Arrow Research search

Author name cluster

Fang Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
1 author row

Possible papers

11

EAAI Journal 2026 Journal Article

Cross-stain knowledge distillation for low-cost lung cancer programmed death ligand-1 assessment with multi-granularity multiple instance learning

  • Yi Shi
  • Chong Ge
  • Fang Zhao
  • Anli Zhang
  • Ao Li
  • Haibo Wu
  • Minghui Wang

Accurately assessing programmed death ligand-1 (PD-L1) status, recognizing patients potentially responsive to immunotherapies. Since immunohistochemistry (IHC) staining is gold standard in identifying molecular biomarker but often expensive and unattainable, routine hematoxylin and eosin (H&E) staining offers a low-cost alternative. However, H&E images primarily reveal morphological knowledge and inherently lack PD-L1-related molecular information, resulting in a severe mono-stain knowledge limitation. Additionally, most existing approaches analyze gigapixel whole-slide images at only a single magnification, which fails to unravel complex pathological information across granularities, leading to a significant uni-granularity information limitation. Therefore, we propose an innovative cross-stain knowledge distillation with multi-granularity framework, namely CroSMuG. First, to alleviate uni-granularity information limitation, a new multi-granularity multiple instance learning framework is introduced. This is based on macro-micro dual branches, which comprises a macro-branch and a micro-branch to extract global and local pathological information. Furthermore, we develop a novel cross-stain knowledge distillation strategy featuring triple-united distillation loss. Specifically, this approach introduces globality-, locality- and task-aware knowledge distillation, enabling the H&E-based predictive network as a student to learn crucial molecular knowledge from an IHC teacher network. Extensive experiments are conducted on diverse real-world datasets from multiple medical centers, and CroSMuG achieves the superior performance with area under the curve (AUC) of 83. 6 % on internal dataset and 81. 2 % on external dataset. These results highlight the generalizability of CroSMuG for accurate H&E-based PD-L1 assessment, offering the potential for practical applications in lung cancer immunotherapy decision-making in clinical practices.

AAAI Conference 2026 Conference Paper

HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment

  • Ruijia Wu
  • Ping Chen
  • Fei Shen
  • Shaoan Zhao
  • Qiang Hui
  • Huanlin Gao
  • Ting Lu
  • Zhaoxiang Liu

Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions.

AAAI Conference 2025 Conference Paper

Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly Detection

  • Fenfang Tao
  • Guo-Sen Xie
  • Fang Zhao
  • Xiangbo Shu

Few-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Existing FSAD methods usually find anomalies by directly designing complex text prompts to align them with visual features under the prevailing large vision-language model paradigm. However, these methods, almost always, neglect intrinsic contextual information in visual features, e.g., the interaction relationships between different vision layers, which is an important clue for detecting anomalies comprehensively. To this end, we propose a kernel-aware graph prompt learning framework, termed as KAG-prompt, by reasoning the cross-layer relations among visual features for FSAD. Specifically, a kernel-aware hierarchical graph is built by taking the different layer features focusing on anomalous regions of different sizes as nodes, meanwhile, the relationships between arbitrary pairs of nodes stand for the edges of the graph. By message passing over this graph, KAG-prompt can capture cross-layer contextual information, thus leading to more accurate anomaly prediction. Moreover, to integrate the information of multiple important anomaly signals in the prediction map, we propose a novel image-level scoring method based on multi-level information fusion. Extensive experiments on MVTecAD and VisA datasets show that KAG-prompt achieves state-of-the-art FSAD results for image-level/pixel-level anomaly detection.

NeurIPS Conference 2025 Conference Paper

LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation

  • Huanlin Gao
  • Ping Chen
  • Fuyuan Shi
  • Chao Tan
  • Zhaoxiang Liu
  • Fang Zhao
  • Kai Wang
  • Shiguo Lian

We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between accelerated and original videos. To address this issue, we formulate cache scheduling as a directed graph with error-weighted edges and introduce a Lexicographic Minimax Path Optimization strategy that explicitly bounds the worst-case path error. This approach substantially improves the consistency of global content and style across generated frames. Extensive experiments on multiple text-to-video benchmarks demonstrate that LeMiCa delivers dual improvements in both inference speed and generation quality. Notably, our method achieves a 2. 9× speedup on the Latte model and reaches an LPIPS score of 0. 05 on Open-Sora, outperforming prior caching techniques. Importantly, these gains come with minimal perceptual quality degradation, making LeMiCa a robust and generalizable paradigm for accelerating diffusion-based video generation. We believe this approach can serve as a strong foundation for future research on efficient and reliable video synthesis.

JBHI Journal 2025 Journal Article

MIF: Multi-Shot Interactive Fusion Model for Cancer Survival Prediction Using Pathological Image and Genomic Data

  • Yi Shi
  • Minghui Wang
  • Honglei Liu
  • Fang Zhao
  • Ao Li
  • Xun Chen

Accurate cancer survival prediction is crucial for oncologists to determine therapeutic plan, which directly influences the treatment efficacy and survival outcome of patient. Recently, multimodal fusion-based prognostic methods have demonstrated effectiveness for survival prediction by fusing diverse cancer-related data from different medical modalities, e. g. , pathological images and genomic data. However, these works still face significant challenges. First, most approaches attempt multimodal fusion by simple one-shot fusion strategy, which is insufficient to explore complex interactions underlying in highly disparate multimodal data. Second, current methods for investigating multimodal interactions face the capability-efficiency dilemma, which is the difficult balance between powerful modeling capability and applicable computational efficiency, thus impeding effective multimodal fusion. In this study, to encounter these challenges, we propose an innovative multi-shot interactive fusion method named MIF for precise survival prediction by utilizing pathological and genomic data. Particularly, a novel multi-shot fusion framework is introduced to promote multimodal fusion by decomposing it into successive fusing stages, thus delicately integrating modalities in a progressive way. Moreover, to address the capacity-efficiency dilemma, various affinity-based interactive modules are introduced to synergize the multi-shot framework. Specifically, by harnessing comprehensive affinity information as guidance for mining interactions, the proposed interactive modules can efficiently generate low-dimensional discriminative multimodal representations. Extensive experiments on different cancer datasets unravel that our method not only successfully achieves state-of-the-art performance by performing effective multimodal fusion, but also possesses high computational efficiency compared to existing survival prediction methods.

NeurIPS Conference 2020 Conference Paper

Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View Consistency

  • Fang Zhao
  • Shengcai Liao
  • Kaihao Zhang
  • Ling Shao

This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation of human image and preserve the spatial distribution of semantic parts. Meanwhile, in order to improve the prediction for textures of invisible parts, we explicitly enforce the consistency across different views of the same subject by exchanging the textures predicted by two views to render images during training. The perception loss and total variation regularization are optimized to maximize the similarity between rendered and input images, which does not necessitate extra 3D texture supervision. Experimental results on pedestrian images and fashion photos demonstrate that our method can produce higher quality textures with convincing details than other texture generation methods.

AAAI Conference 2019 Conference Paper

Look across Elapse: Disentangled Representation Learning and Photorealistic Cross-Age Face Synthesis for Age-Invariant Face Recognition

  • Jian Zhao
  • Yu Cheng
  • Yi Cheng
  • Yang Yang
  • Fang Zhao
  • Jianshu Li
  • Hengzhu Liu
  • Shuicheng Yan

Despite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages still remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intraclass variations. As opposed to current techniques for ageinvariant face recognition, which either directly extract ageinvariant features for recognition, or first synthesize a face that matches target age before feature extraction, we argue that it is more desirable to perform both tasks jointly so that they can leverage each other. To this end, we propose a deep Age-Invariant Model (AIM) for face recognition in the wild with three distinct novelties. First, AIM presents a novel unified deep architecture jointly performing cross-age face synthesis and recognition in a mutual boosting way. Second, AIM achieves continuous face rejuvenation/aging with remarkable photorealistic and identity-preserving properties, avoiding the requirement of paired data and the true age of testing samples. Third, we develop effective and novel training strategies for end-to-end learning the whole deep architecture, which generates powerful age-invariant face representations explicitly disentangled from the age variation. Extensive experiments on several cross-age datasets (MORPH, CACD and FG-NET) demonstrate the superiority of the proposed AIM model over the state-of-the-arts. Benchmarking our model on one of the most popular unconstrained face recognition datasets IJB-C additionally verifies the promising generalizability of AIM in recognizing faces in the wild.

IJCAI Conference 2019 Conference Paper

Multi-Prototype Networks for Unconstrained Set-based Face Recognition

  • Jian Zhao
  • Jianshu Li
  • Xiaoguang Tu
  • Fang Zhao
  • Yuan Xin
  • Junliang Xing
  • Hengzhu Liu
  • Shuicheng Yan

In this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e. g. , varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts.

IS Journal 2018 Journal Article

Activity Recognition for a Smartphone and Web-Based Human Mobility Sensing System

  • Youngsung Kim
  • Ajinkya Ghorpade
  • Fang Zhao
  • Francisco C. Pereira
  • P. Christopher Zegras
  • Moshe Ben-Akiva

Activity-based models in transport modeling and prediction are built from a large number of observed trips and their purposes. However, data acquired through traditional interview-based travel surveys is often inaccurate and insufficient. Recently, a human mobility sensing system, called Future Mobility Survey (FMS), was developed and used to collect travel data from more than 1, 000 participants. FMS combines a smartphone and interactive web interface in order to better infer users activities and patterns. This paper presents a model that infers an activity at a certain location. We propose to generate a set of predictive features based on spatial, temporal, transitional, and environmental contexts with an appropriate quantization. In order to improve the generalization performance of the proposed model, we employ a robust approach with ensemble learning. Empirical results using FMS data demonstrate that the proposed method contributes significantly to providing accurate activity estimates for the user in our travel-sensing application.

NeurIPS Conference 2017 Conference Paper

Dual-Agent GANs for Photorealistic and Identity Preserving Profile Face Synthesis

  • Jian Zhao
  • Lin Xiong
  • Panasonic Karlekar Jayashree
  • Jianshu Li
  • Fang Zhao
  • Zhecan Wang
  • Panasonic Sugiri Pranata
  • Panasonic Shengmei Shen

Synthesizing realistic profile faces is promising for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by populating samples with extreme poses and avoiding tedious annotations. However, learning from synthetic faces may not achieve the desired performance due to the discrepancy between distributions of the synthetic and real face images. To narrow this gap, we propose a Dual-Agent Generative Adversarial Network (DA-GAN) model, which can improve the realism of a face simulator's output using unlabeled real faces, while preserving the identity information during the realism refinement. The dual agents are specifically designed for distinguishing real v. s. fake and identities simultaneously. In particular, we employ an off-the-shelf 3D face model as a simulator to generate profile face images with varying poses. DA-GAN leverages a fully convolutional network as the generator to generate high-resolution images and an auto-encoder as the discriminator with the dual agents. Besides the novel architecture, we make several key modifications to the standard GAN to preserve pose and texture, preserve identity and stabilize training process: (i) a pose perception loss; (ii) an identity perception loss; (iii) an adversarial loss with a boundary equilibrium regularization term. Experimental results show that DA-GAN not only presents compelling perceptual results but also significantly outperforms state-of-the-arts on the large-scale and challenging NIST IJB-A unconstrained face recognition benchmark. In addition, the proposed DA-GAN is also promising as a new approach for solving generic transfer learning problems more effectively.

NeurIPS Conference 2013 Conference Paper

Relevance Topic Model for Unstructured Social Group Activity Recognition

  • Fang Zhao
  • Yongzhen Huang
  • Liang Wang
  • Tieniu Tan

Unstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a relevance topic model" for jointly learning meaningful mid-level representations upon bag-of-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i. e. , Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos. "

v2026.09.13