Arrow Research search

Author name cluster

Chen Huang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
2 author rows

Possible papers

15

AAMAS Conference 2026 Conference Paper

Cross-Domain Alignment with Fine Geometric Perception for Detail-Preserving Point Cloud Completion

  • Chen Huang
  • Haobo Ma
  • Yan Zhang
  • Chao Yang
  • Jianhua Song

Point cloud completion involves inferring and reconstructing the full structure of an object or scene from incomplete 3D point cloud data. Deep learning-based methods typically use encoder-decoder architecturestolearngeometricpriorsfrompartialinputsforreconstruction. However, these methods often prioritize global features over local geometric details, leading to coarse completions lacking high-frequency information. Sequential application of such models can also cause error accumulation and increased computational costs. To address these issues, we propose CAM-FGP, a Cross-domain Alignment Method with Fine Geometric Perception, designed to enhance structural integrity and restore details, especially in regions with missing geometry. CAM-FGP first employs a Fine Geometry Detail Extraction Network (FGDE) to gather highresolution local details from visible point clouds while integrating low-resolution global information to reinforce the missing areas’ structure. Then, aHierarchicalOptimalTransportNetwork(HOTN) aligns multi-source point cloud distributions, improving the transferability of local geometric features. Lastly, CAM-FGP utilizes a multi-stage hidden state completion and fusion strategy to merge local and global features. This approach preserves continuity, reduces memory-induced information loss, and lowers computational costs. CAM-FGP achieves state-of-the-art performance on several benchmark datasets, demonstrating its superiority in point cloud completion.

JBHI Journal 2026 Journal Article

Group Information Guided Smooth Independent Component Analysis Method for Multi-Subject fMRI Data Analysis

  • Yuhui Du
  • Chen Huang
  • Vince D. Calhoun

Group independent component analysis (ICA) has been extensively used to extract brain functional networks (FNs) and associated neuroimaging measures from multi-subject functional magnetic resonance imaging (fMRI) data. However, the inherent noise in fMRI data can adversely affect the performance of ICA, often leading to noisy FNs and hindering the identification of network-level biomarkers. To address this challenge, we propose a novel method called group information guided smooth independent component analysis (GIG-sICA). Our method effectively generates smoother functional networks with reduced noise and enhanced functional coherence, while preserving intra-subject independence and inter-subject correspondence of FN. Importantly, GIG-sICA is capable of handling different types of noise either separately or in combination. To validate the efficacy of our approach, we conducted comprehensive experiments, comparing GIG-sICA with traditional group ICA methods on both simulated and real fMRI datasets. Experiments on five simulated datasets, generated by adding various types of noise, demonstrate that GIG-sICA produces smoother functional networks with enhanced spatial accuracy. Additionally, experiments on real fMRI data from 137 schizophrenia patients and 144 healthy controls demonstrate that GIG-sICA more effectively captures functionally meaningful brain networks and reveals clearer group differences. Overall, GIG-sICA produces smooth and precise network estimations, supporting the discovery of robust biomarkers at the network level for neuroscience research.

AAMAS Conference 2026 Conference Paper

Multimodal Emotion Recognition in Conversation via Large Language Models and Global-Local Cross-Domain Graphs

  • Haobo Ma
  • Chen Huang
  • Yan Zhang
  • Chao Yang
  • Jianhua Song

Multimodal Emotion Recognition in Conversation (MERC) aims to identify emotions in target utterances using multimodal data and has garnered significant interest due to its applications in conversational AI. Recognition accuracy hinges on effectively integrating multimodal cues and contextual information. However, local noise and global outliers often impair performance, while traditional approaches based on simple feature concatenation struggle to capture complex cross-modal interactions. To address these challenges, we propose LLM-EmoGraph, a novel framework that combines large language models (LLMs) with a global-local cross-domain graph architecture. Specifically, LLM-EmoGraph leverages multimodal masking strategies, a large-scale cross-domain multi-graph pretraining to improve transferability across modalities and graph structures. Then, LLM-EmoGraph further introduces an adaptive dual-scale feature fusion strategy to align semantic features across text, speech, and visual inputs. In addition, a weakly supervised hierarchical emotion classification scheme enhanced by LLMs boosts robustness and accuracy. Experiments on two benchmark datasets show that LLM-EmoGraph significantly outperforms existing methods.

ICML Conference 2025 Conference Paper

Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention

  • Dejia Xu
  • Yifan Jiang 0001
  • Chen Huang
  • Liangchen Song
  • Thorsten Gernoth
  • Liangliang Cao
  • Zhangyang Wang
  • Hao Tang 0001

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are often limited to simple trajectories or lack the ability to generate consistent videos from multiple distinct camera paths for the same scene. To address these limitations, we introduce Cavia, a novel framework for camera-controllable, multi-view video generation, capable of converting an input image into multiple spatiotemporally consistent videos. Our framework extends the spatial and temporal attention modules into view-integrated attention modules, improving both viewpoint and temporal consistency. This flexible design allows for joint training with diverse curated data sources, including scene-level static videos, object-level synthetic multi-view dynamic videos, and real-world monocular dynamic videos. To the best of our knowledge, Cavia is the first framework that enables users to generate multiple videos of the same scene with precise control over camera motion, while simultaneously preserving object motion. Extensive experiments demonstrate that Cavia surpasses state-of-the-art methods in terms of geometric consistency and perceptual quality.

AAAI Conference 2025 Conference Paper

CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation

  • Yuxuan Wang
  • Yijun Liu
  • Fei Yu
  • Chen Huang
  • Kexin Li
  • Zhiguo Wan
  • Wanxiang Che
  • Hongyang Chen

Despite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias in the images makes these datasets unsuitable for evaluating VLMs in Chinese culture. To remedy this issue, we present a new Chinese Vision-Language Understanding Evaluation (CVLUE) benchmark dataset, where the selection of object categories and images is entirely driven by Chinese native speakers, ensuring that the source images are representative of Chinese culture. The benchmark contains four distinct VL tasks ranging from image-text retrieval to visual question answering, visual grounding and visual dialogue. We present a detailed statistical analysis of CVLUE and provide a baseline performance analysis with several open-source multilingual VLMs on CVLUE and its English counterparts to reveal their performance gap between English and Chinese. Our in-depth category-level analysis reveals a lack of Chinese cultural knowledge in existing VLMs. We also find that fine-tuning on Chinese culture-related VL datasets effectively enhances VLMs' understanding of Chinese culture.

AAAI Conference 2025 Conference Paper

LEGEND: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets

  • Duanyu Feng
  • Bowen Qin
  • Chen Huang
  • Youcheng Huang
  • Zheng Zhang
  • Wenqiang Lei

The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture the fine-grained nuances of harmful and harmless responses. This motivates the need to develop the datasets involving preference margins, which accurately quantify how harmless one response is compared to another. In this paper, we take the first step to propose an effective and cost-efficient framework to promote the margin-enhanced preference dataset development. Our framework, Legend, Leverages rEpresentation enGineering to annotate preferENce Datasets. It constructs the specific direction within the LLM's embedding space that represents safety. By leveraging this safety direction, Legend can then leverage the semantic distances of paired responses along this direction to annotate margins automatically. We experimentally demonstrate our effectiveness in both reward modeling and harmless alignment for LLMs. Legend also stands out for its efficiency, requiring only the inference time rather than additional training. This efficiency allows for easier implementation and scalability, making Legend particularly valuable for practical applications in aligning LLMs with safe conversations.

TMLR Journal 2025 Journal Article

Pre-Training Representations of Binary Code Using Contrastive Learning

  • Yifan Zhang
  • Chen Huang
  • Yueke Zhang
  • Huajie Shao
  • Kevin Leach
  • Yu Huang

Binary code analysis and comprehension is critical to applications in reverse engineering and computer security tasks where source code is not available. Unfortunately, unlike source code, binary code lacks semantics and is more difficult for human engineers to understand and analyze. In this paper, we present ContraBin, a contrastive learning technique that integrates source code and comment information along with binaries to create an embedding capable of aiding binary analysis and comprehension tasks. Specifically, we present three components in ContraBin: (1) a primary contrastive learning method for initial pre-training, (2) a simplex interpolation method to integrate source code, comments, and binary code, and (3) an intermediate representation learning algorithm to train a binary code embedding. We further analyze the impact of human-written and synthetic comments on binary code comprehension tasks, revealing a significant performance disparity. While synthetic comments provide substantial benefits, human-written comments are found to introduce noise, even resulting in performance drops compared to using no comments. These findings reshape the narrative around the role of comment types in binary code analysis. We evaluate the effectiveness of ContraBin through four indicative downstream tasks related to binary code: algorithmic functionality classification, function name recovery, code summarization, and reverse engineering. The results show that ContraBin considerably improves performance on all four tasks, measured by accuracy, mean of average precision, and BLEU scores as appropriate. ContraBin is the first language representation model to incorporate source code, binary code, and comments into contrastive code representation learning and is intended to contribute to the field of binary code analysis. The dataset used in this study is available for further research.

NeurIPS Conference 2025 Conference Paper

ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG

  • Yikuan Hu
  • Jifeng Zhu
  • Lanrui Tang
  • Chen Huang

Knowledge graphs (KGs), with their structured representation capabilities, offer promising avenue for enhancing Retrieval Augmented Generation (RAG) systems, leading to the development of KG-RAG systems. Nevertheless, existing methods often struggle to achieve effective synergy between system effectiveness and cost efficiency, leading to neither unsatisfying performance nor excessive LLM prompt tokens and inference time. To this end, this paper proposes REMINDRAG, which employs an LLM-guided graph traversal featuring node exploration, node exploitation, and, most notably, memory replay, to improve both system effectiveness and cost efficiency. Specifically, REMINDRAG memorizes traversal experience within KG edge embeddings, mirroring the way LLMs "memorize" world knowledge within their parameters, but in a train-free manner. We theoretically and experimentally confirm the effectiveness of REMINDRAG, demonstrating its superiority over existing baselines across various benchmark datasets and LLM backbones. Our code is available at https: //github. com/kilgrims/ReMindRAG.

NeurIPS Conference 2024 Conference Paper

Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP

  • Chen Huang
  • Skyler Seto
  • Samira Abnar
  • David Grangier
  • Navdeep Jaitly
  • Josh Susskind

Large pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e. g. , satellite imagery) or fine-grained classification (e. g. , car models) where the visual concepts are unseen or under-represented during pretraining. Prompt learning offers a parameter-efficient finetuning framework that can adapt CLIP to downstream tasks even when limited annotation data are available. In this paper, we improve prompt learning by distilling the textual knowledge from natural language prompts (either human- or LLM-generated) to provide rich priors for those under-represented concepts. We first obtain a prompt ``summary'' aligned to each input image via a learned prompt aggregator. Then we jointly train a prompt generator, optimized to produce a prompt embedding that stays close to the aggregated summary while minimizing task loss at the same time. We dub such prompt embedding as Aggregate-and-Adapted Prompt Embedding (AAPE). AAPE is shown to be able to generalize to different downstream data distributions and tasks, including vision-language understanding tasks (e. g. , few-shot classification, VQA) and generation tasks (image captioning) where AAPE achieves competitive performance. We also show AAPE is particularly helpful to handle non-canonical and OOD examples. Furthermore, AAPE learning eliminates LLM-based inference cost as required by baselines, and scales better with data and LLM model size.

NeurIPS Conference 2024 Conference Paper

How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks

  • Etai Littwin
  • Omid Saremi
  • Madhu Advani
  • Vimal Thilak
  • Preetum Nakkiran
  • Chen Huang
  • Joshua Susskind

Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architectures (JEPAs) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes with a lightweight predictor network. This is contrasted with the Masked Auto Encoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in ambient space rather than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative). In this work, we seek to understand the mechanism behind this empirical observation by analyzing deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high influence features, or features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice.

EAAI Journal 2024 Journal Article

Incorporating environmental knowledge embedding and spatial-temporal graph attention networks for inland vessel traffic flow prediction

  • Chen Huang
  • Deshan Chen
  • Tengze Fan
  • Bing Wu
  • Xinping Yan

Accurate prediction of vessel traffic flow is crucial for maritime regulatory authorities and transportation planners. However, existing methods for inland vessel traffic flow prediction often overlook spatial correlation and environmental influences, leading to suboptimal accuracy. To address this issue, we propose an innovative model that incorporates environmental knowledge embedding and a spatial-temporal information extraction module. Our approach involves constructing a vessel traffic knowledge graph, embedding traffic flow through knowledge representation learning. The spatial-temporal information extraction module is leveraged to analyze inherent periodicity and external spatial relationships in vessel traffic flow. Extensive experiments on real-world datasets demonstrate that our approach significantly enhances predictive accuracy. In comparison to the second-ranked model, our approach achieves a decrease of 0. 46 in mean absolute error, a decrease of 0. 64 in root mean squared error, an increase of 3. 06% in accuracy, and an increase of 0. 07 in R-squared. Furthermore, our approach excels in upstream, downstream and long-term prediction, and displays robustness in handling noisy data.

AAAI Conference 2024 Conference Paper

Towards Equipping Transformer with the Ability of Systematic Compositionality

  • Chen Huang
  • Peixin Qin
  • Wenqiang Lei
  • Jiancheng Lv

One of the key factors in language productivity and human cognition is the ability of Systematic Compositionality, which refers to understanding composed, unseen examples of seen primitives. However, recent evidence reveals that the Transformers have difficulty in generalizing the composed context based on the seen primitives. To this end, we take the first step to propose a compositionality-aware Transformer called CAT and two novel pre-training tasks to facilitate the systematic compositionality. We tentatively provide a successful implementation of a multi-layer CAT on the basis of the especially popular BERT. The experimental results demonstrate that CAT outperforms baselines on compositionality-aware tasks with minimal impact on effectiveness on standardized language understanding tasks.

EAAI Journal 2023 Journal Article

Adaptive cylinder vector particle swarm optimization with differential evolution for UAV path planning

  • Chen Huang
  • Xiangbing Zhou
  • Xiaojuan Ran
  • Jiamiao Wang
  • Huayue Chen
  • Wu Deng

Particle swarm optimization (PSO) algorithm has a potential to solve route planning problem for unmanned aerial vehicle (UAV). However, the traditional PSO algorithm is easy to fall into local optimum under the complicated environments with multiple threats. In order to improve the performance in different complicated environments, a novel and effective PSO algorithm with adaptive adjustment of the parameters, cylinder vector and different evolution operator, named ACVDEPSO, is proposed and demonstrated to be effective for route planning problem for UAV. In the proposed ACVDEPSO, the velocity of the particle is converted to its cylinder vector for the convenience of the path search. It is worth highlighting that the parameters of ACVDEPSO algorithm are automatically chosen by the time and the fitness values of the particles. Furthermore, a challenger based on differential evolution operator is introduced to reduce the probability of falling into local optimum and accelerate the algorithm convergence speed. The simulation experiments have been conducted in real digital elevation model (DEM) maps to test the performance of the ACVDEPSO. The experiment results validate that the optimization performance of the ACVDEPSO outperforms the other comparison methods, which can efficiently generate a higher quality path for UAV under the complicated 3D environments.

JBHI Journal 2021 Journal Article

Deep Semantic Segmentation Feature-Based Radiomics for the Classification Tasks in Medical Image Analysis

  • Bingsheng Huang
  • Junru Tian
  • Hongyuan Zhang
  • Zixin Luo
  • Jing Qin
  • Chen Huang
  • Xueping He
  • Yanji Luo

Recently, an emerging trend in medical image classification is to combine radiomics framework with deep learning classification network in an integrated system. Although this combination is efficient in some tasks, the deep learning-based classification network is often difficult to capture an effective representation of lesion regions, and prone to face the challenge of overfitting, leading to unreliable features and inaccurate results, especially when the sizes of the lesions are small or the training dataset is small. In addition, these combinations mostly lack an effective feature selection mechanism, which makes it difficult to obtain the optimal feature selection. In this paper, we introduce a novel and effective deep semantic segmentation feature-based radiomics (DSFR) framework to overcome the above-mentioned challenges, which consists of two modules: the deep semantic feature extraction module and the feature selection module. Specifically, the extraction module is utilized to extract hierarchical semantic features of the lesions from a trained segmentation network. The feature selection module aims to select the most representative features by using a novel feature similarity adaptation algorithm. Experiments are extensively conducted to evaluate our method in two clinical tasks: the pathological grading prediction in pancreatic neuroendocrine neoplasms (pNENs), and the prediction of thrombolytic therapy efficacy in deep venous thrombosis (DVT). Experimental results on both tasks demonstrate that the proposed method consistently outperforms the state-of-the-art approaches by a large margin.

NeurIPS Conference 2016 Conference Paper

Local Similarity-Aware Deep Feature Embedding

  • Chen Huang
  • Chen Change Loy
  • Xiaoou Tang

Existing deep embedding methods in vision tasks are capable of learning a compact Euclidean space from images, where Euclidean distances correspond to a similarity metric. To make learning more effective and efficient, hard sample mining is usually employed, with samples identified through computing the Euclidean feature distance. However, the global Euclidean distance cannot faithfully characterize the true feature similarity in a complex visual feature space, where the intraclass distance in a high-density region may be larger than the interclass distance in low-density regions. In this paper, we introduce a Position-Dependent Deep Metric (PDDM) unit, which is capable of learning a similarity metric adaptive to local feature structure. The metric can be used to select genuinely hard samples in a local neighborhood to guide the deep embedding learning in an online and robust manner. The new layer is appealing in that it is pluggable to any convolutional networks and is trained end-to-end. Our local similarity-aware feature embedding not only demonstrates faster convergence and boosted performance on two complex image retrieval datasets, its large margin nature also leads to superior generalization results under the large and open set scenarios of transfer learning and zero-shot learning on ImageNet 2010 and ImageNet-10K datasets.

v2026.09.13