Arrow Research search

Author name cluster

Xiang Gao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

AAAI Conference 2026 Conference Paper

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

  • Jun Sun
  • Xinxin Zhang
  • Simin Hong
  • Jian Zhu
  • Xiang Gao

Multimodal learning, while contributing to numerous success stories across various fields, faces the challenge of prohibitively expensive manual annotation. To address the scarcity of annotated data, a popular solution is unsupervised domain adaptation, which has been extensively studied in unimodal settings yet remains less explored in multimodal settings. In this paper, we investigate heterogeneous multimodal domain adaptation, where the primary challenge is the varying domain shifts of different modalities from the source to the target domain. We first introduce the information bottleneck method to learn representations for each modality independently, and then match the source and target domains in the representation space with correlation alignment. To balance the domain alignment of all modalities, we formulate the problem as a multi-objective task, aiming for a Pareto optimal solution. By exploiting the properties specific to our model, the problem can be simplified to a quadratic programming problem. Further approximation yields a closed-form solution, leading to an efficient modality-balanced multimodal domain adaptation algorithm. The proposed method features Balanced multi-objective optimization for multimodal domain adaptation, termed Boomda. Extensive empirical results showcase the effectiveness of the proposed approach and demonstrate that Boomda outperforms the competing schemes.

EAAI Journal 2026 Journal Article

Lightweight method of foreign matter detection in coal conveying based on improved you only look once version 8 and embedded equipment

  • Guanfeng Du
  • Hongzheng Zhang
  • Yupeng Luo
  • Zhibo Bao
  • Zhiwei Li
  • Mingxin Zhou
  • Zhelin Liu
  • Shengxian Cao

During the process of conveying pulverized coal, the mixed foreign matter will not only affect the combustion efficiency of pulverized coal, but also cause safety accidents in coal conveying equipment. Therefore, it is very important to monitor the foreign matter in the process of conveying coal. Due to the limited scope of the actual conveying site, embedded equipment is needed for inspection. Aiming at the computing and memory challenges of embedded equipment, an improved lightweight YOLOv8 (you only look once version 8) algorithm is proposed. In the backbone of the algorithm, cross stage partial with 2 convolutions and lightweight PoolFormer (C2f_LPF) module is used to extract lightweight features, and foreign matter information is extracted by using multi-scale concerns in cross stage partial with deformable convolution (CSPDC) module. Then the part of the feature aggregation (PFA) module of the neck is used for lightweight feature fusion. The proposed C2f_LPF+CSPDC+PFA combination realizes a more balanced optimization of lightweight performance and detection accuracy and provides a solution to the contradiction between the limitation of computing resources of embedded equipment and the demand for real-time detection of accuracy in coal conveying. A self-made datasets contain scrap iron, stones, wooden stick and branch are trained and compared with faster region-based convolutional neural network (Faster R-CNN) and YOLO (you only look once) series algorithm on computer and embedded equipment. The datasets consist of 612 images with 4413 examples of foreign matter, which are collected on a lab-scale self-made coal conveying platform. The mean average precision (mAP), Giga floating-point operations per second (GFLOPS), parameters and frame per second (FPS) on embedded equipment are 0. 963, 6. 1, 2. 44 million and 37. 04, respectively. Compared with the original YOLOv8, the computation and parameters are reduced by 24. 7 % and 18. 9 % respectively, and the FPS is improved by 29. 6 %. By contrast, it also has better results than other algorithms, which is well compatible with the configuration requirements of embedded equipment and achieves a good balance between precision and speed. This shows a promising performance on a lab-scale platform and may be extended to real industrial lines after further validation.

EAAI Journal 2026 Journal Article

Rethinking the local constraints: Geometric continuity regularization for image alignment

  • Yinqi Chen
  • Yangting Zheng
  • Peiwen Li
  • Weijian Luo
  • Shuo Kang
  • Xiang Gao
  • Chao Liu
  • Shuo Zhang

In image alignment, existing studies frequently neglect the modeling of featureless areas where reliable features are inherently absent. While indirect strategies, such as adding more geometric features, have been used to reduce such regions, they are limited by the natural variability of scenes. Instead, directly modeling these areas allows local consistency constraints to propagate transformations from feature-rich to featureless regions. However, existing local consistency constraints rely solely on parametric continuity (C1), which can cause excessive smoothness and distortion due to the excessive constraints on parameters. In contrast, geometric continuity (G1) relaxes parameter constraints and ensures visual accuracy, leading to results with lower distorted energy. Thus, this paper, for the first time, rigorously examines the rationale of local constraints, validates their capacity for featureless-region modeling, and theoretically demonstrates that G1 continuity effectively minimizes distortion. Building on these analyses, we introduce G1 continuity regularization; to enforce this property, the regularization term directly penalizes deviations from collinearity at mesh vertices or within network-learned transformations. Compared with existing approaches, our method achieves markedly superior performance.

AAAI Conference 2026 Conference Paper

WIET: Harmonizing Group-aware Model Weighting and Worker Allocation for Ensemble Temporal Prediction MaaS

  • Binbin Feng
  • Shikun He
  • Yingxin Wang
  • Pengwei Wang
  • Xiang Gao
  • Zhijun Ding

Ensemble Temporal Prediction Model-as-a-Service (ETP-MaaS) has become crucial in fields like financial modeling and cloud monitoring. Existing solutions fail to co-optimally address a two-fold challenge of dynamic collaboration and heterogeneity, treating models as independent entities and employing simplistic worker allocation rules. However, at the model level, data volatility means that optimal performance requires identifying and weighting constantly shifting subgroups of base models, not just individual ones; at the system level, these model groups must be efficiently mapped to a pool of heterogeneous and dynamically available workers. To this end, we introduce WIET, an efficient ETP-MaaS system that co-optimizes model weighting and worker allocation. For adaptive weighting, WIET identifies evolving group behaviors among base models and propose a novel group temporal locality-enhanced weighting method. Additionally, WIET develops an efficient, multi-dimensional worker allocation method powered by hybrid heuristic optimization, effectively reducing bottlenecks and resource waste. Experiments show WIET consistently outperforms state-of-the-art methods in terms of accuracy, latency, and resource usage across various workloads and tasks.

TMLR Journal 2025 Journal Article

Enhancing Diversity in Text-to-Image Generation without Compromising Fidelity

  • Jiazhi Li
  • Mi Zhou
  • Mahyar Khayatkhoei
  • Jingyu Shi
  • Xiang Gao
  • Jiageng Zhu
  • Hanchen Xie
  • Xiyun Song

Effective text-to-image generation must synthesize images that are both realistic in appearance (sample fidelity) and have sufficient variations (sample diversity). Diffusion models have achieved promising results in generating high-fidelity images based on textual prompts, and recently, several diversity-focused works have been proposed to improve their demographic diversity by enforcing the generation of samples from various demographic groups. However, another essential aspect of diversity, sample diversity—which enhances prompt reusability to generate creative samples that reflect real-world variability—has been largely overlooked. Specifically, how to generate images that have sufficient demographic and sample diversity while preserving sample fidelity remains an open problem because increasing diversity comes at the cost of reduced fidelity in existing works. To address this problem, we first propose a bimodal low-rank adaptation of pretrained diffusion models, which decouples the text-to-image conditioning, and then propose a lightweight bimodal guidance method that introduces additional diversity to the generation process using reference images retrieved through a fairness strategy by separately controlling the strength of text and image conditioning. We conduct extensive experiments to demonstrate the effectiveness of our method in enhancing demographic diversity (Intersectional Diversity (Shrestha et al., 2024)) by 2.47× and sample diversity (Recall (Kynkäänniemi et al., 2019)) by 1.45× while preserving sample fidelity (Precision (Kynkäänniemi et al., 2019)) compared to the baseline diffusion model.

NeurIPS Conference 2025 Conference Paper

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

  • Siting Li
  • Xiang Gao
  • Simon Du

While an image is worth more than a thousand words, only a few provide crucial information for a given task and thus should be focused on. In light of this, ideal text-to-image (T2I) retrievers should prioritize specific visual attributes relevant to queries. To evaluate current retrievers on handling attribute-focused queries, we build COCO-Facet, a COCO-based benchmark with 9, 112 queries about diverse attributes of interest. We find that CLIP-like retrievers, which are widely adopted due to their efficiency and zero-shot ability, have poor and imbalanced performance, possibly because their image embeddings focus on global semantics and subjects while leaving out other details. Notably, we reveal that even recent Multimodal Large Language Model (MLLM)-based, stronger retrievers with a larger output dimension struggle with this limitation. Hence, we hypothesize that retrieving with general image embeddings is suboptimal for performing such queries. As a solution, we propose to use promptable image embeddings enabled by these multimodal retrievers, which boost performance by highlighting required attributes. Our pipeline for deriving such embeddings generalizes across query types, image pools, and base retriever architectures. To enhance real-world applicability, we offer two acceleration strategies: Pre-processing promptable embeddings and using linear approximations. We show that the former yields a 15\% improvement in Recall@5 when prompts are predefined, while the latter achieves an 8\% improvement when prompts are only available during inference.

NeurIPS Conference 2025 Conference Paper

UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss

  • Zhichao Wang
  • Xinhai Chen
  • Qinglin Wang
  • Xiang Gao
  • Qingyang Zhang
  • Menghan Jia
  • Xiang Zhang
  • Jie Liu

Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-varying regions, enhancing both simulation accuracy and computational efficiency. However, traditional approaches suffer from high computational complexity and geometric inflexibility, limiting their applicability, and existing supervised learning-based approaches face challenges in zero-shot generalization across diverse PDEs and mesh topologies. In this paper, we present an $\textbf{U}$nsupervised and $\textbf{G}$eneralizable $\textbf{M}$esh $\textbf{M}$ovement $\textbf{N}$etwork (UGM2N). We first introduce unsupervised mesh adaptation through localized geometric feature learning, eliminating the dependency on pre-adapted meshes. We then develop a physics-constrained loss function, M-Uniform loss, that enforces mesh equidistribution at the nodal level. Experimental results demonstrate that the proposed network exhibits equation-agnostic generalization and geometric independence in efficient mesh adaptation. It demonstrates consistent superiority over existing methods, including robust performance across diverse PDEs and mesh geometries, scalability to multi-scale resolutions and guaranteed error reduction without mesh tangling.

ICRA Conference 2024 Conference Paper

Bi 2 Lane: Bi-Directional Temporal Refinement with Bi-Level Feature Aggregation for 3D Lane Detection

  • Chengxin Li
  • Yihui Hu
  • Zewen Zheng
  • Xiang Gao
  • Yongqiang Mou
  • Peng Nie
  • Jun Li

Monocular 3D lane detection has recently received increasing research attention in autonomous driving due to its application effectiveness and simplicity. However, depending solely on the limited semantic information from a single image makes current monocular detection methods unable to deal with complex scenarios, such as occluded, blurred, and unaligned scenes. In this study, we introduce an end-to-end framework named Bi 2 Lane which models temporal dependency in a continuous sequence. It recurrently utilizes detected lanes within historical frames as prior information to achieve robust lane detection. Additionally, Bi 2 Lane employs temporal reverse refinement together with temporal forward refinement to achieve bi-directional temporal refinement (BDTR) while maintaining a robust temporal dependency. For the refined features of different frames, we design a bi-level feature aggregation module (BLFA) to fuse them in both point-level and line-level manners, enabling a comprehensive feature representation to deal with complicated road scenes. Extensive experiments conducted on the OpenLane dataset demonstrate the superiority of Bi 2 Lane, achieving a notable F1 score of 63. 8% using a simple ResNet50 backbone, surpassing the performance of existing state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Customizing Language Model Responses with Contrastive In-Context Learning

  • Xiang Gao
  • Kamalika Das

Large language models (LLMs) are becoming increasingly important for machine learning applications. However, it can be challenging to align LLMs with our intent, particularly when we want to generate content that is preferable over others or when we want the LLM to respond in a certain style or tone that is hard to describe. To address this challenge, we propose an approach that uses contrastive examples to better describe our intent. This involves providing positive examples that illustrate the true intent, along with negative examples that show what characteristics we want LLMs to avoid. The negative examples can be retrieved from labeled data, written by a human, or generated by the LLM itself. Before generating an answer, we ask the model to analyze the examples to teach itself what to avoid. This reasoning step provides the model with the appropriate articulation of the user's need and guides it towards generting a better answer. We tested our approach on both synthesized and real-world datasets, including StackExchange and Reddit, and found that it significantly improves performance compared to standard few-shot prompting.

AAAI Conference 2024 Conference Paper

Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation

  • Xiang Gao
  • Zhengbo Xu
  • Junhan Zhao
  • Jiaying Liu

Recently, text-to-image diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing flexible image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework contributing a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequency-domain filtering module based on Discrete Cosine Transform, which extracts image features carrying different DCT spectral bands to control the text-to-image generation process of the Latent Diffusion Model, realizing versatile I2I applications including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related methods, FCDiffusion establishes a unified text-driven I2I framework suiting diverse I2I application scenarios simply by switching among different frequency control branches. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/.

AAAI Conference 2024 Conference Paper

PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature Alignment

  • Zewen Zheng
  • Xuemin Zhang
  • Yongqiang Mou
  • Xiang Gao
  • Chengxin Li
  • Guoheng Huang
  • Chi-Man Pun
  • Xiaochen Yuan

Monocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. However, the perspective geometric distortion between FV and 3D space in this FV-based approach introduces extremely dense anchor designs, which ultimately leads to confusing lane representations. In this paper, we introduce a novel prior-guided perspective on lane detection and propose an end-to-end framework named PVALane, which utilizes 2D prior knowledge to achieve precise and efficient 3D lane detection. Since 2D lane predictions can provide strong priors for lane existence, PVALane exploits FV features to generate sparse prior anchors with potential lanes in 2D space. These dynamic prior anchors help PVALane to achieve distinct lane representations and effectively improve the precision of PVALane due to the reduced lane search space. Additionally, by leveraging these prior anchors and representing lanes in both FV and bird-eye-viewed (BEV) spaces, we effectively align and merge semantic and geometric information from FV and BEV features. Extensive experiments conducted on the OpenLane and ONCE-3DLanes datasets demonstrate the superior performance of our method compared to existing state-of-the-art approaches and exhibit excellent robustness.

IJCAI Conference 2023 Conference Paper

Divide Rows and Conquer Cells: Towards Structure Recognition for Large Tables

  • Huawen Shen
  • Xiang Gao
  • Jin Wei
  • Liang Qiao
  • Yu Zhou
  • Qiang Li
  • Zhanzhan Cheng

Recent advanced Table Structure Recognition (TSR) models adopt image-to-text solutions to parse table structure. These methods can be formulated as image caption problem, i. e. , input a single-table image and output table structure description in a specific text format, e. g. , HTML. With the impressive success of Transformer in text generation tasks, these methods use Transformer architecture to predict HTML table text in an autoregressive manner. However, tables always emerge with a large variety of shapes and sizes. Autoregressive models usually suffer from the error accumulation problem as the length of predicted text increases, which results in unsatisfactory performance for large tables. In this paper, we propose a novel image-to-text based TSR method that relieves error accumulation problems and improves performance noticeably. At the core of our method is a cascaded two-step decoder architecture with the former decoder predicting HTML table row tags non-autoregressively and the latter predicting HTML table cell tags of each row in a semi-autoregressive manner. Compared with existing methods that predict HTML text autoregressively, the superiority of our row-to-cell progressive table parsing is twofold: (1) it generates an HTML tag sequence with a vertical-and-horizontal two-step `scanning', which better fits the inherent 2D structure of image data, (2) it performs substantially better for large tables (long sequence prediction) since it alleviates error accumulation problem specific to autoregressive models. Extensive experiments demonstrate that our method achieves competitive performance on three public benchmarks.

NeurIPS Conference 2023 Conference Paper

Structured Neural Networks for Density Estimation and Causal Inference

  • Asic Chen
  • Ruian (Ian) Shi
  • Xiang Gao
  • Ricardo Baptista
  • Rahul G. Krishnan

Injecting structure into neural networks enables learning functions that satisfy invariances with respect to subsets of inputs. For instance, when learning generative models using neural networks, it is advantageous to encode the conditional independence structure of observed variables, often in the form of Bayesian networks. We propose the Structured Neural Network (StrNN), which injects structure through masking pathways in a neural network. The masks are designed via a novel relationship we explore between neural network architectures and binary matrix factorization, to ensure that the desired independencies are respected. We devise and study practical algorithms for this otherwise NP-hard design problem based on novel objectives that control the model architecture. We demonstrate the utility of StrNN in three applications: (1) binary and Gaussian density estimation with StrNN, (2) real-valued density estimation with Structured Autoregressive Flows (StrAFs) and Structured Continuous Normalizing Flows (StrCNF), and (3) interventional and counterfactual analysis with StrAFs for causal inference. Our work opens up new avenues for learning neural networks that enable data-efficient generative modeling and the use of normalizing flows for causal effect estimation.

NeurIPS Conference 2022 Conference Paper

Human-Robotic Prosthesis as Collaborating Agents for Symmetrical Walking

  • Ruofan Wu
  • Junmin Zhong
  • Brent Wallace
  • Xiang Gao
  • He Huang
  • Jennie Si

This is the first attempt at considering human influence in the reinforcement learning control of a robotic lower limb prosthesis toward symmetrical walking in real world situations. We propose a collaborative multi-agent reinforcement learning (cMARL) solution framework for this highly complex and challenging human-prosthesis collaboration (HPC) problem. The design of an automatic controller of the robot within the HPC context is based on accessible physical features or measurements that are known to affect walking performance. Comparisons are made with the current state-of-the-art robot control designs, which are single-agent based, as well as existing MARL solution approaches tailored to the problem, including multi-agent deep deterministic policy gradient (MADDPG) and counterfactual multi-agent policy gradient (COMA). Results show that, when compared to these approaches, treating the human and robot as coupled agents and using estimated human adaption in robot control design can achieve lower stage cost, peak error, and symmetry value to ensure better human walking performance. Additionally, our approach accelerates learning of walking tasks and increases learning success rate. The proposed framework can potentially be further developed to examine how human and robotic lower limb prosthesis interact, an area that little is known about. Advancing cMARL toward real world applications such as HPC for normative walking sets a good example of how AI can positively impact on people’s lives.

ICML Conference 2022 Conference Paper

Learning to Incorporate Texture Saliency Adaptive Attention to Image Cartoonization

  • Xiang Gao
  • Yuqi Zhang
  • Yingjie Tian 0001

Image cartoonization is recently dominated by generative adversarial networks (GANs) from the perspective of unsupervised image-to-image translation, in which an inherent challenge is to precisely capture and sufficiently transfer characteristic cartoon styles (e. g. , clear edges, smooth color shading, vivid colors, etc.). Existing advanced models try to enhance cartoonization effect by learning to promote edges adversarially, introducing style transfer loss, or learning to align style from multiple representation space. This paper demonstrates that more distinct and vivid cartoonization effect could be easily achieved with only basic adversarial loss. Observing that cartoon style is more evident in cartoon-texture-salient local image regions, we build a region-level adversarial learning branch in parallel with the normal image-level one, which constrains adversarial learning on cartoon-texture-salient local patches for better perceiving and transferring cartoon texture features. To this end, a novel cartoon-texture-saliency-sampler (CTSS) module is proposed to adaptively sample cartoon-texture-salient patches from training data. We present that such texture saliency adaptive attention is of significant importance in facilitating and enhancing cartoon stylization, which is a key missing ingredient of related methods. The superiority of our model in promoting cartoonization effect, especially for high-resolution input images, are fully demonstrated with extensive experiments.

AAAI Conference 2022 Conference Paper

RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling

  • Yizhe Zhang
  • Siqi Sun
  • Xiang Gao
  • Yuwei Fang
  • Chris Brockett
  • Michel Galley
  • Jianfeng Gao
  • Bill Dolan

Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation.

AAAI Conference 2021 Conference Paper

A Controllable Model of Grounded Response Generation

  • Zeqiu Wu
  • Michel Galley
  • Chris Brockett
  • Yizhe Zhang
  • Xiang Gao
  • Chris Quirk
  • Rik Koncel-Kedziorski
  • Jianfeng Gao

Current end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models’ propensity to “hallucinate” facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines.

IJCAI Conference 2021 Conference Paper

Multi-view Feature Augmentation with Adaptive Class Activation Mapping

  • Xiang Gao
  • YingJie Tian
  • Zhiquan Qi

We propose an end-to-end-trainable feature augmentation module built for image classification that extracts and exploits multi-view local features to boost model performance. Different from using global average pooling (GAP) to extract vectorized features from only the global view, we propose to sample and ensemble diverse multi-view local features to improve model robustness. To sample class-representative local features, we incorporate a simple auxiliary classifier head (comprising only one 1x1 convolutional layer) which efficiently and adaptively attends to class-discriminative local regions of feature maps via our proposed AdaCAM (Adaptive Class Activation Mapping). Extensive experiments demonstrate consistent and noticeable performance gains achieved by our multi-view feature augmentation module.

v2026.09.13