Arrow Research search

Author name cluster

Gong Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

AAAI Conference 2026 Conference Paper

CoRA: A Collaborative Robust Architecture with Hybrid Fusion for Efficient Perception

  • Gong Chen
  • Chaokun Zhang
  • Pengcheng Lv
  • Xiaohui Xie

Collaborative perception has garnered significant attention as a crucial technology to overcome the perceptual limitations of single-agent systems. Many state-of-the-art (SOTA) methods have achieved communication efficiency and high performance via intermediate fusion. However, they share a critical vulnerability: their performance degrades under adverse communication conditions due to the misalignment induced by data transmission, which severely hampers their practical deployment. To bridge this gap, we re-examine different fusion paradigms, and recover that the strengths of intermediate and late fusion are not a trade-off, but a complementary pairing. Based on this key insight, we propose CoRA, a novel collaborative robust architecture with a hybrid approach to decouple performance from robustness with low communication. It is composed of two components: a feature-level fusion branch and an object-level correction branch. Its first branch selects critical features and fuses them efficiently to ensure both performance and scalability. The second branch leverages semantic relevance to correct spatial displacements, guaranteeing resilience against pose errors. Experiments demonstrate the superiority of CoRA. Under extreme scenarios, CoRA improves upon its baseline performance by approximately 19% in AP@0.7 with more than 5x less communication volume, which makes it a promising solution for robust collaborative perception.

AAAI Conference 2026 Conference Paper

Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMs

  • Zhejing Hu
  • Yan Liu
  • Zhi Zhang
  • Aiwei Zhang
  • Sheng-hua Zhong
  • Bruce X.B. Yu
  • Gong Chen

Large Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studies validate that symbolic music can be modeled as a human language, yet composing structured music from partial symbolic inputs through natural language interaction remains underexplored. Even LLMs struggle to generate structurally coherent compositions in such hybrid input-output scenarios, highlighting a fundamental gap that calls for a domain-specific learning paradigm. To this end, we propose Inspiration-to-Structure (IoS), a cognitively inspired framework that enables LLMs to generate structured musical sections from melodic ideas. IoS employs a three-phase process—semantic, structural, and collaborative cognition—and is supported by two key components: (1) a new dataset and construction protocol called Structured Triplet Data (STD), and (2) a training method, Dual-Instance Structural Contrastive Optimization (DiSCO), designed to enhance structural awareness. Experiments show that IoS improves structural coherence by 47.8% and artistic creativity by 21.8% compared to conventional language modeling paradigm, supervised fine-tuning, and even enables smaller LLMs to surpass larger LLMs. These results suggest that symbolic music, while language-like, demands specialized modeling beyond standard language modeling paradigms. IoS enables LLMs to transform music theory knowledge into structured composition, empowering users to compose music interactively via language and advancing toward general creative AI.

EAAI Journal 2026 Journal Article

Theory-guided data-driven based on the learning curve for fracturing performance prediction

  • Yunjin Wang
  • Leyi Zheng
  • Gong Chen
  • Jianlong Zhang
  • Hao Bai
  • Hanxuan Song
  • Tingxue Jiang
  • Fujian Zhou

Accurate and robust prediction of fracturing performance is essential for optimizing fracturing strategies. Here, a fracturing learning curve is proposed based on the fracturing characteristics in Gimsar shale oil, and is used as a theoretical guide to build a theory-guided data-driven (TgDD) model to predict the fracturing performance. The fracturing learning curve is further decomposed into dimensionless trends and local fluctuations. Convolutional neural network (CNN) and gated recurrent unit (GRU) were combined to build a CNN-GRU to predict the dimensionless trend. Using adaptive boosting (AdaBoost) integrated random forest (RF) to build an AdaBoost-RF to predict the local fluctuations. The results show that dimensionless trend has time series characteristics. CNN-GRU can extract and select the features, and its prediction ability is 28. 1 % and 12. 9 % higher than that of CNN and GRU. AdaBoost-RF can dynamically adjust the weights, and its prediction ability is about 37% higher than that of the RF. TgDD is more sensitive to engineering parameters. Relative to the direct prediction, the prediction accuracy of the TgDD is improved by 47. 6 %. There are two main reasons for the higher prediction accuracy of TgDD. One is that the dimensionless trend belongs to the time series data, for which the established CNN-GRU model has an extremely strong prediction ability. The second is that the fluctuation amplitude of local fluctuations is reduced, which improves the data quality. The engineering parameters of the newly fractured wells were optimized using TgDD, and its estimated ultimate recovery was improved from 0. 4847 to 0. 4917.

IJCAI Conference 2025 Conference Paper

CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation

  • Zhejing Hu
  • Yan Liu
  • Gong Chen
  • Bruce X. B. Yu

Generative artificial intelligence in music has made significant strides, yet it still falls short of the substantial achievements seen in natural language processing, primarily due to the limited availability of music data. Knowledge-informed approaches have been shown to enhance the performance of music generation models, even when only a few pieces of musical knowledge are integrated. This paper seeks to leverage comprehensive music theory in AI-driven music generation tasks, such as algorithmic composition and style transfer, which traditionally require significant manual effort with existing techniques. We introduce a novel automatic music lexicon construction model that generates a lexicon, named CompLex, comprising 37, 432 items derived from just 9 manually input category keywords and 5 sentence prompt templates. A new multi-agent algorithm is proposed to automatically detect and mitigate hallucinations. CompLex demonstrates impressive performance improvements across three state-of-the-art text-to-music generation models, encompassing both symbolic and audio-based methods. Furthermore, we evaluate CompLex in terms of completeness, accuracy, non-redundancy, and executability, confirming that it possesses the key characteristics of an effective lexicon.

AAAI Conference 2025 Conference Paper

Compose with Me: Collaborative Music Inpainter for Symbolic Music Infilling

  • Zhejing Hu
  • Yan Liu
  • Gong Chen
  • Bruce X.B. Yu

The field of music generation has seen a surge of interest from both academia and industry, with innovative platforms such as Suno, Udio, and SkyMusic earning widespread recognition. However, the challenge of music infilling—modifying specific music segments without reconstructing the entire piece—remains a significant hurdle for both audio-based and symbolic-based models, limiting their adaptability and practicality. In this paper, we address symbolic music infilling by introducing the Collaborative Music Inpainter (CMI), an advanced human-in-the-loop (HITL) model for music infilling. The CMI features the Joint Embedding Predictive Autoregressive Generative Architecture (JEP-AGA), which learns the high-level predictive representations of the masked part that needs to be infilled during the autoregressive generative process, akin to how humans perceive and interpret music. The newly developed Dynamic Interaction Learner (DIL) achieves HITL by iteratively refining the infilled output based on user interactions alone, significantly reducing the interaction cost without requiring further input. Experimental results confirm CMI’s superior performance in music infilling, demonstrating its efficiency in producing high-quality music.

AAAI Conference 2025 Conference Paper

Mixture of Knowledge Minigraph Agents for Literature Review Generation

  • Zhi Zhang
  • Yan Liu
  • Sheng-hua Zhong
  • Gong Chen
  • Yu Yang
  • Jiannong Cao

Literature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel framework, collaborative knowledge minigraph agents (CKMAs), to automate scholarly literature reviews. A novel prompt-based algorithm, the knowledge minigraph construction agent (KMCA), is designed to identify relations between concepts from academic literature and automatically constructs knowledge minigraphs. By leveraging the capabilities of large language models on constructed knowledge minigraphs, the multiple path summarization agent (MPSA) efficiently organizes concepts and relations from different viewpoints to generate literature review paragraphs. We evaluate CKMAs on three benchmark datasets. Experimental results show the effectiveness of the proposed method, further revealing promising applications of LLMs in scientific research.

ICML Conference 2025 Conference Paper

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

  • Haotian Ni
  • Yake Wei
  • Hang Liu
  • Gong Chen
  • Chong Peng
  • Hao Lin
  • Di Hu 0001

Multimodal learning faces challenges in effectively fusing information from diverse modalities, especially when modality quality varies across samples. Dynamic fusion strategies, such as attention mechanism in Transformers, aim to address such challenge by adaptively emphasizing modalities based on the characteristics of input data. However, through amounts of carefully designed experiments, we surprisingly observed that the dynamic adaptability of widely-used self-attention models diminishes. Model tends to prefer one modality regardless of data characteristics. This bias triggers a self-reinforcing cycle that progressively overemphasizes the favored modality, widening the distribution gap in attention keys across modalities and deactivating attention mechanism’s dynamic properties. To revive adaptability, we propose a simple yet effective method Rolling Query (RollingQ), which balances attention allocation by rotating the query to break the self-reinforcing cycle and mitigate the key distribution gap. Extensive experiments on various multimodal scenarios validate the effectiveness of RollingQ and the restoration of cooperation dynamics is pivotal for enhancing the broader capabilities of widely deployed multimodal Transformers. The source code is available at https: //github. com/GeWu-Lab/RollingQ_ICML2025.

EAAI Journal 2024 Journal Article

Expert's experience-informed hierarchical kriging method for aerodynamic data modeling

  • Chen-Zhou Xu
  • Zhong-Hua Han
  • Bo-Wen Zan
  • Ke-Shi Zhang
  • Gong Chen
  • Wen-Zheng Wang

Data-driven models, such as kriging, have gained popularity in aerospace engineering due to their capability of predicting multidimensional, nonlinear aerodynamic characteristics of an aircraft. However, they are still suffering from the problem associated with poor extrapolation capability and physical interpretability, which in turn has great impacts on aerodynamic performance and flight safety. To address this problem, this article proposes to incorporate an empirical aerodynamic model obtained from expert's understanding of aerodynamics in the fitting process of a kriging model. First, empirical aerodynamic models based on expert's understanding and experience are derived. Second, the regression term of a kriging model is replaced by the expert's experience-informed model so that the global trend can be consistent with physical laws and provides extra knowledge in the subregion(s) without training data, especially for the extrapolation region(s). Finally, the expert's experience-informed hierarchical kriging (EEI-HK) model is built in a sequential way. The proposed method is validated with analytical test examples and demonstrated by aerodynamic data modeling of an AGARD-B missile and an FDL-5A hypersonic flight vehicle. Results show that, compared with ordinary and universal kriging models, the proposed EEI-HK model can dramatically improve the prediction accuracy in the extrapolation domain and slightly enhance the interpolation accuracy, with assistance from physical information provided by an empirical model. In consequence, it can be a promising approach for aerodynamic data extrapolation and saving the cost of establishing an aerodynamic database.

AAAI Conference 2024 Conference Paper

Responding to the Call: Exploring Automatic Music Composition Using a Knowledge-Enhanced Model

  • Zhejing Hu
  • Yan Liu
  • Gong Chen
  • Xiao Ma
  • Shenghua Zhong
  • Qianwen Luo

Call-and-response is a musical technique that enriches the creativity of music, crafting coherent musical ideas that mirror the back-and-forth nature of human dialogue with distinct musical characteristics. Although this technique is integral to numerous musical compositions, it remains largely uncharted in automatic music composition. To enhance the creativity of machine-composed music, we first introduce the Call-Response Dataset (CRD) containing 19,155 annotated musical pairs and crafted comprehensive objective evaluation metrics for musical assessment. Then, we design a knowledge-enhanced learning-based method to bridge the gap between human and machine creativity. Specifically, we train the composition module using the call-response pairs, supplementing it with musical knowledge in terms of rhythm, melody, and harmony. Our experimental results underscore that our proposed model adeptly produces a wide variety of creative responses for various musical calls.

IJCAI Conference 2022 Conference Paper

EGCN: An Ensemble-based Learning Framework for Exploring Effective Skeleton-based Rehabilitation Exercise Assessment

  • Bruce X. B. Yu
  • Yan Liu
  • Xiang Zhang
  • Gong Chen
  • Keith C. C. Chan

Recently, some skeleton-based physical therapy systems have been attempted to automatically evaluate the correctness or quality of an exercise performed by rehabilitation subjects. However, in terms of algorithms and evaluation criteria, the task remains not fully explored regarding making full use of different skeleton features. To advance the prior work, we propose a learning framework called Ensemble-based Graph Convolutional Network (EGCN) for skeleton-based rehabilitation exercise assessment. As far as we know, this is the first attempt that utilizes both two skeleton feature groups and investigates different ensemble strategies for the task. We also examine the properness of existing evaluation criteria and focus on evaluating the prediction ability of our proposed method. We then conduct extensive cross-validation experiments on two latest public datasets: UI-PRMD and KIMORE. Results indicate that the model-level ensemble scheme of our EGCN achieves better performance than existing methods. Code is available: https: //github. com/bruceyo/EGCN.

AAAI Conference 2020 Conference Paper

Social Influence Does Matter: User Action Prediction for In-Feed Advertising

  • Hongyang Wang
  • Qingfei Meng
  • Ju Fan
  • Yuchen Li
  • Laizhong Cui
  • Xiaoman Zhao
  • Chong Peng
  • Gong Chen

Social in-feed advertising delivers ads that seamlessly fit inside a user’s feed, and allows users to engage in social actions (likes or comments) with the ads. Many businesses pay higher attention to “engagement marketing” that maximizes social actions, as social actions can effectively promote brand awareness. This paper studies social action prediction for infeed advertising. Most existing works overlook the social in- fluence as a user’s action may be affected by her friends’ actions. This paper introduces an end-to-end approach that leverages social influence for action prediction, and focuses on addressing the high sparsity challenge for in-feed ads. We propose to learn influence structure that models who tends to be influenced. We extract a subgraph with the near neighbors a user interacts with, and learn topological features of the subgraph by developing structure-aware graph encoding methods. We also introduce graph attention networks to learn influence dynamics that models how a user is influenced by neighbors’ actions. We conduct extensive experiments on real datasets from the commercial advertising platform of WeChat and a public dataset. The experimental results demonstrate that social influence learned by our approach can significantly boost performance of social action prediction.

v2026.09.13