Arrow Research search

Author name cluster

Congcong Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

EAAI Journal 2026 Journal Article

Trend-aware forecasting of pipeline displacement in thermal power plants via variational mode decomposition-aware temporal transformer

  • Minglu Dai
  • Qi Liu
  • Shuaiming Niu
  • Congcong Li
  • Jingdan Yuan
  • Xinyao Yu

Accurate forecasting of pipeline displacement in thermal power plants is essential for deep peak regulation, operational safety, and structural health monitoring. However, hybrid-driven methods based on mathematical decomposition often involve high computational costs, while conventional models that rely on single loss functions struggle to capture long-term trends in multi-step forecasts. To address these challenges, this study proposes a Variational Mode Decomposition-inspired forecasting network that integrates a time–frequency-aware embedding module with a trend-aware loss function. The network employs a patch-based linear embedding to efficiently extract multi-scale temporal features and introduces a customized loss to improve trend alignment across forecasting horizons. The model is evaluated using real-world displacement data from a thermal power plant. The results demonstrate that it achieves higher forecasting accuracy, stronger trend preservation, and substantially better computational efficiency than both decomposition-based approaches and purely data-driven neural networks.

ICRA Conference 2024 Conference Paper

STT: Stateful Tracking with Transformers for Autonomous Driving

  • Longlong Jing
  • Ruichi Yu
  • Xu Chen
  • Zhengli Zhao
  • Shiwei Sheng
  • Colin Graber
  • Qi Chen
  • Qinru Li

Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model’s performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTP S that address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset.

YNIMG Journal 2024 Journal Article

Wireless optically pumped magnetometer MEG

  • Hao Cheng
  • Kaiyan He
  • Congcong Li
  • Xiao Ma
  • Fufu Zheng
  • Wei Xu
  • Pan Liao
  • Rui Yang

The current magnetoencephalography (MEG) systems, which rely on cables for control and signal transmission, do not fully realize the potential of wearable optically pumped magnetometers (OPM). This study presents a significant advancement in wireless OPM-MEG by reducing magnetization in the electronics and developing a tailored wireless communication protocol. Our protocol effectively eliminates electromagnetic interference, particularly in the critical frequency bands of MEG signals, and accurately synchronizes the acquisition and stimulation channels with the host computer's clock. We have successfully achieved single-channel wireless OPM-MEG measurement and demonstrated its reliability by replicating three well-established experiments: The alpha rhythm, auditory evoked field, and steady-state visual evoked field in the human brain. Our prototype wireless OPM-MEG system not only streamlines the measurement process but also represents a major step forward in the development of wearable OPM-MEG applications in both neuroscience and clinical research.

EAAI Journal 2023 Journal Article

Exploring 2-rank strategic weight manipulation in multiple attribute decision making and its applications in project review and university ranking

  • Yating Liu
  • Siqi Wu
  • Congcong Li
  • Yucheng Dong

In some real multiple attribute decision making (MADM) problems, sometimes, it is time-consuming and unnecessary to obtain a complete ranking of alternatives, thus, a decision maker would classify the alternatives into two ordered categories, forming a 2-rank MADM problem. Occasionally, a decision maker can manipulate the desired 2-rank results by strategically setting attribute weights. This process is called 2-rank strategic weight manipulation (2RSWM). First, this study defines the 2-rank range of alternatives. Subsequently, several mixed 0–1 linear programming models (MLPMs) are constructed to obtain the 2-rank range and the strategic attribute weight vector of the desired 2-rank result of the alternative(s) of the decision maker. Furthermore, we provide conditions for the existence of the strategic attribute weight vector based on the 2-rank range of the alternatives and the proposed MLPMs. Finally, two illustrative examples and two simulation experiments are conducted to validate the effectiveness of our proposed models. Due to the ordered weighted averaging (OWA) operator having smaller average width of the 2-rank range, and a larger minimum distance between the impersonal and strategic attribute weight vectors, we argue that the OWA operator has a better performance than the weighted averaging (WA) operator in defending against 2RSWM.

ICRA Conference 2023 Conference Paper

Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints

  • Jiachen Li
  • Xinwei Shi
  • Feiyu Chen
  • Jonathan Stroud
  • Zhishuai Zhang
  • Tian Lan
  • Junhua Mao
  • Jeonhyung Kang

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory pre-diction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study.

ICRA Conference 2022 Conference Paper

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

  • Longlong Jing
  • Ruichi Yu
  • Henrik Kretzschmar
  • Kang Li
  • Charles R. Qi
  • Hang Zhao 0021
  • Alper Ayvaci
  • Xu Chen

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior performance when compared to LiDAR-based techniques. Through systematic analysis, we identified that per-object depth estimation accuracy is a major factor bounding the performance. Motivated by this observation, we propose a multi-level fusion method that combines different representations (RGB and pseudo-LiDAR) and temporal information across multiple frames for objects (tracklets) to enhance per-object depth estimation. Our proposed fusion method achieves the state-of-the-art performance of per-object depth estimation on the Waymo Open Dataset, the KITTI detection dataset, and the KITTI MOT dataset. We further demonstrate that by simply replacing estimated depth with fusion-enhanced depth, we can achieve significant improvements in monocular 3D perception tasks, including detection and tracking.

EAAI Journal 2022 Journal Article

Multi-camera joint spatial self-organization for intelligent interconnection surveillance

  • Congcong Li
  • Jing Li
  • Yuguang Xie
  • Jiayang Nie
  • Tao Yang
  • Zhaoyang Lu

The construction of smart city makes information interconnection play an increasingly important role in intelligent surveillance systems. Especially the interconnection among massive cameras is the key to realizing the evolution from current fragmented monitoring to interconnection surveillance. However, it remains a challenging problem in practical systems due to large sensor quantity, various camera types, and complex spatial layout. Aimed at this problem, this paper proposes a novel multi-camera joint spatial self-organization approach, which realizes interconnection surveillance by unifying cameras into one imaging space. Differing from existing back-end data association strategy, our method takes front-end data calibration as a breakthrough to relate surveillance data. Specifically, this paper first initials camera spatial parameter by sequence complementary feature integration. Through integrating complementarity and redundancy among sequence features, our method has robustness under scene dynamic changes and noise. Then, we propose a multi-camera joint optimization method based on common monitoring coverage correlation analysis to estimate a more accurate relative relationship. By leveraging the two strategies, the spatial relationship and visual data association across monitoring cameras are returned finally. Our system organizes all cameras into a unified imaging space by itself. Extensive experimental evaluations on an actual campus environment demonstrate our method achieves remarkable performance.

YNIMG Journal 2022 Journal Article

Multimodal neuroimaging with optically pumped magnetometers: A simultaneous MEG-EEG-fNIRS acquisition system

  • Xingyu Ru
  • Kaiyan He
  • Bingjiang Lyu
  • Dongxu Li
  • Wei Xu
  • Wenyu Gu
  • Xiao Ma
  • Jiayi Liu

Multimodal neuroimaging plays an important role in neuroscience research. Integrated noninvasive neuroimaging modalities, such as magnetoencephalography (MEG), electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS), allow neural activity and related physiological processes in the brain to be precisely and comprehensively depicted, providing an effective and advanced platform to study brain function. Noncryogenic optically pumped magnetometer (OPM) MEG has high signal power due to its on-scalp sensor layout and enables more flexible configurations than traditional commercial superconducting MEG. Here, we integrate OPM-MEG with EEG and fNIRS to develop a multimodal neuroimaging system that can simultaneously measure brain electrophysiology and hemodynamics. We conducted a series of experiments to demonstrate the feasibility and robustness of our MEG-EEG-fNIRS acquisition system. The complementary neural and physiological signals simultaneously collected by our multimodal imaging system provide opportunities for a wide range of potential applications in neurovascular coupling, wearable neuroimaging, hyperscanning and brain-computer interfaces.

NeurIPS Conference 2011 Conference Paper

$\theta$-MRF: Capturing Spatial and Semantic Structure in the Parameters for Scene Understanding

  • Congcong Li
  • Ashutosh Saxena
  • Tsuhan Chen

For most scene understanding tasks (such as object detection or depth estimation), the classifiers need to consider contextual information in addition to the local features. We can capture such contextual information by taking as input the features/attributes from all the regions in the image. However, this contextual dependence also varies with the spatial location of the region of interest, and we therefore need a different set of parameters for each spatial location. This results in a very large number of parameters. In this work, we model the independence properties between the parameters for each location and for each task, by defining a Markov Random Field (MRF) over the parameters. In particular, two sets of parameters are encouraged to have similar values if they are spatially close or semantically close. Our method is, in principle, complementary to other ways of capturing context such as the ones that use a graphical model over the labels instead. In extensive evaluation over two different settings, of multi-class object detection and of multiple scene understanding tasks (scene categorization, depth estimation, geometric labeling), our method beats the state-of-the-art methods in all the four tasks.

ICRA Conference 2011 Conference Paper

FeCCM for scene understanding: Helping the robot to learn multiple tasks

  • Congcong Li
  • T. P. Wong
  • Norris Xu
  • Ashutosh Saxena

Helping a robot to understand a scene can include many sub-tasks, such as scene categorization, object detection, geometric labeling, etc. Each sub-task is notoriously hard, and state-of-art classifiers exist for many sub-tasks. It is desirable to have an algorithm that can capture such correlation without requiring to make any changes to the inner workings of any classifier, and therefore make the perception for a robot better. We have recently proposed a generic model (Feedback Enabled Cascaded Classification Model) that enables us to easily take state-of-art classifiers as black-boxes and improve performance. In this video, we show that we can use our FeCCM model to quickly combine existing classifiers for various sub-tasks, and build a shoe finder robot in a day. The video shows our robot using FeCCM to find a shoe on request.

NeurIPS Conference 2010 Conference Paper

Towards Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models

  • Congcong Li
  • Adarsh Kowdle
  • Ashutosh Saxena
  • Tsuhan Chen

In many machine learning domains (such as scene understanding), several related sub-tasks (such as scene categorization, depth estimation, object detection) operate on the same raw data and provide correlated outputs. Each of these tasks is often notoriously hard, and state-of-the-art classifiers already exist for many sub-tasks. It is desirable to have an algorithm that can capture such correlation without requiring to make any changes to the inner workings of any classifier. We propose Feedback Enabled Cascaded Classification Models (FE-CCM), that maximizes the joint likelihood of the sub-tasks, while requiring only a ‘black-box’ interface to the original classifier for each sub-task. We use a two-layer cascade of classifiers, which are repeated instantiations of the original ones, with the output of the first layer fed into the second layer as input. Our training method involves a feedback step that allows later classifiers to provide earlier classifiers information about what error modes to focus on. We show that our method significantly improves performance in all the sub-tasks in two different domains: (i) scene understanding, where we consider depth estimation, scene categorization, event categorization, object detection, geometric labeling and saliency detection, and (ii) robotic grasping, where we consider grasp point detection and object classification.

v2026.09.13