Arrow Research search

Author name cluster

Jie He

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
1 author row

Possible papers

12

AAAI Conference 2026 Conference Paper

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

  • Yijie Zhu
  • Rui Shao
  • Ziyang Liu
  • Jie He
  • Jizhihui Liu
  • Jiuru Wang
  • Zitong Yu

Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how interactions shape the environment. However, most existing approaches treat observation and action generation in a monolithic and goal-agnostic manner, often leading to semantically misaligned predictions and incoherent behaviors. To this end, we propose H-GAR, a Hierarchical interaction framework via Goal-driven observation-Action Refinement. To anchor prediction to the task objective, H-GAR first produces a goal observation and a coarse action sketch that outline a high-level route toward the goal. To enable explicit interaction between observation and action under the guidance of the goal observation for more coherent decision-making, we devise two synergistic modules. (1) Goal-Conditioned Observation Synthesizer (GOS) synthesizes intermediate observations based on the coarse-grained actions and the predicted goal observation. (2) Interaction-Aware Action Refiner (IAAR) refines coarse actions into fine-grained, goal-consistent actions by leveraging feedback from the intermediate observations and a Historical Action Memory Bank that encodes prior actions to ensure temporal consistency. By integrating goal grounding with explicit action-observation interaction in a coarse-to-fine manner, H-GAR enables more accurate manipulation. Extensive experiments on both simulation and real-world robotic manipulation tasks demonstrate that H-GAR achieves state-of-the-art performance.

EAAI Journal 2026 Journal Article

Identification of freeway crash hotspots: A novel methodological framework using non-crash data for identification and errors reduction

  • Yuntao Ye
  • Jie He
  • Changjian Zhang
  • Xintong Yan
  • Xiazhi Zhang
  • Pengcheng Qin
  • Zhiming Fang

Identifying crash hotspots is essential for freeway safety, but most existing methods are crash-based and suffer from limitations such as passive analysis, long data collection periods, and regression-to-the-mean (RTM) bias. To address these limitations, this study proposes a novel three-module framework to identify crash hotspots using both crash data and non-crash data, while mitigating RTM bias. Specifically, in Module 1, a deep convolutional generative adversarial network (DCGAN) is utilized to learn the temporal-spatial patterns of vehicle kinematic parameters under ideal conditions. By comparing actual samples with DCGAN-generated samples, driving instability is quantified as a surrogate measure for crash risk estimation. In Module 2, an adaptive multivariate kernel density estimation method was introduced to identify hotspots using crash data and macro-level factors. Module 3 employs machine learning models to integrate the Modules 1 and 2. The proposed framework was validated through a case study. The results demonstrate that the framework outperforms existing methods, including negative binomial models and kernel density estimation, achieving average reductions of 0. 0070 in mean absolute error and 0. 0057 in root mean square error, together with an average increase of 0. 0255 in R2. Furthermore, Module 1 is demonstrated to estimate the mean crash risk without crash data. The analysis of Module 3 suggests that Module 1 exerts a corrective effect on Module 2, reducing errors caused by RTM bias. In practical applications, Module 1 can be applied to newly constructed freeways. For freeways with sufficient crash data, the comprehensive framework can be utilized to achieve more accurate identification.

NeurIPS Conference 2025 Conference Paper

CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification

  • Wei Li
  • Renshan Zhang
  • Rui Shao
  • Jie He
  • Liqiang Nie

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits scalability and deployment. Existing sparsification strategies—such as Mixture-of-Depths, layer skipping, and early exit—fall short by neglecting the semantic coupling across vision-language-action modalities, and focusing narrowly on intra-LLM computation while overlooking end-to-end coherence from perception to control. To address these challenges, we propose **CogVLA**, a Cognition-Aligned Vision-Language-Action framework that leverages instruction-driven routing and sparsification to improve both efficiency and performance. CogVLA draws inspiration from human multimodal coordination and introduces a 3-stage progressive architecture. 1) **Encoder-FiLM based Aggregation Routing (EFA-Routing)** injects instruction information into the vision encoder to selectively aggregate and compress dual-stream visual tokens, forming a instruction-aware latent representation. 2) Building upon this compact visual encoding, **LLM-FiLM based Pruning Routing (LFP-Routing)** introduces action intent into the language model by pruning instruction-irrelevant visually grounded tokens, thereby achieving token-level sparsity. 3) To ensure that compressed perception inputs can still support accurate and coherent action generation, we introduce **V‑L‑A Coupled Attention (CAtten)**, which combines causal vision-language attention with bidirectional action parallel decoding. Extensive experiments on the LIBERO benchmark and real-world robotic tasks demonstrate that CogVLA achieves state-of-the-art performance with success rates of 97. 4\% and 70. 0\%, respectively, while reducing training costs by 2. 5$\times$ and decreasing inference latency by 2. 8$\times$ compared to OpenVLA.

EAAI Journal 2025 Journal Article

Sliding window regression method for mechanical meter reading recognition in time-series images

  • Hao Xiu
  • Siran Hu
  • Yuanxin Cui
  • Jie He
  • Xiaotong Zhang
  • Yue Qi

Remote meter reading is an important part of the intelligent transformation of old meters. In addition to smart meters, installing cameras outside traditional meters and transmitting images to cloud servers via Narrowband Internet of Things technology can make the intelligent transformation process more convenient. Against this backdrop, this paper proposes a sliding window-based joint recognition network to address the significant numerical errors and misrecognition issues in current meter reading methods. It combines maximum probability decoding and maximum average probability decoding to classify and recognize complete instrument images, effectively reducing numerical errors and improving accuracy. After the instrument data acquisition is systematized, combining access to historical data and current images can effectively improve the recognition accuracy rate. The network collaborates with a regression prediction model, performing regression analysis on historical data to establish equations for small-range visual-assisted regression recognition. Experiments on instrument datasets show that the algorithm achieves a 99. 57% recognition accuracy, which is 4% higher than holistic methods and nearly 1% higher than single-digit methods. Moreover, it reduces the maximum error from an average of 68. 5 to 1. 06, and the mean absolute error and mean square error by nearly 20 times. The model is now ready for production use. This method is highly significant for industrial users, especially in old-style electricity meter reading scenarios in factories, such as in power monitoring systems of large-scale manufacturing enterprises, strongly supporting industrial automation.

EAAI Journal 2025 Journal Article

Vertically invariant network for vertical and horizontal scene text recognition

  • Chengli Zhu
  • Zhao Feng
  • Jie He
  • Xiaohui Xiao

Vertical scene text is widely used for traffic signs, ShopSigns, and ship container codes. Compared to horizontal or curved text, vertical text contains different layout and orientation features between adjacent characters. Therefore, most existing text recognizers struggle to handle vertical samples or require more parameters to model the orientation representation. In this paper, we propose a vertically invariant network to address these issues, which explicitly extracts vertically equivariant and vertically invariant features from vertical and horizontal text. Specifically, we introduce a four-rotation convolution and a vertically equivariant backbone based on the characteristics of vertical and horizontal text images, which can encode vertical equivariance with significantly fewer network parameters. Then, a reconstruction-based vertically invariant module is developed to adaptively transform the vertically equivariant features into vertically invariant features. Besides, we construct a Vertical and Horizontal Text Recognition dataset to boost research in vertical text recognition applications, including 26308 images collected from traffic signs, ShopSigns, and ship container codes. Extensive experiments on our dataset and two public scene text datasets show that our proposal can achieve state-of-the-art vertical and horizontal text recognition accuracy. On the Wuhan University vertical and horizontal text recognition (WHU-VHTR) dataset, our model outperforms the second-best method (Character Image Reconstruction Network, CIRN) while using only 35. 2% of its parameters and reducing the inference time by 12. 4%, while achieving 2. 6% higher recognition accuracy. The code is available online at https: //github. com/ChengliZhu777/VINet.

EAAI Journal 2024 Journal Article

Driving risk prediction of urban arterial and collector roads using multi-dimensional real-time data

  • Xintong Yan
  • Jie He
  • Guanhe Wu
  • Chenwei Wang
  • Changjian Zhang
  • Yuntao Ye

With the objective of predicting driving risks on urban arterial and collector roads, this paper primarily collects and processes vehicle trajectory data and driver-vehicle-road-environment data. Using Entropy-Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) and K-means, risk scores and levels for the target roads are determined. An original dataset comprising 10, 320 samples is established, with risk levels as labels and multi-dimensional data as features. SHapley Additive exPlanations (SHAP) analysis is utilized to extract important features. The performance of four ensemble models are compared, including Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM) and Stacking. Three data resampling techniques are integrated to address data imbalance. The combination of LightGBM and Deep Convolutional Generative Adversarial Network (DCGAN) outperforms other approaches, showing potential for real-time risk prediction. Besides, DCGAN is employed to update the multi-dimensional features. The models' predictive performance is evaluated under various update ratios using F1-score and training time, with LightGBM demonstrating superior accuracy and efficiency compared to its counterparts (RF, XGBoost, and Stacking). This study develops a driving risk prediction system based on multi-dimensional features, offering a theoretical foundation and technical approach to enhance urban road traffic safety management.

AAMAS Conference 2024 Conference Paper

JDRec: Practical Actor-Critic Framework for Online Combinatorial Recommender System

  • Xin Zhao
  • Jiaxin Li
  • Zhiwei Fang
  • Yuchen Guo
  • Jinyuan Zhao
  • Jie He
  • Wenlong Chen
  • Changping Peng

In the realm of online recommendation systems, the Combinatorial Recommender (CR) system stands out for its unique approach. It presents users with a list of items on a result page, where user behavior is simultaneously influenced by contextual information and the items listed. Formulated as a combinatorial optimization problem, the objective of the CR system is to maximize the recommendation reward across the entire list of items. Despite the significant potential of CR systems, developing a practical and efficient model remains substantial challenges. These challenges stem from the dynamic nature of online environments and the pressing need for personalized recommendations. To tackle these challenges, we decompose the overarching problem into two sub-problems: list generation and list evaluation. We propose novel and pragmatic model architectures for each sub-problem aiming to concurrently enhance both effectiveness and efficiency. To further adapt the CR system to online scenarios, we integrate a bootstrap algorithm into an actor-critic reinforcement framework. This innovative approach called JD Recommender System (JDRec) is designed to continuously refine the recommendation mode through sustained user interaction, ensuring the system’s adaptability and relevance. The proposed JDRec framework, tested through rigorous offline and online experiments, has shown promising results. It has been successfully deployed in online JD recommendation systems, yielding a notable improvement in click-through rate by 2. 6% and augmenting the total value of the platform by 5. 03%. Besides, we release the large scale dataset used in our work to facilitate further research. This work is licensed under a Creative Commons Attribution International 4. 0 License. *Equal contribution. Proc. of the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024), N. Alechina, V. Dignum, M. Dastani, J. S. Sichman (eds.), May 6 – 10, 2024, Auckland, New Zealand. © 2024 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org).

AAAI Conference 2024 Conference Paper

Parallel Ranking of Ads and Creatives in Real-Time Advertising Systems

  • Zhiguang Yang
  • Liufang Sang
  • Haoran Wang
  • Wenlong Chen
  • Lu Wang
  • Jie He
  • Changping Peng
  • Zhangang Lin

Creativity is the heart and soul of advertising services. Effective creatives can create a win-win scenario: advertisers each target users and achieve marketing objectives more effectively, users more quickly find products of interest, and platforms generate more advertising revenue. With the advent of AI-Generated Content, advertisers now can produce vast amounts of creative content at a minimal cost. The current challenge lies in how advertising systems can select the most pertinent creative in real-time for each user personally. Existing methods typically perform serial ranking of ads or creatives, limiting the creative module in terms of both effectiveness and efficiency. In this paper, we propose for the first time a novel architecture for online parallel estimation of ads and creatives ranking, as well as the corresponding offline joint optimization model. The online architecture enables sophisticated personalized creative modeling while reducing overall latency. The offline joint model for CTR estimation allows mutual awareness and collaborative optimization between ads and creatives. Additionally, we optimize the offline evaluation metrics for the implicit feedback sorting task involved in ad creative ranking. We conduct extensive experiments to compare ours with two state-of-the-art approaches. The results demonstrate the effectiveness of our approach in both offline evaluations and real-world advertising platforms online in terms of response time, CTR, and CPM.

AAAI Conference 2024 Conference Paper

S2CycleDiff: Spatial-Spectral-Bilateral Cycle-Diffusion Framework for Hyperspectral Image Super-resolution

  • Jiahui Qu
  • Jie He
  • Wenqian Dong
  • Jingyu Zhao

Hyperspectral image super-resolution (HISR) is a technique that can break through the limitation of imaging mechanism to obtain the hyperspectral image (HSI) with high spatial resolution. Although some progress has been achieved by existing methods, most of them directly learn the spatial-spectral joint mapping between the observed images and the target high-resolution HSI (HrHSI), failing to fully reserve the spectral distribution of low-resolution HSI (LrHSI) and the spatial distribution of high-resolution multispectral imagery (HrMSI). To this end, we propose a spatial-spectral-bilateral cycle-diffusion framework (S2CycleDiff) for HISR, which can step-wise generate the HrHSI with high spatial-spectral fidelity by learning the conditional distribution of spatial and spectral super-resolution processes bilaterally. Specifically, a customized conditional cycle-diffusion framework is designed as the backbone to achieve the spatial-spectral-bilateral super-resolution by repeated refinement, wherein the spatial/spectral guided pyramid denoising (SGPD) module seperately takes HrMSI and LrHSI as the guiding factors to achieve the spatial details injection and spectral correction. The outputs of the conditional cycle-diffusion framework are fed into a complementary fusion block to integrate the spatial and spectral details to generate the desired HrHSI. Experiments have been conducted on three widely used datasets to demonstrate the superiority of the proposed method over state-of-the-art HISR methods. The code is available at https://github.com/Jiahuiqu/S2CycleDiff.

JBHI Journal 2020 Journal Article

A Novel MKL Method for GBM Prognosis Prediction by Integrating Histopathological Image and Multi-Omics Data

  • Ya Zhang
  • Ao Li
  • Jie He
  • Minghui Wang

Glioblastoma multiforme (GBM) is one of the most malignant brain tumors with very short prognosis expectation. To improve patients’ clinical treatment and their life quality after surgery, researches have developed tremendous in silico models and tools for predicting GBM prognosis based on molecular datasets and have earned great success. However, pathology still plays the most critical role in cancer diagnosis and prognosis in the clinic at present. Recent advancement of storing and processing histopathological images has drawn attention of researchers. Models based on histopathological images are developed, which show great potential for computer-aided pathological diagnoses. But models based on both molecular and histopathological images that could predict GBM prognosis with high accuracy are not present yet. In our previous research, we used the simple MKL method to integrate multi-omics data to improve GBM prognosis prediction successfully. In this paper, we have developed a novel multiple kernel learning (MKL) method, named histopathological integrating multiple kernel learning (HI-MKL), that could integrate both histopathological images and multi-omics data efficiently. By using datasets from The Cancer Genome Atlas project, we have built a system that could predict the GBM prognosis with high accuracy. Our research shows that HI-MKL is an accurate, robust, and generalized MKL method, which performs well in a GBM prognosis task.

NeurIPS Conference 2019 Conference Paper

Joint Optimization of Tree-based Index and Deep Model for Recommender Systems

  • Han Zhu
  • Daqing Chang
  • Ziru Xu
  • Pengye Zhang
  • Xiang Li
  • Jie He
  • Han Li
  • Jian Xu

Large-scale industrial recommender systems are usually confronted with computational problems due to the enormous corpus size. To retrieve and recommend the most relevant items to users under response time limits, resorting to an efficient index structure is an effective and practical solution. The previous work Tree-based Deep Model (TDM) \cite{zhu2018learning} greatly improves recommendation accuracy using tree index. By indexing items in a tree hierarchy and training a user-node preference prediction model satisfying a max-heap like property in the tree, TDM provides logarithmic computational complexity w. r. t. the corpus size, enabling the use of arbitrary advanced models in candidate retrieval and recommendation. In tree-based recommendation methods, the quality of both the tree index and the user-node preference prediction model determines the recommendation accuracy for the most part. We argue that the learning of tree index and preference model has interdependence. Our purpose, in this paper, is to develop a method to jointly learn the index structure and user preference prediction model. In our proposed joint optimization framework, the learning of index and user preference prediction model are carried out under a unified performance measure. Besides, we come up with a novel hierarchical user preference representation utilizing the tree index hierarchy. Experimental evaluations with two large-scale real-world datasets show that the proposed method improves recommendation accuracy significantly. Online A/B test results at a display advertising platform also demonstrate the effectiveness of the proposed method in production environments.

EAAI Journal 2012 Journal Article

Recognition of driving postures by multiwavelet transform and multilayer perceptron classifier

  • Chihang Zhao
  • Yongsheng Gao
  • Jie He
  • Jie Lian

To develop Human-centric Driver Assistance Systems (HDAS) for automatic understanding and charactering of driver behaviors, an efficient feature extraction of driving postures based on Geronimo–Hardin–Massopust (GHM) multiwavelet transform is proposed, and Multilayer Perceptron (MLP) classifiers with three layers are then exploited in order to recognize four pre-defined classes of driving postures. With features extracted from a driving posture dataset created at Southeast University (SEU), the holdout and cross-validation experiments on driving posture classification are conducted by MLP classifiers, compared with the Intersection Kernel Support Vector Machines (IKSVMs), the k-Nearest Neighbor (kNN) classifier and the Parzen classifier. The experimental results show that feature extraction based on GHM multwavelet transform and MLP classifier, using softmax activation function in the output layer and hyperbolic tangent activation function in the hidden layer, offer the best classification performance compared to IKSVMs, kNN and Parzen classifiers. The experimental results also show that talking on a cellular phone is the most difficult one to classify among four predefined classes, which are 83. 01% and 84. 04% in the holdout and cross-validation experiments respectively. These results show the effectiveness of the feature extraction approach using GHM multiwavelet transform and MLP classifier in automatically understanding and characterizing driver behaviors towards Human-centric Driver Assistance Systems (HDAS).

v2026.09.13