Arrow Research search

Author name cluster

Ling Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

52 papers
2 author rows

Possible papers

52

YNIMG Journal 2026 Journal Article

Distinct frontal lobe subregions mediate the emergence and reporting of visual consciousness

  • Yan Zhang
  • Xiarong Li
  • Zhenlan Jin
  • Junjun Zhang
  • Ling Li

Persistent debate surrounds whether the frontal lobe supports the emergence or reporting of consciousness, raising the hypothesis that distinct frontal subregions may support these processes. We addressed this by combining electroencephalography (EEG) with eye-tracking in Report and No-Report paradigms. Eye-movement features distinguished conscious and unconscious trials in the no-report task. Event-related potential analyses showed that the Visual Awareness Negativity (VAN) was independent of reporting, whereas P3b occurred only with explicit reports. Importantly, the frontal Dorsal Attention Network (DAN) supported the emergence of consciousness, independent of post-perceptual reporting, as shown by multivoxel pattern analysis showing that a classifier's ability to decode visual consciousness generalized bidirectionally between report and no-report tasks. In contrast, frontal components of the Default Mode Network (DMN) and Frontoparietal Control Network (FPN) encoded visual consciousness only when explicit reports were required, indicating roles in reporting. These findings demonstrate a functional dissociation within the frontal lobe and refine the anatomical framework for the neural basis of visual consciousness.

AAAI Conference 2026 Conference Paper

Efficient Diffusion Planning with Temporal Diffusion

  • Jiaming Guo
  • Rui Zhang
  • Zerun Li
  • Yunkai Gao
  • Shaohui Peng
  • Siming Lan
  • Xing Hu
  • Zidong Du

Diffusion planning is a promising method for learning high-performance policies from offline data. To avoid the impact of discrepancies between planning and reality on performance, previous works generate new plans at each time step. However, this incurs significant computational overhead and leads to lower decision frequencies, and frequent plan switching may also affect performance. In contrast, humans might create detailed short-term plans and more general, sometimes vague, long-term plans, and adjust them over time. Inspired by this, we propose the Temporal Diffusion Planner (TDP) which improves decision efficiency by distributing the denoising steps across the time dimension. TDP begins by generating an initial plan that becomes progressively more vague over time. At each subsequent time step, rather than generating an entirely new plan, TDP updates the previous one with a small number of denoising steps. This reduces the average number of denoising steps, improving decision efficiency. Additionally, we introduce an automated replanning mechanism to prevent significant deviations between the plan and reality. Experiments on D4RL show that, compared to previous works that generate new plans every time step, TDP significantly improves the decision-making frequency by 11-24.8 times while achieving higher or comparable performance.

YNIMG Journal 2026 Journal Article

Gray matter volume predicts decision speed and reveals stage-specific contributions of large-scale brain networks in gambling tasks

  • Tingting Zhang
  • Qiuzhu Zhang
  • Ronglong Xiong
  • Junjun Zhang
  • Zhenlan Jin
  • Ling Li

Large-scale brain networks are well-established in resting-state research and are increasingly being used in task-based functional magnetic resonance imaging (fMRI) studies. However, the mechanisms by which brain networks dynamically reorganize across the various stages of decision-making remain unclear. Here, we investigated the neural basis of decision-making by integrating voxel-based morphometry and fMRI within a modified "Wheel of Fortune" gambling task. Stage-specific brain activation was characterized using the Yeo-7 network atlas to delineate large-scale network dynamics across task stages. We found that: (1) Reaction time (RTs) were significantly longer during choose conditions compared to follow conditions; (2) Gray matter volume correlated with individual variability in RT and predicted RT during choose conditions using multivariate pattern analysis with a Kernel Ridge Regression model, effects absent during follow conditions; (3) A negative correlation was observed between RT and activation in the right superior temporal gyrus and left mid-cingulate cortex; (4) Choice stage involved more extensive network engagement than the result and rating stages, with the rating stage showing the lowest overall activation. Network-specific fractional contributions revealed dominant engagement of the ventral attention network, default mode network, and somato-motor network during the choice stage; the frontoparietal network (FPN), dorsal attention network (DAN), and visual network during the result stage; and the DAN and FPN during the rating stage. These findings provide structural and functional explanations for individual differences in decision speed within a gambling paradigm, revealing the distinct and dynamic roles of brain networks across decision stages and offering mechanistic insights into the neural architecture of this process.

AAAI Conference 2026 Conference Paper

QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation

  • Xinguo Zhu
  • Shaohui Peng
  • Jiaming Guo
  • Yunji Chen
  • Qi Guo
  • Yuanbo Wen
  • Hang Qin
  • Ruizhi Chen

Developing high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While large language models (LLMs) offer promise for automation, both general-purpose and finetuned LLMs suffer from two fundamental and conflicting limitations: correctness and efficiency. The key reason is that existing LLM-based approaches directly generate the entire optimized low-level programs, requiring exploration of an extremely vast space encompassing both optimization policies and implementation codes. To address the challenge of exploring an intractable space, we propose Macro Thinking Micro Coding (MTMC), a hierarchical framework inspired by the staged optimization strategy of human experts. It decouples optimization strategy from implementation details, ensuring efficiency through high-level strategy and correctness through low-level implementation. Specifically, Macro Thinking employs reinforcement learning to guide lightweight LLMs in efficiently exploring and learning semantic optimization strategies that maximize hardware utilization. Micro Coding leverages general-purpose LLMs to incrementally implement the stepwise optimization proposals from Macro Thinking, avoiding full-kernel generation errors. Together, they effectively navigate the vast optimization space and intricate implementation details, enabling LLMs for high-performance GPU kernel generation. Comprehensive results on widely adopted benchmarks demonstrate the superior performance of MTMC on GPU kernel generation in both accuracy and running time. On KernelBench, MTMC achieves near 100% and 70% accuracy at Levels 1-2 and 3, over 50% than SOTA general-purpose and domain-finetuned LLMs, with up to 7.3× speedup over LLMs, and 2.2× over expert-optimized PyTorch Eager kernels. On the more challenging TritonBench, MTMC attains up to 59.64% accuracy and 34× speedup. All models and datasets will be made publicly available.

YNIMG Journal 2026 Journal Article

The Erlangen Program in lateral occipital cortex: Hierarchical encoding of emergent features

  • Junjun Zhang
  • Shi Zeng
  • Baochen Wang
  • Jingyu He
  • Zhenlan Jin
  • Ling Li

Emergent features are fundamental concepts in Gestalt psychology, yet the neural encoding of these features, particularly a quantitative understanding of their relative superiority, remains elusive. This study bridges this gap by conceptualizing emergent features through geometric transformations within the Erlangen Program, which provides a principled framework to quantify their hierarchical relationships. We propose that the lateral occipital cortex (LOC) encodes these emergent features in accordance with the geometric hierarchies defined by this program. Using fMRI and multivariate pattern analysis, we demonstrate that LOC reliably discriminates between distinct geometric transformations (Euclidean, affine, projective, and topology). Critically, representational similarity analysis reveals that neural dissimilarities in LOC align with the relative stability of geometries predicted by the Erlangen Program. However, the LOC exhibits similar representational structures for lower-order transformations like Euclidean and affine geometries, suggesting a potential collapse of these distinctions in the region's global geometric hierarchy. Furthermore, transfer learning confirms hierarchical nesting relationships among the geometries: classifiers trained on specific geometric distinctions generalize to others in a manner consistent with the Erlangen hierarchy. These findings establish LOC as the neural substrate where emergent features are organized hierarchically by geometric stability, revealing how the visual system prioritizes invariant global structures to optimize perceptual efficiency.

EAAI Journal 2025 Journal Article

A hyperparameter-fusion neural networks for deposition prediction

  • Li Ding
  • Kun Pang
  • Junjie Li
  • Hua Shao
  • Nan Liu
  • Rui Chen
  • Zhiqiang Li
  • Zhenjie Yao

As integrated circuit manufacturing processes develop into the nanometer scale, precise control and prediction of the deposition process have become crucial. Nanoscale manufacturing imposes unprecedentedly high demands on film quality, uniformity, and consistency, presenting significant challenges to traditional control and prediction methodologies. This study proposes a novel approach that, for the first time, formulates the thin-film deposition process as a video prediction task, enabling the use of deep learning for morphological forecasting under varying process conditions, and introduces a novel hyperparameter-fusion neural network, referred to as DepositionNet (DepoNet). Unlike conventional video prediction models, DepoNet specifically accounts for the influence of deposition parameters on the entire simulation process. We have incorporated a novel Hyper Projector that allows the model to flexibly adapt to varying deposition conditions and material characteristics. Through comprehensive comparative experimental analyses, we demonstrate that DepoNet significantly outperforms existing deep-learning models and achieves a mean squared error of 17. 34, representing a 3. 67% improvement over the second best model and a 1, 435 × speedup over physics-based methods, thereby validating its exceptional generalization capability. Extensive experiments reveal that the model maintains high performance even under conditions of limited training data, for instance, achieving a peak signal-to-noise ratio (PSNR) of 41. 516 decibels (dB) when trained with only 20% of the available data.

IROS Conference 2025 Conference Paper

A Safety-Enhanced Autonomous Resection Method for Precision Laparoscopic Surgery amid Tissue Deformation

  • Yudong Shi
  • Hangjie Mo
  • Xilin Xiao
  • Ruiming Duan
  • Ling Li
  • Xiaojian Li

Resection of pathological tissue is a common procedure in surgical oncology for treating tumors. In robot-assisted electrosurgery, the use of predefined markers to guide autonomous robotic resection is gaining traction. Accurate tracking of these markers and minimizing electrocautery damage are critical for the safe and effective autonomous resection of tumors. This paper introduces a safety enhanced autonomous resection method for laparoscopic surgery, designed to mitigate the risks posed by tissue deformation during the resection process. Initially, we pre-plan the cutting path and design a switching strategy for navigation waypoints based on a preview tracking mechanism. Then, we develop a depth-fused navigation controller and a safe withdrawal motion controller. Next, an inertial tracking mechanism is established to evaluate tissue deformation over short periods. Finally, we develop a confidence generator to fuse the two controllers, ensuring that tissue deformation during the resection process does not cause additional electrocautery damage. Simulation and phantom experiments were conducted, demonstrating the effectiveness of our proposed method. This work represents a significant step toward achieving autonomous robotic resection.

IJCAI Conference 2025 Conference Paper

Automated Superscalar Processor Design by Learning Data Dependencies

  • Shuyao Cheng
  • Rui Zhang
  • Wenkai He
  • Pengwei Jin
  • Chongxiao Li
  • Zidong Du
  • Xing Hu
  • Yifan Hao

Automated processor design, which can significantly reduce human efforts and accelerate design cycles, has received considerable attention. While recent advancements have automatically designed single-cycle processors that execute one instruction per cycle, their performance cannot compete with modern superscalar processors that execute multiple instructions per cycle. Previous methods fail on superscalar processor design because they cannot address inter-instruction data dependencies, leading to inefficient sequential instruction execution. This paper proposes a novel approach to automatically designing superscalar processors using a hardware-friendly model called the Stateful Binary Speculation Diagram (State-BSD). We observe that processor parallelism can be enhanced through on-the-fly inter-instruction dependent data predictors, reusing the processor's internal states to learn the data dependency. To meet the challenge of both hardware-resource limitation and design functional correctness, State-BSD consists of two components: 1) a lightweight state-selector trained by simulated annealing method to detect the most reusable processor states and store them in a small buffer; and 2) a highly precise state-speculator trained by BSD expansion method to predict the inter-instruction dependent data using the selected states. It is the first work to achieve the automated superscalar processor design, i. e. QiMeng-CPU-v2, which improves the performance by about 380x than the state-of-the-art automated design and is comparable to human-designed superscalar processors such as ARM Cortex A53.

NeurIPS Conference 2025 Conference Paper

EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization

  • Yize Wu
  • Ke Gao
  • Ling Li
  • Yanjun Wu

Speculative decoding is an effective and lossless method for Large Language Model (LLM) inference acceleration. It employs a smaller model to generate a draft token sequence, which is then verified by the original base model. In multi-GPU systems, inference latency can be further reduced through tensor parallelism (TP), while the optimal TP size of the draft model is typically smaller than that of the base model, leading to GPU idling during the drafting stage. We observe that such inefficiency stems from the sequential execution of layers, which is seemingly natural but actually unnecessary. Therefore, we propose EasySpec, a layer-parallel speculation strategy that optimizes the efficiency of multi-GPU utilization. EasySpec breaks the inter-layer data dependencies in the draft model, enabling multiple layers to run simultaneously across multiple devices as ``fuzzy'' speculation. After each drafting-and-verification iteration, the draft model’s key-value cache is calibrated in a single forward pass, preventing long-term fuzzy-error accumulation at minimal additional latency. EasySpec is a training-free and plug-in method. We evaluated EasySpec on several mainstream open-source LLMs, using smaller versions of models from the same series as drafters. The results demonstrate that EasySpec can achieve a peak speedup of 4. 17x compared to vanilla decoding, while preserving the original distributions of the base LLMs. Specifically, the drafting stage can be accelerated by up to 1. 62x with a maximum speculation accuracy drop of only 7\%. The code is available at https: //github. com/Yize-Wu/EasySpec.

ICML Conference 2025 Conference Paper

Equivalence is All: A Unified View for Self-supervised Graph Learning

  • Yejiang Wang
  • Yuhai Zhao
  • Zhengkui Wang
  • Ling Li
  • Jiapu Wang
  • Fangting Li
  • Miaomiao Huang
  • Shirui Pan

Node equivalence is common in graphs, such as computing networks, encompassing automorphic equivalence (preserving adjacency under node permutations) and attribute equivalence (nodes with identical attributes). Despite their importance for learning node representations, these equivalences are largely ignored by existing graph models. To bridge this gap, we propose a GrAph self-supervised Learning framework with Equivalence (GALE) and analyze its connections to existing techniques. Specifically, we: 1) unify automorphic and attribute equivalence into a single equivalence class; 2) enforce the equivalence principle to make representations within the same class more similar while separating those across classes; 3) introduce approximate equivalence classes with linear time complexity to address the NP-hardness of exact automorphism detection and handle node-feature variation; 4) analyze existing graph encoders, noting limitations in message passing neural networks and graph transformers regarding equivalence constraints; 5) show that graph contrastive learning are a degenerate form of equivalence constraint; and 6) demonstrate that GALE achieves superior performance over baselines.

EAAI Journal 2025 Journal Article

Etching process prediction based on cascade recurrent neural network

  • Zhenjie Yao
  • Ziyi Hu
  • Panpan Lai
  • Fengling Qin
  • Wenrui Wang
  • Zhicheng Wu
  • Lingfei Wang
  • Hua Shao

Etching is one of the most critical processes in semiconductor manufacturing. Etch models have been developed to reveal the underlying etch mechanisms, which employs rigorous physical and chemical process simulation. Traditional simulation is very time consuming. The data-driven artificial intelligence model provides an alternative modeling approach. In this paper, a Cascade Recurrent Neural Networks (CRNN) is proposed to model and predict etching profiles. The etching profile is represented by polar coordinates and modeled by the recurrent neural networks, the corresponding etching parameters (e. g. , pressure, power, temperature, and voltage) are integrated into the network through cascade combination layers. Experimental results on a dataset of 10, 000 simulated etching profiles demonstrated the effectiveness of our method: compared with traditional etching simulation methods, CRNN can speedup 21, 000 × with an average error of less than 0. 7 nm for 1 step prediction. Furthermore, compared to simple deep neural networks, the Mean Absolute Errors (MAE) could be reduced from 1. 7329 nm to 1. 3845 nm for 10 steps prediction. Finally, the effectiveness and accuracy of CRNN etching predictor is validated through fine-tuning on experimental data.

JBHI Journal 2025 Journal Article

Exploring Neural Mechanisms of Visual Working Memory for Real-World Stimuli Categories: Insights from the Fusiform Gyrus

  • Ronglong Xiong
  • Xiaotong Wei
  • Junjun Zhang
  • Zhenlan Jin
  • Ling Li

Objective: To dissect the neural mechanisms underlying visual working memory (VWM) processing of real-world stimuli (Body, Face, Place, Tool). Methods: This study leveraged task-fMRI data from the Human Connectome Project (HCP) n-back paradigm. The Neurofunctional Integration and Specificity Analysis (NISA) framework was proposed to synergistically combine Representational Similarity Analysis (RSA) and multivoxel machine learning classification and regression, enabling distinct characterization of visual perception and VWM processes. Functional connectivity (FC) patterns of NISA-selected regions of interest were further integrated with transcriptomic data to probe molecular substrates. Results: Bilateral fusiform gyrus (FFG) voxel patterns showed maximal stimulus representation fidelity ( r = -0. 43 to -0. 42, q r = 0. 42 to, 0. 56, q ⁻7 ). Transcriptomic decoding revealed associations between bilateral FFG FC profiles and genes implicated in mental and psychiatric disorders ( q < 0. 05). Conclusion: The FFG operates as a dual-process hub, concurrently mediating visual perceptual categorization and working memory maintenance. Its synergistic excitation-inhibition in FFG may optimize the behavior performance through dynamic resource allocation, while FC-transcriptome coupling further revealed gene networks implicated in cognitive vulnerability across VWM categories.

AAAI Conference 2025 Conference Paper

Multi-View 3D Human Pose Estimation with Weakly Synchronized Images

  • Ling Li
  • Ruiwen Gu
  • Chongyang Wang
  • Junliang Xing
  • Xinchun Yu
  • Xiao-Ping Zhang

Multi-view 3D human pose estimation (MHPE) is an important research task in computer vision. To maintain consistency during the data collection, hardware synchronization devices are commonly used to connect cameras, ensuring that images from different views are captured simultaneously. However, synchronizing with extra devices has two apparent limitations: the hardware is i) usually expensive and ii) less flexible for deployment in outdoor open scenarios. Suppose the model can improve its tolerance for the time differences in multi-view image capture. In that case, the difficulty and cost of deployment will be greatly reduced, and MHPE will become more widespread. In this paper, we try to answer how to build a model that performs pose estimation directly using ''weakly synchronized images" from multiple views, where the captured images shift from each other within a frame. To this end, we introduce a new multi-view 3D human pose estimation task given weakly synchronized image inputs. Apart from existing well-synchronized datasets, we present the first weakly synchronized dataset comprising 800k images. Thereon, we propose SyncDiffPose, a novel model based on the diffusion method for pose estimation to denoise the error in such data. By combining simple synchronization strategies, e.g., the timer method, our approach can perform pose estimation without hardware calibration.

ICML Conference 2025 Conference Paper

N2GON: Neural Networks for Graph-of-Net with Position Awareness

  • Yejiang Wang
  • Yuhai Zhao
  • Zhengkui Wang
  • Wen Shan
  • Ling Li
  • Qian Li 0043
  • Miaomiao Huang
  • Meixia Wang

Graphs, fundamental in modeling various research subjects such as computing networks, consist of nodes linked by edges. However, they typically function as components within larger structures in real-world scenarios, such as in protein-protein interactions where each protein is a graph in a larger network. This study delves into the Graph-of-Net (GON), a structure that extends the concept of traditional graphs by representing each node as a graph itself. It provides a multi-level perspective on the relationships between objects, encapsulating both the detailed structure of individual nodes and the broader network of dependencies. To learn node representations within the GON, we propose a position-aware neural network for Graph-of-Net which processes both intra-graph and inter-graph connections and incorporates additional data like node labels. Our model employs dual encoders and graph constructors to build and refine a constraint network, where nodes are adaptively arranged based on their positions, as determined by the network’s constraint system. Our model demonstrates significant improvements over baselines in empirical evaluations on various datasets.

YNIMG Journal 2025 Journal Article

Noise and artifact suppression in SQUID and wearable OPM-MEG: A systematic review of background, physiological, and Technical interference

  • Ruonan Wang
  • Yujie Ma
  • Ruochen Zhao
  • Jin Ding
  • Ling Li
  • Yanfei Yang
  • Fulong Wang
  • Zhiqiang Cao

Magnetoencephalography (MEG) is a non-invasive imaging technique that captures neural activity with high spatio-temporal resolution. In recent years, novel wearable devices based on Optically Pumped Magnetometer (OPM) have emerged as a new driving force for advancing MEG due to their cost-effectiveness, portability, and mobility. In practical applications, MEG signals are frequently influenced by various interference sources, resulting in degradation of signal quality. Consequently, numerous suppression techniques have been proposed to overcome these challenges. This manuscript presents a comprehensive review of the most advanced methods for suppressing MEG noise or artifacts, with a specific focus on mitigating background noise, physiological artifacts (such as those caused by heartbeat, eye movements, and muscle contractions), as well as technical artifacts (including system-related artifacts associated with devices, motion-induced artifacts, and metal-induced artifacts). Additionally, the current limitations and challenges of these approaches in real-world scenarios are highlighted. Reviewing nearly a decade of research, there is an urgent need for a lightweight noise analysis framework in the complex measurement environment of wearable OPM-MEG devices. This framework should be capable of effectively detecting, classifying, and suppressing individual and combined MEG interference. By addressing this need, we can enhance the reliability and practicality of MEG signals while advancing brain science research.

AAAI Conference 2025 Conference Paper

QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models

  • Qirui Zhou
  • Yuanbo Wen
  • Ruizhi Chen
  • Ke Gao
  • Weiqiang Xiong
  • Ling Li
  • Qi Guo
  • Yanjun Wu

As a crucial operator in numerous scientific and engineering computing applications, the automatic optimization of General Matrix Multiplication (GEMM) with full utilization of ever-evolving hardware architectures (e.g. GPUs and RISC-V) is of paramount importance. While Large Language Models (LLMs) can generate functionally correct code for simple tasks, they have yet to produce high-performance code. The key challenge resides in deeply understanding diverse hardware architectures and crafting prompts that effectively unleash the potential of LLMs to generate high-performance code. In this paper, we propose a novel prompt mechanism called QiMeng-GEMM which enables LLMs to comprehend the architectural characteristics of different hardware platforms and automatically search for the optimization combinations for GEMM. The key of QiMeng-GEMM is a set of informative, adaptive, and iterative meta-prompts. Based on this, a searching strategy for optimal combinations of meta-prompts is used to iteratively generate high-performance code. Extensive experiments conducted on 4 leading LLMs, various paradigmatic hardware platforms, and representative matrix dimensions unequivocally demonstrate QiMeng-GEMM’s superior performance in auto-generating optimized GEMM code. Compared to vanilla prompts, our method achieves a performance enhancement of up to 113×. Even when compared to human experts, our method can reach 115% of cuBLAS on NVIDIA GPUs and 211% of OpenBLAS on RISC-V CPUs. Notably, while human experts often take months to optimize GEMM, our approach reduces the development cost by over 240×.

NeurIPS Conference 2025 Conference Paper

QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation

  • Changxin Ke
  • Rui Zhang
  • Shuo Wang
  • Li Ding
  • Guangli Li
  • Yuanbo Wen
  • Shuoming Zhang
  • Ruiyuan Xu

The rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches. However, data scarcity poses a significant challenge for machine learning-based sequential-to-parallel code translation. Although recent back-translation methods show promise, they still fail to ensure functional equivalence in the translated code. In this paper, we propose \textbf{QiMeng-MuPa}, a novel \textbf{Mu}tual-Supervised Learning framework for Sequential-to-\textbf{Pa}rallel code translation, to address the functional equivalence issue. QiMeng-MuPa consists of two models, a Translator and a Tester. Through an iterative loop consisting of Co-verify and Co-evolve steps, the Translator and the Tester mutually generate data for each other and improve collectively. The Tester generates unit tests to verify and filter functionally equivalent translated code, thereby evolving the Translator, while the Translator generates translated code as augmented input to evolve the Tester. Experimental results demonstrate that QiMeng-MuPa significantly enhances the performance of the base models: when applied to Qwen2. 5-Coder, it not only improves Pass@1 by up to 28. 91\% and boosts Tester performance by 68. 90\%, but also outperforms the previous state-of-the-art method CodeRosetta by 1. 56 and 6. 92 in BLEU and CodeBLEU scores, while achieving performance comparable to DeepSeek-R1 and GPT-4. 1. Our code is available at \url{https: //github. com/kcxain/mupa}.

IJCAI Conference 2025 Conference Paper

QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives

  • Xuzhi Zhang
  • Shaohui Peng
  • Qirui Zhou
  • Yuanbo Wen
  • Qi Guo
  • Ruizhi Chen
  • Xinguo Zhu
  • Weiqiang Xiong

Computation-intensive tensor operators constitute over 90% of the computations in Large Language Models (LLMs) and Deep Neural Networks. Automatically and efficiently generating high-performance tensor operators with hardware primitives is crucial for diverse and ever-evolving hardware architectures like RISC-V, ARM, and GPUs, as manually optimized implementation takes at least months and lacks portability. LLMs excel at generating high-level language codes, but they struggle to fully comprehend hardware characteristics and produce high-performance tensor operators. We introduce a tensor-operator auto-generation framework with a one-line user prompt (QiMeng-TensorOp), which enables LLMs to automatically exploit hardware characteristics to generate tensor operators with hardware primitives, and tune parameters for optimal performance across diverse hardware. Experimental results on various hardware platforms, SOTA LLMs, and typical tensor operators demonstrate that QiMeng-TensorOp effectively unleashes the computing capability of various hardware platforms, and automatically generates tensor operators of superior performance. Compared with vanilla LLMs, QiMeng-TensorOp achieves up to 1291× performance improvement. Even compared with human experts, QiMeng-TensorOp could reach 251% of OpenBLAS on RISC-V CPUs, and 124% of cuBLAS on NVIDIA GPUs. Additionally, QiMeng-TensorOp also significantly reduces development costs by 200× compared with human experts.

NeurIPS Conference 2025 Conference Paper

Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models

  • Ling Li
  • Yao Zhou
  • Yuxuan Liang
  • Fugee Tsung
  • Jiaheng Wei

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models (LVLMs) has enabled a rethinking of geo-localization as a reasoning-driven task grounded in visual cues. However, two major challenges persist. On the data side, existing reasoning-focused datasets are primarily based on street-view imagery, offering limited scene diversity and constrained viewpoints. On the modeling side, current approaches predominantly rely on supervised fine-tuning, which yields only marginal improvements in reasoning capabilities. To address these challenges, we propose a novel pipeline that constructs a reasoning-oriented geo-localization dataset, $\textit{MP16-Reason}$, using diverse social media images. We introduce $\textit{GLOBE}$, $\textbf{G}$roup-relative policy optimization for $\textbf{L}$ocalizability assessment and $\textbf{O}$ptimized visual-cue reasoning, yielding $\textbf{B}$i-objective geo-$\textbf{E}$nhancement for the VLM in recognition and reasoning. $\textit{GLOBE}$ incorporates task-specific rewards that jointly enhance localizability assessment, visual-cue reasoning, and geolocation accuracy. Both qualitative and quantitative results demonstrate that $\textit{GLOBE}$ outperforms state-of-the-art open-source LVLMs on geo-localization tasks, particularly in diverse visual scenes, while also generating more insightful and interpretable reasoning trajectories. The data and code are available at https: //github. com/lingli1996/GLOBE.

YNIMG Journal 2025 Journal Article

Study on individual differences in visual working memory tasks based on spatiotemporal brain functional metrics and biological perspectives

  • Ronglong Xiong
  • Qiuzhu Zhang
  • Junjun Zhang
  • Zhenlan Jin
  • Ling Li

Visual working memory (VWM) is a critical area of study in cognitive neuroscience, yet the neural and genetic foundations of individual differences in VWM remain unclear. This study investigates individual differences in VWM performance across four types of visual stimuli (Body, Face, Place, Tool) under 0-back and 2-back conditions by integrating gene expression data and spatiotemporal brain function metrics. First, multiple spatiotemporal brain function metrics were extracted, and Sequential Backward Selection (SBS) and Leave-One-Subject-Out Cross-Validation (LOSO-CV) linear regression were applied to predict behavioral performance under VWM conditions. Model performance was evaluated using RMSE. Next, the Working Memory Individual Differences Map (WMIDM) was constructed based on Pearson correlation coefficients between actual and predicted behavioral performance. Finally, WMIDM was integrated with Allen Human Brain Atlas (AHBA) gene expression data to explore its genetic underpinnings. Notably, the gene analysis is exploratory, providing a preliminary framework for future investigations into the molecular basis of working memory. The results demonstrated that under the 2 vs. 0-back condition, spatiotemporal metrics outperformed static metrics ( r spa = 0. 40, q = 8. 9 × 1 0 − 28, RMSE = 0. 928 vs. r sta = 0. 28, q = 2. 7 × 1 0 − 14, RMSE = 0. 966 ). Brain regions contributing to the WMIDM were primarily located in the frontal lobe. Furthermore, genes associated with WMIDM were significantly enriched in pathways linked to intellectual disability and mental disorders, as well as related biological processes and cell types. This study highlights the neural and potential genetic foundations of individual differences in working memory through the lens of spatiotemporal multidimensional brain function and gene expression. These findings provide valuable insights for future neuroscience research and pave the way for personalized cognitive interventions.

ICRA Conference 2024 Conference Paper

A Force-driven and Vision-driven Hybrid Control Method of Autonomous Laparoscope-Holding Robot

  • Jin Fang
  • Ling Li
  • Xiaojian Li
  • Hangjie Mo
  • Pengxin Guo
  • Xilin Xiao
  • Yanwei Qu

Laparoscope-holding robots significantly enhance the stability and precision of visualization in minimally invasive surgeries. Most existing robots of this kind depend on visual servo systems and struggle with efficient, rapid adjustments in the field-of-view (FOV), especially when identifying organs and needles outside the FOV. This paper presents a laparoscope-holding robot system capable of employing both vision-driven and force-driven mechanisms for continuous and large-scale FOV adjustments, respectively. The system features an integrated tactile handle, enabling the reception of human-robot interaction forces during surgical navigation. We propose a hybrid control method that leverages both force and vision inputs for laparoscopic FOV adjustments. This approach integrates a virtual wrench, generated from visual information, and an interaction wrench, obtained from the tactile handle, into the robot's dynamic model, which complies with remote center of motion constraints. The interaction wrench's gain is adjusted with the gripping force on the integrated tactile handle, ensuring that unintended movements caused by accidental contacts are prevented, thus safeguarding operational safety. The proposed method eliminates the need to switch control modes, enabling simultaneous visual tracking and tactile interaction guidance. Experimental results demonstrate that the proposed method not only allows for FOV adjustments with surgical instrument guiding but also adapts well to large-scale FOV adjustment tasks.

NeurIPS Conference 2024 Conference Paper

DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object Detection

  • Haochen Li
  • Rui Zhang
  • Hantao Yao
  • Xin Zhang
  • Yifan Hao
  • Xinkai Song
  • Xiaqing Li
  • Yongwei Zhao

Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. As the visual-language models (VLMs) can provide essential general knowledge on unseen images, freezing the visual encoder and inserting a domain-agnostic adapter can learn domain-invariant knowledge for DAOD. However, the domain-agnostic adapter is inevitably biased to the source domain. It discards some beneficial knowledge discriminative on the unlabelled domain, \ie domain-specific knowledge of the target domain. To solve the issue, we propose a novel Domain-Aware Adapter (DA-Ada) tailored for the DAOD task. The key point is exploiting domain-specific knowledge between the essential general knowledge and domain-invariant knowledge. DA-Ada consists of the Domain-Invariant Adapter (DIA) for learning domain-invariant knowledge and the Domain-Specific Adapter (DSA) for injecting the domain-specific knowledge from the information discarded by the visual encoder. Comprehensive experiments over multiple DAOD tasks show that DA-Ada can efficiently infer a domain-aware visual encoder for boosting domain adaptive object detection. Our code is available at https: //github. com/Therock90421/DA-Ada.

IROS Conference 2024 Conference Paper

DESectBot: Design and Validation of a Novel Two-Segment Decoupled Continuum Robotic System for Endoscopic Submucosal Dissection

  • Wenjie Liu
  • Yuancheng Shao
  • Yao Zhang 0029
  • Zixi Chen
  • Di Wu 0053
  • Yuqiao Chen
  • Cesare Stefanini
  • Ling Li

Endoscopic Submucosal Dissection (ESD) is a minimally invasive procedure designed to remove precancerous and cancerous lesions from the gastrointestinal (GI) tract. Given the GI tract’s tortuous and narrow shape, along with the need for varied movements during dissection, this requires highly flexible and compact instruments, making flexible continuum robots suitable candidates. In this paper, we propose a novel two-segment continuum robot system named DESectBot, featuring a diameter of 5. 5 mm and a total length of the active bending module of 48 mm, while the robot’s total length exceeds 1 m. We designed a novel joint combination structure called the spatial cross-curved disk skeleton for the robot, which addresses the mechanical coupling problem between flexible robot actuators. The DESectBot boasts six degrees of freedom, and its kinematic modeling has been derived and utilized in the closed-loop control of the DESectBot. The validation of the DESectBot was conducted through a two-stage test: first, the decoupling performance of the DESectBot was validated. The results show that when one active bending segment bends, the other segment remains almost uninfluenced, with a maximum variation of 1. 15 degrees, demonstrating the robot’s effective decoupling capability. Secondly, the accuracy of DESectBot was validated through trajectory-following experiments. The results reveal that the average tracking error for both trajectories is less than 2 mm, and the maximum tracking error is below 2. 5 mm. Taking marking, one of the ESD procedures with a 5mm tolerance, as an example, the DESectBot has the potential to be utilized for ESD procedure.

ICML Conference 2024 Conference Paper

GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model

  • Ling Li
  • Yu Ye 0002
  • Bingchuan Jiang
  • Wei Zeng 0004

This work tackles the problem of geo-localization with a new paradigm using a large vision-language model (LVLM) augmented with human inference knowledge. A primary challenge here is the scarcity of data for training the LVLM - existing street-view datasets often contain numerous low-quality images lacking visual clues, and lack any reasoning inference. To address the data-quality issue, we devise a CLIP-based network to quantify the degree of street-view images being locatable, leading to the creation of a new dataset comprising highly locatable street views. To enhance reasoning inference, we integrate external knowledge obtained from real geo-localization games, tapping into valuable human inference capabilities. The data are utilized to train GeoReasoner, which undergoes fine-tuning through dedicated reasoning and location-tuning stages. Qualitative and quantitative evaluations illustrate that GeoReasoner outperforms counterpart LVLMs by more than 25% at country-level and 38% at city-level geo-localization tasks, and surpasses StreetCLIP performance while requiring fewer training resources. The data and code are available at https: //github. com/lingli1996/GeoReasoner.

AAAI Conference 2024 Conference Paper

Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill Learning

  • Shaohui Peng
  • Xing Hu
  • Qi Yi
  • Rui Zhang
  • Jiaming Guo
  • Di Huang
  • Zikang Tian
  • Ruizhi Chen

Large language models (LLMs) show their powerful automatic reasoning and planning capability with a wealth of semantic knowledge about the human world. However, the grounding problem still hinders the applications of LLMs in the real-world environment. Existing studies try to fine-tune the LLM or utilize pre-defined behavior APIs to bridge the LLMs and the environment, which not only costs huge human efforts to customize for every single task but also weakens the generality strengths of LLMs. To autonomously ground the LLM onto the environment, we proposed the Hypothesis, Verification, and Induction (HYVIN) framework to automatically and progressively ground the LLM with self-driven skill learning. HYVIN first employs the LLM to propose the hypothesis of sub-goals to achieve tasks and then verify the feasibility of the hypothesis via interacting with the underlying environment. Once verified, HYVIN can then learn generalized skills with the guidance of these successfully grounded subgoals. These skills can be further utilized to accomplish more complex tasks that fail to pass the verification phase. Verified in the famous instruction following task set, BabyAI, HYVIN achieves comparable performance in the most challenging tasks compared with imitation learning methods that cost millions of demonstrations, proving the effectiveness of learned skills and showing the feasibility and efficiency of our framework.

YNIMG Journal 2024 Journal Article

Identifying individual's distractor suppression using functional connectivity between anatomical large-scale brain regions

  • Lei Zhuo
  • Zhenlan Jin
  • Ke Xie
  • Simeng Li
  • Feng Lin
  • Junjun Zhang
  • Ling Li

Distractor suppression (DS) is crucial in goal-oriented behaviors, referring to the ability to suppress irrelevant information. Current evidence points to the prefrontal cortex as an origin region of DS, while subcortical, occipital, and temporal regions are also implicated. The present study aimed to examine the contribution of communications between these brain regions to visual DS. To do it, we recruited two independent cohorts of participants for the study. One cohort participated in a visual search experiment where a salient distractor triggering distractor suppression to measure their DS and the other cohort filled out a Cognitive Failure Questionnaire to assess distractibility in daily life. Both cohorts collected resting-state functional magnetic resonance imaging (rs-fMRI) data to investigate function connectivity (FC) underlying DS. First, we generated predictive models of the DS measured in visual search task using resting-state functional connectivity between large anatomical regions. It turned out that the models could successfully predict individual's DS, indicated by a significant correlation between the actual and predicted DS (r = 0.32, p < 0.01). Importantly, Prefrontal-Temporal, Insula-Limbic and Parietal-Occipital connections contributed to the prediction model. Furthermore, the model could also predict individual's daily distractibility in the other independent cohort (r = -0.34, p < 0.05). Our findings showed the efficiency of the predictive models of distractor suppression encompassing connections between large anatomical regions and highlighted the importance of the communications between attention-related and visual information processing regions in distractor suppression. Current findings may potentially provide neurobiological markers of visual distractor suppression.

AAAI Conference 2024 Conference Paper

OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement Learning

  • Fan Wu
  • Rui Zhang
  • Qi Yi
  • Yunkai Gao
  • Jiaming Guo
  • Shaohui Peng
  • Siming Lan
  • Husheng Han

Model-based offline reinforcement learning (RL) algorithms have emerged as a promising paradigm for offline RL. These algorithms usually learn a dynamics model from a static dataset of transitions, use the model to generate synthetic trajectories, and perform conservative policy optimization within these trajectories. However, our observations indicate that policy optimization methods used in these model-based offline RL algorithms are not effective at exploring the learned model and induce biased exploration, which ultimately impairs the performance of the algorithm. To address this issue, we propose Offline Conservative ExplorAtioN (OCEAN), a novel rollout approach to model-based offline RL. In our method, we incorporate additional exploration techniques and introduce three conservative constraints based on uncertainty estimation to mitigate the potential impact of significant dynamic errors resulting from exploratory transitions. Our work is a plug-in method and can be combined with classical model-based RL algorithms, such as MOPO, COMBO, and RAMBO. Experiment results of our method on the D4RL MuJoCo benchmark show that OCEAN significantly improves the performance of existing algorithms.

YNIMG Journal 2024 Journal Article

Task functional networks predict individual differences in the speed of emotional facial discrimination

  • Toluwani Joan Amos
  • Bishal Guragai
  • Qianru Rao
  • Wenjuan Li
  • Zhenlan Jin
  • Junjun Zhang
  • Ling Li

Every individual experiences negative emotions, such as fear and anger, significantly influencing how external information is perceived and processed. With the gradual rise in brain-behavior relationship studies, analyses investigating individual differences in negative emotion processing and a more objective measure such as the response time (RT) remain unexplored. This study aims to address this gap by establishing that the individual differences in the speed of negative facial emotion discrimination can be predicted from whole-brain functional connectivity when participants were performing a face discrimination task. Employing the connectome predictive modeling (CPM) framework, we demonstrated this in the young healthy adult group from the Human Connectome Project-Young Adults (HCP-YA) dataset and the healthy group of the Boston Adolescent Neuroimaging of Depression and Anxiety (BANDA) dataset. We identified distinct network contributions in the adult and adolescent predictive models. The highest represented brain networks involved in the adult model predictions included representations from the motor, visual association, salience, and medial frontal networks. Conversely, the adolescent predictive models showed substantial contributions from the cerebellum-frontoparietal network interactions. Finally, we observed that despite the successful within-dataset prediction in healthy adults and adolescents, the predictive models failed in the cross-dataset generalization. In conclusion, our study shows that individual differences in the speed of emotional facial discrimination can be predicted in healthy adults and adolescent samples using their functional connectivity during negative facial emotion processing. Future research is needed in the derivation of more generalizable models.

IJCAI Conference 2023 Conference Paper

ALL-E: Aesthetics-guided Low-light Image Enhancement

  • Ling Li
  • Dong Liang
  • Yuanhang Gao
  • Sheng-Jun Huang
  • Songcan Chen

Evaluating the performance of low-light image enhancement (LLE) is highly subjective, thus making integrating human preferences into image enhancement a necessity. Existing methods fail to consider this and present a series of potentially valid heuristic criteria for training enhancement models. In this paper, we propose a new paradigm, i. e. , aesthetics-guided low-light image enhancement (ALL-E), which introduces aesthetic preferences to LLE and motivates training in a reinforcement learning framework with an aesthetic reward. Each pixel, functioning as an agent, refines itself by recursive actions, i. e. , its corresponding adjustment curve is estimated sequentially. Extensive experiments show that integrating aesthetic assessment improves both subjective experience and objective evaluation. Our results on various benchmarks demonstrate the superiority of ALL-E over state-of-the-art methods. Source code: https: //dongl-group. github. io/project pages/ALLE. html

AAAI Conference 2023 Conference Paper

Conceptual Reinforcement Learning for Language-Conditioned Tasks

  • Shaohui Peng
  • Xing Hu
  • Rui Zhang
  • Jiaming Guo
  • Qi Yi
  • Ruizhi Chen
  • Zidong Du
  • Ling Li

Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recently, the language-conditioned policy is proposed to facilitate policy transfer through learning the joint representation of observation and text that catches the compact and invariant information across various environments. Existing studies of language-conditioned RL methods often learn the joint representation as a simple latent layer for the given instances (episode-specific observation and text), which inevitably includes noisy or irrelevant information and cause spurious correlations that are dependent on instances, thus hurting generalization performance and training efficiency. To address the above issue, we propose a conceptual reinforcement learning (CRL) framework to learn the concept-like joint representation for language-conditioned policy. The key insight is that concepts are compact and invariant representations in human cognition through extracting similarities from numerous instances in real-world. In CRL, we propose a multi-level attention encoder and two mutual information constraints for learning compact and invariant concepts. Verified in two challenging environments, RTFM and Messenger, CRL significantly improves the training efficiency (up to 70%) and generalization ability (up to 30%) to the new environment dynamics.

NeurIPS Conference 2023 Conference Paper

Context Shift Reduction for Offline Meta-Reinforcement Learning

  • Yunkai Gao
  • Rui Zhang
  • Jiaming Guo
  • Fan Wu
  • Qi Yi
  • Shaohui Peng
  • Siming Lan
  • Ruizhi Chen

Offline meta-reinforcement learning (OMRL) utilizes pre-collected offline datasets to enhance the agent's generalization ability on unseen tasks. However, the context shift problem arises due to the distribution discrepancy between the contexts used for training (from the behavior policy) and testing (from the exploration policy). The context shift problem leads to incorrect task inference and further deteriorates the generalization ability of the meta-policy. Existing OMRL methods either overlook this problem or attempt to mitigate it with additional information. In this paper, we propose a novel approach called Context Shift Reduction for OMRL (CSRO) to address the context shift problem with only offline datasets. The key insight of CSRO is to minimize the influence of policy in context during both the meta-training and meta-test phases. During meta-training, we design a max-min mutual information representation learning mechanism to diminish the impact of the behavior policy on task representation. In the meta-test phase, we introduce the non-prior context collection strategy to reduce the effect of the exploration policy. Experimental results demonstrate that CSRO significantly reduces the context shift and improves the generalization ability, surpassing previous methods across various challenging domains.

NeurIPS Conference 2023 Conference Paper

Contrastive Modules with Temporal Attention for Multi-Task Reinforcement Learning

  • Siming Lan
  • Rui Zhang
  • Qi Yi
  • Jiaming Guo
  • Shaohui Peng
  • Yunkai Gao
  • Fan Wu
  • Ruizhi Chen

In the field of multi-task reinforcement learning, the modular principle, which involves specializing functionalities into different modules and combining them appropriately, has been widely adopted as a promising approach to prevent the negative transfer problem that performance degradation due to conflicts between tasks. However, most of the existing multi-task RL methods only combine shared modules at the task level, ignoring that there may be conflicts within the task. In addition, these methods do not take into account that without constraints, some modules may learn similar functions, resulting in restricting the model's expressiveness and generalization capability of modular methods. In this paper, we propose the Contrastive Modules with Temporal Attention(CMTA) method to address these limitations. CMTA constrains the modules to be different from each other by contrastive learning and combining shared modules at a finer granularity than the task level with temporal attention, alleviating the negative transfer within the task and improving the generalization ability and the performance for multi-task RL. We conducted the experiment on Meta-World, a multi-task RL benchmark containing various robotics manipulation tasks. Experimental results show that CMTA outperforms learning each task individually for the first time and achieves substantial performance improvements over the baselines.

NeurIPS Conference 2023 Conference Paper

Decompose a Task into Generalizable Subtasks in Multi-Agent Reinforcement Learning

  • Zikang Tian
  • Ruizhi Chen
  • Xing Hu
  • Ling Li
  • Rui Zhang
  • Fan Wu
  • Shaohui Peng
  • Jiaming Guo

In recent years, Multi-Agent Reinforcement Learning (MARL) techniques have made significant strides in achieving high asymptotic performance in single task. However, there has been limited exploration of model transferability across tasks. Training a model from scratch for each task can be time-consuming and expensive, especially for large-scale Multi-Agent Systems. Therefore, it is crucial to develop methods for generalizing the model across tasks. Considering that there exist task-independent subtasks across MARL tasks, a model that can decompose such subtasks from the source task could generalize to target tasks. However, ensuring true task-independence of subtasks poses a challenge. In this paper, we propose to \textbf{d}ecompose a \textbf{t}ask in\textbf{to} a series of \textbf{g}eneralizable \textbf{s}ubtasks (DT2GS), a novel framework that addresses this challenge by utilizing a scalable subtask encoder and an adaptive subtask semantic module. We show that these components endow subtasks with two properties critical for task-independence: avoiding overfitting to the source task and maintaining consistent yet scalable semantics across tasks. Empirical results demonstrate that DT2GS possesses sound zero-shot generalization capability across tasks, exhibits sufficient transferability, and outperforms existing methods in both multi-task and single-task problems.

NeurIPS Conference 2023 Conference Paper

Efficient Symbolic Policy Learning with Differentiable Symbolic Expression

  • Jiaming Guo
  • Rui Zhang
  • Shaohui Peng
  • Qi Yi
  • Xing Hu
  • Ruizhi Chen
  • Zidong Du
  • Xishan Zhang

Deep reinforcement learning (DRL) has led to a wide range of advances in sequential decision-making tasks. However, the complexity of neural network policies makes it difficult to understand and deploy with limited computational resources. Currently, employing compact symbolic expressions as symbolic policies is a promising strategy to obtain simple and interpretable policies. Previous symbolic policy methods usually involve complex training processes and pre-trained neural network policies, which are inefficient and limit the application of symbolic policies. In this paper, we propose an efficient gradient-based learning method named Efficient Symbolic Policy Learning (ESPL) that learns the symbolic policy from scratch in an end-to-end way. We introduce a symbolic network as the search space and employ a path selector to find the compact symbolic policy. By doing so we represent the policy with a differentiable symbolic expression and train it in an off-policy manner which further improves the efficiency. In addition, in contrast with previous symbolic policies which only work in single-task RL because of complexity, we expand ESPL on meta-RL to generate symbolic policies for unseen tasks. Experimentally, we show that our approach generates symbolic policies with higher performance and greatly improves data efficiency for single-task RL. In meta-RL, we demonstrate that compared with neural network policies the proposed symbolic policy achieves higher performance and efficiency and shows the potential to be interpretable.

NeurIPS Conference 2023 Conference Paper

GALOPA: Graph Transport Learning with Optimal Plan Alignment

  • Yejiang Wang
  • Yuhai Zhao
  • Daniel Zhengkui Wang
  • Ling Li

Self-supervised learning on graph aims to learn graph representations in an unsupervised manner. While graph contrastive learning (GCL - relying on graph augmentation for creating perturbation views of anchor graphs and maximizing/minimizing similarity for positive/negative pairs) is a popular self-supervised method, it faces challenges in finding label-invariant augmented graphs and determining the exact extent of similarity between sample pairs to be achieved. In this work, we propose an alternative self-supervised solution that (i) goes beyond the label invariance assumption without distinguishing between positive/negative samples, (ii) can calibrate the encoder for preserving not only the structural information inside the graph, but the matching information between different graphs, (iii) learns isometric embeddings that preserve the distance between graphs, a by-product of our objective. Motivated by optimal transport theory, this scheme relays on an observation that the optimal transport plans between node representations at the output space, which measure the matching probability between two distributions, should be consistent to the plans between the corresponding graphs at the input space. The experimental findings include: (i) The plan alignment strategy significantly outperforms the counterpart using the transport distance; (ii) The proposed model shows superior performance using only node attributes as calibration signals, without relying on edge information; (iii) Our model maintains robust results even under high perturbation rates; (iv) Extensive experiments on various benchmarks validate the effectiveness of the proposed method.

NeurIPS Conference 2023 Conference Paper

Learning Domain-Aware Detection Head with Prompt Tuning

  • Haochen Li
  • Rui Zhang
  • Hantao Yao
  • Xinkai Song
  • Yifan Hao
  • Yongwei Zhao
  • Ling Li
  • Yunji Chen

Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. However, existing methods focus on reducing the domain bias of the detection backbone by inferring a discriminative visual encoder, while ignoring the domain bias in the detection head. Inspired by the high generalization of vision-language models (VLMs), applying a VLM as the robust detection backbone following a domain-aware detection head is a reasonable way to learn the discriminative detector for each domain, rather than reducing the domain bias in traditional methods. To achieve the above issue, we thus propose a novel DAOD framework named Domain-Aware detection head with Prompt tuning (DA-Pro), which applies the learnable domain-adaptive prompt to generate the dynamic detection head for each domain. Formally, the domain-adaptive prompt consists of the domain-invariant tokens, domain-specific tokens, and the domain-related textual description along with the class label. Furthermore, two constraints between the source and target domains are applied to ensure that the domain-adaptive prompt can capture the domains-shared and domain-specific knowledge. A prompt ensemble strategy is also proposed to reduce the effect of prompt disturbance. Comprehensive experiments over multiple cross-domain adaptation tasks demonstrate that using the domain-adaptive prompt can produce an effectively domain-related detection head for boosting domain-adaptive object detection. Our code is available at https: //github. com/Therock90421/DA-Pro.

JBHI Journal 2023 Journal Article

Selective Enhancement of Frontal-Posterior Functional Connectivity by Anodal tDCS Over the Right Posterior Parietal Cortex During Temporal Attention

  • Bo Tan
  • Qian Liao
  • Ping Xu
  • Junjun Zhang
  • Zhenlan Jin
  • Ling Li

Temporal attention is the concentration of perceptual resources at a specific point in time, which can help individuals get prepared to improve their behavioral performance, whereas the neural mechanism of temporal attention is yet to be well understood. In this study, behavioral measurement, transcranial direct current stimulation (tDCS), and electroencephalography (EEG) were combined to explore the effects of task performance and whole-brain functional connectivities (FCs) during temporal attention with different time intervals after applying anodal and sham tDCS over the right posterior parietal cortex (PPC). Although anodal tDCS, compared with sham tDCS, did not induce a significant effect on the task performance of temporal attention, it could effectively increase long-range FCs of gamma rhythms between the right frontal and parieto-occipital regions during temporal attention, and most of the increased FCs were in the right hemisphere with certain hemispheric laterality. Meanwhile, there were intensively more increased long-range FCs at short-time intervals than those at long-time intervals, and the increased FCs at neutral long-time intervals were the least and mainly inter-hemispheric FCs. The current study not only further enriched the evidence on the key role of the right PPC during temporal attention but also proved that anodal tDCS could indeed enhance whole-brain functional connectivity architecture involving intra- and inter-hemispheric long-range FCs, which would provide ideas and references for subsequent studies of temporal attention as well as attention deficit disorder.

EAAI Journal 2023 Journal Article

Self-Attention Causal Dilated Convolutional Neural Network for Multivariate Time Series Classification and Its Application

  • Wenbiao Yang
  • Kewen Xia
  • Zhaocheng Wang
  • Shurui Fan
  • Ling Li

Time Series Classification (TSC) in data mining is gradually developing as an important research direction. Many researchers have developed an extensive interest in Multivariate Time Series Classification (MTSC). The Self-Attention Causal Dilated Convolutional Neural Network (SACDCNN) is proposed to address the limitations of existing models that perform poorly on classification tasks. It designs the residual and dense blocks based on Causal Dilated Convolution based on the traditional residual and dense networks that still have superior performance after deepening the network hierarchy and the dependence of time series on long-range information. A Self-Attention mechanism (SA) is also incorporated to extract the internal autocorrelation of time series features. Comparison experiments on 20 benchmark University of California, Riverside (UCR) and University of California, Irvine (UCI) datasets with eight high-performance classification models show that the method can improve the classification accuracy of time series datasets. Finally, it was applied to petroleum logging reservoir recognition, and a comparison experiment was conducted on two wells. The results show that SACDCNN is effective and significantly superior. It overcomes the shortcomings of traditional logging interpretation techniques and improves the efficiency and success rate of oil and gas exploration.

NeurIPS Conference 2022 Conference Paper

Causality-driven Hierarchical Structure Discovery for Reinforcement Learning

  • Shaohui Peng
  • Xing Hu
  • Rui Zhang
  • Ke Tang
  • Jiaming Guo
  • Qi Yi
  • Ruizhi Chen
  • Xishan Zhang

Hierarchical reinforcement learning (HRL) has been proven to be effective for tasks with sparse rewards, for it can improve the agent's exploration efficiency by discovering high-quality hierarchical structures (e. g. , subgoals or options). However, automatically discovering high-quality hierarchical structures is still a great challenge. Previous HRL methods can only find the hierarchical structures in simple environments, as they are mainly achieved through the randomness of agent's policies during exploration. In complicated environments, such a randomness-driven exploration paradigm can hardly discover high-quality hierarchical structures because of the low exploration efficiency. In this paper, we propose CDHRL, a causality-driven hierarchical reinforcement learning framework, to build high-quality hierarchical structures efficiently in complicated environments. The key insight is that the causalities among environment variables are naturally fit for modeling reachable subgoals and their dependencies; thus, the causality is suitable to be the guidance in building high-quality hierarchical structures. Roughly, we build the hierarchy of subgoals based on causality autonomously, and utilize the subgoal-based policies to unfold further causality efficiently. Therefore, CDHRL leverages a causality-driven discovery instead of a randomness-driven exploration for high-quality hierarchical structure construction. The results in two complex environments, 2D-Minecraft and Eden, show that CDHRL can discover high-quality hierarchical structures and significantly enhance exploration efficiency.

AAAI Conference 2022 Conference Paper

Conditional Local Convolution for Spatio-Temporal Meteorological Forecasting

  • Haitao Lin
  • Zhangyang Gao
  • Yongjie Xu
  • Lirong Wu
  • Ling Li
  • Stan Z. Li

Spatio-temporal forecasting is challenging attributing to the high nonlinearity in temporal dynamics as well as complex location-characterized patterns in spatial domains, especially in fields like weather forecasting. Graph convolutions are usually used for modeling the spatial dependency in meteorology to handle the irregular distribution of sensors' spatial location. In this work, a novel graph-based convolution for imitating the meteorological flows is proposed to capture the local spatial patterns. Based on the assumption of smoothness of location-characterized patterns, we propose conditional local convolution whose shared kernel on nodes' local space is approximated by feedforward networks, with local representations of coordinate obtained by horizon maps into cylindrical-tangent space as its input. The established united standard of local coordinate system preserves the orientation on geography. We further propose the distance and orientation scaling terms to reduce the impacts of irregular spatial distribution. The convolution is embedded in a Recurrent Neural Network architecture to model the temporal dynamics, leading to the Conditional Local Convolution Recurrent Network (CLCRN). Our model is evaluated on real-world weather benchmark datasets, achieving state-of-the-art performance with obvious improvements. We conduct further analysis on local pattern visualization, model's framework choice, advantages of horizon maps and etc. The source code is available at https://github.com/BIRD-TAO/CLCRN.

AAAI Conference 2022 Conference Paper

Semantically Contrastive Learning for Low-Light Image Enhancement

  • Dong Liang
  • Ling Li
  • Mingqiang Wei
  • Shuo Yang
  • Liyan Zhang
  • Wenhan Yang
  • Yun Du
  • Huiyu Zhou

Low-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weakvisibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question – if leveraging both accessible unpaired over/underexposed images and high-level semantic guidance, can improve the performance of cutting-edge LLE models? Here, we propose an effective semantically contrastive learning paradigm for LLE (namely SCL-LLE). Beyond the existing LLE wisdom, it casts the image enhancement task as multi-task joint learning, where LLE is converted into three constraints of contrastive learning, semantic brightness consistency, and feature preservation for simultaneously ensuring the exposure, texture, and color consistency. SCL-LLE allows the LLE model to learn from unpaired positives (normal-light)/negatives (over/underexposed), and enables it to interact with the scene semantics to regularize the image enhancement network, yet the interaction of high-level semantic knowledge and the lowlevel signal prior is seldom investigated in previous methods. Training on readily available open data, extensive experiments demonstrate that our method surpasses the state-of-thearts LLE models over six independent cross-scenes datasets. Moreover, SCL-LLE’s potential to benefit the downstream semantic segmentation under extremely dark conditions is discussed. Source Code: https: //github. com/LingLIx/SCL-LLE.

YNIMG Journal 2022 Journal Article

Shared and distinct structure-function substrates of heterogenous distractor suppression ability between high and low working memory capacity individuals

  • Ke Xie
  • Zhenlan Jin
  • Dong-Gang Jin
  • Junjun Zhang
  • Ling Li

Salient stimuli can capture attention in a bottom-up manner; however, this attentional capture can be suppressed in a top-down manner. It has been shown that individuals with high working memory capacity (WMC) can suppress salient‑but-irrelevant distractors better than those with low WMC; however, neural substrates underlying this difference remain unclear. To examine this, participants with high or low WMC (high-/low-WMC, n = 44/44) performed a visual search task wherein a color singleton item served as a salient distractor, and underwent structural and resting-state functional magnetic resonance imaging scans. Behaviorally, the color singleton distractor generally reduced the reaction time (RT). This RT benefit (ΔRT) was higher in the high-WMC group relative to the low-WMC group, indicating the superior distractor suppression ability of the high-WMC group. Moreover, leveraging voxel-based morphometry analysis, gray matter morphology (volume and deformation) in the ventral attention network (VAN) was found to show the same, positive associations with ΔRT in both WMC groups. However, correlations of the opposite sign were found between ΔRT and gray matter morphology in the frontoparietal (FPN)/default mode network (DMN) in the two WMC groups. Furthermore, resting-state functional connectivity analysis centering on regions with a structural-behavioral relationship found that connections between the left orbital and right superior frontal gyrus (hubs of DMN and VAN, respectively) was correlated with ΔRT in the high-WMC group (but not in the low-WMC group). Collectively, our work present shared and distinct neuroanatomical substrates of distractor suppression in high- and low-WMC individuals. Furthermore, intrinsic connectivity of the brain network hubs in high-WMC individuals may account for their superior ability in suppressing salient distractors.

YNICL Journal 2021 Journal Article

Disturbed temporal dynamics of episodic retrieval activity with preserved spatial activity pattern in amnestic mild cognitive impairment: A simultaneous EEG-fMRI study

  • Hao Shu
  • Lihua Gu
  • Ping Yang
  • Molly V. Lucas
  • Lijuan Gao
  • Hongxing Zhang
  • Haisan Zhang
  • Zhan Xu

Episodic memory (EM) deficit is the core cognitive dysfunction of amnestic mild cognitive impairment (aMCI). However, the episodic retrieval pattern detected by functional MRI (fMRI) appears preserved in aMCI subjects. To address this discrepancy, simultaneous electroencephalography (EEG)-fMRI recording was employed to determine whether temporal dynamics of brain episodic retrieval activity were disturbed in patients with aMCI. Twenty-six aMCI and 29 healthy control (HC) subjects completed a word-list memory retrieval task during simultaneous EEG-fMRI. The retrieval success activation pattern was detected with fMRI analysis, and the familiarity- and recollection-related components of episodic retrieval activity were identified using event-related potential (ERP) analysis. The fMRI-constrained ERP analysis explored the temporal dynamics of brain activity in the retrieval success pattern, and the ERP-informed fMRI analysis detected fMRI correlates of the ERP components related to familiarity and recollection processes. The two groups exhibited similar retrieval success patterns in the bilateral posteromedial parietal cortex, the left inferior parietal lobule (IPL), and the left lateral prefrontal cortex (LPFC). The fMRI-constrained ERP analysis showed that the aMCI group did not exhibit old/new effects in the IPL and LPFC that were observed in the HC group. In addition, the aMCI group showed disturbed fMRI correlate of ERP recollection component that was associated with inferior EM performance. Therefore, in this study, we identified disturbed temporal dynamics in episodic retrieval activity with a preserved spatial activity pattern in aMCI. Taken together, the simultaneous EEG-fMRI technique demonstrated the potential to identify individuals with a high risk of cognitive deterioration.

YNICL Journal 2021 Journal Article

Hippocampal subfield and anterior-posterior segment volumes in patients with sporadic amyotrophic lateral sclerosis

  • Shuangwu Liu
  • Qingguo Ren
  • Gaolang Gong
  • Yuan Sun
  • Bing Zhao
  • Xiaotian Ma
  • Na Zhang
  • Suyu Zhong

Neuroimaging studies of hippocampal volumes in patients with amyotrophic lateral sclerosis (ALS) have reported inconsistent results. Our aims were to demonstrate that such discrepancies are largely due to atrophy of different regions of the hippocampus that emerge in different disease stages of ALS and to explore the existence of co-pathology in ALS patients. We used the well-validated King's clinical staging system for ALS to classify patients into different disease stages. We investigated in vivo hippocampal atrophy patterns across subfields and anterior-posterior segments in different King's stages using structural MRI in 76 ALS patients and 94 health controls (HCs). The thalamus, corticostriatal tract and perforant path were used as structural controls to compare the sequence of alterations between these structures and the hippocampal subfields. Compared with HCs, ALS patients at King's stage 1 had lower volumes in the bilateral posterior subiculum and presubiculum; ALS patients at King's stage 2 exhibited lower volumes in the bilateral posterior subiculum, left anterior presubiculum and left global hippocampus; ALS patients at King's stage 3 showed significantly lower volumes in the bilateral posterior subiculum, dentate gyrus and global hippocampus. Thalamic atrophy emerged at King's stage 3. White matter tracts remained normal in a subset of ALS patients. Our study demonstrated that the pattern of hippocampal atrophy in ALS patients varies greatly across King's stages. Future studies in ALS patients that focus on the hippocampus may help to further clarify possible co-pathologies in ALS.

YNIMG Journal 2021 Journal Article

Where there is no object formation, there is no perceptual organization: Evidence from the configural superiority effect

  • Junjun Zhang
  • Xiaoyan Yang
  • Zhenlan Jin
  • Ling Li

Object formation is considered the aim of perceptual organization, but such a proposition has been neglected in empirical studies. In the current study, we investigated the role of object formation in configural superiority. Essentially, discrimination on bar orientations was enhanced by adding a right angle to each of the bars. Such facilitation is due to the emergent feature (EF) of closure formed by combining the bars with right angles. To study object formation, visual stimuli were generated by random dot stereograms to form objects or holes in 3D. Behaviorally, we found that the EF of closure facilitated oddball discrimination on objects, as demonstrated by previous studies, but did not facilitate oddball discrimination on holes with the same shape as objects. Multivariate pattern analysis of functional magnetic resonance imaging (fMRI) data showed that the EF of closure increased the object classification accuracy compared to the holes in the lateral occipital cortex (LOC), where object information is encoded, but not in the early visual cortex (EVC). The neural representations of objects and holes with and without EFs were further investigated using representational similarity analysis. The results demonstrate that in the LOC, the neural representations of objects with EFs showed a greater difference than those of the other three, that is, objects without EFs and holes with or without EFs. However, the uniqueness of objects with EFs was not observed in the EVC. Thus, our results suggest that the EF of closure, which leads to the configural superiority effect, only emerges for objects but not for holes, and only in the LOC but not the EVC. Our study provides the first empirical evidence suggesting that object formation plays an indispensable role in perceptual organization.

YNIMG Journal 2020 Journal Article

Distinct neural substrates underlying target facilitation and distractor suppression: A combined voxel-based morphometry and resting-state functional connectivity study

  • Ke Xie
  • Zhenlan Jin
  • Xuejin Ni
  • Junjun Zhang
  • Ling Li

Selective attention, the ability to filter relevant from a sea of sensory information, relies on the prioritization of goal-relevant information (target facilitation) and the suppression of goal-irrelevant information (distractor suppression). Although several lines of evidence have shown that target facilitation and distractor suppression were mediated by distinct mechanisms, the underlying neural substrates remain unclear. To address this question, we acquired structural and resting-state magnetic resonance imaging scans, as well as behavioral data from a modified Posner cueing task. Specifically, the location of a target (Target Cue, TC) and a distractor (Distractor Cue, DC) was either cued in advance to separately trigger target facilitation and distractor suppression, or no predictive information was provided, serving as a baseline. We combined voxel-based morphometry (VBM) and resting-state functional connectivity (rsFC) analyses to explore the neural correlates of behavioral benefits, yielding the following results. First, behavioral data showed faster responses to TC and DC conditions compared to baseline, the benefits of which were named TC-benefit and DC-benefit. Second, the VBM analysis revealed that the gray matter volume (GMV) in the superior frontal (SFG) and postcentral gyrus inversely correlated with individual TC-benefit, while the GMV in the superior parietal lobe, middle frontal gyrus, and angular gyrus inversely correlated with individual DC-benefit, indicating that target facilitation and distractor suppression was associated with the GMV of distinct and distributed regions in the frontoparietal cortex. Third, the rsFC analysis with the SFG as a seed region further found distinct patterns of rsFC for target facilitation and distractor suppression. Specifically, individual TC-benefit were positively correlated with distributed connections between the SFG and brain regions, mainly within the ventral attention and somato-motor network; but individual DC-benefit were positively correlated with centralized connections between the SFG and brain regions, mainly within the frontoparietal, dorsal attention and ventral attention network. Finally, a multiple linear regression analysis showed that the GMV and rsFC could jointly explain individual differences in TC- and DC-benefit. Taken together, these results provided neural evidence for different structural and functional substrates underlying target facilitation and distractor suppression.

YNICL Journal 2020 Journal Article

Reproducible metabolic topographies associated with multiple system atrophy: Network and regional analyses in Chinese and American patient cohorts

  • Bo Shen
  • Sidi Wei
  • Jingjie Ge
  • Shichun Peng
  • Fengtao Liu
  • Ling Li
  • Sisi Guo
  • Ping Wu

PURPOSE: Multiple system atrophy (MSA) is an atypical parkinsonian syndrome and often difficult to discriminate clinically from progressive supranuclear palsy (PSP) and Parkinson's disease (PD) in early stages. Although a characteristic metabolic brain network has been reported for MSA, it is unknown whether this network can provide a clinically useful biomarker in different centers. This study was aimed to identify and cross-validate MSA-related brain network and assess its ability for differential diagnosis and clinical correlations in Chinese and American patient cohorts. METHODS: F-FDG PET scans retrospectively from 128 clinically diagnosed parkinsonian patients (34 MSA, 34 PSP and 60 PD) and 40 normal subjects in China and in the USA. Using PET images from 20 moderate-stage MSA patients of parkinsonian subtype and 20 normal subjects in both centers, we reproduced MSA-related pattern (MSAPRP) of spatial covariance and estimated its reliability. MSAPRP scores were evaluated in assessing differential diagnosis among moderate- and early-stage MSA, PSP or PD patients and clinical correlations with disease severity. Regional metabolic differences were detected using statistical parameter mapping analysis. MSA-related network and regional topographies of metabolic abnormality were cross-validated between the Chinese and American cohorts. RESULTS: We generated a highly reliable MSAPRP characterized by decreased loading in inferior frontal cortex, striatum and cerebellum, and increased loading in sensorimotor, parietal and occipital cortices. MSAPRP scores discriminated between normal, MSA, PSP and PD subjects and correlated with standardized ratings of clinical stages and motor symptoms in MSA. High similarities in MSAPRPs, network scores and corresponding maps of metabolic abnormality were observed between two different cohorts. CONCLUSION: We have demonstrated reproducible metabolic topographies associated with MSA at both network and regional levels in two independent patient cohorts. Moreover, MSAPRP scores are sensitive for evaluating disease discrimination and clinical correlates. This study supports differential diagnosis of MSA regardless of different patient populations, PET scanners and imaging protocols.

YNIMG Journal 2017 Journal Article

The effects of changes in object location on object identity detection: A simultaneous EEG-fMRI study

  • Ping Yang
  • Chenggui Fan
  • Min Wang
  • Noa Fogelson
  • Ling Li

Object identity and location are bound together to form a unique integration that is maintained and processed in visual working memory (VWM). Changes in task-irrelevant object location have been shown to impair the retrieval of memorial representations and the detection of object identity changes. However, the neural correlates of this cognitive process remain largely unknown. In the present study, we aim to investigate the underlying brain activation during object color change detection and the modulatory effects of changes in object location and VWM load. To this end we used simultaneous electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) recordings, which can reveal the neural activity with both high temporal and high spatial resolution. Subjects responded faster and with greater accuracy in the repeated compared to the changed object location condition, when a higher VWM load was utilized. These results support the spatial congruency advantage theory and suggest that it is more pronounced with higher VWM load. Furthermore, the spatial congruency effect was associated with larger posterior N1 activity, greater activation of the right inferior frontal gyrus (IFG) and less suppression of the right supramarginal gyrus (SMG), when object location was repeated compared to when it was changed. The ERP-fMRI integrative analysis demonstrated that the object location discrimination-related N1 component is generated in the right SMG.

TIST Journal 2013 Journal Article

Effective and efficient microprocessor design space exploration using unlabeled design configurations

  • Tianshi Chen
  • Yunji Chen
  • Qi Guo
  • Zhi-Hua Zhou
  • Ling Li
  • Zhiwei Xu

Ever-increasing design complexity and advances of technology impose great challenges on the design of modern microprocessors. One such challenge is to determine promising microprocessor configurations to meet specific design constraints, which is called Design Space Exploration (DSE). In the computer architecture community, supervised learning techniques have been applied to DSE to build regression models for predicting the qualities of design configurations. For supervised learning, however, considerable simulation costs are required for attaining the labeled design configurations. Given limited resources, it is difficult to achieve high accuracy. In this article, inspired by recent advances in semisupervised learning and active learning, we propose the COAL approach which can exploit unlabeled design configurations to significantly improve the models. Empirical study demonstrates that COAL significantly outperforms a state-of-the-art DSE technique by reducing mean squared error by 35% to 95%, and thus, promising architectures can be attained more efficiently.

JMLR Journal 2008 Journal Article

Support Vector Machinery for Infinite Ensemble Learning

  • Hsuan-Tien Lin
  • Ling Li

Ensemble learning algorithms such as boosting can achieve better performance by averaging over the predictions of some base hypotheses. Nevertheless, most existing algorithms are limited to combining only a finite number of hypotheses, and the generated ensemble is usually sparse. Thus, it is not clear whether we should construct an ensemble classifier with a larger or even an infinite number of hypotheses. In addition, constructing an infinite ensemble itself is a challenging task. In this paper, we formulate an infinite ensemble learning framework based on the support vector machine (SVM). The framework can output an infinite and nonsparse ensemble through embedding infinitely many hypotheses into an SVM kernel. We use the framework to derive two novel kernels, the stump kernel and the perceptron kernel. The stump kernel embodies infinitely many decision stumps, and the perceptron kernel embodies infinitely many perceptrons. We also show that the Laplacian radial basis function kernel embodies infinitely many decision trees, and can thus be explained through infinite ensemble learning. Experimental results show that SVM with these kernels is superior to boosting with the same base hypothesis set. In addition, SVM with the stump kernel or the perceptron kernel performs similarly to SVM with the Gaussian radial basis function kernel, but enjoys the benefit of faster parameter selection. These properties make the novel kernels favorable choices in practice. [abs] [ pdf ][ bib ] &copy JMLR 2008. ( edit, beta )

NeurIPS Conference 2006 Conference Paper

Ordinal Regression by Extended Binary Classification

  • Ling Li
  • Hsuan-Tien Lin

We present a reduction framework from ordinal regression to binary classification based on extended examples. The framework consists of three steps: extracting extended examples from the original examples, learning a binary classifier on the extended examples with any binary classification algorithm, and constructing a ranking rule from the binary classifier. A weighted 0/1 loss of the binary classifier would then bound the mislabeling cost of the ranking rule. Our framework allows not only to design good ordinal regression algorithms based on well-tuned binary classification approaches, but also to derive new generalization bounds for ordinal regression from known bounds for binary classification. In addition, our framework unifies many existing ordinal regression algorithms, such as perceptron ranking and support vector ordinal regression. When compared empirically on benchmark data sets, some of our newly designed algorithms enjoy advantages in terms of both training speed and generalization performance over existing algorithms, which demonstrates the usefulness of our framework.

v2026.09.13