Arrow Research search

Author name cluster

Qi Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

145 papers
2 author rows

Possible papers

145

TAAS Journal 2026 Journal Article

A Unified Framework for Noisy Image Super-Resolution

  • Ziang Wu
  • Qi Zhang
  • Xiaoli Sun
  • Yonglin Tian
  • Qi Zhu
  • Chunwei Tian

High-quality images are essential for human-computer interaction in industrial systems; however, captured images are often degraded by device vibrations and moving objects. The single-image super-resolution (SISR) task aims to reconstruct high-quality images from low-quality inputs, where deep networks have demonstrated significant success. Nevertheless, noisy image super-resolution remains challenging due to the difficulty of separating noise from genuine structural details. Unlike traditional methods that struggle to distinguish noise from actual image details, the proposed unified framework for noisy image super-resolution (UFNet) combines discriminative learning with a degradation model to enhance both noise suppression and detail recovery. UFNet employs two parallel networks to extract informative features for reconstructing high-quality images. The upper branch utilizes a discriminative learning strategy to remove noise, while the lower branch applies the concept of a degradation model to recover structural details. To restore lost details while maintaining naturalness and structural consistency, a Feature Distillation and Refinement Block (FDRB) is embedded in the lower network. Furthermore, a refinement network is employed to eliminate redundant information introduced during the fusion operation, thereby unifying the results of the two branches and enhancing the final super-resolution outcomes. Extensive experiments demonstrate that the proposed UFNet achieves excellent performance in noisy image super-resolution. The source code of UFNet is available at https://github.com/WuZiang73/UFNet.

AAAI Conference 2026 Conference Paper

DFRec: Dual Fluctuation Modeling of Multi-level Intent Evolution for Next-Item Recommendation

  • Nengjun Zhu
  • Lingdan Sun
  • Qi Zhang
  • Jian Cao
  • Hang Yu

User sequential behaviors are driven by a variety of complex and evolving intents. Capturing the dynamic change of user intents has become critical yet challenging in the next-item recommendation. Existing studies usually model the transition relationships among multiple intents within a session or integrate temporal information to capture the dynamic evolution of user intents. However, they struggle to identify the precise timing and magnitudes of intent changes, leading to ambiguity in providing consistent or violated recommendations and ultimately yielding subpar performance. To this end, we propose a novel framework called Dual Fluctuation Modeling of Multi-level Intent Evolution for Next-Item Recommendation (DFRec) in this paper. DFRec explicitly identifies the user intent changes and further quantifies the magnitude of the changes. Specifically, we assume that a user's intent fluctuates around an inherent intent, with the magnitude of fluctuations indicating the extent of changes in user intents. Thus, we design an Emerging Intent Generation Module that employs a normal distribution with dynamic variance to capture intent fluctuations at each time step. Furthermore, we introduce a dual-layer dynamic variance update mechanism to capture fluctuation characteristics at different temporal levels, enhancing the representation of possible emergent intents. Extensive experiments on three real-world datasets verify DFRec's superiority over state-of-the-art baselines.

EAAI Journal 2026 Journal Article

Feature-driven machine learning for prediction and inverse design of bistable twist metastructures with highly nonlinear responses

  • Peiyuan Zheng
  • Bin Han
  • Hao Wang
  • Zhipeng Liu
  • Qinze Wang
  • Qi Zhang

Mechanical metastructures offer promising applications in energy-absorbing devices and soft actuators. However, their highly nonlinear responses, especially sudden force drop within a quite narrow displacement range, pose a significant challenge in accurate prediction and inverse optimization. This study introduces a feature-driven machine learning approach that significantly enhances the predictive accuracy and efficiency of highly nonlinear responses. Herein, bistable twist metastructures with highly nonlinear sudden force drop responses are utilized as the platform. By employing feature-driven classification and reconstruction, force-displacement curves are modularized according to the feature types. The feature-driven sampling strategy increases the sampling density in highly nonlinear modules while reducing the redundant points in approximately linear modules. This facilitates the optimized distribution of sampling points in each module, ensuring the integrity of highly nonlinear features. Moreover, the feature-driven machine learning approach reduces the severe reliance on large data amount. Utilizing only 1500 data groups, a mapping paradigm between 9 geometric parameters and highly nonlinear responses is established through machine learning within only 1∼3 min, achieving a significantly higher accuracy. Additionally, inverse design is accomplished through the integration of trained network and search algorithms. Targeted on the pronounced structural hysteresis, which is advantageous to cushioning performance, desired mechanical responses are identified, exhibiting an enhancement exceeding 30%, with the corresponding geometric parameters effectively determined. It is believed that this feature-driven modular approach holds immense potentials for both precise prediction and effective optimization of mechanical metastructures with highly nonlinear responses.

AAAI Conference 2026 Conference Paper

Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths

  • Rui Yao
  • Qi Chai
  • Jinhai Yao
  • Siyuan Li
  • Junhao Chen
  • Qi Zhang
  • Hao Wang

"Fedspeak", the stylized and often nuanced language used by the U.S. Federal Reserve, encodes implicit policy signals and strategic stances. The Federal Open Market Committee strategically employs Fedspeak as a communication tool to shape market expectations and influence both domestic and global economic conditions. As such, automatically parsing and interpreting Fedspeak presents a high-impact challenge, with significant implications for financial forecasting, algorithmic trading, and data-driven policy analysis. Technically, to enrich the semantic and contextual representation of Fedspeak texts, we incorporate domain-specific reasoning grounded in the monetary policy transmission mechanism. We further introduce a dynamic uncertainty decoding module to assess the confidence of model predictions, thereby enhancing both classification accuracy and model reliability. Experimental results demonstrate that our framework achieves state-of-the-art performance on the policy stance analysis task. Moreover, statistical analysis reveals a significant positive correlation between perceptual uncertainty and model error rates, validating the effectiveness of perceptual uncertainty as a diagnostic signal.

AAAI Conference 2026 Conference Paper

MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning

  • Zhiheng Xi
  • Yuhui Wang
  • Yiwen Ding
  • Guanyu Li
  • Senjie Jin
  • Shichun Liu
  • Jixuan Huang
  • Dingwen Yang

Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B.

AAAI Conference 2026 Conference Paper

One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion

  • Yitong Dong
  • Qi Zhang
  • Minchao Jiang
  • Zhiqiang Wu
  • Qingnan Fan
  • Ying Feng
  • Huaqi Zhang
  • Hujun Bao

We present a novel framework for high-fidelity novel view synthesis (NVS) from sparse images, addressing key limitations in recent feed-forward 3D Gaussian Splatting (3DGS) methods built on Vision Transformer (ViT) backbones. While ViT-based pipelines offer strong geometric priors, they are often constrained by low-resolution inputs due to computational costs. Moreover, existing generative enhancement methods tend to be 3D-agnostic, resulting in inconsistent structures across views, especially in unseen regions. To overcome these challenges, we design a Dual-Domain Detail Perception Module, which enables handling high-resolution images without being limited by the ViT backbone, and endows Gaussians with additional features to store high-frequency details. We develop a feature-guided diffusion network, which can preserve high-frequency details during the restoration process. We introduce a unified training strategy that enables joint optimization of the ViT-based geometric backbone and the diffusion-based refinement module. Experiments demonstrate that our method can maintain superior generation quality across multiple datasets.

AAAI Conference 2026 Conference Paper

ORTCL: Towards Continual Learning of Time Series Foundation Models on Streaming Data via Orthogonal Rotation

  • Li Lin
  • Xinrui Zhang
  • Qi Zhang
  • Shuai Wang
  • Kaiwen Xia

Time Series Foundation Models (TSFMs) have emerged as a promising approach in time series analysis. Due to the large-scale parameters of TSFMs and pretraining cost, how to adapt TDFMs in streaming data is always the key factor constraining their application effectiveness. Because streaming data often experiences data distribution and task drifts, which cannot be learnt by offline training. Existing methods typically address streaming data modeling with continuous learning through model fine-tuning or model editing. However, fine-tuning incurs significant computational costs, while editing methods can lead to shifts in the original feature space during streaming updates. To address these limitations, we propose a novel Orthogonal Rotation Transformation-based Continuous Learning method, called ORTCL, for TSFMs. Our key insight is to apply orthogonal matrix rotations to the input and output feature spaces of the TSFMs during model editing. This preserves the metric structure of the original feature space and enables new data to be directly mapped into the existing feature space of the TSFMs. Specifically, we obtain the orthogonal matrix for the input layer via singular value decomposition and derive the corresponding transformation matrix for the output layer through least squares optimization. Extensive experimental results demonstrate that ORTCL outperforms existing methods in both single-domain and cross-domain streaming time series forecasting tasks, effectively mitigating catastrophic forgetting.

JBHI Journal 2026 Journal Article

Pathology-Guided AI System for Accurate Segmentation and Diagnosis of Cervical Spondylosis

  • Qi Zhang
  • Xiuyuan Chen
  • Ziyi He
  • Lianming Wu
  • Kun Wang
  • Jianqi Sun
  • Hongxing Shen

Cervical spondylosis, a complex and prevalent condition, demands precise and efficient diagnostic techniques for accurate assessment. While MRI offers detailed visualization of cervical spine anatomy, manual interpretation remains labor-intensive and prone to error. To address this, we developed an innovative AI-assisted Expert-based Diagnosis System 1 that automates both segmentation and diagnosis of cervical spondylosis using MRI. Leveraging multi-center datasets of cervical MRI images from patients with cervical spondylosis, our system features a pathology-guided segmentation model capable of accurately segmenting key cervical anatomical structures. The segmentation is followed by an expert-based diagnostic framework that automates the calculation of critical clinical indicators. Our segmentation model achieved an impressive average Dice coefficient exceeding 0. 90 across four cervical spinal anatomies and demonstrated enhanced accuracy in herniation areas. Diagnostic evaluation further showcased the system’s precision, with the lowest mean average errors (MAE) for the C2-C7 Cobb angle and the Maximum Spinal Cord Compression (MSCC) coefficient. In addition, our method delivered high accuracy, precision, recall, and F1 scores in herniation localization, K-line status assessment, T2 hyperintensity detection, and Kang grading. Comparative analysis and external validation demonstrate that our system outperforms existing methods, establishing a new benchmark for segmentation and diagnostic tasks for cervical spondylosis.

TCS Journal 2026 Journal Article

Penalty-enhanced quantum approximate optimization algorithm framework for maximization and minimization problems

  • Hao Zhong
  • Qi Zhang

The Quantum Approximate Optimization Algorithm (QAOA for short) has demonstrated great potential in solving NP-hard combinatorial optimization problems. This study proposes a penalty-enhanced QAOA framework for addressing both maximization and minimization problems. By uniformly setting penalty coefficients, the framework provides general support for both types of problems. It ensures the feasibility of output solutions and improves the quality of approximate solutions by adjusting the objective function and the construction of the Hamiltonian. We apply this framework to the Minimum Vertex Cover problem (as a minimization task) and the Maximum Independent Set problem (as a maximization task), designing corresponding quantum Hamiltonians and penalty terms.

AAAI Conference 2026 Conference Paper

Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination

  • Mingqi Wu
  • Zhihao Zhang
  • Qiaole Dong
  • Zhiheng Xi
  • Jun Zhao
  • Senjie Jin
  • Xiaoran Fan
  • Yuhao Zhou

Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods.

EAAI Journal 2026 Journal Article

Research and traceability assessment of thermal radiation damage of glass based on artificial neural network

  • Yidan Cui
  • Qi Zhang

Glass softening under intense thermal radiation is a critical failure mechanism in fires, explosions, and nuclear detonations. However, the diversity of transient thermal radiation scenarios makes it difficult for traditional databases and interpolation methods to provide accurate, generalizable assessments, while numerical simulations, though precise, are often computationally prohibitive for rapid use. To address these challenges, this study proposes an artificial intelligence framework based on artificial neural networks for rapid prediction and inverse analysis of glass softening thickness under multi-factor transient thermal radiation. Key findings include: (1) A physics-based five-dimensional thermal radiation–softening database with 1340 datasets was constructed. Outliers were identified and corrected using a joint Isolation Forest and Local Outlier Factor approach, revealing the nonlinear coupling between heat source parameters and softening sensitivity. (2) A four-layer fully connected artificial neural network was developed to predict softening thickness under high thermal radiation. The model achieved high accuracy (validation loss 0. 0094, test set prediction–truth correlation coefficient of 0. 9944) with a response time of 1. 2 ms. In nuclear explosion scenarios, it maintained over 75 % prediction accuracy within a thermal impulse range from 3. 54 × 106 to 5. 73 × 107 J, with error below 11 % for yields between 1 and 50 kilotons. (3) A two-stage hybrid optimization algorithm combining particle swarm optimization and adaptive moment estimation enabled efficient, accurate inverse analysis of key parameters, achieving error below 0. 5 % in multi-parameter inversion. This framework offers a novel, scalable approach for intelligent evaluation of thermal radiation-induced glass damage, supporting applications in disaster assessment, nuclear safety, and thermal protection design.

AAAI Conference 2026 Conference Paper

Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model

  • Rong Bao
  • Bo Wang
  • Xiao Wang
  • Hongyu Li
  • Rui Zheng
  • Leszek Rutkowski
  • Qi Zhang
  • Liang Ding

Long Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions: 1) The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient.

AAAI Conference 2026 Conference Paper

UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations

  • Wei Feng
  • Chi Huang
  • Qi Zhang
  • Qian Zhang
  • Nan Li

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical in real-world scenarios. In this paper, we propose a novel learning from Unposed views under Varied illuminations Relightable 3D Gaussian Splatting (dubbed UV-RGS), to address this challenge by jointly optimizing camera poses, 3DGS representations, surface materials, and environment illuminations (i.e., unknown and varied lighting conditions in training) using only unposed views under varied lightings. Firstly, UV-RGS presents a viewpoint dividing strategy to group inputs into constituent units, enabling each unit can perform similar poses and illuminations. Next, for each unit, to get the constituent model, UV-RGS establishes an incrementally pose learning module to estimate coarse camera parameters, which also enjoy a proxy-view refinement to alleviate the sparse view learning. Additionally, for all constituent unit models, we introduce a holistic model learning strategy that integrates progressive unit aggregation component and the 3DGS coupled with camera poses joint optimization, which realizes the scene high-fidelity perception by the physical-based rendering. Extensive experiments on both real-world and synthetic challenging datasets demonstrate the effectiveness of UV-RGS, achieving the state-of-the-art performance for scene inverse rendering by learning 3DGS from only unposed views under varied illuminations.

AIIM Journal 2026 Journal Article

Weakly-supervised ultrasound image segmentation with elliptical shape prior constraint

  • Changyan Wang
  • Yehua Cai
  • Ruyi Yang
  • Haobo Chen
  • Jiang Shang
  • Hong Ding
  • Qi Zhang

Accurate pixel-level segmentation of ultrasound (US) images is vital for computer-aided disease screening, diagnosis, and treatment response evaluation. The weakly supervised methods have the potential to reduce the time-consuming and labor-intensive workload for radiologists, paving the way for further automation in the quantitative analysis of US images. Among these methods, the multiple instance learning (MIL) has proven effective and is often applied to prediction tasks with insufficiently labeled data. In US examinations, the elliptical region formed by intersecting lines used by radiologists for target annotation serves as a crucial prior information. Therefore, we propose a novel weakly supervised method called elliptical shape prior constraint MIL (ESPC-MIL) for pixel-level segmentation of US images. ESPC-MIL incorporates an elliptical shape prior constraint into the MIL framework, delivering more accurate foreground and background candidate regions for MIL, which enhances its predictive performance for tissues and organs with approximately elliptical shapes. Furthermore, the method utilizes elliptical shape prior information for global supervision, improving edge segmentation and localization accuracy. Compared to other weakly supervised methods, ESPC-MIL achieves state-of-the-art results on four US image datasets: Achilles tendon dataset, median nerve dataset, private breast tumor dataset, and public breast ultrasound image dataset, with Dice similarity coefficients of 0. 855, 0. 849, 0. 876, and 0. 748, respectively. It demonstrates performance comparable to fully supervised segmentation methods while significantly reducing annotation requirements. Notably, the method demonstrates a more significant performance improvement in segmenting objects with approximately elliptical shapes compared to those with complex shapes. Source codes and models are available at https: //github. com/CYWang-kayla/ESPC-MIL-Model.

IJCAI Conference 2025 Conference Paper

2D Gaussian Splatting for Outdoor Scene Decomposition and Relighting

  • Wei Feng
  • Kangrui Ye
  • Qi Zhang
  • Qian Zhang
  • Nan Li

Gaussian splatting techniques have recently revolutionized outdoor scene decomposition and relighting through multi-view images. However, achieving high rendering quality still requires a fixed lighting condition among all input views, which is costly or even impractical to capture in outdoor scenes. In this paper, we propose outdoor scene decomposition and relighting with 2D Gaussian splatting (OSDR-GS), a novel inverse rendering strategy under outdoor changing and unknown lighting conditions. Firstly, we present a lighting-based group learning framework that categorizes input images into multiple lighting groups, to learn the separate lighting from each group individually. Secondly, OSDR-GS introduces a fine-grained outdoor lighting component to represent sun-light and sky-light, respectively, which are also adjusted via the correlative exposure factors adaptively. Finally, we construct a visibility-driven shadow module to characterize the nuanced interplay of light and occlusion realistically, for eliminating the uncertainty of dark pixels on lighting-based group learning. Extensive experiments on multiple challenging outdoor datasets validate the effectiveness of OSDR-GS, which achieves the state-of-the-art performance in changing lighting scene inverse rendering.

IROS Conference 2025 Conference Paper

A Fast-moving Underwater Wall-climbing Robotic Fish Inspired by Rock-climbing Fish

  • Hengshen Qin
  • Chuang Zhang
  • Wenjun Tan
  • Ruiqian Wang
  • Yiwei Zhang 0012
  • Lianchao Yang
  • Qi Zhang
  • Lianqing Liu

The rock-climbing fish is a benthic organism that can move rapidly and flexibly on rock surfaces in complex underwater environments. Studies have shown that this unique adhesion-sliding movement mechanism of the rock-climbing fish relies on the anisotropic friction exhibited by its sucker structure, which helps to reduce friction in the forward direction and defend against the impact of the flow field. In this work, inspired by the anisotropic friction phenomenon of the sucker of the rock-climbing fish, we designed the absorption module and pectoral and pelvic fin flapping module of the robotic fish to realize the contact switching from low friction to high friction. Meanwhile the propulsion module adopts a novel design of wire-driven caudal fin that can oscillate at high frequency (exceeding ~5 Hz). Resulting robotic fish can realize different motion modes such as adhesion-sliding movement (~0. 5 BL/s) and wall-stabilized adsorption. This work will provide a new solution for the design pattern of underwater wall robots.

EAAI Journal 2025 Journal Article

A lightweight hybrid neural network for remote sensing image super-resolution reconstruction

  • Qi Zhang
  • Fuzhen Zhu
  • Bing Zhu
  • Puying Li
  • Yvshuo Han

To address the challenges of large model sizes and high computational complexity in existing deep learning-based super-resolution neural networks, we design a lightweight hybrid neural network for better used in applications of resource constraints or high real-time requirements, such as edge computing devices, autonomous driving, intelligent agriculture, natural disaster monitoring, and so on. The developed method can reduce the model complexity by combining lightweight convolutional neural network and transformer, while maintaining more super-resolution details. Firstly, the lightweight convolutional ghost modules replace traditional convolutions to build a lightweighted network structure, because of using the convolution with fewer channels and a cheap operation. Secondly, the high-frequency filtering multi-distillation block refines features and extracts high-frequency details through average pooling and upsampling, thereby enriching the details of the reconstructed images. And then, the shifted window transformer is used to capture global information to enhance the overall visual coherence of the reconstructed image by using the shifted window-based self-attention mechanism. Finally, a high-resolution image is obtained by feature fusion and sampling. Experimental results show that the developed method improved both objective evaluation metrics and visual effects. Compared with the state-of-the-art method, the peak signal-to-noise ratio of the developed method is increased about 0. 054 dB, the structural similarity is increased by 0. 003, while the learned perceptual image patch similarity is decreased by 0. 029. The number of model parameters is reduced by 1. 71× 10 5, and the number of multiplication and addition operations is decreased by 8. 2× 10 9.

IROS Conference 2025 Conference Paper

Accelerating Inverse Kinematic Solutions for a Cable-Driven Soft Robotic Manipulator via Physics-Informed Neural Network

  • Rui Lin
  • Shuyou He
  • Ming Xu
  • Kangjia Fu
  • Xuesong Wu
  • Xiucong Sun
  • Qi Zhang
  • Sunquan Yu

Cable-driven soft manipulators, with inherent compliance and hyper-redundancy, offer significant advantages in unstructured environments but present formidable challenges in modeling of inverse kinematics due to nonlinear deformations and underactuation. In this paper, building on a modified forward kinematic model, a physics-informed neural networks (PINN) framework based on spatiotemporal data is proposed for efficient inverse kinematics computation of cable-driven soft robotic manipulators. A geometrically exact forward kinematic model is constructed under the Piecewise Constant Curvature (PCC) assumption, extended to multi-section configurations, and enhanced by cable deflection compensation to account for practical routing constraints. Experimental validation shows a 40. 11% reduction in end-effector positioning error (average 15. 98 mm) when deflection effects are included. The proposed PINN architecture takes time and section count as inputs and outputs the corresponding manipulator configuration, enabling unified spatiotemporal trajectory tracking by minimizing elastic energy while satisfying kinematic constraints. Compared to particle swarm optimization (PSO), which requires iterative computation for each trajectory sample, the proposed method reduces computational time by over 71. 9%, demonstrating superior efficiency in solving redundant inverse kinematics problems. This work bridges data-driven and mechanics-based approaches, offering a scalable solution for real-time control of soft manipulators.

AAAI Conference 2025 Conference Paper

Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning

  • Shihan Dou
  • Yan Liu
  • Enyu Zhou
  • Songyang Gao
  • Tianlong Li
  • Limao Xiong
  • Xin Zhao
  • Haoxiang Jia

The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions.

AAAI Conference 2025 Conference Paper

Amplifier: Bringing Attention to Neglected Low-Energy Components in Time Series Forecasting

  • Jingru Fei
  • Kun Yi
  • Wei Fan
  • Qi Zhang
  • Zhendong Niu

We propose an energy amplification technique to address the issue that existing models easily overlook low-energy components in time series forecasting. This technique comprises an energy amplification block and an energy restoration block. The energy amplification block enhances the energy of low-energy components to improve the model's learning efficiency for these components, while the energy restoration block returns the energy to its original level. Moreover, considering that the energy-amplified data typically displays two distinct energy peaks in the frequency spectrum, we integrate the energy amplification technique with a seasonal-trend forecaster to model the temporal relationships of these two peaks independently, serving as the backbone for our proposed model, Amplifier. Additionally, we propose a semi-channel interaction temporal relationship enhancement block for Amplifier, which enhances the model's ability to capture temporal relationships from the perspective of the commonality and specificity of each channel in the data. Extensive experiments on eight time series forecasting benchmarks consistently demonstrate our model's superiority in both effectiveness and efficiency compared to state-of-the-art methods.

JMLR Journal 2025 Journal Article

An Augmentation Overlap Theory of Contrastive Learning

  • Qi Zhang
  • Yifei Wang
  • Yisen Wang

Recently, self-supervised contrastive learning has achieved great success on various tasks. However, its underlying working mechanism is yet unclear. In this paper, we first provide the tightest bounds based on the widely adopted assumption of conditional independence. Further, we relax the conditional independence assumption to a more practical assumption of augmentation overlap and derive the asymptotically closed bounds for the downstream performance. Our proposed augmentation overlap theory hinges on the insight that the support of different intra-class samples will become more overlapped under aggressive data augmentations, thus simply aligning the positive samples (augmented views of the same sample) could make contrastive learning cluster intra-class samples together. Moreover, from the newly derived augmentation overlap perspective, we develop an unsupervised metric for the representation evaluation of contrastive learning, which aligns well with the downstream performance almost without relying on additional modules. Code is available at https://github.com/PKU-ML/GARC. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2025. ( edit, beta )

AAAI Conference 2025 Conference Paper

Autonomous Goal Detection and Cessation in Reinforcement Learning: A Case Study on Source Term Estimation

  • Yiwei Shi
  • Muning Wen
  • Qi Zhang
  • Weinan Zhang
  • Cunjia Liu
  • Weiru Liu

Reinforcement Learning has revolutionized decision-making processes in dynamic environments, yet it often struggles with autonomously detecting and achieving goals without clear feedback signals. For example, in a Source Term Estimation problem, the lack of precise environmental information makes it challenging to provide clear feedback signals and to define and evaluate how the source's location is determined. To address this challenge, the Autonomous Goal Detection and Cessation (AGDC) module was developed, enhancing various RL algorithms by incorporating a self-feedback mechanism for autonomous goal detection and cessation upon task completion. Our method effectively identifies and ceases undefined goals by approximating the agent's belief, significantly enhancing the capabilities of RL algorithms in environments with limited feedback. To validate effectiveness of our approach, we integrated AGDC with deep Q-Network, proximal policy optimization, and deep deterministic policy gradient algorithms, and evaluated its performance on the Source Term Estimation problem. The experimental results showed that AGDC-enhanced RL algorithms significantly outperformed traditional statistical methods such as infotaxis, entrotaxis, and dual control for exploitation and exploration, as well as a non-statistical random action selection method. These improvements were evident in terms of success rate, mean traveled distance, and search time, highlighting AGDC's effectiveness and efficiency in complex, real-world scenarios.

NeurIPS Conference 2025 Conference Paper

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

  • Zhiheng Xi
  • Guanyu Li
  • Yutao Fan
  • Honglin Guo
  • Yufang Liu
  • Xiaoran Fan
  • Jiaqi Liu
  • Wangmeng Zuo

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community.

JBHI Journal 2025 Journal Article

CoDPN: An Unsupervised Collaborative Dual-path Network For Contactless Remote Physiological Measurement

  • Yimou Lv
  • Xinyuan Jiao
  • Lin Qi
  • Qi Zhang
  • Lisheng Xu
  • Wei Qian

Remote photoplethysmography (rPPG) provides a convenient solution for contactless physiological measurement, and is a promising technique for daily health monitoring and clinical application. However, traditional supervised methods rely heavily on labeled data, incurring substantial annotation costs. Current unsupervised contrastive learning methods encounter challenges under challenging conditions, including variations in illumination and motion of the head or limbs. In this paper, we propose a novel unsupervised end-to-end framework, called CoDPN for physiological measurement. CoDPN employs a dual-path architecture consisting of the Short-term Information Extraction Path (SIEP) and Long-term Information Extraction Path (LIEP), which capture short-term contextual relevance and long-term periodic dependence of rPPG signals, respectively. Next, we propose a collaborative learning strategy to integrate the latent features from both the SIEP and LIEP, facilitating the exchange of complementary information. Furthermore, our unsupervised learning strategy leverages rPPG features in both the frequency and time domains, guiding the CoDPN to extract rPPG signals consistent with physiological information while reducing dependence on labeled datasets. We conduct extensive experiments on three benchmark datasets (UBFC-rPPG, PURE, and UBFC-phys) to implement physiological measurement, including heart rate (HR), heart rate variability (HRV) and respiratory frequency (RF). The experimental results demonstrate that our CoDPN outperforms other state-of-the-art methods under complex conditions.

IJCAI Conference 2025 Conference Paper

COLUR: Confidence-Oriented Learning, Unlearning and Relearning with Noisy-Label Data for Model Restoration and Refinement

  • Zhihao Sui
  • Liang Hu
  • Jian Cao
  • Usman Naseem
  • Zhongyuan Lai
  • Qi Zhang

Large deep learning models have achieved significant success in various tasks. However, the performance of a model can significantly degrade if it is needed to train on datasets with noisy labels with misleading or ambiguous information. To date, there are limited investigations on how to restore performance when model degradation has been incurred by noisy label data. Inspired by the "forgetting mechanism" in neuroscience, which enables accelerating the relearning of correct knowledge by unlearning the wrong knowledge, we propose a robust model restoration and refinement (MRR) framework COLUR, namely Confidence-Oriented Learning, Unlearning and Relearning. Specifically, we implement COLUR with an efficient co-training architecture to unlearn the influence of label noise, and then refine model confidence on each label for relearning. Extensive experiments are conducted on four real datasets and all evaluation results show that COLUR consistently outperforms other SOTA methods after MRR.

TMLR Journal 2025 Journal Article

Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance

  • Qi Zhang
  • Yi Zhou
  • Shaofeng Zou

This paper provides the first tight convergence analyses for RMSProp and Adam for non-convex optimization under the most relaxed assumptions of coordinate-wise generalized smoothness and affine noise variance. RMSProp is firstly analyzed, which is a special case of Adam with adaptive learning rates but without first-order momentum. Specifically, to solve the challenges due to the dependence among adaptive update, unbounded gradient estimate and Lipschitz constant, we demonstrate that the first-order term in the descent lemma converges and its denominator is upper bounded by a function of gradient norm. Based on this result, we show that RMSProp with proper hyperparameters converges to an $\epsilon$-stationary point with an iteration complexity of $\mathcal O(\epsilon^{-4})$. We then generalize our analysis to Adam, where the additional challenge is due to a mismatch between the gradient and the first-order momentum. We develop a new upper bound on the first-order term in the descent lemma, which is also a function of the gradient norm. We show that Adam with proper hyperparameters converges to an $\epsilon$-stationary point with an iteration complexity of $\mathcal O(\epsilon^{-4})$. Our complexity results for both RMSProp and Adam match with the complexity lower bound established in Arjevani et al. (2023).

AAAI Conference 2025 Conference Paper

COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting Mechanism

  • Jianing He
  • Qi Zhang
  • Hongyun Zhang
  • Xuanjing Huang
  • Usman Naseem
  • Duoqian Miao

Early exiting is an effective paradigm for improving the inference efficiency of pre-trained language models (PLMs) by dynamically adjusting the number of executed layers for each sample. However, in most existing works, easy and hard samples are treated equally by each classifier during training, which neglects the test-time early exiting behavior, leading to inconsistency between training and testing. Although some methods have tackled this issue under a fixed speed-up ratio, the challenge of flexibly adjusting the speed-up ratio while maintaining consistency between training and testing is still under-explored. To bridge the gap, we propose a novel Consistency-Oriented Signal-based Early Exiting (COSEE) framework, which leverages a calibrated sample weighting mechanism to enable each classifier to emphasize the samples that are more likely to exit at that classifier under various acceleration scenarios. Extensive experiments on the GLUE benchmark demonstrate the effectiveness of our COSEE across multiple exiting signals and backbones, yielding a better trade-off between performance and efficiency.

TMLR Journal 2025 Journal Article

Decision-Focused Surrogate Modeling for Mixed-Integer Linear Optimization

  • Shivi Dixit
  • Rishabh Gupta
  • Qi Zhang

Mixed-integer optimization is at the core of many online decision-making systems that demand frequent updates of decisions in real time. However, due to their combinatorial nature, mixed-integer linear programs (MILPs) can be difficult to solve, rendering them often unsuitable for time-critical online applications. To address this challenge, we develop a data-driven approach for constructing surrogate optimization models in the form of linear programs (LPs) that can be solved much more efficiently than the corresponding MILPs. We train these surrogate LPs in a decision-focused manner such that for different model inputs, they achieve the same or close to the same optimal solutions as the original MILPs. One key advantage of the proposed method is that it allows the incorporation of all of the original MILP’s linear constraints, which significantly increases the likelihood of obtaining feasible predicted solutions. Results from two computational case studies indicate that this decision-focused surrogate modeling approach is highly data-efficient and provides very accurate predictions of the optimal solutions. In these examples, the resulting surrogate LPs outperform state-of-the-art neural-network-based optimization proxies.

RLJ Journal 2025 Journal Article

Efficient Information Sharing for Training Decentralized Multi-Agent World Models

  • Xiaoling Zeng
  • Qi Zhang

World models, which were originally developed for single-agent reinforcement learning, have recently been extended to multi-agent settings. Due to unique challenges in multi-agent reinforcement learning, agents' independently training of their world models often leads to underperforming policies, and therefore existing work has largely been limited to the centralized training framework that requires excessive communication. As communication is key, we ask the question of how the agents should communicate efficiently to train and learn policies from their decentralized world models. We address this question progressively. We first allow the agents to communicate with unlimited bandwidth to identify which algorithmic components would benefit the most from what types of communication. Then, we restrict the inter-agent communication with a predetermined bandwidth limit to challenge the agents to communicate efficiently. Our algorithmic innovations develop a scheme that prioritizes important information to share while respecting the bandwidth limit. The resulting method yields superior sample efficiency, sometimes even over centralized training baselines, in a range of cooperative multi-agent reinforcement learning benchmarks.

RLC Conference 2025 Conference Paper

Efficient Information Sharing for Training Decentralized Multi-Agent World Models

  • Xiaoling Zeng
  • Qi Zhang

World models, which were originally developed for single-agent reinforcement learning, have recently been extended to multi-agent settings. Due to unique challenges in multi-agent reinforcement learning, agents' independently training of their world models often leads to underperforming policies, and therefore existing work has largely been limited to the centralized training framework that requires excessive communication. As communication is key, we ask the question of how the agents should communicate efficiently to train and learn policies from their decentralized world models. We address this question progressively. We first allow the agents to communicate with unlimited bandwidth to identify which algorithmic components would benefit the most from what types of communication. Then, we restrict the inter-agent communication with a predetermined bandwidth limit to challenge the agents to communicate efficiently. Our algorithmic innovations develop a scheme that prioritizes important information to share while respecting the bandwidth limit. The resulting method yields superior sample efficiency, sometimes even over centralized training baselines, in a range of cooperative multi-agent reinforcement learning benchmarks.

IJCAI Conference 2025 Conference Paper

Empowering Multimodal Road Traffic Profiling with Vision Language Models and Frequency Spectrum Fusion

  • Haolong Xiang
  • Xiaolong Xu
  • Guangdong Wang
  • Xuyun Zhang
  • Xiaoyong Li
  • Qi Zhang
  • Amin Beheshti
  • Wei Fan

With the rapid urbanization in the modern era, smart traffic profiling based on multimodal sources of data has been playing a significant role in ensuring safe travel, reducing traffic congestion and optimizing urban mobility. Most existing methods for traffic profiling on the road level usually utilize single-modality data, i. e. , they mainly focus on image processing with deep vision models or auxiliary analysis on the textual data. However, the joint modeling and multimodal fusion of the textual and visual modalities have been rarely studied in road traffic profiling, which largely hinders the accurate prediction or classification of traffic conditions. To address this issue, we propose a novel multimodal learning and fusion framework for road traffic profiling, named TraffiCFUS. Specifically, given the traffic images, our TraffiCFUS framework first introduces Vision Language Models (VLMs) to generate text and then creates tailored prompt instructions for refining this text according to the specific scene requirements of road traffic profiling. Next, we apply the discrete Fourier transform to convert multimodal data from the spatial domain to the frequency domain and perform a cross-modal spectrum transform to filter out irrelevant information for traffic profiling. Furthermore, the processed spatial multimodal data is combined to generate fusion loss and interaction loss with contrastive learning. Finally, extensive experiments on four real-world datasets illustrate superior performance compared with the state-of-the-art approaches.

NeurIPS Conference 2025 Conference Paper

Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning

  • Yu Zhang
  • Jialei Zhou
  • Xinchen Li
  • Qi Zhang
  • Zhongwei Wan
  • Duoqian Miao
  • Changwei Wang
  • Longbing Cao

Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or causes semantic confusion by simultaneously modeling diverse semantic primitive types. To mitigate this defect of DiTs, we propose a novel split-text conditioning framework named DiT-ST. This framework converts a complete-text caption into a split-text caption, a collection of simplified sentences, to explicitly express various semantic primitives and their interconnections. The split-text caption is then injected into different denoising stages of DiT-ST in a hierarchical and incremental manner. Specifically, DiT-ST leverages Large Language Models to parse captions, extracting diverse primitives and hierarchically sorting out and constructing these primitives into a split-text input. Moreover, we partition the diffusion denoising process according to its differential sensitivities to diverse semantic primitive types and determine the appropriate timesteps to incrementally inject tokens of diverse semantic primitive types into input tokens via cross-attention. In this way, DiT-ST enhances the representation learning of specific semantic primitive types across different stages. Extensive experiments validate the effectiveness of our proposed DiT-ST in mitigating the complete-text comprehension defect. Datasets and models are available.

NeurIPS Conference 2025 Conference Paper

EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving

  • Shihan Dou
  • Ming Zhang
  • Chenhao Huang
  • Jiayi Chen
  • Feng Chen
  • Shichun Liu
  • Yan Liu
  • Chenxiao Liu

We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3. 7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available in the supplementary materials.

JBHI Journal 2025 Journal Article

HealthiVert-GAN: A Novel Framework of Pseudo-Healthy Vertebral Image Synthesis for Interpretable Compression Fracture Grading

  • Qi Zhang
  • Cheng Chuang
  • Shunan Zhang
  • Ziqi Zhao
  • Kun Wang
  • Jun Xu
  • Jianqi Sun

Osteoporotic vertebral compression fractures (OVCFs) are prevalent in the elderly population, typically assessed on computed tomography (CT) scans by evaluating vertebral height loss. This assessment helps determine the fracture's impact on spinal stability and the need for surgical intervention. However, the absence of pre-fracture CT scans and standardized vertebral references leads to measurement errors and inter-observer variability, while irregular compression patterns further challenge the precise grading of fracture severity. While deep learning methods have shown promise in aiding OVCFs screening, they often lack interpretability and sufficient sensitivity, limiting their clinical applicability. To address these challenges, we introduce a novel vertebra synthesis-height loss quantification-OVCFs grading framework. Our proposed model, HealthiVert-GAN 1, utilizes a coarse-to-fine synthesis network designed to generate pseudo-healthy vertebral images that simulate the pre-fracture state of fractured vertebrae. This model integrates three auxiliary modules that leverage the morphology and height information of adjacent healthy vertebrae to ensure anatomical consistency. Additionally, we introduce the Relative Height Loss of Vertebrae (RHLV) as a quantification metric, which divides each vertebra into three sections to measure height loss between pre-fracture and post-fracture states, followed by fracture severity classification using a Support Vector Machine (SVM). Our approach achieves state-of-the-art classification performance on both the Verse2019 dataset and in-house dataset, and it provides cross-sectional distribution maps of vertebral height loss. This practical tool enhances diagnostic accuracy in clinical settings and assisting in surgical decision-making.

IJCAI Conference 2025 Conference Paper

Improving Prediction Certainty Estimation for Reliable Early Exiting via Null Space Projection

  • Jianing He
  • Qi Zhang
  • Duoqian Miao
  • Kun Yi
  • Shufeng Hao
  • Hongyun Zhang
  • Zhihua Wei

Early exiting has demonstrated great potential in accelerating the inference of pre-trained language models (PLMs) by enabling easy samples to exit at shallow layers, eliminating the need for executing deeper layers. However, existing early exiting methods primarily rely on class-relevant logits to formulate their exiting signals for estimating prediction certainty, neglecting the detrimental influence of class-irrelevant information in the features on prediction certainty. This leads to an overestimation of prediction certainty, causing premature exiting of samples with incorrect early predictions. To remedy this, we define an NSP score to estimate prediction certainty by considering the proportion of class-irrelevant information in the features. On this basis, we propose a novel early exiting method based on the Certainty-Aware Probability (CAP) score, which integrates insights from both logits and the NSP score to enhance prediction certainty estimation, thus enabling more reliable exiting decisions. The experimental results on the GLUE benchmark show that our method can achieve an average speed-up ratio of 2. 19× across all tasks with negligible performance degradation, surpassing the state-of-the-art (SOTA) ConsistentEE by 28%, yielding a better trade-off between task performance and inference efficiency. The code is available at https: //github. com/He-Jianing/NSP. git.

IROS Conference 2025 Conference Paper

Interdigitated Electrodes for Selective Stimulation of Skeletal Muscle Actuators in Biosyncretic Robots

  • Lianchao Yang
  • Chuang Zhang
  • Qi Zhang
  • Yiwei Zhang 0012
  • Hengshen Qin
  • Lianqing Liu

Engineered skeletal muscle tissue (SMT) is the ideal driving units for achieving fine movements in biosyncretic robots due to their excellent controllability and potentially large driving force. However, the selective stimulation of SMTs continues to pose a significant technical challenge. In this study, we propose a method for the selective stimulation of 3D SMT using thin-film interdigitated electrodes (IDEs). By optimizing the IDEs geometry through finite element simulations, electrical field intensity of fingertip is effectively reduced. The thin-film IDEs are fabricated on a Polyester (PET) substrate using screen printing technology and successfully enable selective activation and controlled contraction of the SMTs. Compared to conventional parallel-plate electrodes (PPEs) and rod-shaped electrodes (RSEs), the IDEs significantly improve the electrical field distribution and enhance spatial resolution. This advancement provides a promising new approach for achieving high-precision motion control in biosyncretic robots (or biohybrid robots).

TMLR Journal 2025 Journal Article

Large Action Models: From Inception to Implementation

  • Lu Wang
  • Fangkai Yang
  • Chaoyun Zhang
  • Junting Lu
  • Jiaxu Qian
  • Shilin He
  • Pu Zhao
  • Bo Qiao

As AI continues to advance, there is a growing demand for systems that go beyond language-based assistance and move toward intelligent agents capable of performing real-world actions. This evolution requires the transition from traditional Large Language Models (LLMs), which excel at generating textual responses, to Large Action Models (LAMs), designed for action generation and execution within dynamic environments. Enabled by agent systems, LAMs hold the potential to transform AI from passive language understanding to active task completion, marking a significant milestone in the progression toward artificial general intelligence. In this paper, we present a comprehensive framework for developing LAMs, offering a systematic approach to their creation, from inception to deployment. We begin with an overview of LAMs, highlighting their unique characteristics and delineating their differences from LLMs. Using a Windows OS-based agent as a case study, we provide a detailed, step-by-step guide on the key stages of LAM development, including data collection, model training, environment integration, grounding, and evaluation. This generalizable workflow can serve as a blueprint for creating functional LAMs in various application domains. We conclude by identifying the current limitations of LAMs and discussing directions for future research and industrial deployment, emphasizing the challenges and opportunities that lie ahead in realizing the full potential of LAMs in real-world applications.

TMLR Journal 2025 Journal Article

Large Language Model-Brained GUI Agents: A Survey

  • Chaoyun Zhang
  • Shilin He
  • Jiaxu Qian
  • Bowen Li
  • Liqun Li
  • Si Qin
  • Yu Kang
  • Minghua Ma

Graphical User Interfaces (GUIs) have long been central to human-computer interaction, providing an intuitive and visually-driven way to access and interact with digital systems. Traditionally, automating GUI interactions relied on script-based or rule-based approaches, which, while effective for fixed workflows, lacked the flexibility and adaptability required for dynamic, real-world applications. The advent of Large Language Models (LLMs), particularly multimodal models, has ushered in a new era of GUI automation. They have demonstrated exceptional capabilities in natural language understanding, code generation, task generalization, and visual processing. This has paved the way for a new generation of ''LLM-brained'' GUI agents capable of interpreting complex GUI elements and autonomously executing actions based on natural language instructions. These agents represent a paradigm shift, enabling users to perform intricate, multi-step tasks through simple conversational commands. Their applications span across web navigation, mobile app interactions, and desktop automation, offering a transformative user experience that revolutionizes how individuals interact with software. This emerging field is rapidly advancing, with significant progress in both research and industry. To provide a structured understanding of this trend, this paper presents a comprehensive survey of LLM-brained GUI agents, exploring their historical evolution, core components, and advanced techniques. We address critical research questions such as existing GUI agent frameworks, the collection and utilization of data for training specialized GUI agents, the development of large action models tailored for GUI tasks, and the evaluation metrics and benchmarks necessary to assess their effectiveness. Additionally, we examine emerging applications powered by these agents. Through a detailed analysis, this survey identifies key research gaps and outlines a roadmap for future advancements in the field. By consolidating foundational knowledge and state-of-the-art developments, this work aims to guide both researchers and practitioners in overcoming challenges and unlocking the full potential of LLM-brained GUI agents. We anticipate that this survey will serve both as a practical cookbook for constructing LLM-powered GUI agents, and as a definitive reference for advancing research in this rapidly evolving domain.

ICLR Conference 2025 Conference Paper

Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

  • Zhongyi Shui
  • Jianpeng Zhang
  • Weiwei Cao
  • Sinuo Wang
  • Ruizhe Guo
  • Le Lu 0001
  • Lin Yang 0002
  • Xianghua Ye

Artificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities, leading to ambiguous patient-level pairings. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease (including several most deadly cancers) diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively. Additionally, on the publicly available CT-RATE and Rad-ChestCT benchmarks, our fVLM outperformed the current state-of-the-art methods with absolute AUC gains of 7.4% and 4.8%, respectively.

AAAI Conference 2025 Conference Paper

MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction

  • Zixuan Gong
  • Qi Zhang
  • Guangyin Bao
  • Lei Zhu
  • Rongtao Xu
  • Ke Liu
  • Liang Hu
  • Duoqian Miao

Decoding natural visual scenes from brain activity has flourished, with extensive research in single-subject tasks and, however, less in cross-subject tasks. Reconstructing high-quality images in cross-subject tasks is a challenging problem due to profound individual differences between subjects and the scarcity of data annotation. In this work, we proposed MindTuner for cross-subject visual decoding, which achieves high-quality and rich semantic reconstructions using only 1 hour of fMRI training data benefiting from the phenomena of visual fingerprint in the human visual system and a novel fMRI-to-text alignment paradigm. Firstly, we pre-train a multi-subject model among 7 subjects and fine-tune it with scarce data on new subjects, where LoRAs with Skip-LoRAs are utilized to learn the visual fingerprint. Then, we take the image modality as the intermediate pivot modality to achieve fMRI-to-text alignment, which achieves impressive fMRI-to-text retrieval performance and corrects fMRI-to-image reconstruction with fine-tuned semantics. The results of both qualitative and quantitative analyses demonstrate that MindTuner surpasses state-of-the-art cross-subject visual decoding models on the Natural Scenes Dataset (NSD), whether using training data of 1 hour or 40 hours.

AAAI Conference 2025 Conference Paper

MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning

  • Yaming Yang
  • Dilxat Muhtar
  • Yelong Shen
  • Yuefeng Zhan
  • Jianfeng Liu
  • Yujing Wang
  • Hao Sun
  • Weiwei Deng

Parameter-efficient fine-tuning (PEFT) has been widely employed for domain adaptation, with LoRA being one of the most prominent methods due to its simplicity and effectiveness. However, in multi-task learning (MTL) scenarios, LoRA tends to obscure the distinction between tasks by projecting sparse high-dimensional features from different tasks into the same dense low-dimensional intrinsic space. This leads to task interference and suboptimal performance for LoRA and its variants. To tackle this challenge, we propose MTL-LoRA, which retains the advantages of low-rank adaptation while significantly enhancing MTL capabilities. MTL-LoRA augments LoRA by incorporating additional task-adaptive parameters that differentiate task-specific information and capture shared knowledge across various tasks within low-dimensional spaces. This approach enables pretrained models to jointly adapt to different target domains with a limited number of trainable parameters. Comprehensive experimental results, including evaluations on public academic benchmarks for natural language understanding, commonsense reasoning, and image-text understanding, as well as real-world industrial text Ads relevance datasets, demonstrate that MTL-LoRA outperforms LoRA and its various variants with comparable or even fewer learnable parameters in MTL setting.

AAAI Conference 2025 Conference Paper

Position-Aware Guided Point Cloud Completion with CLIP Model

  • Feng Zhou
  • Qi Zhang
  • Ju Dai
  • Lei Li
  • Qing Fan
  • Junliang Xing

Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.

NeurIPS Conference 2025 Conference Paper

Pre-Trained Policy Discriminators are General Reward Models

  • Shihan Dou
  • Shichun Liu
  • Yuming Yang
  • Yicheng Zou
  • Yunhua Zhou
  • Shuhao Xing
  • Chenhao Huang
  • Qiming Ge

We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1. 8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54. 8% to 81. 0% on STEM tasks and from 57. 9% to 85. 5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3. 1-8B from an average of 47. 36% to 56. 33% and Qwen2. 5-32B from 64. 49% to 70. 47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0. 99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.

IJCAI Conference 2025 Conference Paper

Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy

  • Zhihao Sui
  • Liang Hu
  • Jian Cao
  • Dora D. Liu
  • Usman Naseem
  • Zhongyuan Lai
  • Qi Zhang

Machine Unlearning (MU) technology facilitates the removal of the influence of specific data instances from trained models on request. Despite rapid advancements in MU technology, its vulnerabilities are still underexplored, posing potential risks of privacy breaches through leaks of ostensibly unlearned information. Current limited research on MU attacks requires access to original models containing privacy data, which violates the critical privacy-preserving objective of MU. To address this gap, we initiate the innovative study on recalling the forgotten class memberships from unlearned models (ULMs) without requiring access to the original one. Specifically, we implement a Membership Recall Attack (MRA) framework with a teacher-student knowledge distillation architecture, where ULMs serve as noisy labelers to transfer knowledge to student models. Then, it is translated into a Learning with Noisy Labels (LNL) problem for inferring correct labels of the forgetting instances. Extensive experiments on state-of-the-art MU methods with multiple real datasets demonstrate that the proposed MRA strategy exhibits high efficacy in recovering class memberships of unlearned instances. As a result, our study and evaluation have established a benchmark for future research on MU vulnerabilities.

NeurIPS Conference 2025 Conference Paper

Revealing Multimodal Causality with Large Language Models

  • Jin Li
  • Shoujin Wang
  • Qi Zhang
  • Feng Liu
  • Tongliang Liu
  • Longbing Cao
  • Shui Yu
  • Fang Chen

Uncovering cause-and-effect mechanisms from data is fundamental to scientific progress. While large language models (LLMs) show promise for enhancing causal discovery (CD) from unstructured data, their application to the increasingly prevalent multimodal setting remains a critical challenge. Even with the advent of multimodal LLMs (MLLMs), their efficacy in multimodal CD is hindered by two primary limitations: (1) difficulty in exploring intra- and inter-modal interactions for comprehensive causal variable identification; and (2) insufficiency to handle structural ambiguities with purely observational data. To address these challenges, we propose MLLM-CD, a novel framework for multimodal causal discovery from unstructured data. It consists of three key components: (1) a novel contrastive factor discovery module to identify genuine multimodal factors based on the interactions explored from contrastive sample pairs; (2) a statistical causal structure discovery module to infer causal relationships among discovered factors; and (3) an iterative multimodal counterfactual reasoning module to refine the discovery outcomes iteratively by incorporating the world knowledge and reasoning capabilities of MLLMs. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed MLLM-CD in revealing genuine factors and causal relationships among them from multimodal unstructured data. The implementation code and data are available at https: //github. com/JinLi-i/MLLM-CD.

IJCAI Conference 2025 Conference Paper

RF-DTR: A Multi-Stage DCT Token Regression Network for Progressive Rib Fracture Mask Refinement

  • ShouYu Chen
  • Liang Hu
  • Juntao Wang
  • Usman Naseem
  • Zhongyuan Lai
  • Qi Zhang

Rib fracture patterns are key indicators of trauma severity. Detecting and locating these fractures is a critical yet time-consuming task, especially in 3D imaging, due to their minute size and irregular geometries. Existing voxel-based spatial methods fail to capture frequency-domain variations inherent in imaging and do not replicate the progressive refinement process used by clinicians during manual annotation, leading to suboptimal results. We propose a novel regression network, RF-DTR, incorporating a gated regressor mechanism and operating entirely in the frequency domain to address these challenges. Specifically, we present an innovative spatial-frequency transform applied to volumes and corresponding masks. Furthermore, we introduce a Mahalanobis regularization technique to enhance the model and learn high-frequency DCT components relevant to clinical tasks. Finally, a hierarchical penalty is proposed to improve the confidence of the prediction. Extensive experiments confirm our method's superiority in handling complex, sparsely annotated medical imaging datasets.

NeurIPS Conference 2025 Conference Paper

SEMPO: Lightweight Foundation Models for Time Series Forecasting

  • Hui He
  • Kun Yi
  • Yuanchi Ma
  • Qi Zhang
  • Zhendong Niu
  • Guansong Pang

The recent boom of large pre-trained models witnesses remarkable success in developing foundation models (FMs) for time series forecasting. Despite impressive performance across diverse downstream forecasting tasks, existing time series FMs possess massive network architectures and require substantial pre-training on large-scale datasets, which significantly hinders their deployment in resource-constrained environments. In response to this growing tension between versatility and affordability, we propose SEMPO, a novel lightweight foundation model that requires pretraining on relatively small-scale data, yet exhibits strong general time series forecasting. Concretely, SEMPO comprises two key modules: 1) energy-aware S p E ctral decomposition module, that substantially improves the utilization of pre-training data by modeling not only the high-energy frequency signals but also the low-energy yet informative frequency signals that are ignored in current methods; and 2) M ixture-of- P r O mpts enabled Transformer, that learns heterogeneous temporal patterns through small dataset-specific prompts and adaptively routes time series tokens to prompt-based experts for parameter-efficient model adaptation across different datasets and domains. Equipped with these modules, SEMPO significantly reduces both pre-training data scale and model size, while achieving strong generalization. Extensive experiments on two large-scale benchmarks covering 16 datasets demonstrate the superior performance of SEMPO in both zero-shot and few-shot forecasting scenarios compared with state-of-the-art methods. Code and data are available at https: //github. com/mala-lab/SEMPO.

IJCAI Conference 2025 Conference Paper

Sentiment-enhanced Multi-hop Connected Graph Attention Network for Multimodal Aspect-Based Sentiment Analysis

  • Linlin Zhu
  • Heli Sun
  • Xiaoyong Huang
  • Qi Zhang
  • Ruichen Cao
  • Liang He

Multimodal aspect-based sentiment analysis aims to extract aspects from different data sources and recognize the corresponding sentiments. While current research has broadly focused on syntax relation-driven semantic comprehension, the impact of the importance of different syntactic relations on semantic understanding has not been adequately investigated. To address this issue, we propose a Sentiment-enhanced Multi-hop Connected Graph Attention Network (MCG), aiming to enhance the discriminative capability of model for sentiments and to delve into the syntactic relationships within the text. Firstly, we design a contrastive sentiment-enhanced pre-training task that expands the diversity and complexity of training samples to improve the recognition of multiple sentiments. Secondly, we construct a multi-hop connected syntactic dependency graph to deeply explore the rich syntactic dependencies in the text and to reveal the differences among syntactic relations. Moreover, we develop a multi-hop connected graph attention mechanism that enables the model to focus on the key syntactic relations within the syntactic structure, thereby enhancing the comprehension and predictive capabilities of model in multimodal sentiment analysis. Experimental results on two benchmark datasets demonstrate that our method outperforms state-of-the-art methods. The source code is provided in the supplementary materials.

ICRA Conference 2025 Conference Paper

Trustworthy Robot Behavior Tree Generation Based on Multi-Source Heterogeneous Knowledge Graph

  • Jianchao Yuan
  • Shuo Yang
  • Qi Zhang
  • Ge Li
  • Jianping Tang

In robotics, the design of robot behavior trees generally requires roboticists to comprehensively and customizable consider all the relevant factors including the robot hardware capabilities, task descriptions, etc, posing great challenges for design quality and efficiency. The mainstream practice of BT design paradigm has been utilizing the BT component framework to develop task-specific BT structures manually. In contrast, the latest advances in Generative Pretrained Transformers (GPTs) have also opened up the possibility of BT design automation. However, these approaches generally show low efficiency or are less trustworthy for complex robot task goals due to time-consuming manual design and unreliable GPT reasoning. To solve the above limitations, this paper proposes a novel knowledge-driven approach that develops a specialized knowledge graph from multi-sourced and heterogeneous highquality robot knowledge to reason out a trustworthy robot plan for achieving complex task goals. Then we present the plan transformation and BT merging algorithms to automatically generate the plan-level BT structure. The comparative experiment results have shown that our approach can generate highquality and trustworthy BT structure regarding the task plan accuracy and consistency, as well as the BT generation time, compared with the manual design and GPT-based approaches.

NeurIPS Conference 2025 Conference Paper

Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models

  • Jun Zhao
  • Yongzhuo Yang
  • Xiang Hu
  • Jingqi Tong
  • Yi Lu
  • Wei Wu
  • Tao Gui
  • Qi Zhang

Retrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However, the internal mechanisms by which LLMs utilize these knowledge remain unclear. We propose modeling the forward propagation of knowledge as an entity flow, employing this framework to trace LLMs' internal behaviors when processing mixed-source knowledge. Linear probing utilizes a trainable linear classifier to detect specific attributes in hidden layers. However, once trained, a probe cannot adapt to dynamically specified entities. To address this challenge, we construct an entity-aware probe, which introduces special tokens to mark probing targets and employs a small trainable rank-8 lora update to process these special markers. We first verify this approach through an attribution experiment, demonstrating that it can accurately detect information about ad-hoc entities from complex hidden states. Next, we trace entity flows across layers to understand how LLMs reconcile conflicting knowledge internally. Our probing results reveal that contextual and parametric knowledge are routed between tokens through distinct sets of attention heads, supporting attention competition only within knowledge types. While conflicting knowledge maintains a residual presence across layers, aligned knowledge from multiple sources gradually accumulates, with the magnitude of this accumulation directly determining its influence on final outputs.

AAAI Conference 2025 Conference Paper

View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection

  • Qi Zhang
  • Zhouhang Luo
  • Tao Yu
  • Hui Huang

View transformation robustness (VTR) is critical for deep-learning-based multi-view 3D object reconstruction models, which indicates the methods' stability under inputs with various view transformations. However, existing research seldom focused on view transformation robustness in multi-view 3D object reconstruction. One direct way to improve the models' VTR is to produce data with more view transformations and add them to model training. Recent progress on large vision models, particularly Stable Diffusion models, has provided great potential for generating 3D models or synthesizing novel view images with only a single image input. Directly deploying these models at inference consumes heavy computation resources and their robustness to view transformations is not guaranteed either. To fully utilize the power of Stable Diffusion models without extra inference computation burdens, we propose to generate novel views with Stable Diffusion models for better view transformation robustness. Instead of synthesizing random views, we propose a reconstruction error-guided view selection method, which considers the reconstruction errors' spatial distribution of the 3D predictions and chooses the views that could cover the reconstruction errors as much as possible. The methods are trained and tested on sets with large view transformations to validate the 3D reconstruction models' robustness to view transformations. Extensive experiments demonstrate that the proposed method can outperform state-of-the-art 3D reconstruction methods and other view transformation robustness comparison methods.

AAAI Conference 2025 Conference Paper

Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding

  • Guangyin Bao
  • Qi Zhang
  • Zixuan Gong
  • Jialei Zhou
  • Wei Fan
  • Kun Yi
  • Usman Naseem
  • Liang Hu

Decoding visual information from human brain activity has seen remarkable advancements in recent research. However, the diversity in cortical parcellation and fMRI patterns across individuals has prompted the development of deep learning models tailored to each subject. The personalization limits the broader applicability of brain visual decoding in real-world scenarios. To address this issue, we introduce Wills Aligner, a novel approach designed to achieve multi-subject collaborative brain visual decoding. Wills Aligner begins by aligning the fMRI data from different subjects at the anatomical level. It then employs delicate mixture-of-brain-expert adapters and a meta-learning strategy to account for individual fMRI pattern differences. Additionally, Wills Aligner leverages the semantic relation of visual stimuli to guide the learning of inter-subject commonality, enabling visual decoding for each subject to draw insights from other subjects' data. We rigorously evaluate our Wills Aligner across various visual decoding tasks, including classification, cross-modal retrieval, and image reconstruction. The experimental results demonstrate that Wills Aligner achieves promising performance.

TMLR Journal 2025 Journal Article

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

  • Jiaxu Qian
  • Chendong Wang
  • Yifan Yang
  • Chaoyun Zhang
  • Huiqiang Jiang
  • Xufang Luo
  • Yu Kang
  • Qingwei Lin

Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real-world scenarios—especially when small objects or fine spatial context are involved. We pinpoint two core causes of this failure: the absence of region-adaptive attention and inflexible token budgets that force uniform downsampling, leading to critical information loss. To overcome these limitations, we introduce Zoomer a visual prompting framework that delivers token-efficient, detail-preserving image representations for black-box MLLMs. Zoomer integrates (1) a prompt-aware emphasis module to highlight semantically relevant regions, (2) a spatial-preserving orchestration schema to maintain object relationships, and (3) a budget-aware strategy to optimally allocate tokens between global context and local details. Extensive experiments on nine benchmarks and three commercial MLLMs demonstrate that Zoomer boosts accuracy by up to 27% while cutting image token usage by up to 67\%. Our approach establishes a principled methodology for robust, resource-aware multimodal understanding in settings where model internals are inaccessible.

IROS Conference 2024 Conference Paper

A Low-Texture Robust Hybrid Feature Based Visual Odometry

  • He Wang 0046
  • Qi Zhang
  • Zhiwen Zheng
  • Xiaoli Li 0001
  • Hongye Tan
  • Ru Li 0001

In low-texture scenes, Visual Odometry (VO) algorithms often encounter challenges stemming from sparse feature sets and reduced accuracy in feature matching. To overcome this, integrating plane features and vanishing point characteristics can provide additional constraints for refining camera poses. Optical flow-based tracking methods may also offer improved matching precision compared to traditional feature-based approaches. Motivated by these challenges, we present a robust Visual Odometry system tailored for low-texture environments. Our system combines a vanishing point-based approach for camera pose optimization with a Manhattan-aided algorithm for matching line segments using optical flow. By incorporating planes and vanishing points as supplementary features for pose estimation, we enhance overall accuracy without significant time overhead. We utilize detected line features to compute vanishing points, improving accuracy without compromising efficiency. In addition, our Manhattan-aided optical flow technique supplements and refines the results of line feature matching, further enhancing the accuracy of vanishing points. Evaluation on various public datasets demonstrates the superior accuracy and robustness of our system compared to state-of-the-art Simultaneous Localization And Mapping (SLAM) and VO methods. Notably, our method effectively addresses issues of failure in low-texture scenes and improves the accuracy of line feature matching compared to baseline methods. We will release our source code upon paper acceptance.

AAAI Conference 2024 Conference Paper

A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields

  • Yiyu Zhuang
  • Qi Zhang
  • Xuan Wang
  • Hao Zhu
  • Ying Feng
  • Xiaoyu Li
  • Ying Shan
  • Xun Cao

Recent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes.

IJCAI Conference 2024 Conference Paper

Deep Frequency Derivative Learning for Non-stationary Time Series Forecasting

  • Wei Fan
  • Kun Yi
  • Hangting Ye
  • Zhiyuan Ning
  • Qi Zhang
  • Ning An

While most time series are non-stationary, it is inevitable for models to face the distribution shift issue in time series forecasting. Existing solutions manipulate statistical measures (usually mean and std. ) to adjust time series distribution. However, these operations can be theoretically seen as the transformation towards zero frequency component of the spectrum which cannot reveal full distribution information and would further lead to information utilization bottleneck in normalization, thus hindering forecasting performance. To address this problem, we propose to utilize the whole frequency spectrum to transform time series to make full use of data distribution from the frequency perspective. We present a deep frequency derivative learning framework, DERITS, for non-stationary time series forecasting. Specifically, DERITS is built upon a novel reversible transformation, namely Frequency Derivative Transformation (FDT) that makes signals derived in the frequency domain to acquire more stationary frequency representations. Then, we propose the Order-adaptive Fourier Convolution Network to conduct adaptive frequency filtering and learning. Furthermore, we organize DERITS as a parallel-stacked architecture for the multi-order derivation and fusion for forecasting. Finally, we conduct extensive experiments on several datasets which show the consistent superiority in both time series forecasting and shift alleviation.

YNIMG Journal 2024 Journal Article

Detection of individual brain tau deposition in Alzheimer's disease based on latent feature-enhanced generative adversarial network

  • Jiehui Jiang
  • Rong Shi
  • Jiaying Lu
  • Min Wang
  • Qi Zhang
  • Shuoyan Zhang
  • Luyao Wang
  • Ian Alberts

OBJECTIVE: The conventional methods for interpreting tau PET imaging in Alzheimer's disease (AD), including visual assessment and semi-quantitative analysis of fixed hallmark regions, are insensitive to detect individual small lesions because of the spatiotemporal neuropathology's heterogeneity. In this study, we proposed a latent feature-enhanced generative adversarial network model for the automatic extraction of individual brain tau deposition regions. METHODS: The latent feature-enhanced generative adversarial network we propose can learn the distribution characteristics of tau PET images of cognitively normal individuals and output the abnormal distribution regions of patients. This model was trained and validated using 1131 tau PET images from multiple centres (with distinct races, i.e., Caucasian and Mongoloid) with different tau PET ligands. The overall quality of synthetic imaging was evaluated using structural similarity (SSIM), peak signal to noise ratio (PSNR), and mean square error (MSE). The model was compared to the fixed templates method for diagnosing and predicting AD. RESULTS: The reconstructed images archived good quality, with SSIM = 0.967 ± 0.008, PSNR = 31.377 ± 3.633, and MSE = 0.0011 ± 0.0007 in the independent test set. The model showed higher classification accuracy (AUC = 0.843, 95 % CI = 0.796-0.890) and stronger correlation with clinical scales (r = 0.508, P < 0.0001). The model also achieved superior predictive performance in the survival analysis of cognitive decline, with a higher hazard ratio: 3.662, P < 0.001. INTERPRETATION: The LFGAN4Tau model presents a promising new approach for more accurate detection of individualized tau deposition. Its robustness across tracers and races makes it a potentially reliable diagnostic tool for AD in practice.

NeurIPS Conference 2024 Conference Paper

FilterNet: Harnessing Frequency Filters for Time Series Forecasting

  • Kun Yi
  • Jingru Fei
  • Qi Zhang
  • Hui He
  • Shufeng Hao
  • Defu Lian
  • Wei Fan

Given the ubiquitous presence of time series data across various domains, precise forecasting of time series holds significant importance and finds widespread real-world applications such as energy, weather, healthcare, etc. While numerous forecasters have been proposed using different network architectures, the Transformer-based models have state-of-the-art performance in time series forecasting. However, forecasters based on Transformers are still suffering from vulnerability to high-frequency signals, efficiency in computation, and bottleneck in full-spectrum utilization, which essentially are the cornerstones for accurately predicting time series with thousands of points. In this paper, we explore a novel perspective of enlightening signal processing for deep time series forecasting. Inspired by the filtering process, we introduce one simple yet effective network, namely FilterNet, built upon our proposed learnable frequency filters to extract key informative temporal patterns by selectively passing or attenuating certain components of time series signals. Concretely, we propose two kinds of learnable filters in the FilterNet: (i) Plain shaping filter, that adopts a universal frequency kernel for signal filtering and temporal modeling; (ii) Contextual shaping filter, that utilizes filtered frequencies examined in terms of its compatibility with input signals fordependency learning. Equipped with the two filters, FilterNet can approximately surrogate the linear and attention mappings widely adopted in time series literature, while enjoying superb abilities in handling high-frequency noises and utilizing the whole frequency spectrum that is beneficial for forecasting. Finally, we conduct extensive experiments on eight time series forecasting benchmarks, and experimental results have demonstrated our superior performance in terms of both effectiveness and efficiency compared with state-of-the-art methods. Our code is available at$^1$.

ECAI Conference 2024 Conference Paper

FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding

  • Huitong Pan
  • Qi Zhang
  • Cornelia Caragea
  • Eduard C. Dragut
  • Longin Jan Latecki

Flowcharts are graphical tools for representing complex concepts in concise visual representations. This paper introduces the FlowLearn dataset, a resource tailored to enhance the understanding of flowcharts. FlowLearn contains complex scientific flowcharts and simulated flowcharts. The scientific subset contains 3, 858 flowcharts sourced from scientific literature and the simulated subset contains 10, 000 flowcharts created using a customizable script. The dataset is enriched with annotations for visual components, OCR, Mermaid code representation, and VQA question-answer pairs. Despite the proven capabilities of Large Vision-Language Models (LVLMs) in various visual understanding tasks, their effectiveness in decoding flowcharts—a crucial element of scientific communication—has yet to be thoroughly investigated. The FlowLearn test set is crafted to assess the performance of LVLMs in flowchart comprehension. Our study thoroughly evaluates state-of-the-art LVLMs, identifying existing limitations and establishing a foundation for future enhancements in this relatively underexplored domain. For instance, in tasks involving simulated flowcharts, GPT-4V achieved the highest accuracy (58%) in counting the number of nodes, while Claude recorded the highest accuracy (83%) in OCR tasks. Notably, no single model excels in all tasks within the FlowLearn framework, highlighting significant opportunities for further development.

AAAI Conference 2024 Conference Paper

Frequency Spectrum Is More Effective for Multimodal Representation and Fusion: A Multimodal Spectrum Rumor Detector

  • An Lao
  • Qi Zhang
  • Chongyang Shi
  • Longbing Cao
  • Kun Yi
  • Liang Hu
  • Duoqian Miao

Multimodal content, such as mixing text with images, presents significant challenges to rumor detection in social media. Existing multimodal rumor detection has focused on mixing tokens among spatial and sequential locations for unimodal representation or fusing clues of rumor veracity across modalities. However, they suffer from less discriminative unimodal representation and are vulnerable to intricate location dependencies in the time-consuming fusion of spatial and sequential tokens. This work makes the first attempt at multimodal rumor detection in the frequency domain, which efficiently transforms spatial features into the frequency spectrum and obtains highly discriminative spectrum features for multimodal representation and fusion. A novel Frequency Spectrum Representation and fUsion network (FSRU) with dual contrastive learning reveals the frequency spectrum is more effective for multimodal representation and fusion, extracting the informative components for rumor detection. FSRU involves three novel mechanisms: utilizing the Fourier transform to convert features in the spatial domain to the frequency domain, the unimodal spectrum compression, and the cross-modal spectrum co-selection module in the frequency domain. Substantial experiments show that FSRU achieves satisfactory multimodal rumor detection performance.

IJCAI Conference 2024 Conference Paper

HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy Preservation in Multimodal Sentiment Analysis

  • Zhuojia Wu
  • Qi Zhang
  • Duoqian Miao
  • Kun Yi
  • Wei Fan
  • Liang Hu

Multimodal Sentiment Analysis (MSA) aims to identify speakers' sentiment tendencies in multimodal video content, raising serious concerns about privacy risks associated with multimodal data, such as voiceprints and facial images. Recent distributed collaborative learning has been verified as an effective paradigm for privacy preservation in multimodal tasks. However, they often overlook the privacy distinctions among different modalities, struggling to strike a balance between performance and privacy preservation. Consequently, it poses an intriguing question of maximizing multimodal utilization to improve performance while simultaneously protecting necessary modalities. This paper forms the first attempt at modality-specified (i. e. , audio and visual) privacy preservation in MSA tasks. We propose a novel Hybrid Distributed cross-modality cGAN framework (HyDiscGAN), which learns multimodality alignment to generate fake audio and visual features conditioned on shareable de-identified textual data. The objective is to leverage the fake features to approximate real audio and visual content to guarantee privacy preservation while effectively enhancing performance. Extensive experiments show that compared with the state-of-the-art MSA model, HyDiscGAN can achieve superior or competitive performance while preserving privacy.

EAAI Journal 2024 Journal Article

Integrated dynamic multi-threshold pattern recognition with graph attention long short-term neural memory network for water distribution network losses prediction: An automated expert system

  • Minglei Fu
  • Qi Zhang
  • Kezhen Rong
  • Zaher Mundher Yaseen
  • Lejin Zheng
  • Jianfeng Zheng

Water loss is a common and critical problem in water distribution networks, resulting in a decrease in wastewater and user experience. In this research, prediction-classification framework based on deep learning (i. e. , graph attention long short-term neural memory network (GA-LSTM)) is proposed. Dynamic multi-threshold pattern recognition (DMTPR) was developed to identify and classify losses. To improve the model's reliability, we performed hyperparameter optimisation, trained and verified using one year's operation data of a real pipe network, and compared it with the previous methods. In addition, the influence of different aggregation ranges on the accuracy and stability of the model was analysed. The results showed that the GA-LSTM-DMTPR framework can be used as a reliable tool for practical applications because it provides high-precision and high-stability water demand prediction and realises accurate loss identification and classification. This is owing the potential of the model to aggregate the spatial and temporal multi-dimensional information.

AAAI Conference 2024 Conference Paper

Large-Scale Non-convex Stochastic Constrained Distributionally Robust Optimization

  • Qi Zhang
  • Yi Zhou
  • Ashley Prater-Bennette
  • Lixin Shen
  • Shaofeng Zou

Distributionally robust optimization (DRO) is a powerful framework for training robust models against data distribution shifts. This paper focuses on constrained DRO, which has an explicit characterization of the robustness level. Existing studies on constrained DRO mostly focus on convex loss function, and exclude the practical and challenging case with non-convex loss function, e.g., neural network. This paper develops a stochastic algorithm and its performance analysis for non-convex constrained DRO. The computational complexity of our stochastic algorithm at each iteration is independent of the overall dataset size, and thus is suitable for large-scale applications. We focus on the general Cressie-Read family divergence defined uncertainty set which includes chi^2-divergences as a special case. We prove that our algorithm finds an epsilon-stationary point with an improved computational complexity than existing methods. Our method also applies to the smoothed conditional value at risk (CVaR) DRO.

AAAI Conference 2024 Conference Paper

LLMEval: A Preliminary Study on How to Evaluate Large Language Models

  • Yue Zhang
  • Ming Zhang
  • Haipeng Yuan
  • Shichun Liu
  • Yongyao Shi
  • Tao Gui
  • Qi Zhang
  • Xuanjing Huang

Recently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first two questions, which are basically what tasks to give the LLM during testing and what kind of knowledge it should deal with. As for the third question, which is about what standards to use, the types of evaluators, how to score, and how to rank, there hasn't been much discussion. In this paper, we analyze evaluation methods by comparing various criteria with both manual and automatic evaluation, utilizing onsite, crowd-sourcing, public annotators and GPT-4, with different scoring methods and ranking systems. We propose a new dataset, LLMEval and conduct evaluations on 20 LLMs. A total of 2,186 individuals participated, leading to the generation of 243,337 manual annotations and 57,511 automatic evaluation results. We perform comparisons and analyses of different settings and conduct 10 conclusions that can provide some insights for evaluating LLM in the future. The dataset and the results are publicly available at https://github.com/llmeval. The version with the appendix are publicly available at https://arxiv.org/abs/2312.07398.

AAAI Conference 2024 Conference Paper

Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting

  • Qi Zhang
  • Yunfei Gong
  • Daijie Chen
  • Antoni B. Chan
  • Hui Huang

Recent deep learning-based multi-view people detection (MVD) methods have shown promising results on existing datasets. However, current methods are mainly trained and evaluated on small, single scenes with a limited number of multi-view frames and fixed camera views. As a result, these methods may not be practical for detecting people in larger, more complex scenes with severe occlusions and camera calibration errors. This paper focuses on improving multi-view people detection by developing a supervised view-wise contribution weighting approach that better fuses multi-camera information under large scenes. Besides, a large synthetic dataset is adopted to enhance the model's generalization ability and enable more practical evaluation and comparison. The model's performance on new testing scenes is further improved with a simple domain adaptation technique. Experimental results demonstrate the effectiveness of our approach in achieving promising cross-scene multi-view people detection performance.

NeurIPS Conference 2024 Conference Paper

NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction

  • Zixuan Gong
  • Guangyin Bao
  • Qi Zhang
  • Zhongwei Wan
  • Duoqian Miao
  • Shoujin Wang
  • Lei Zhu
  • Changwei Wang

Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e. g. , a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https: //github. com/gongzix/NeuroClips.

EAAI Journal 2024 Journal Article

Reference-based super-resolution reconstruction of remote sensing images based on a coarse-to-fine feature matching transformer

  • Chen Wang
  • Fuzhen Zhu
  • Bing Zhu
  • Qi Zhang
  • Hongbin Ma

Remote sensing image super-resolution reconstruction technology mines the deep details of remote sensing image, which has been widely used in the fields of intelligent and precise agriculture, intelligent transportation, earth surface object recognition, and so on. To obtain more detailed information, we design a reference-based super-resolution reconstruction network. Firstly, the coarse-to-fine feature matching strategy is adopted for the features of the input image. Global coarse matching is performed on the center patch of each block, and then pixel-level local fine matching is performed on the edge patches of the block. This both reduces the amount of computation and improves the matching accuracy. A threshold is set to determine whether the feature matching results meet the criteria for feature transfer. Finally, different scale features are fused through several convolutional layers and sampling operations, obtaining the reconstruction features after a fourfold increase in resolution. The ultimate super-resolution image is generated through a decoder. We have performed training and testing on remote sensing datasets. Compared to the current state-of-the-art methods, our proposed method is visually more details and outperforms other methods in terms of objective evaluation metrics.

NeurIPS Conference 2024 Conference Paper

Reinforcement Learning with Euclidean Data Augmentation for State-Based Continuous Control

  • Jinzhu Luo
  • Dingyang Chen
  • Qi Zhang

Data augmentation creates new data points by transforming the original ones for an reinforcement learning (RL) agent to learn from, which has been shown to be effective for the objective of improving data efficiency of RL for continuous control. Prior work towards this objective has been largely restricted to perturbation-based data augmentation where new data points are created by perturbing the original ones, which has been impressively effective for tasks where the RL agent observe control states as images with perturbations including random cropping, shifting, etc. This work focuses on state-based control, where the RL agent can directly observe raw kinematic and task features, and considers an alternative data augmentation applied to these features based on Euclidean symmetries under transformations like rotations. We show that the default state features used in exiting benchmark tasks that are based on joint configurations are not amenable to Euclidean transformations. We therefore advocate using state features based on configurations of the limbs (i. e. , rigid bodies connected by joints) that instead provides rich augmented data under Euclidean transformations. With minimal hyperparameter tuning, we show this new Euclidean data augmentation strategy significantly improve both data efficiency and asymptotic performance of RL on a wide range of continuous control tasks.

AAAI Conference 2024 Conference Paper

Selective Focus: Investigating Semantics Sensitivity in Post-training Quantization for Lane Detection

  • Yunqian Fan
  • Xiuying Wei
  • Ruihao Gong
  • Yuqing Ma
  • Xiangguo Zhang
  • Qi Zhang
  • Xianglong Liu

Lane detection (LD) plays a crucial role in enhancing the L2+ capabilities of autonomous driving, capturing widespread attention. The Post-Processing Quantization (PTQ) could facilitate the practical application of LD models, enabling fast speeds and limited memories without labeled data. However, prior PTQ methods do not consider the complex LD outputs that contain physical semantics, such as offsets, locations, etc., and thus cannot be directly applied to LD models. In this paper, we pioneeringly investigate semantic sensitivity to post-processing for lane detection with a novel Lane Distortion Score. Moreover, we identify two main factors impacting the LD performance after quantization, namely intra-head sensitivity and inter-head sensitivity, where a small quantization error in specific semantics can cause significant lane distortion. Thus, we propose a Selective Focus framework deployed with Semantic Guided Focus and Sensitivity Aware Selection modules, to incorporate post-processing information into PTQ reconstruction. Based on the observed intra-head sensitivity, Semantic Guided Focus is introduced to prioritize foreground-related semantics using a practical proxy. For inter-head sensitivity, we present Sensitivity Aware Selection, efficiently recognizing influential prediction heads and refining the optimization objectives at runtime. Extensive experiments have been done on a wide variety of models including keypoint-, anchor-, curve-, and segmentation-based ones. Our method produces quantized models in minutes on a single GPU and can achieve 6.4\% F1 Score improvement on the CULane dataset. Code and supplementary statement can be found at https://github.com/PannenetsF/SelectiveFocus.

JBHI Journal 2024 Journal Article

SSCRB: Predicting circRNA-RBP Interaction Sites Using a Sequence and Structural Feature-Based Attention Model

  • Liwei Liu
  • Yuxiao Wei
  • Qi Zhang
  • Qi Zhao

The prediction of interaction sites between circular RNA (circRNA) and RNA binding proteins (RBPs) is crucial for regulating diseases and discovering new treatment approaches. Computational models have been widely used to predict circRNA-RBP interaction sites due to the availability of genome-wide circRNA binding event data. However, efficiently obtaining multi-scale circRNA features to improve prediction accuracy remains a challenging problem. In this study, we propose SSCRB, a lightweight model for predicting circRNA-RBP interaction sites. Our model extracts both sequence and structural features of circRNA and incorporates multi-scale features through the attention mechanism. Furthermore, we develop an ensemble model by combining multiple submodels to enhance predictive performance and generalizability. We evaluate SSCRB on 37 circRNA datasets and compare it with other state-of-the-art methods. The average AUC of SSCRB is 97. 66%, demonstrating its efficiency and robustness. SSCRB outperforms other methods in terms of prediction accuracy while requiring significantly fewer computational resources.

AAAI Conference 2024 Conference Paper

Text Diffusion with Reinforced Conditioning

  • Yuxuan Liu
  • Tianchi Yang
  • Shaohan Huang
  • Zihan Zhang
  • Haizhen Huang
  • Furu Wei
  • Weiwei Deng
  • Feng Sun

Diffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models still fall short in their performance due to a challenge in handling the discreteness of language. This paper thoroughly analyzes text diffusion models and uncovers two significant limitations: degradation of self-conditioning during training and misalignment between training and sampling. Motivated by our findings, we propose a novel Text Diffusion model called TReC, which mitigates the degradation with Reinforced Conditioning and the misalignment by Time-Aware Variance Scaling. Our extensive experiments demonstrate the competitiveness of TReC against autoregressive, non-autoregressive, and diffusion baselines. Moreover, qualitative analysis shows its advanced ability to fully utilize the diffusion process in refining samples.

IROS Conference 2024 Conference Paper

Visual Perception System for Autonomous Driving

  • Qi Zhang
  • Siyuan Gou
  • Wenbin Li

The recent surge in interest in autonomous driving is fueled by its rapidly developing capacity to enhance safety, efficiency, and convenience. A key component of autonomous driving technology lies in its perceptual systems, where advancements have led to more precise algorithms applicable to autonomous driving, such as vision-based Simultaneous Localization and Mapping (SLAM), object detection, and tracking algorithms. This work introduces a visual-based perception system for autonomous driving that integrates trajectory tracking and prediction of moving objects to prevent collisions while addressing the localization and mapping needs of autonomous driving. The system leverages motion cues from pedestrians to monitor and forecast their movements while simultaneously mapping the environment. This integrated approach resolves camera localization and tracks other moving objects in the scene, ultimately generating a sparse map to facilitate vehicle navigation. The performance, efficiency, and resilience of this approach are demonstrated through comprehensive evaluations of both simulated and real-world datasets.

NeurIPS Conference 2023 Conference Paper

A Comprehensive Study on Text-attributed Graphs: Benchmarking and Rethinking

  • Hao Yan
  • Chaozhuo Li
  • Ruosong Long
  • Chao Yan
  • Jianan Zhao
  • Wenwen Zhuang
  • Jun Yin
  • Peiyan Zhang

Text-attributed graphs (TAGs) are prevalent in various real-world scenarios, where each node is associated with a text description. The cornerstone of representation learning on TAGs lies in the seamless integration of textual semantics within individual nodes and the topological connections across nodes. Recent advancements in pre-trained language models (PLMs) and graph neural networks (GNNs) have facilitated effective learning on TAGs, garnering increased research interest. However, the absence of meaningful benchmark datasets and standardized evaluation procedures for TAGs has impeded progress in this field. In this paper, we propose CS-TAG, a comprehensive and diverse collection of challenging benchmark datasets for TAGs. The CS-TAG datasets are notably large in scale and encompass a wide range of domains, spanning from citation networks to purchase graphs. In addition to building the datasets, we conduct extensive benchmark experiments over CS-TAG with various learning paradigms, including PLMs, GNNs, PLM-GNN co-training methods, and the proposed novel topological pre-training of language models. In a nutshell, we provide an overview of the CS-TAG datasets, standardized evaluation procedures, and present baseline experiments. The entire CS-TAG project is publicly accessible at \url{https: //github. com/sktsherlock/TAG-Benchmark}.

AAAI Conference 2023 Conference Paper

A Dynamics and Task Decoupled Reinforcement Learning Architecture for High-Efficiency Dynamic Target Intercept

  • Dora D. Liu
  • Liang Hu
  • Qi Zhang
  • Tangwei Ye
  • Usman Naseem
  • Zhong Yuan Lai

Due to the flexibility and ease of control, unmanned aerial vehicles (UAVs) have been increasingly used in various scenarios and applications in recent years. Training UAVs with reinforcement learning (RL) for a specific task is often expensive in terms of time and computation. However, it is known that the main effort of the learning process is made to fit the low-level physical dynamics systems instead of the high-level task itself. In this paper, we study to apply UAVs in the dynamic target intercept (DTI) task, where the dynamics systems equipped by different UAV models are correspondingly distinct. To this end, we propose a dynamics and task decoupled RL architecture to address the inefficient learning procedure, where the RL module focuses on modeling the DTI task without involving physical dynamics, and the design of states, actions, and rewards are completely task-oriented while the dynamics control module can adaptively convert actions from the RL module to dynamics signals to control different UAVs without retraining the RL module. We show the efficiency and efficacy of our results in comparison and ablation experiments against state-of-the-art methods.

TCS Journal 2023 Journal Article

A unified greedy approximation for several dominating set problems

  • Hao Zhong
  • Yong Tang
  • Qi Zhang
  • Ronghua Lin
  • Weisheng Li

Minimum Dominating Set and Minimum Connected Dominating Set are classic graph problems that have been studied extensively in the literature. These two problems and their various variants are NP-hard in a general graph, and for some of them greedy approximation algorithms have been proposed. In this paper, by designing two potential functions that enjoy submodularity or a weak submodularity, we propose a unified O( ln ⁡ δ )-approximation algorithm for a generalized Minimum (Connected) Dominating Set that includes Minimum (Connected) Dominating Set, Minimum (Connected) Total Dominating Set, Minimum (Connected) *-Dominating Set and Minimum (Connected) Positive Influence Dominating Set, where δ is the maximum node degree of the input graph. For each specific version of the generalized Minimum (Connected) Dominating Set, the unified algorithm either matches the best one of existing approximation algorithms, or gives the first approximation solution.

EAAI Journal 2023 Journal Article

Automatic topology optimization of echo state network based on particle swarm optimization

  • Yu Xue
  • Qi Zhang
  • Adam Slowik

The task of time series forecasting is to predict the future trend of data based on the collected historical data, providing theoretical and data support for human judgment and decision making. Randomization-based echo state networks (ESNs) are widely used in the research and application field of time series analysis for their simple structure and fast training speed. The core of the ESN is its dynamic reservoir, which the original reservoir are randomly generated and controlled only by parameter sparsity, often performing poorly on complex tasks and affecting the performance of networks. Manual design of the topology of reservoir is difficult, time-consuming and inconvenient to operate, which is not conducive to the development of ESNs. The construction of a suitable reservoir topology for practical application problems to enrich reservoir dynamics is a hot research point for researchers. In this paper, an automatic optimization method is introduced into the topology optimization of ESN (TP-ESN), and the particle swarm optimization algorithm is used to optimize the topological construction of the ESN. The connection structure between the reservoir neurons is first encoded and then iteratively optimized. The optimized structure is decoded and then the reservoir is initialized for ESN training. Prediction results on Mackey–Glass benchmark time series and two electroencephalogram (EEG) datasets demonstrate that TP-ESN method can have better adaptability, stronger prediction ability and stability than several other manually designed ESN reservoir topologies in the case of relatively complex tasks.

JBHI Journal 2023 Journal Article

CCT-Unet: A U-Shaped Network Based on Convolution Coupled Transformer for Segmentation of Peripheral and Transition Zones in Prostate MRI

  • Yifei Yan
  • Rongzong Liu
  • Haobo Chen
  • Limin Zhang
  • Qi Zhang

The accurate segmentation of prostate region in magnetic resonance imaging (MRI) can provide reliable basis for artificially intelligent diagnosis of prostate cancer. Transformer-based models have been increasingly used in image analysis due to their ability to acquire long-term global contextual features. Although Transformer can provide feature representations of the overall appearance and contour representations at long distance, it does not perform well on small-scale datasets of prostate MRI due to its insensitivity to local variation such as the heterogeneity of the grayscale intensities in the peripheral zone and transition zone across patients; meanwhile, the convolutional neural network (CNN) could retain these local features well. Therefore, a robust prostate segmentation model that can aggregate the characteristics of CNN and Transformer is desired. In this work, a U-shaped network based on the convolution coupled Transformer is proposed for segmentation of peripheral and transition zones in prostate MRI, named the convolution coupled Transformer U-Net (CCT-Unet). The convolutional embedding block is first designed for encoding high-resolution input to retain the edge detail of the image. Then the convolution coupled Transformer block is proposed to enhance the ability of local feature extraction and capture long-term correlation that encompass anatomical information. The feature conversion module is also proposed to alleviate the semantic gap in the process of jumping connection. Extensive experiments have been conducted to compare our CCT-Unet with several state-of-the-art methods on both the ProstateX open dataset and the self-bulit Huashan dataset, and the results have consistently shown the accuracy and robustness of our CCT-Unet in MRI prostate segmentation.

NeurIPS Conference 2023 Conference Paper

FourierGNN: Rethinking Multivariate Time Series Forecasting from a Pure Graph Perspective

  • Kun Yi
  • Qi Zhang
  • Wei Fan
  • Hui He
  • Liang Hu
  • Pengyang Wang
  • Ning An
  • Longbing Cao

Multivariate time series (MTS) forecasting has shown great importance in numerous industries. Current state-of-the-art graph neural network (GNN)-based forecasting methods usually require both graph networks (e. g. , GCN) and temporal networks (e. g. , LSTM) to capture inter-series (spatial) dynamics and intra-series (temporal) dependencies, respectively. However, the uncertain compatibility of the two networks puts an extra burden on handcrafted model designs. Moreover, the separate spatial and temporal modeling naturally violates the unified spatiotemporal inter-dependencies in real world, which largely hinders the forecasting performance. To overcome these problems, we explore an interesting direction of directly applying graph networks and rethink MTS forecasting from a pure graph perspective. We first define a novel data structure, hypervariate graph, which regards each series value (regardless of variates or timestamps) as a graph node, and represents sliding windows as space-time fully-connected graphs. This perspective considers spatiotemporal dynamics unitedly and reformulates classic MTS forecasting into the predictions on hypervariate graphs. Then, we propose a novel architecture Fourier Graph Neural Network (FourierGNN) by stacking our proposed Fourier Graph Operator (FGO) to perform matrix multiplications in Fourier space. FourierGNN accommodates adequate expressiveness and achieves much lower complexity, which can effectively and efficiently accomplish {the forecasting}. Besides, our theoretical analysis reveals FGO's equivalence to graph convolutions in the time domain, which further verifies the validity of FourierGNN. Extensive experiments on seven datasets have demonstrated our superior performance with higher efficiency and fewer parameters compared with state-of-the-art methods. Code is available at this repository: https: //github. com/aikunyi/FourierGNN.

NeurIPS Conference 2023 Conference Paper

Frequency-domain MLPs are More Effective Learners in Time Series Forecasting

  • Kun Yi
  • Qi Zhang
  • Wei Fan
  • Shoujin Wang
  • Pengyang Wang
  • Hui He
  • Ning An
  • Defu Lian

Time series forecasting has played the key role in different industrial, including finance, traffic, energy, and healthcare domains. While existing literatures have designed many sophisticated architectures based on RNNs, GNNs, or Transformers, another kind of approaches based on multi-layer perceptrons (MLPs) are proposed with simple structure, low complexity, and superior performance. However, most MLP-based forecasting methods suffer from the point-wise mappings and information bottleneck, which largely hinders the forecasting performance. To overcome this problem, we explore a novel direction of applying MLPs in the frequency domain for time series forecasting. We investigate the learned patterns of frequency-domain MLPs and discover their two inherent characteristic benefiting forecasting, (i) global view: frequency spectrum makes MLPs own a complete view for signals and learn global dependencies more easily, and (ii) energy compaction: frequency-domain MLPs concentrate on smaller key part of frequency components with compact signal energy. Then, we propose FreTS, a simple yet effective architecture built upon Frequency-domain MLPs for Time Series forecasting. FreTS mainly involves two stages, (i) Domain Conversion, that transforms time-domain signals into complex numbers of frequency domain; (ii) Frequency Learning, that performs our redesigned MLPs for the learning of real and imaginary part of frequency components. The above stages operated on both inter-series and intra-series scales further contribute to channel-wise and time-wise dependency learning. Extensive experiments on 13 real-world benchmarks (including 7 benchmarks for short-term forecasting and 6 benchmarks for long-term forecasting) demonstrate our consistent superiority over state-of-the-art methods. Code is available at this repository: https: //github. com/aikunyi/FreTS.

NeurIPS Conference 2023 Conference Paper

Identifiable Contrastive Learning with Automatic Feature Importance Discovery

  • Qi Zhang
  • Yifei Wang
  • Yisen Wang

Existing contrastive learning methods rely on pairwise sample contrast $z_x^\top z_{x'}$ to learn data representations, but the learned features often lack clear interpretability from a human perspective. Theoretically, it lacks feature identifiability and different initialization may lead to totally different features. In this paper, we study a new method named tri-factor contrastive learning (triCL) that involves a 3-factor contrast in the form of $z_x^\top S z_{x'}$, where $S=\text{diag}(s_1, \dots, s_k)$ is a learnable diagonal matrix that automatically captures the importance of each feature. We show that by this simple extension, triCL can not only obtain identifiable features that eliminate randomness but also obtain more interpretable features that are ordered according to the importance matrix $S$. We show that features with high importance have nice interpretability by capturing common classwise features, and obtain superior performance when evaluated for image retrieval using a few features. The proposed triCL objective is general and can be applied to different contrastive learning methods like SimCLR and CLIP. We believe that it is a better alternative to existing 2-factor contrastive learning by improving its identifiability and interpretability with minimal overhead. Code is available at https: //github. com/PKU-ML/Tri-factor-Contrastive-Learning.

NeurIPS Conference 2023 Conference Paper

MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision Transformers

  • Yu Zhang
  • Yepeng Liu
  • Duoqian Miao
  • Qi Zhang
  • Yiwei Shi
  • Liang Hu

Vision Transformer (ViT) faces obstacles in wide application due to its huge computational cost. Almost all existing studies on compressing ViT adopt the manner of splitting an image with a single granularity, with very few exploration of splitting an image with multi-granularity. As we know, important information often randomly concentrate in few regions of an image, necessitating multi-granularity attention allocation to an image. Enlightened by this, we introduce the multi-granularity strategy to compress ViT, which is simple but effective. We propose a two-stage multi-granularity framework, MG-ViT, to balance ViT’s performance and computational cost. In single-granularity inference stage, an input image is split into a small number of patches for simple inference. If necessary, multi-granularity inference stage will be instigated, where the important patches are further subsplit into multi-finer-grained patches for subsequent inference. Moreover, prior studies on compression only for classification, while we extend the multi-granularity strategy to hierarchical ViT for downstream tasks such as detection and segmentation. Extensive experiments Prove the effectiveness of the multi-granularity strategy. For instance, on ImageNet, without any loss of performance, MG-ViT reduces 47\% FLOPs of LV-ViT-S and 56\% FLOPs of DeiT-S.

NeurIPS Conference 2023 Conference Paper

Model-enhanced Vector Index

  • Hailin Zhang
  • Yujing Wang
  • Qi Chen
  • Ruiheng Chang
  • Ting Zhang
  • Ziming Miao
  • Yingyan Hou
  • Yang Ding

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions offer better model quality, but are hindered by unacceptable serving latency and the inability to support document updates. In this paper, we aim to enhance the vector index with end-to-end deep generative models, leveraging the differentiable advantages of deep retrieval models while maintaining desirable serving efficiency. We propose Model-enhanced Vector Index (MEVI), a differentiable model-enhanced index empowered by a twin-tower representation model. MEVI leverages a Residual Quantization (RQ) codebook to bridge the sequence-to-sequence deep retrieval and embedding-based models. To substantially reduce the inference time, instead of decoding the unique document ids in long sequential steps, we first generate some semantic virtual cluster ids of candidate documents in a small number of steps, and then leverage the well-adapted embedding vectors to further perform a fine-grained search for the relevant documents in the candidate virtual clusters. We empirically show that our model achieves better performance on the commonly used academic benchmarks MSMARCO Passage and Natural Questions, with comparable serving latency to dense retrieval solutions.

AIJ Journal 2023 Journal Article

Risk-aware analysis for interpretations of probabilistic achievement and maintenance commitments

  • Qi Zhang
  • Edmund H. Durfee
  • Satinder Singh

Probabilistic commitments provide a computational framework for multi-agent coordination, where one autonomous agent (the commitment provider), commits to a future course of action that probabilistically influences the local state of another agent (the commitment recipient) in ways that the recipient desires. Conventionally, a probabilistic commitment is specified abstractly so as to give the provider latitude at run time about how to achieve it. Unfortunately, as we analyze in this article, this abstraction incurs a risk of suboptimal performance for the recipient. For (achievement) commitments by the provider to achieve conditions that the recipient prefers but that do not initially hold, we prove that the recipient can make modeling choices that bound its risk of suboptimality. Somewhat surprisingly, however, for (maintenance) commitments by the provider to maintain conditions whose initial values are already ones the recipient prefers, we prove that no such bounds on suboptimality risk are possible. We study the two types of commitments empirically to measure the suboptimality they incur under different conditions, and based on our theoretical and empirical results suggest that adding selective details when specifying probabilistic maintenance commitments can be beneficial.

AAAI Conference 2023 Conference Paper

Self-Supervised Learning for Multilevel Skeleton-Based Forgery Detection via Temporal-Causal Consistency of Actions

  • Liang Hu
  • Dora D. Liu
  • Qi Zhang
  • Usman Naseem
  • Zhong Yuan Lai

Skeleton-based human action recognition and analysis have become increasingly attainable in many areas, such as security surveillance and anomaly detection. Given the prevalence of skeleton-based applications, tampering attacks on human skeletal features have emerged very recently. In particular, checking the temporal inconsistency and/or incoherence (TII) in the skeletal sequence of human action is a principle of forgery detection. To this end, we propose an approach to self-supervised learning of the temporal causality behind human action, which can effectively check TII in skeletal sequences. Especially, we design a multilevel skeleton-based forgery detection framework to recognize the forgery on frame level, clip level, and action level in terms of learning the corresponding temporal-causal skeleton representations for each level. Specifically, a hierarchical graph convolution network architecture is designed to learn low-level skeleton representations based on physical skeleton connections and high-level action representations based on temporal-causal dependencies for specific actions. Extensive experiments consistently show state-of-the-art results on multilevel forgery detection tasks and superior performance of our framework compared to current competing methods.

AAAI Conference 2022 Conference Paper

A Multi-Agent Reinforcement Learning Approach for Efficient Client Selection in Federated Learning

  • Sai Qian Zhang
  • Jieyu Lin
  • Qi Zhang

Federated learning (FL) is a training technique that enables client devices to jointly learn a shared model by aggregating locally computed models without exposing their raw data. While most of the existing work focuses on improving the FL model accuracy, in this paper, we focus on the improving the training efficiency, which is often a hurdle for adopting FL in real world applications. Specifically, we design an efficient FL framework which jointly optimizes model accuracy, processing latency and communication efficiency, all of which are primary design considerations for real implementation of FL. Inspired by the recent success of Multi Agent Reinforcement Learning (MARL) in solving complex control problems, we present FedMarl, a federated learning framework that relies on trained MARL agents to perform efficient client selection. Experiments show that FedMarl can significantly improve model accuracy with much lower processing latency and communication cost.

NeurIPS Conference 2022 Conference Paper

A Neural Corpus Indexer for Document Retrieval

  • Yujing Wang
  • Yingyan Hou
  • Haonan Wang
  • Ziming Miao
  • Shibin Wu
  • Qi Chen
  • Yuqing Xia
  • Chengmin Chi

Current state-of-the-art document retrieval solutions mainly follow an index-retrieve paradigm, where the index is hard to be directly optimized for the final retrieval target. In this paper, we aim to show that an end-to-end deep neural network unifying training and indexing stages can significantly improve the recall performance of traditional methods. To this end, we propose Neural Corpus Indexer (NCI), a sequence-to-sequence network that generates relevant document identifiers directly for a designated query. To optimize the recall performance of NCI, we invent a prefix-aware weight-adaptive decoder architecture, and leverage tailored techniques including query generation, semantic document identifiers, and consistency-based regularization. Empirical studies demonstrated the superiority of NCI on two commonly used academic benchmarks, achieving +21. 4% and +16. 8% relative enhancement for Recall@1 on NQ320k dataset and R-Precision on TriviaQA dataset, respectively, compared to the best baseline method.

IJCAI Conference 2022 Conference Paper

A Probabilistic Code Balance Constraint with Compactness and Informativeness Enhancement for Deep Supervised Hashing

  • Qi Zhang
  • Liang Hu
  • Longbing Cao
  • Chongyang Shi
  • Shoujin Wang
  • Dora D. Liu

Building on deep representation learning, deep supervised hashing has achieved promising performance in tasks like similarity retrieval. However, conventional code balance constraints (i. e. , bit balance and bit uncorrelation) imposed on avoiding overfitting and improving hash code quality are unsuitable for deep supervised hashing owing to their inefficiency and impracticality of simultaneously learning deep data representations and hash functions. To address this issue, we propose probabilistic code balance constraints on deep supervised hashing to force each hash code to conform to a discrete uniform distribution. Accordingly, a Wasserstein regularizer aligns the distribution of generated hash codes to a uniform distribution. Theoretical analyses reveal that the proposed constraints form a general deep hashing framework for both bit balance and bit uncorrelation and maximizing the mutual information between data input and their corresponding hash codes. Extensive empirical analyses on two benchmark datasets further demonstrate the enhancement of compactness and informativeness of hash codes for deep supervised hash to improve retrieval performance (code available at: https: //github. com/mumuxi/dshwr).

EAAI Journal 2022 Journal Article

An improved brain storm optimization algorithm with new solution generation strategies for classification

  • Yu Xue
  • Qi Zhang
  • Yan Zhao

In recent years, brain storm optimization (BSO) algorithm has received much attention in solving classical optimization problems and is used to implement evolutionary classification models. However, in practical applications, large-scale datasets complicate the structure of the classification model, which can have a great impact on the classification performance. In the optimization process, the traditional single-strategy BSO cannot preserve the information of dominant solution well, and its generation strategy is inefficient in solving various complex practical problems. To solve this problem, we introduce feature selection to improve the optimization model structure. Meanwhile, in order to enhance the search capability of BSO, three new generation strategy are embedded in the BSO algorithm in this paper. With the three generation methods of global optimal, local optimal and nearest neighbor, the information of the dominant solution can be better preserved and the search efficiency can be improved. The performance of the proposed generation strategy in solving classification problems is demonstrated on ten datasets with different sizes and dimensions. The experimental results reveal that the new generation strategy can enhance the performance of BSO algorithm for solving classification problems.

AAAI Conference 2022 Conference Paper

CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting

  • Hui He
  • Qi Zhang
  • Simeng Bai
  • Kun Yi
  • Zhendong Niu

Modeling complex hierarchical and grouped feature interaction in the multivariate time series data is indispensable to comprehending the data dynamics and predicting the future condition. The implicit feature interaction and highdimensional data make multivariate forecasting very challenging. Many existing works did not put more emphasis on exploring explicit correlation among multiple time-series data, and complicated models are designed to capture longand short-range patterns with the aid of attention mechanisms. In this work, we think that a pre-defined graph or a general learning method is difficult due to its irregular structure. Hence, we present CATN, an end-to-end model of Cross Attentive Tree-aware Network to jointly capture the interseries correlation and intra-series temporal patterns. We first construct a tree structure to learn hierarchical and grouped correlation and design an embedding approach that can pass a dynamic message to generalize implicit but interpretable cross features among multiple time series. Next in the temporal aspect, we propose a multi-level dependency learning mechanism including global&local learning and cross attention mechanism, which can combine long-range dependencies, short-range dependencies as well as cross dependencies at different time steps. The extensive experiments on different datasets from real-world show the effectiveness and robustness of the method we proposed when compared with existing state-of-the-art methods.

YNIMG Journal 2022 Journal Article

Distinct networks coupled with parietal cortex for spatial representations inside and outside the visual field

  • Bo Zhang
  • Fan Wang
  • Qi Zhang
  • Yuji Naya

Our mental representation of egocentric space is influenced by the disproportionate sensory perception of the body. Previous studies have focused on the neural architecture for egocentric representations within the visual field. However, the space representation underlying the body is still unclear. To address this problem, we applied both functional Magnitude Resonance Imaging (fMRI) and Magnetoencephalography (MEG) to a spatial-memory paradigm by using a virtual environment in which human participants remembered a target location left, right, or back relative to their own body. Both experiments showed larger involvement of the frontoparietal network in representing a retrieved target on the left/right side than on the back. Conversely, the medial temporal lobe (MTL)-parietal network was more involved in retrieving a target behind the participants. The MEG data showed an earlier activation of the MTL-parietal network than that of the frontoparietal network during retrieval of a target location. These findings suggest that the parietal cortex may represent the entire space around the self-body by coordinating two distinct brain networks.

NeurIPS Conference 2022 Conference Paper

How Mask Matters: Towards Theoretical Understandings of Masked Autoencoders

  • Qi Zhang
  • Yifei Wang
  • Yisen Wang

Masked Autoencoders (MAE) based on a reconstruction task have risen to be a promising paradigm for self-supervised learning (SSL) and achieve state-of-the-art performance across different benchmark datasets. However, despite its impressive empirical success, there is still limited theoretical understanding of it. In this paper, we propose a theoretical understanding of how masking matters for MAE to learn meaningful features. We establish a close connection between MAE and contrastive learning, which shows that MAE implicit aligns the mask-induced positive pairs. Built upon this connection, we develop the first downstream guarantees for MAE methods, and analyze the effect of mask ratio. Besides, as a result of the implicit alignment, we also point out the dimensional collapse issue of MAE, and propose a Uniformity-enhanced MAE (U-MAE) loss that can effectively address this issue and bring significant improvements on real-world datasets, including CIFAR-10, ImageNet-100, and ImageNet-1K. Code is available at https: //github. com/zhangq327/U-MAE.

NeurIPS Conference 2022 Conference Paper

Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models

  • Xiuying Wei
  • Yunchen Zhang
  • Xiangguo Zhang
  • Ruihao Gong
  • Shanghang Zhang
  • Qi Zhang
  • Fengwei Yu
  • Xianglong Liu

Transformer architecture has become the fundamental element of the widespread natural language processing~(NLP) models. With the trends of large NLP models, the increasing memory and computation costs hinder their efficient deployment on resource-limited devices. Therefore, transformer quantization attracts wide research interest. Recent work recognizes that structured outliers are the critical bottleneck for quantization performance. However, their proposed methods increase the computation overhead and still leave the outliers there. To fundamentally address this problem, this paper delves into the inherent inducement and importance of the outliers. We discover that $\boldsymbol \gamma$ in LayerNorm (LN) acts as a sinful amplifier for the outliers, and the importance of outliers varies greatly where some outliers provided by a few tokens cover a large area but can be clipped sharply without negative impacts. Motivated by these findings, we propose an outlier suppression framework including two components: Gamma Migration and Token-Wise Clipping. The Gamma Migration migrates the outlier amplifier to subsequent modules in an equivalent transformation, contributing to a more quantization-friendly model without any extra burden. The Token-Wise Clipping takes advantage of the large variance of token range and designs a token-wise coarse-to-fine pipeline, obtaining a clipping range with minimal final quantization loss in an efficient way. This framework effectively suppresses the outliers and can be used in a plug-and-play mode. Extensive experiments prove that our framework surpasses the existing works and, for the first time, pushes the 6-bit post-training BERT quantization to the full-precision (FP) level. Our code is available at https: //github. com/wimh966/outlier_suppression.

ICLR Conference 2021 Conference Paper

BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

  • Yuhang Li 0001
  • Ruihao Gong
  • Xu Tan
  • Yang Yang
  • Peng Hu
  • Qi Zhang
  • Fengwei Yu
  • Wei Wang 0059

We study the challenging task of neural network quantization without end-to-end retraining, called Post-training Quantization (PTQ). PTQ usually requires a small subset of training data but produces less powerful quantized models than Quantization-Aware Training (QAT). In this work, we propose a novel PTQ framework, dubbed BRECQ, which pushes the limits of bitwidth in PTQ down to INT2 for the first time. BRECQ leverages the basic building blocks in neural networks and reconstructs them one-by-one. In a comprehensive theoretical study of the second-order error, we show that BRECQ achieves a good balance between cross-layer dependency and generalization error. To further employ the power of quantization, the mixed precision technique is incorporated in our framework by approximating the inter-layer and intra-layer sensitivity. Extensive experiments on various handcrafted and searched neural architectures are conducted for both image classification and object detection tasks. And for the first time we prove that, without bells and whistles, PTQ can attain 4-bit ResNet and MobileNetV2 comparable with QAT and enjoy 240 times faster production of quantized models. Codes are available at https://github.com/yhhhli/BRECQ.

AAAI Conference 2021 Conference Paper

Efficient Querying for Cooperative Probabilistic Commitments

  • Qi Zhang
  • Edmund H. Durfee
  • Satinder Singh

Multiagent systems can use commitments as the core of a general coordination infrastructure, supporting both cooperative and non-cooperative interactions. Agents whose objectives are aligned, and where one agent can help another achieve greater reward by sacrificing some of its own reward, should choose a cooperative commitment to maximize their joint reward. We present a solution to the problem of how cooperative agents can efficiently find an (approximately) optimal commitment by querying about carefully-selected commitment choices. We prove structural properties of the agents’ values as functions of the parameters of the commitment specification, and develop a greedy method for composing a query with provable approximation bounds, which we empirically show can find nearly optimal commitments in a fraction of the time methods that lack our insights require.

JBHI Journal 2021 Journal Article

Graph Based Multichannel Feature Fusion for Wrist Pulse Diagnosis

  • Qi Zhang
  • Jianhang Zhou
  • Bob Zhang

It is well known in Traditional Chinese Medicine (TCM) that a person's wrist pulse signal can reflect their health condition. Recently, many computerized wrist pulse AI systems have been proposed to simulate a practitioner's three fingers in order to acquire the wrist pulse signals (three positions/channels) from a candidate's wrist dynamically, before evaluating their health status based on the various feature extraction and detection methods. However, few works have investigated the correlation of the extracted features from the three wrist channels and comprehensively fused the various features together, which can improve the performance of wrist pulse diagnosis. In this paper, we propose a graph based multichannel feature fusion (GBMFF) method to utilize the multichannel features of the wrist pulse signals effectively. In detail, two different sensors, i. e. , pressure and photoelectricity are used to capture the three channels of the wrist pulse signals. These are used to generate two different features by applying the stacked sparse autoencoder and wavelet scattering. Each feature of one wrist pulse sample is regarded as a node associated with its corresponding feature vector, and used to construct a graph for one candidate. A novel algorithm is implemented to construct different graphs for different candidates, which are used for wrist pulse diagnosis by developing graph convolutional networks. Experimental results indicate that our proposed AI-based method can obtain superior performances compared to other state-of-the-art approaches.

JBHI Journal 2021 Journal Article

Method of Tumor Pathological Micronecrosis Quantification Via Deep Learning From Label Fuzzy Proportions

  • Qiancheng Ye
  • Qi Zhang
  • Yu Tian
  • Tianshu Zhou
  • Hongbin Ge
  • Jiajun Wu
  • Na Lu
  • Xueli Bai

The presence of necrosis is associated with tumor progression and patient outcomes in many cancers, but existing analyses rarely adopt quantitative methods because the manual quantification of histopathological features is too expensive. We aim to accurately identify necrotic regions on hematoxylin and eosin (HE)–stained slides and to calculate the ratio of necrosis with minimal annotations on the images. An adaptive method named Learning from Label Fuzzy Proportions (LLFP) was introduced to histopathological image analysis. Two datasets of liver cancer HE slides were collected to verify the feasibility of the method by training on the internal set using cross validation and performing validation on the external set, along with ensemble learning to improve performance. The models from cross validation performed relatively stably in identifying necrosis, with a Concordance Index of the Slide Necrosis Score (CISNS) of 0. 9165±0. 0089 in the internal test set. The integration model improved the CISNS to 0. 9341 and achieved a CISNS of 0. 8278 on the external set. There were significant differences in survival (p = 0. 0060) between the three groups divided according to the calculated necrosis ratio. The proposed method can build an integration model good at distinguishing necrosis and capable of clinical assistance as an automatic tool to stratify patients with different risks or as a cluster tool for the quantification of histopathological features. We presented a method effective for identifying histopathological features and suggested that the extent of necrosis, especially micronecrosis, in liver cancer is related to patient outcomes.

NeurIPS Conference 2021 Conference Paper

MQBench: Towards Reproducible and Deployable Model Quantization Benchmark

  • Yuhang Li
  • Mingzhu Shen
  • Jian Ma
  • Yan Ren
  • Mingxin Zhao
  • Qi Zhang
  • Ruihao Gong
  • Fengwei Yu

Model quantization has emerged as an indispensable technique to accelerate deep learning inference. Although researchers continue to push the frontier of quantization algorithms, existing quantization work is often unreproducible and undeployable. This is because researchers do not choose consistent training pipelines and ignore the requirements for hardware deployments. In this work, we propose Model Quantization Benchmark (MQBench), a first attempt to evaluate, analyze, and benchmark the reproducibility and deployability for model quantization algorithms. We choose multiple different platforms for real-world deployments, including CPU, GPU, ASIC, DSP, and evaluate extensive state-of-the-art quantization algorithms under a unified training pipeline. MQBench acts like a bridge to connect the algorithm and the hardware. We conduct a comprehensive analysis and find considerable intuitive or counter-intuitive insights. By aligning up the training settings, we find existing algorithms have about-the-same performance on the conventional academic track. While for the hardware-deployable quantization, there is a huge accuracy gap and still a long way to go. Surprisingly, no existing algorithm wins every challenge in MQBench, and we hope this work could inspire future research directions.

AAAI Conference 2021 Conference Paper

Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic Modeling

  • Yicheng Zou
  • Lujun Zhao
  • Yangyang Kang
  • Jun Lin
  • Minlong Peng
  • Zhuoren Jiang
  • Changlong Sun
  • Qi Zhang

In a customer service system, dialogue summarization can boost service efficiency by automatically creating summaries for long spoken dialogues in which customers and agents try to address issues about specific topics. In this work, we focus on topic-oriented dialogue summarization, which generates highly abstractive summaries that preserve the main ideas from dialogues. In spoken dialogues, abundant dialogue noise and common semantics could obscure the underlying informative content, making the general topic modeling approaches difficult to apply. In addition, for customer service, role-specific information matters and is an indispensable part of a summary. To effectively perform topic modeling on dialogues and capture multi-role information, in this work we propose a novel topic-augmented two-stage dialogue summarizer (TDS) jointly with a saliency-aware neural topic model (SATM) for topic-oriented summarization of customer service dialogues. Comprehensive studies on a real-world Chinese customer service dataset demonstrated the superiority of our method against several strong baselines.

AAAI Conference 2021 Conference Paper

Tripartite Collaborative Filtering with Observability and Selection for Debiasing Rating Estimation on Missing-Not-at-Random Data

  • Qi Zhang
  • Longbing Cao
  • Chongyang Shi
  • Liang Hu

Most collaborative filtering (CF) models estimate missing ratings with an implicit assumption that the ratings are missingat-random, which may cause the biased rating estimation and degraded performance since recent deep exploration shows that ratings may likely be missing-not-at-random (MNAR). To debias MNAR rating estimation, we introduce item observability and user selection to depict the generation of MNAR ratings and propose a tripartite CF (TCF) framework to jointly model the triple aspects of rating generation: item observability, user selection, and ratings, and to estimate the MNAR ratings. An item observability variable is introduced to a complete observability model to infer whether an item is observable to a user. TCF also conducts a complete rating model for rating generation and utilizes a user selection model dependent on the item observability and rating values to model user selection of the observable items. We further elaborately instantiate TCF as a Tripartite Probabilistic Matrix Factorization model (TPMF) by leveraging the probabilistic matrix factorization. Besides, TPMF introduces multifaceted dependency between user selection and ratings to model the influence of user selection on ratings. Extensive experiments on synthetic and real-world datasets show that modeling item observability and user selection effectively debias MNAR rating estimation, and TPMF outperforms the state-of-the-art methods in estimating the MNAR ratings.

IJCAI Conference 2021 Conference Paper

UNBERT: User-News Matching BERT for News Recommendation

  • Qi Zhang
  • Jingjie Li
  • Qinglin Jia
  • Chuyuan Wang
  • Jieming Zhu
  • Zhaowei Wang
  • Xiuqiang He

Nowadays, news recommendation has become a popular channel for users to access news of their interests. How to represent rich textual contents of news and precisely match users' interests and candidate news lies in the core of news recommendation. However, existing recommendation methods merely learn textual representations from in-domain news data, which limits their generalization ability to new news that are common in cold-start scenarios. Meanwhile, many of these methods represent each user by aggregating the historically browsed news into a single vector and then compute the matching score with the candidate news vector, which may lose the low-level matching signals. In this paper, we explore the use of the successful BERT pre-training technique in NLP for news recommendation and propose a BERT-based user-news matching model, called UNBERT. In contrast to existing research, our UNBERT model not only leverages the pre-trained model with rich language knowledge to enhance textual representation, but also captures multi-grained user-news matching signals at both word-level and news-level. Extensive experiments on the Microsoft News Dataset (MIND) demonstrate that our approach constantly outperforms the state-of-the-art methods.

AAAI Conference 2021 Conference Paper

Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-Encoders

  • Yicheng Zou
  • Jun Lin
  • Lujun Zhao
  • Yangyang Kang
  • Zhuoren Jiang
  • Changlong Sun
  • Qi Zhang
  • Xuanjing Huang

Automatic chat summarization can help people quickly grasp important information from numerous chat messages. Unlike conventional documents, chat logs usually have fragmented and evolving topics. In addition, these logs contain a quantity of elliptical and interrogative sentences, which make the chat summarization highly context dependent. In this work, we propose a novel unsupervised framework called RankAE to perform chat summarization without employing manually labeled data. RankAE consists of a topic-oriented ranking strategy that selects topic utterances according to centrality and diversity simultaneously, as well as a denoising auto-encoder that is carefully designed to generate succinct but contextinformative summaries based on the selected utterances. To evaluate the proposed method, we collect a large-scale dataset of chat logs from a customer service environment and build an annotated set only for model evaluation. Experimental results show that RankAE significantly outperforms other unsupervised methods and is able to generate high-quality summaries in terms of relevance and topic coverage.

AAAI Conference 2020 Conference Paper

3D Crowd Counting via Multi-View Fusion with 3D Gaussian Kernels

  • Qi Zhang
  • Antoni B. Chan

Crowd counting has been studied for decades and a lot of works have achieved good performance, especially the DNNs-based density map estimation methods. Most existing crowd counting works focus on single-view counting, while few works have studied multi-view counting for large and wide scenes, where multiple cameras are used. Recently, an end-to-end multi-view crowd counting method called multiview multi-scale (MVMS) has been proposed, which fuses multiple camera views using a CNN to predict a 2D scenelevel density map on the ground-plane. Unlike MVMS, we propose to solve the multi-view crowd counting task through 3D feature fusion with 3D scene-level density maps, instead of the 2D ground-plane ones. Compared to 2D fusion, the 3D fusion extracts more information of the people along zdimension (height), which helps to solve the scale variations across multiple views. The 3D density maps still preserve the 2D density maps property that the sum is the count, while also providing 3D information about the crowd density. We also explore the projection consistency among the 3D prediction and the ground-truth in the 2D views to further enhance the counting performance. The proposed method is tested on 3 multi-view counting datasets and achieves better or comparable counting performance to the state-of-the-art.

AAAI Conference 2020 Conference Paper

Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means Features

  • Tao Gui
  • Lizhi Qing
  • Qi Zhang
  • Jiacheng Ye
  • Hang Yan
  • Zichu Fei
  • Xuanjing Huang

Multi-task learning (MTL) has received considerable attention, and numerous deep learning applications benefit from MTL with multiple objectives. However, constructing multiple related tasks is difficult, and sometimes only a single task is available for training in a dataset. To tackle this problem, we explored the idea of using unsupervised clustering to construct a variety of auxiliary tasks from unlabeled data or existing labeled data. We found that some of these newly constructed tasks could exhibit semantic meanings corresponding to certain human-specific attributes, but some were non-ideal. In order to effectively reduce the impact of non-ideal auxiliary tasks on the main task, we further proposed a novel meta-learning-based multi-task learning approach, which trained the shared hidden layers on auxiliary tasks, while the meta-optimization objective was to minimize the loss on the main task, ensuring that the optimizing direction led to an improvement on the main task. Experimental results across five image datasets demonstrated that the proposed method significantly outperformed existing single task learning, semi-supervised learning, and some data augmentation methods, including an improvement of more than 9% on the Omniglot dataset.

IJCAI Conference 2020 Conference Paper

Leveraging Document-Level Label Consistency for Named Entity Recognition

  • Tao Gui
  • Jiacheng Ye
  • Qi Zhang
  • Yaqian zhou
  • Yeyun Gong
  • Xuanjing Huang

Document-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types. Previous work focused on better context representations and used the CRF for label decoding. However, CRF-based methods are inadequate for modeling document-level label consistency. This work introduces a novel two-stage label refinement approach to handle document-level label consistency, where a key-value memory network is first used to record draft labels predicted by the base model, and then a multi-channel Transformer makes refinements on these draft predictions based on the explicit co-occurrence relationship derived from the memory network. In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing the incorrect refinement of correct draft labels. The experimental results on three named entity recognition benchmarks demonstrated that the proposed method significantly outperformed the state-of-the-art methods.

AAAI Conference 2020 Conference Paper

Modeling Probabilistic Commitments for Maintenance Is Inherently Harder than for Achievement

  • Qi Zhang
  • Edmund Durfee
  • Satinder Singh

Most research on probabilistic commitments focuses on commitments to achieve enabling preconditions for other agents. Our work reveals that probabilistic commitments to instead maintain preconditions for others are surprisingly harder to use well than their achievement counterparts, despite strong semantic similarities. We isolate the key difference as being not in how the commitment provider is constrained, but rather in how the commitment recipient can locally use the commitment specification to approximately model the provider’s effects on the preconditions of interest. Our theoretic analyses show that we can more tightly bound the potential suboptimality due to approximate modeling for achievement than for maintenance commitments. We empirically evaluate alternative approximate modeling strategies, confirming that probabilistic maintenance commitments are qualitatively more challenging for the recipient to model well, and indicating the need for more detailed specifications that can sacrifice some of the agents’ autonomy.

AAAI Conference 2020 Conference Paper

Neural Question Generation with Answer Pivot

  • Bingning Wang
  • Xiaochuan Wang
  • Ting Tao
  • Qi Zhang
  • Jingfang Xu

Neural question generation (NQG) is the task of generating questions from the given context with deep neural networks. Previous answer-aware NQG methods suffer from the problem that the generated answers are focusing on entity and most of the questions are trivial to be answered. The answeragnostic NQG methods reduce the bias towards named entities and increasing the model’s degrees of freedom, but sometimes result in generating unanswerable questions which are not valuable for the subsequent machine reading comprehension system. In this paper, we treat the answers as the hidden pivot for question generation and combine the question generation and answer selection process in a joint model. We achieve the state-of-the-art result on the SQuAD dataset according to automatic metric and human evaluation.

AAAI Conference 2020 Conference Paper

ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion

  • Bingning Wang
  • Ting Yao
  • Qi Zhang
  • Jingfang Xu
  • Xiaochuan Wang

This paper presents the ReCO, a human-curated Chinese Reading Comprehension dataset on Opinion. The questions in ReCO are opinion based queries issued to commercial search engine. The passages are provided by the crowdworkers who extract the support snippet from the retrieved documents. Finally, an abstractive yes/no/uncertain answer was given by the crowdworkers. The release of ReCO consists of 300k questions that to our knowledge is the largest in Chinese reading comprehension. A prominent characteristic of ReCO is that in addition to the original context paragraph, we also provided the support evidence that could be directly used to answer the question. Quality analysis demonstrates the challenge of ReCO that it requires various types of reasoning skills such as causal inference, logical reasoning, etc. Current QA models that perform very well on many question answering problems, such as BERT (Devlin et al. 2018), only achieves 77% accuracy on this dataset, a large margin behind humans nearly 92% performance, indicating ReCO present a good challenge for machine reading comprehension. The codes, dataset and leaderboard will be freely available at https: //github. com/benywon/ReCO.

AAAI Conference 2020 Conference Paper

Rethinking Generalization of Neural Models: A Named Entity Recognition Case Study

  • Jinlan Fu
  • Pengfei Liu
  • Qi Zhang

While neural network-based models have achieved impressive performance on a large body of NLP tasks, the generalization behavior of different models remains poorly understood: Does this excellent performance imply a perfect generalization model, or are there still some limitations? In this paper, we take the NER task as a testbed to analyze the generalization behavior of existing models from different perspectives and characterize the differences of their generalization abilities through the lens of our proposed measures, which guides us to better design models and training methods. Experiments with in-depth analyses diagnose the bottleneck of existing neural NER models in terms of breakdown performance analysis, annotation errors, dataset bias, and category relationships, which suggest directions for improvement. We have released the datasets: (ReCoNLL, PLONER) for the future research at our project page: http: //pfliu. com/InterpretNER/.

JAAMAS Journal 2020 Journal Article

Semantics and algorithms for trustworthy commitment achievement under model uncertainty

  • Qi Zhang
  • Edmund H. Durfee
  • Satinder Singh

Abstract We focus on how an agent can exercise autonomy while still dependably fulfilling commitments it has made to another, despite uncertainty about outcomes of its actions and how its own objectives might evolve. Our formal semantics treats a probabilistic commitment as constraints on the actions an autonomous agent can take, rather than as promises about states of the environment it will achieve. We have developed a family of commitment-constrained (iterative) lookahead algorithms that provably respect the semantics, and that support different tradeoffs between computation and plan quality. Our empirical results confirm that our algorithms’ ability to balance (selfish) autonomy and (unselfish) dependability outperforms optimizing either alone, that our algorithms can effectively handle uncertainty about both what actions do and which states are rewarding, and that our algorithms can solve more computationally-demanding problems through judicious parameter choices for how far our algorithms should lookahead and how often they should iterate.

AAAI Conference 2020 Conference Paper

Spatio-Temporal Graph Structure Learning for Traffic Forecasting

  • Qi Zhang
  • Jianlong Chang
  • Gaofeng Meng
  • Shiming Xiang
  • Chunhong Pan

As an indispensable part in Intelligent Traffic System (ITS), the task of traffic forecasting inherently subjects to the following three challenging aspects. First, traffic data are physically associated with road networks, and thus should be formatted as traffic graphs rather than regular grid-like tensors. Second, traffic data render strong spatial dependence, which implies that the nodes in the traffic graphs usually have complex and dynamic relationships between each other. Third, traffic data demonstrate strong temporal dependence, which is crucial for traffic time series modeling. To address these issues, we propose a novel framework named Structure Learning Convolution (SLC) that enables to extend the traditional convolutional neural network (CNN) to graph domains and learn the graph structure for traffic forecasting. Technically, SLC explicitly models the structure information into the convolutional operation. Under this framework, various non-Euclidean CNN methods can be considered as particular instances of our formulation, yielding a flexible mechanism for learning on the graph. Along this technical line, two SLC modules are proposed to capture the global and local structures respectively and they are integrated to construct an endto-end network for traffic forecasting. Additionally, in this process, Pseudo three Dimensional convolution (P3D) networks are combined with SLC to capture the temporal dependencies in traffic data. Extensively comparative experiments on six real-world datasets demonstrate our proposed approach significantly outperforms the state-of-the-art ones.

NeurIPS Conference 2020 Conference Paper

Spectral Temporal Graph Neural Network for Multivariate Time-series Forecasting

  • Defu Cao
  • Yujing Wang
  • Juanyong Duan
  • Ce Zhang
  • Xia Zhu
  • Congrui Huang
  • Yunhai Tong
  • Bixiong Xu

Multivariate time-series forecasting plays a crucial role in many real-world applications. It is a challenging problem as one needs to consider both intra-series temporal correlations and inter-series correlations simultaneously. Recently, there have been multiple works trying to capture both correlations, but most, if not all of them only capture temporal correlations in the time domain and resort to pre-defined priors as inter-series relationships. In this paper, we propose Spectral Temporal Graph Neural Network (StemGNN) to further improve the accuracy of multivariate time-series forecasting. StemGNN captures inter-series correlations and temporal dependencies jointly in the spectral domain. It combines Graph Fourier Transform (GFT) which models inter-series correlations and Discrete Fourier Transform (DFT) which models temporal dependencies in an end-to-end framework. After passing through GFT and DFT, the spectral representations hold clear patterns and can be predicted effectively by convolution and sequential learning modules. Moreover, StemGNN learns inter-series correlations automatically from the data without using pre-defined priors. We conduct extensive experiments on ten real-world datasets to demonstrate the effectiveness of StemGNN.

AAAI Conference 2020 Conference Paper

Storytelling from an Image Stream Using Scene Graphs

  • Ruize Wang
  • Zhongyu Wei
  • Piji Li
  • Qi Zhang
  • Xuanjing Huang

Visual storytelling aims at generating a story from an image stream. Most existing methods tend to represent images directly with the extracted high-level features, which is not intuitive and difficult to interpret. We argue that translating each image into a graph-based semantic representation, i. e. , scene graph, which explicitly encodes the objects and relationships detected within image, would benefit representing and describing images. To this end, we propose a novel graph-based architecture for visual storytelling by modeling the two-level relationships on scene graphs. In particular, on the within-image level, we employ a Graph Convolution Network (GCN) to enrich local fine-grained region representations of objects on scene graphs. To further model the interaction among images, on the cross-images level, a Temporal Convolution Network (TCN) is utilized to refine the region representations along the temporal dimension. Then the relation-aware representations are fed into the Gated Recurrent Unit (GRU) with attention mechanism for story generation. Experiments are conducted on the public visual storytelling dataset. Automatic and human evaluation results indicate that our method achieves state-of-the-art.

NeurIPS Conference 2020 Conference Paper

Succinct and Robust Multi-Agent Communication With Temporal Message Control

  • Sai Qian Zhang
  • Qi Zhang
  • Jieyu Lin

Recent studies have shown that introducing communication between agents can significantly improve overall performance in cooperative Multi-agent reinforcement learning (MARL). However, existing communication schemes often require agents to exchange an excessive number of messages at run-time under a reliable communication channel, which hinders its practicality in many real-world situations. In this paper, we present \textit{Temporal Message Control} (TMC), a simple yet effective approach for achieving succinct and robust communication in MARL. TMC applies a temporal smoothing technique to drastically reduce the amount of information exchanged between agents. Experiments show that TMC can significantly reduce inter-agent communication overhead without impacting accuracy. Furthermore, TMC demonstrates much better robustness against transmission loss than existing approaches in lossy networking environments.

IJCAI Conference 2019 Conference Paper

CNN-Based Chinese NER with Lexicon Rethinking

  • Tao Gui
  • Ruotian Ma
  • Qi Zhang
  • Lujun Zhao
  • Yu-Gang Jiang
  • Xuanjing Huang

Character-level Chinese named entity recognition (NER) that applies long short-term memory (LSTM) to incorporate lexicons has achieved great success. However, this method fails to fully exploit GPU parallelism and candidate lexicons can conflict. In this work, we propose a faster alternative to Chinese NER: a convolutional neural network (CNN)-based method that incorporates lexicons using a rethinking mechanism. The proposed method can model all the characters and potential words that match the sentence in parallel. In addition, the rethinking mechanism can address the word conflict by feeding back the high-level features to refine the networks. Experimental results on four datasets show that the proposed method can achieve better performance than both word-level and character-level baseline methods. In addition, the proposed method performs up to 3. 21 times faster than state-of-the-art methods, while realizing better performance.

AAAI Conference 2019 Conference Paper

Cooperative Multimodal Approach to Depression Detection in Twitter

  • Tao Gui
  • Liang Zhu
  • Qi Zhang
  • Minlong Peng
  • Xu Zhou
  • Keyu Ding
  • Zhigang Chen

The advent of social media has presented a promising new opportunity for the early detection of depression. To do so effectively, there are two challenges to overcome. The first is that textual and visual information must be jointly considered to make accurate inferences about depression. The second challenge is that due to the variety of content types posted by users, it is difficult to extract many of the relevant indicator texts and images. In this work, we propose the use of a novel cooperative multi-agent model to address these challenges. From the historical posts of users, the proposed method can automatically select related indicator texts and images. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods by a large margin (over 30% error reduction). In several experiments and examples, we also verify that the selected posts can successfully indicate user depression, and our model can obtained a robust performance in realistic scenarios.

NeurIPS Conference 2019 Conference Paper

Efficient Communication in Multi-Agent Reinforcement Learning via Variance Based Control

  • Sai Qian Zhang
  • Qi Zhang
  • Jieyu Lin

Multi-agent reinforcement learning (MARL) has recently received considerable attention due to its applicability to a wide range of real-world applications. However, achieving efficient communication among agents has always been an overarching problem in MARL. In this work, we propose Variance Based Control (VBC), a simple yet efficient technique to improve communication efficiency in MARL. By limiting the variance of the exchanged messages between agents during the training phase, the noisy component in the messages can be eliminated effectively, while the useful part can be preserved and utilized by the agents for better performance. Our evaluation using multiple MARL benchmarks indicates that our method achieves $2-10\times$ lower in communication overhead than state-of-the-art MARL algorithms, while allowing agents to achieve better overall performance.

IJCAI Conference 2019 Conference Paper

GSTNet: Global Spatial-Temporal Network for Traffic Flow Prediction

  • Shen Fang
  • Qi Zhang
  • Gaofeng Meng
  • Shiming Xiang
  • Chunhong Pan

Predicting traffic flow on traffic networks is a very challenging task, due to the complicated and dynamic spatial-temporal dependencies between different nodes on the network. The traffic flow renders two types of temporal dependencies, including short-term neighboring and long-term periodic dependencies. What's more, the spatial correlations over different nodes are both local and non-local. To capture the global dynamic spatial-temporal correlations, we propose a Global Spatial-Temporal Network (GSTNet), which consists of several layers of spatial-temporal blocks. Each block contains a multi-resolution temporal module and a global correlated spatial module in sequence, which can simultaneously extract the dynamic temporal dependencies and the global spatial correlations. Extensive experiments on the real world datasets verify the effectiveness and superiority of the proposed method on both the public transportation network and the road network.

IJCAI Conference 2019 Conference Paper

Learning Task-Specific Representation for Novel Words in Sequence Labeling

  • Minlong Peng
  • Qi Zhang
  • Xiaoyu Xing
  • Tao Gui
  • Jinlan Fu
  • Xuanjing Huang

Word representation is a key component in neural-network-based sequence labeling systems. However, representations of unseen or rare words trained on the end task are usually poor for appreciable performance. This is commonly referred to as the out-of-vocabulary (OOV) problem. In this work, we address the OOV problem in sequence labeling using only training data of the task. To this end, we propose a novel method to predict representations for OOV words from their surface-forms (e. g. , character sequence) and contexts. The method is specifically designed to avoid the error propagation problem suffered by existing approaches in the same paradigm. To evaluate its effectiveness, we performed extensive empirical studies on four part-of-speech tagging (POS) tasks and four named entity recognition (NER) tasks. Experimental results show that the proposed method can achieve better or competitive performance on the OOV problem compared with existing state-of-the-art methods.

AAAI Conference 2019 Conference Paper

Learning to Communicate and Solve Visual Blocks-World Tasks

  • Qi Zhang
  • Richard Lewis
  • Satinder Singh
  • Edmund Durfee

We study emergent communication between speaker and listener recurrent neural-network agents that are tasked to cooperatively construct a blocks-world target image sampled from a generative grammar of blocks configurations. The speaker receives the target image and learns to emit a sequence of discrete symbols from a fixed vocabulary. The listener learns to construct a blocks-world image by choosing block placement actions as a function of the speaker’s full utterance and the image of the ongoing construction. Our contributions are (a) the introduction of a task domain for studying emergent communication that is both challenging and affords useful analyses of the emergent protocols; (b) an empirical comparison of the interpolation and extrapolation performance of training via supervised, (contextual) Bandit, and reinforcement learning; and (c) evidence for the emergence of interesting linguistic properties in the RL agent protocol that are distinct from the other two.

AAAI Conference 2019 Conference Paper

Long Short-Term Memory with Dynamic Skip Connections

  • Tao Gui
  • Qi Zhang
  • Lujun Zhao
  • Yaosong Lin
  • Minlong Peng
  • Jingjing Gong
  • Xuanjing Huang

In recent years, long short-term memory (LSTM) has been successfully used to model sequential data of variable length. However, LSTM can still experience difficulty in capturing long-term dependencies. In this work, we tried to alleviate this problem by introducing a dynamic skip connection, which can learn to directly connect two dependent words. Since there is no dependency information in the training data, we propose a novel reinforcement learning-based method to model the dependency relationship and connect dependent words. The proposed model computes the recurrent transition functions based on the skip connections, which provides a dynamic skipping advantage over RNNs that always tackle entire sentences sequentially. Our experimental results on three natural language processing tasks demonstrate that the proposed method can achieve better performance than existing methods. In the number prediction experiment, the proposed model outperformed LSTM with respect to accuracy by nearly 20%.

JBHI Journal 2019 Journal Article

MR Image Super-Resolution via Wide Residual Networks With Fixed Skip Connection

  • Jun Shi
  • Zheng Li
  • Shihui Ying
  • Chaofeng Wang
  • Qingping Liu
  • Qi Zhang
  • Pingkun Yan

Spatial resolution is a critical imaging parameter in magnetic resonance imaging. The image super-resolution (SR) is an effective and cost efficient alternative technique to improve the spatial resolution of MR images. Over the past several years, the convolutional neural networks (CNN)-based SR methods have achieved state-of-the-art performance. However, CNNs with very deep network structures usually suffer from the problems of degradation and diminishing feature reuse, which add difficulty to network training and degenerate the transmission capability of details for SR. To address these problems, in this work, a progressive wide residual network with a fixed skip connection (named FSCWRN) based SR algorithm is proposed to reconstruct MR images, which combines the global residual learning and the shallow network based local residual learning. The strategy of progressive wide networks is adopted to replace deeper networks, which can partially relax the above-mentioned problems, while a fixed skip connection helps provide rich local details at high frequencies from a fixed shallow layer network to subsequent networks. The experimental results on one simulated MR image database and three real MR image databases show the effectiveness of the proposed FSCWRN SR algorithm, which achieves improved reconstruction performance compared with other algorithms.

EAAI Journal 2019 Journal Article

Online probabilistic goal recognition and its application in dynamic shortest-path local network interdiction

  • Kai Xu
  • Yunxiu Zeng
  • Qi Zhang
  • Quanjun Yin
  • Lin Sun
  • Kaiming Xiao

Goal recognition is the task of inferring an agent’s goals given some or all of the agent’s observed actions. However, few research focuses on how to improve the usage effectiveness of knowledge produced by a goal recognition system. In this work, we propose a probabilistic goal recognition approach tailored to a dynamic shortest-path network interdiction problem. Apart from inferring a probabilistic distribution over the possible goals of an agent, our work has another four key novelties: (i) a dynamic shortest-path local network interdiction model that allocates resources locally per step using goal recognition information; (ii) two behavior modeling approaches, including a data-driven learning method based on Inverse Reinforcement Learning as well as a heuristic method taking advantage of the network information, to help solve both the data-intensive and no available data situations; (iii) a heuristic named Subjective Confidence that uses variance in particle system for flexible resource allocation adjustment. The empirical test results show the effectiveness of our goal recognition method, and also verify the practical implications of these methods in solving scalable multi-terminus network interdiction problem.

AAAI Conference 2019 Conference Paper

Trainable Undersampling for Class-Imbalance Learning

  • Minlong Peng
  • Qi Zhang
  • Xiaoyu Xing
  • Tao Gui
  • Xuanjing Huang
  • Yu-Gang Jiang
  • Keyu Ding
  • Zhigang Chen

Undersampling has been widely used in the class-imbalance learning area. The main deficiency of most existing undersampling methods is that their data sampling strategies are heuristic-based and independent of the used classifier and evaluation metric. Thus, they may discard informative instances for the classifier during the data sampling. In this work, we propose a meta-learning method built on the undersampling to address this issue. The key idea of this method is to parametrize the data sampler and train it to optimize the classification performance over the evaluation metric. We solve the non-differentiable optimization problem for training the data sampler via reinforcement learning. By incorporating evaluation metric optimization into the data sampling process, the proposed method can learn which instance should be discarded for the given classifier and evaluation metric. In addition, as a data level operation, this method can be easily applied to arbitrary evaluation metric and classifier, including non-parametric ones (e. g. , C4. 5 and KNN). Experimental results on both synthetic and realistic datasets demonstrate the effectiveness of the proposed method.

AAAI Conference 2018 Conference Paper

Adaptive Co-attention Network for Named Entity Recognition in Tweets

  • Qi Zhang
  • Jinlan Fu
  • Xiaoyu Liu
  • Xuanjing Huang

In this study, we investigate the problem of named entity recognition for tweets. Named entity recognition is an important task in natural language processing and has been carefully studied in recent decades. Previous named entity recognition methods usually only used the textual content when processing tweets. However, many tweets contain not only textual content, but also images. Such visual information is also valuable in the name entity recognition task. To make full use of textual and visual information, this paper proposes a novel method to process tweets that contain multimodal information. We extend a bi-directional long short term memory network with conditional random fields and an adaptive co-attention network to achieve this task. To evaluate the proposed methods, we constructed a large scale labeled dataset that contained multimodal tweets. Experimental results demonstrated that the proposed method could achieve a better performance than the previous methods in most cases.

JBHI Journal 2018 Journal Article

Multimodal Neuroimaging Feature Learning With Multimodal Stacked Deep Polynomial Networks for Diagnosis of Alzheimer's Disease

  • Jun Shi
  • Xiao Zheng
  • Yan Li
  • Qi Zhang
  • Shihui Ying

The accurate diagnosis of Alzheimer's disease (AD) and its early stage, i. e. , mild cognitive impairment, is essential for timely treatment and possible delay of AD. Fusion of multimodal neuroimaging data, such as magnetic resonance imaging (MRI) and positron emission tomography (PET), has shown its effectiveness for AD diagnosis. The deep polynomial networks (DPN) is a recently proposed deep learning algorithm, which performs well on both large-scale and small-size datasets. In this study, a multimodal stacked DPN (MM-SDPN) algorithm, which MM-SDPN consists of two-stage SDPNs, is proposed to fuse and learn feature representation from multimodal neuroimaging data for AD diagnosis. Specifically speaking, two SDPNs are first used to learn high-level features of MRI and PET, respectively, which are then fed to another SDPN to fuse multimodal neuroimaging information. The proposed MM-SDPN algorithm is applied to the ADNI dataset to conduct both binary classification and multiclass classification tasks. Experimental results indicate that MM-SDPN is superior over the state-of-the-art multimodal feature-learning-based algorithms for AD diagnosis.

AAAI Conference 2018 Conference Paper

Neural Networks Incorporating Dictionaries for Chinese Word Segmentation

  • Qi Zhang
  • Xiaoyu Liu
  • Jinlan Fu

In recent years, deep neural networks have achieved significant success in Chinese word segmentation and many other natural language processing tasks. Most of these algorithms are end-to-end trainable systems and can effectively process and learn from large scale labeled datasets. However, these methods typically lack the capability of processing rare words and data whose domains are different from training data. Previous statistical methods have demonstrated that human knowledge can provide valuable information for handling rare cases and domain shifting problems. In this paper, we seek to address the problem of incorporating dictionaries into neural networks for the Chinese word segmentation task. Two different methods that extend the bi-directional long short-term memory neural network are proposed to perform the task. To evaluate the performance of the proposed methods, state-of-the-art supervised models based methods and domain adaptation approaches are compared with our methods on nine datasets from different domains. The experimental results demonstrate that the proposed methods can achieve better performance than other state-of-the-art neural network methods and domain adaptation approaches in most cases.

IJCAI Conference 2018 Conference Paper

Neural Networks Incorporating Unlabeled and Partially-labeled Data for Cross-domain Chinese Word Segmentation

  • Lujun Zhao
  • Qi Zhang
  • Peng Wang
  • Xiaoyu Liu

Most existing Chinese word segmentation (CWS) methods are usually supervised. Hence, large-scale annotated domain-specific datasets are needed for training. In this paper, we seek to address the problem of CWS for the resource-poor domains that lack annotated data. A novel neural network model is proposed to incorporate unlabeled and partially-labeled data. To make use of unlabeled data, we combine a bidirectional LSTM segmentation model with two character-level language models using a gate mechanism. These language models can capture co-occurrence information. To make use of partially-labeled data, we modify the original cross entropy loss function of RNN. Experimental results demonstrate that the method performs well on CWS tasks in a series of domains.

NeurIPS Conference 2017 Conference Paper

A Learning Error Analysis for Structured Prediction with Approximate Inference

  • Yuanbin Wu
  • Man Lan
  • Shiliang Sun
  • Qi Zhang
  • Xuanjing Huang

In this work, we try to understand the differences between exact and approximate inference algorithms in structured prediction. We compare the estimation and approximation error of both underestimate and overestimate models. The result shows that, from the perspective of learning errors, performances of approximate inference could be as good as exact inference. The error analyses also suggest a new margin for existing learning algorithms. Empirical evaluations on text classification, sequential labelling and dependency parsing witness the success of approximate inference and the benefit of the proposed margin.

IJCAI Conference 2017 Conference Paper

Hashtag Recommendation for Multimodal Microblog Using Co-Attention Network

  • Qi Zhang
  • Jiawen Wang
  • Haoran Huang
  • Xuanjing Huang
  • Yeyun Gong

In microblogging services, authors can use hashtags to mark keywords or topics. Many live social media applications (e. g. , microblog retrieval, classification) can gain great benefits from these manually labeled tags. However, only a small portion of microblogs contain hashtags inputed by users. Moreover, many microblog posts contain not only textual content but also images. These visual resources also provide valuable information that may not be included in the textual content. So that it can also help to recommend hashtags more accurately. Motivated by the successful use of the attention mechanism, we propose a co-attention network incorporating textual and visual information to recommend hashtags for multimodal tweets. Experimental result on the data collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods using textual information only.

JBHI Journal 2017 Journal Article

Histopathological Image Classification With Color Pattern Random Binary Hashing-Based PCANet and Matrix-Form Classifier

  • Jun Shi
  • Jinjie Wu
  • Yan Li
  • Qi Zhang
  • Shihui Ying

The computer-aided diagnosis for histopathological images has attracted considerable attention. Principal component analysis network (PCANet) is a novel deep learning algorithm for feature learning with the simple network architecture and parameters. In this study, a color pattern random binary hashing-based PCANet (C-RBH-PCANet) algorithm is proposed to learn an effective feature representation from color histopathological images. The color norm pattern and angular pattern are extracted from the principal component images of R, G, and B color channels after cascaded PCA networks. The random binary encoding is then performed on both color norm pattern images and angular pattern images to generate multiple binary images. Moreover, we rearrange the pooled local histogram features by spatial pyramid pooling to a matrix-form for reducing the dimension of feature and preserving spatial information. Therefore, a C-RBH-PCANet and matrix-form classifier-based feature learning and classification framework is proposed for diagnosis of color histopathological images. The experimental results on three color histopathological image datasets show that the proposed C-RBH-PCANet algorithm is superior to the original PCANet and other conventional unsupervised deep learning algorithms, while the best performance is achieved by the proposed feature learning and classification framework that combines C-RBH-PCANet and matrix-form classifier.

ICRA Conference 2017 Conference Paper

In vivo tracking and measurement of pollen tube vesicle motion

  • Chengzhi Hu
  • Qi Zhang
  • Tobias Meyer
  • Hannes Vogler
  • Jan T. Burri
  • Naveen Shamsudhin
  • Ueli Grossniklaus
  • Bradley J. Nelson

Particle tracking has emerged as a powerful tool for investigating the swarm control of microrobots and the dynamic biological processes in the life sciences. In seed plants, pollen tubes, a part of the male gametophyte, are excellent models for understanding plant growth and cellular behavior, because vesicle motion within pollen tubes reveals important information about vesicle function and interactions. Conventional vesicle tracking is based on spatiotemporal image analysis, which requires high-quality images and vesicles with constant velocity. For in vivo tracking, vesicles may disappear in some frames, and image sequences may have spatial and temporal distortions, which hamper vesicle tracking for broader applications. In this paper, we studied intracellular motion during pollen tube growth with an optical flow method. Streaming images from confocal and optical microscopes were recorded to study the intracellular motion of vesicles of different size. Local motion for each vesicle was detected using a local displacement vector field. The displacement from two adjacent frames was then calculated. The flow field shows information such as the dynamics of vesicle secretion, endocytosis, exocytosis, and cytoskeletal stability. Vesicles from different regions inside the tube were tracked simultaneously with a Kanade-Lucas-Tomasi (KLT) feature matching algorithm. The spatial and temporal characteristics of intracellular vesicles were evaluated. The proposed methods can be of great use for studying the dynamics of fluorescently tagged particles in biological systems.

IJCAI Conference 2017 Conference Paper

Mention Recommendation for Twitter with End-to-end Memory Network

  • Haoran Huang
  • Qi Zhang
  • Xuanjing Huang

In this study, we investigated the problem of recommending usernames when people attempt to use the ``@'' sign to mention other people in twitter-like social media. With the extremely rapid development of social networking services, this problem has received considerable attention in recent years. Previous methods have studied the problem from different aspects. Because most of Twitter-like microblogging services limit the length of posts, statistical learning methods may be affected by the problems of word sparseness and synonyms. Although recent progress in neural word embedding methods have advanced the state-of-the-art in many natural language processing tasks, the benefits of word embedding have not been taken into consideration for this problem. In this work, we proposed a novel end-to-end memory network architecture to perform this task. We incorporated the interests of users with external memory. A hierarchical attention mechanism was also applied to better consider the interests of users. The experimental results on a dataset we collected from Twitter demonstrated that the proposed method could outperform state-of-the-art approaches.

EAAI Journal 2016 Journal Article

A novel distance function of D numbers and its application in product engineering

  • Meizhu Li
  • Yong Hu
  • Qi Zhang
  • Yong Deng

The Dempster–Shafer theory is widely applied in uncertainty modelling and knowledge reasoning due to its ability of expressing uncertain information. A distance between two basic probability assignments (BPAs) presents a measure of performance for identification of algorithms based on the evidential theory of Dempster–Shafer. However, some conditions lead to limitations in practical application for the Dempster–Shafer theory, such as exclusiveness hypothesis and completeness constraint. To overcome these shortcomings, a novel theory called D numbers theory is proposed. A distance function of D numbers is proposed to measure the distance between two D numbers. The distance function of D numbers is a generalization of distance between two BPAs, which inherits the advantage of Dempster–Shafer theory and strengthens the capability of uncertainty modeling. An illustrative case about product engineering is provided to demonstrate the effectiveness of the proposed function.

IJCAI Conference 2016 Conference Paper

Commitment Semantics for Sequential Decision Making under Reward Uncertainty

  • Qi Zhang
  • Edmund Durfee
  • Satinder Singh
  • Anna Chen
  • Stefan Witwicki

Cooperating agents can make commitments to help each other, but commitments might have to be probabilistic when actions have stochastic outcomes. We consider the additional complication in cases where an agent might prefer to change its policy as it learns more about its reward function from experience. How should such an agent be allowed to change its policy while still faithfully pursuing its commitment in a principled decision-theoretic manner? We address this question by defining a class of Dec-POMDPs with Bayesian reward uncertainty, and by developing a novel Commitment Constrained Iterative Mean Reward algorithm that implements the semantics of faithful commitment pursuit while still permitting the agent's response to the evolving understanding of its rewards. We bound the performance of our algorithm theoretically, and evaluate empirically how it effectively balances solution quality and computation cost.

AAAI Conference 2016 Conference Paper

Discourse Relations Detection via a Mixed Generative-Discriminative Framework

  • Jifan Chen
  • Qi Zhang
  • Pengfei Liu
  • Xuanjing Huang

Word embeddings, which can better capture the fine-grained semantics of words, have proven to be useful for a variety of natural language processing tasks. However, because discourse structures describe the relationships between segments of discourse, word embeddings cannot be directly integrated to perform the task. In this paper, we introduce a mixed generative-discriminative framework, in which we use vector offsets between embeddings of words to represent the semantic relations between text segments and Fisher kernel framework to convert a variable number of vector offsets into a fixed length vector. In order to incorporate the weights of these offsets into the vector, we also propose the Weighted Fisher Vector. Experimental results on two different datasets show that the proposed method without using manually designed features can achieve better performance on recognizing the discourse level relations in most cases.

IJCAI Conference 2016 Conference Paper

Hashtag Recommendation Using Attention-Based Convolutional Neural Network

  • Yuyun Gong
  • Qi Zhang

Along with the increasing requirements, the hashtag recommendation task for microblogs has been receiving considerable attention in recent years. Various researchers have studied the problem from different aspects. However, most of these methods usually need handcrafted features. Motivated by the successful use of convolutional neural networks (CNNs) for many natural language processing tasks, in this paper, we adopt CNNs to perform the hashtag recommendation problem. To incorporate the trigger words whose effectiveness have been experimentally evaluated in several previous works, we propose a novel architecture with an attention mechanism. The results of experiments on the data collected from a real world microblogging service demonstrated that the proposed model outperforms state-of-the-art methods. By incorporating trigger words into the consideration, the relative improvement of the proposed method over the state-of-the-art method is around 9. 4% in the F1-score.

JBHI Journal 2016 Journal Article

Multimodality Neurological Data Visualization With Multi-VOI-Based DTI Fiber Dynamic Integration

  • Qi Zhang
  • Murray Alexander
  • Lawrence Ryner

Brain lesions are usually located adjacent to critical spinal structures, so it is a challenging task for neurosurgeons to precisely plan a surgical procedure without damaging healthy tissues and nerves. The advancement of medical imaging technologies produces a large amount of neurological data, which are capable of showing a wide variety of brain properties. Advanced algorithms of medical data computing and visualization are critically helpful in efficiently utilizing the acquired data for disease diagnosis and brain function and structure exploration, which is helpful for treatment planning. In this paper, we describe new algorithms and a software framework for multiple volume of interest specified diffusion tensor imaging (DTI) fiber dynamic visualization. The displayed results have been integrated with a volume rendering pipeline for multimodality neurological data exploration. A depth texture indexing algorithm is used to detect DTI fiber tracts in graphics process units (GPUs), which makes fibers to be displayed and interactively manipulated with brain data acquired from functional magnetic resonance imaging, T 1 and T 2 -weighted anatomic imaging, and angiographic imaging. The developed software platform is built on an object-oriented structure, which is transparent and extensible. It provides a comprehensive human-computer interface for data exploration and information extraction. The GPU-accelerated high-performance computing kernels have been implemented to enable our software to dynamically visualize neurological data. The developed techniques will be useful in computer-aided neurological disease diagnosis, brain structure exploration, and general cognitive neuroscience.

ICML Conference 2015 Conference Paper

\(\ell_{1, p}\)-Norm Regularization: Error Bounds and Convergence Rate Analysis of First-Order Methods

  • Zirui Zhou
  • Qi Zhang
  • Anthony Man-Cho So

Recently, \ell_1, p-regularization has been widely used to induce structured sparsity in the solutions to various optimization problems. Motivated by the desire to analyze the convergence rate of first-order methods, we show that for a large class of \ell_1, p-regularized problems, an error bound condition is satisfied when p∈[1, 2] or p=∞but fails to hold for any p∈(2, ∞). Based on this result, we show that many first-order methods enjoy an asymptotic linear rate of convergence when applied to \ell_1, p-regularized linear or logistic regression with p∈[1, 2] or p=∞. By contrast, numerical experiments suggest that for the same class of problems with p∈(2, ∞), the aforementioned methods may not converge linearly.

AAAI Conference 2015 Conference Paper

Retweet Behavior Prediction Using Hierarchical Dirichlet Process

  • Qi Zhang
  • Yeyun Gong
  • Ya Guo
  • Xuanjing Huang

The task of predicting retweet behavior is an important and essential step for various social network applications, such as business intelligence, popular event prediction, and so on. Due to the increasing requirements, in recent years, the task has attracted extensive attentions. In this work, we propose a novel method using non-parametric statistical models to combine structural, textual, and temporal information together to predict retweet behavior. To evaluate the proposed method, we collect a large number of microblogs and their corresponding social networks from a real microblog service. Experimental results on the constructed dataset demonstrate that the proposed method can achieve better performance than state-of-the-art methods. The relative improvement of the the proposed over the method using only textual information is more than 38. 5% in terms of F1-Score.

JMLR Journal 2014 Journal Article

Optimality of Graphlet Screening in High Dimensional Variable Selection

  • Jiashun Jin
  • Cun-Hui Zhang
  • Qi Zhang

Consider a linear model $Y = X \beta + \sigma z$, where $X$ has $n$ rows and $p$ columns and $z \sim N(0, I_n)$. We assume both $p$ and $n$ are large, including the case of $p \gg n$. The unknown signal vector $\beta$ is assumed to be sparse in the sense that only a small fraction of its components is nonzero. The goal is to identify such nonzero coordinates (i.e., variable selection). We are primarily interested in the regime where signals are both rare and weak so that successful variable selection is challenging but is still possible. We assume the Gram matrix $G = X'X$ is sparse in the sense that each row has relatively few large entries (diagonals of $G$ are normalized to $1$). The sparsity of $G$ naturally induces the sparsity of the so-called Graph of Strong Dependence (GOSD). The key insight is that there is an interesting interplay between the signal sparsity and graph sparsity: in a broad context, the signals decompose into many small-size components of GOSD that are disconnected to each other. We propose Graphlet Screening for variable selection. This is a two-step Screen and Clean procedure, where in the first step, we screen subgraphs of GOSD with sequential $\chi^2$-tests, and in the second step, we clean with penalized MLE. The main methodological innovation is to use GOSD to guide both the screening and cleaning processes. For any variable selection procedure $\hat{\beta}$, we measure its performance by the Hamming distance between the sign vectors of $\hat{\beta}$ and $\beta$, and assess the optimality by the minimax Hamming distance. Compared with more stringent criteria such as exact support recovery or oracle property, which demand strong signals, the Hamming distance criterion is more appropriate for weak signals since it naturally allows a small fraction of errors. We show that in a broad class of situations, Graphlet Screening achieves the optimal rate of convergence in terms of the Hamming distance. Unlike Graphlet Screening, well- known procedures such as the $L^0/L^1$-penalization methods do not utilize local graphic structure for variable selection, so they generally do not achieve the optimal rate of convergence, even in very simple settings and even if the tuning parameters are ideally set. The the presented algorithm is implemented as R-CRAN package ScreenClean and in matlab (available at stat.cmu.edu ). [abs] [ pdf ][ bib ] &copy JMLR 2014. ( edit, beta )

IJCAI Conference 2013 Conference Paper

Learning Topical Translation Model for Microblog Hashtag Suggestion

  • Zhuoye Ding
  • Xipeng Qiu
  • Qi Zhang
  • Xuanjing Huang

Hashtags can be viewed as an indication to the context of the tweet or as the core idea expressed in the tweet. They provide valuable information for many applications, such as information retrieval, opinion mining, text classification, and so on. However, only a small number of microblogs are manually tagged. To address this problem, in this work, we propose a topical translation model for microblog hashtag suggestion. We assume that the content and hashtags of the tweet are talking about the same themes but written in different languages. Under the assumption, hashtag suggestion is modeled as a translation process from content to hashtags. Moreover, in order to cover the topic of tweets, the proposed model regards the translation probability to be topic-specific. It uses topic-specific word trigger to bridge the vocabulary gap between the words in tweets and hashtags, and discovers the topics of tweets by a topic model designed for microblogs. Experimental results on the dataset crawled from real world microblogging service demonstrate that the proposed method outperforms state-of-the-art methods.

JMLR Journal 2002 Journal Article

Multiple-Instance Learning of Real-Valued Data

  • Daniel R. Dooly
  • Qi Zhang
  • Sally A. Goldman
  • Robert A. Amar

The multiple-instance learning model has received much attention recently with a primary application area being that of drug activity prediction. Most prior work on multiple-instance learning has been for concept learning, yet for drug activity prediction, the label is a real-valued affinity measurement giving the binding strength. We present extensions of k -nearest neighbors ( k -NN), Citation- k NN, and the diverse density algorithm for the real-valued setting and study their performance on Boolean and real-valued data. We also provide a method for generating chemically realistic artificial data.

NeurIPS Conference 2001 Conference Paper

EM-DD: An Improved Multiple-Instance Learning Technique

  • Qi Zhang
  • Sally Goldman

We present a new multiple-instance (MI) learning technique (EM(cid: 173) DD) that combines EM with the diverse density (DD) algorithm. EM-DD is a general-purpose MI algorithm that can be applied with boolean or real-value labels and makes real-value predictions. On the boolean Musk benchmarks, the EM-DD algorithm without any tuning significantly outperforms all previous algorithms. EM-DD is relatively insensitive to the number of relevant attributes in the data set and scales up well to large bag sizes. Furthermore, EM(cid: 173) DD provides a new framework for MI learning, in which the MI problem is converted to a single-instance setting by using EM to estimate the instance responsible for the label of the bag.

v2026.09.13