Arrow Research search

Author name cluster

Ming Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

58 papers
2 author rows

Possible papers

58

AAAI Conference 2026 Conference Paper

CogniTrust: Cognitive Memory-Driven Verifiable Supervision for Robust Hashing

  • Yiyang Gu
  • Bohan Wu
  • Yifang Qin
  • Jiaru Tang
  • Rong-Cheng Tu
  • Zhiping Xiao
  • Taian Guo
  • Junyu Luo

In this paper, we study the problem of robust multi-label hashing, where label noise hinders the learning of a reliable semantic structure from data. Many existing methods rely on heuristic sample selection or consistency-based training, but lack a unified mechanism to validate and refine supervision across structural and semantic levels. Inspired by cognitive theories of human memory, we propose a novel framework called CogniTrust that unifies verifiable supervision with a triadic memory model: a) In episodic memory, feature activations are decomposed into spatial patterns that support the assessment of structural evidence and the estimation of label reliability; b) Semantic memory keeps track of class-level prototypes from structurally attentive regions to estimate the semantic plausibility of labels; c) Reconstructive memory simulates memory recall through interpolation between images using a diffusion-based mixup process, which enriches the training signals for semantically uncertain regions. These components work together, allowing supervision to be refined through the joint consideration of spatial structure and semantic information. Extensive experiments on noisy hashing benchmarks demonstrate that CogniTrust consistently outperforms a range of state-of-the-art baselines. Our results show that cognitive memory mechanisms offer a principled basis for more reliable label denoising and robust hashing.

JBHI Journal 2026 Journal Article

IMRadar: Bidirectional Velocity Mamba for Contactless Human Behavior Sensing

  • Ming Zhang
  • Jiong Liang
  • Jiayao Li
  • Shaolin Liao
  • Henry Soekmadji
  • Chengpei Tang

In recent years, intelligent human behavior sensing based on channel state information (CSI) has garnered significant attention from researchers, serving as a pivotal application of contactless health monitoring. However, the feature extraction networks used in existing perception schemes have significant limitations in terms of global context perception, computational complexity, and only consider features in one direction. To address these issues, this article proposes a novel bidirectional velocity Mamba (BVMamba) model and constructs an intelligent behavior sensing system, named IMRadar. The system first analyzes the velocity information that better characterizes the human motion state from CSI data, and uses the BVMamba model to extract global deep behavioral features from both forward and reverse directions. The BVMamba model includes forward velocity Mamba block (FVMamba), reverse velocity Mamba block (RVMamba), and bidirectional velocity feature fusion block (FUBlock), which can comprehensively capture the dynamic characteristics of complex behaviors. Experiments have shown that IMRadar exhibits excellent recognition performance on both publicly available datasets (ARIL, Widar) and self-built dataset (IM-HAR), with accuracy rates exceeding 98% for all datasets, providing an efficient and robust solution for non-contact behavior perception technology.

AAAI Conference 2026 Conference Paper

MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning

  • Zhiheng Xi
  • Yuhui Wang
  • Yiwen Ding
  • Guanyu Li
  • Senjie Jin
  • Shichun Liu
  • Jixuan Huang
  • Dingwen Yang

Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B.

AAAI Conference 2026 Conference Paper

Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination

  • Mingqi Wu
  • Zhihao Zhang
  • Qiaole Dong
  • Zhiheng Xi
  • Jun Zhao
  • Senjie Jin
  • Xiaoran Fan
  • Yuhao Zhou

Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods.

YNIMG Journal 2026 Journal Article

System-Level Reconfiguration of the Aging Brain: Linking Dynamics, Morphology and Micro-architectures

  • Liming Fan
  • Youjun Li
  • Yutong Wu
  • Simeng An
  • Nan Yao
  • Qian Zhu
  • Yueye Zhao
  • Daqing Guo

Healthy aging involves complex neural reconfigurations across both structural and functional domains. While resting-state functional magnetic resonance imaging (rs-fMRI) has linked static functional connectivity alterations to aging, the whole-brain dynamics of functional activity and their covariance with structural changes remain poorly characterized. To address this gap, we integrated three data-driven approaches to profile functional dynamics in the aging brain and decode their association with structural atrophy. Using rs-fMRI data from 252 participants-145 young adults (22.7 ± 3.4 years) and 107 older adults (68.7 ± 6.5 years)-we made several key observations. First, normalized Shannon entropy revealed a significant reduction in spatiotemporal complexity among older individuals. Second, phase synchronization analysis of BOLD signals indicated enhanced global integration and metastability in older adults, particularly within the dorsal attention (DAN), ventral attention (VAN), and frontoparietal networks (FPN). Third, temporal asymmetry analysis demonstrated increased nonreversibility and a heightened functional hierarchy in the aging brain, again most evident in the FPN. Morphometric analyses confirmed widespread structural atrophy in older participants. Crucially, partial least squares (PLS) analysis uncovered significant covariance between morphometric patterns and dynamic functional metrics, underscoring a tight structure-dynamics coupling in aging. Furthermore, structural atrophy correlated significantly with variations in micro-architecture maps. Finally, we evaluated the behavioral relevance of these dynamics through correlations with cognitive performance. Our findings offer an integrative, multiscale perspective on neural decline in aging, emphasizing the interplay between dynamic functional reorganization and structural atrophy.

AAAI Conference 2026 Conference Paper

What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study

  • Xiaoran Fan
  • Zhichao Sun
  • Yangfan Gao
  • Jingfei Xiong
  • Hang Yan
  • Yifei Cao
  • Jiajun Sun
  • Shuo Li

Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.

AAAI Conference 2025 Conference Paper

Attention Bootstrapping for Multi-Modal Test-Time Adaptation

  • Yusheng Zhao
  • Junyu Luo
  • Xiao Luo
  • Jinsheng Huang
  • Jingyang Yuan
  • Zhiping Xiao
  • Ming Zhang

Test-time adaptation aims to adapt a well-trained model to potential distribution shifts at test time using only unlabeled test data, without access to the original training data. While previous efforts mainly focus on a single modality, test-time distribution shift in the multi-modal setting is more complex and calls for new solutions. This paper tackles the problem of multi-modal test-time adaptation by proposing a novel method named Attention Bootstrapping with Principal Entropy Minimization (ABPEM). We observe that test-time distribution shift causes misalignment across modalities, leading to a large gap between intra-modality discrepancies (measured by self-attention) and inter-modality discrepancies (measured by cross-attention). We name this the attention gap. This attention gap widens with more severe distribution shifts, hindering effective modality fusion. To mitigate this attention gap and encourage better modality fusion, we propose attention bootstrapping that promotes cross-attention with the guidance of self-attention. Moreover, to reduce the gradient noise in the commonly-used entropy minimization, we adopt principal entropy minimization, a refinement of entropy minimization that reduces gradient noise by focusing on the principal parts of entropy, excluding less reliable gradient information. Extensive experiments on the benchmarks validate the effectiveness of the proposed ABPEM in comparison with competing baselines.

AAAI Conference 2025 Conference Paper

Cluster-guided Contrastive Class-imbalanced Graph Classification

  • Wei Ju
  • Zhengyang Mao
  • Siyu Yi
  • Yifang Qin
  • Yiyang Gu
  • Zhiping Xiao
  • Jianhao Shen
  • Ziyue Qiao

This paper studies the problem of class-imbalanced graph classification, which aims at effectively classifying the graph categories in scenarios with imbalanced class distributions. While graph neural networks (GNNs) have achieved remarkable success, their modeling ability on imbalanced graph-structured data remains suboptimal, which typically leads to predictions biased towards the majority classes. On the other hand, existing class-imbalanced learning methods in vision may overlook the rich graph semantic substructures of the majority classes and excessively emphasize learning from the minority classes. To address these challenges, we propose a simple yet powerful approach called C3GNN that integrates the idea of clustering into contrastive learning to enhance class-imbalanced graph classification. Technically, C3GNN clusters graphs from each majority class into multiple subclasses, with sizes comparable to the minority class, mitigating class imbalance. It also employs the Mixup technique to generate synthetic samples, enriching the semantic diversity of each subclass. Furthermore, supervised contrastive learning is used to hierarchically learn effective graph representations, enabling the model to thoroughly explore semantic substructures in majority classes while avoiding excessive focus on minority classes. Extensive experiments on real-world graph benchmark datasets verify the superior performance of our proposed method against competitive baselines.

NeurIPS Conference 2025 Conference Paper

Deeper with Riemannian Geometry: Overcoming Oversmoothing and Oversquashing for Graph Foundation Models

  • Li Sun
  • Zhenhao Huang
  • Ming Zhang
  • Philip S Yu

Message Passing Neural Networks (MPNNs) are the building block of graph foundation models, but fundamentally suffer from oversmoothing and oversquashing. There has recently been a surge of interest in fixing both issues. Existing efforts primarily adopt global approaches, which may be beneficial in some regions but detrimental in others, ultimately leading to the suboptimal expressiveness. In this paper, we begin by revisiting oversquashing through a global measure -- spectral gap $\lambda$ -- and prove that the increase of $\lambda$ leads to gradient vanishing with respect to the input features, thereby undermining the effectiveness of message passing. Motivated by such theoretical insights, we propose a local approach that adaptively adjusts message passing based on local structures. To achieve this, we connect local Riemannian geometry with MPNNs, and establish a novel nonhomogeneous boundary condition to address both oversquashing and oversmoothing. Building on the Robin condition, we design a GBN network with local bottleneck adjustment, coupled with theoretical guarantees. Extensive experiments on homophilic and heterophilic graphs show the expressiveness of GBN. Furthermore, GBN does not exhibit performance degradation even when the network depth exceeds $256$ layers.

TMLR Journal 2025 Journal Article

DELTA: Dual Consistency Delving with Topological Uncertainty for Active Graph Domain Adaptation

  • Pengyun Wang
  • Yadi Cao
  • Chris Russell
  • Yanxin Shen
  • Junyu Luo
  • Ming Zhang
  • Siyu Heng
  • Xiao Luo

Graph domain adaptation has recently enabled knowledge transfer across different graphs. However, without the semantic information on target graphs, the performance on target graphs is still far from satisfactory. To address the issue, we study the problem of active graph domain adaptation, which selects a small quantitative of informative nodes on the target graph for extra annotation. This problem is highly challenging due to the complicated topological relationships and the distribution discrepancy across graphs. In this paper, we propose a novel approach named Dual Consistency Delving with Topological Uncertainty (DELTA) for active graph domain adaptation. Our DELTA consists of an edge-oriented graph subnetwork and a path-oriented graph subnetwork, which can explore topological semantics from complementary perspectives. In particular, our edge-oriented graph subnetwork utilizes the message passing mechanism to learn neighborhood information, while our path-oriented graph subnetwork explores high-order relationships from substructures. To jointly learn from two subnetworks, we roughly select informative candidate nodes with the consideration of consistency across two subnetworks. Then, we aggregate local semantics from its K-hop subgraph based on node degrees for topological uncertainty estimation. To overcome potential distribution shifts, we compare target nodes and their corresponding source nodes for discrepancy scores as an additional component for fine selection. Extensive experiments on benchmark datasets demonstrate that DELTA outperforms various state-of-the-art approaches. The code implementation of DELTA is available at https://github.com/goose315/DELTA.

AAAI Conference 2025 Conference Paper

DisCo: Graph-Based Disentangled Contrastive Learning for Cold-Start Cross-Domain Recommendation

  • Hourun Li
  • Yifan Wang
  • Zhiping Xiao
  • Jia Yang
  • Changling Zhou
  • Ming Zhang
  • Wei Ju

Recommender systems are widely used in various real-world applications, but they often encounter the persistent challenge of the user cold-start problem. Cross-domain recommendation (CDR), which leverages user interactions from one domain to improve prediction performance in another, has emerged as a promising solution. However, users with similar preferences in the source domain may exhibit different interests in the target domain. Therefore, directly transferring embeddings may introduce irrelevant source-domain collaborative information. In this paper, we propose a novel graph-based disentangled contrastive learning framework to capture fine-grained user intent and filter out irrelevant collaborative information, thereby avoiding negative transfer. Specifically, for each domain, we use a multi-channel graph encoder to capture diverse user intents. We then construct the affinity graph in the embedding space and perform multi-step random walks to capture high-order user similarity relationships. Treating one domain as the target, we propose a disentangled intent-wise contrastive learning approach, guided by user similarity, to refine the bridging of user intents across domains.Extensive experiments on four benchmark CDR datasets demonstrate that DisCo consistently outperforms existing state-of-the-art baselines, thereby validating the effectiveness of both DisCo and its components.

NeurIPS Conference 2025 Conference Paper

Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed Graphs

  • Yusheng Zhao
  • Qixin Zhang
  • Xiao Luo
  • Weizhi Zhang
  • Zhiping Xiao
  • Wei Ju
  • Philip S Yu
  • Ming Zhang

Large language models (LLMs) have been used in many zero-shot learning problems, with their strong generalization ability. Recently, adopting LLMs in text-attributed graphs (TAGs) has drawn increasing attention. However, the adoption of LLMs faces two major challenges: limited information on graph structure and unreliable responses. LLMs struggle with text attributes isolated from the graph topology. Worse still, they yield unreliable predictions due to both information insufficiency and the inherent weakness of LLMs (e. g. , hallucination). Towards this end, this paper proposes a novel method named Dynamic Text Bundling Supervision (DENSE) that queries LLMs with bundles of texts to obtain bundle-level labels and uses these labels to supervise graph neural networks. Specifically, we sample a set of bundles, each containing a set of nodes with corresponding texts of close proximity. We then query LLMs with the bundled texts to obtain the label of each bundle. Subsequently, the bundle labels are used to supervise the optimization of graph neural networks, and the bundles are further refined to exclude noisy items. To justify our design, we also provide theoretical analysis of the proposed method. Extensive experiments across ten datasets validate the effectiveness of the proposed method.

NeurIPS Conference 2025 Conference Paper

EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving

  • Shihan Dou
  • Ming Zhang
  • Chenhao Huang
  • Jiayi Chen
  • Feng Chen
  • Shichun Liu
  • Yan Liu
  • Chenxiao Liu

We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3. 7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available in the supplementary materials.

AAAI Conference 2025 Conference Paper

GeoMamba: Towards Multi-granular POI Recommendation with Geographical State Space Model

  • Yifang Qin
  • Jiaxuan Xie
  • Zhiping Xiao
  • Ming Zhang

Point-of-Interest (POI) recommendation plays an important role in a wide range of location-based social network ap- plications, aiming to accurately predicting users’ next visits based on their historical check-in records. Previous efforts have primarily focused on the modifications of existing sequential models, neglecting the fact that POI visiting sequences typically involve continuous state transformation of geographical and intention signals. Additionally, the diverse time span between check-ins require the model to prop- erly recognize user’s multi-granular preference. While recent advances of State Space Model (SSM) have revealed their potential in handling intricate temporal signals, we propose a state-based model that is tailored for spatio-temporal POI sequences. On top of traditional SSMs that are typically limited to linear sequences like Mamba, we propose GeoMamba, which customizes the model states to accommodate the spatio-temporal sequences, especially fitting for POI recommendations. Specifically, while the approximation operator HiPPO sets the foundation of linear SSMs, we introduce a novel GaPPO operator that extends the model’s state space into graph-represented geographical domains. This innovation allows us to construct locational SSM encoders that seamlessly integrate users’ spatio-temporal characteristics. The sequence-aware outputs of GeoMamba are further processed to generate multi-scale behavior representations. Extensive experimental results illustrate the superiority of GeoMamba over several state-of-the-art baselines.

AIJ Journal 2025 Journal Article

MATE: Masked optimal transport with dynamic selection for partial label graph learning

  • Yiyang Gu
  • Binqi Chen
  • Zihao Chen
  • Ziyue Qiao
  • Xiao Luo
  • Junyu Luo
  • Zhiping Xiao
  • Wei Ju

This paper investigates the problem of partial label graph learning, in which every graph is associated with a set of candidate labels. Previous methods for weakly supervised graph classification often provide pseudo-labels for graph samples that could be overconfident and biased towards the dominant classes, thus resulting in substantial error accumulation. In this paper, we introduce a new framework named Masked Optimal Transport with Dynamic Selection (MATE) for partial label graph learning, which improves the quality of graph assignments from the perspectives of class balancing and uncertainty mining. In particular, our MATE masks probabilities out of candidate sets and then adopts optimal transport to optimize the assignments without class biases. This design is based on the assumption that the true label distribution is class-balanced or nearly balanced, which is common in various training datasets and real-world scenarios. To further reduce potential noise, we propose a novel scoring metric termed partial energy discrepancy (PED) to evaluate the uncertainty of assignments, and then introduce a dynamic selection strategy that modifies the sample-specific thresholds via momentum updating. Finally, these samples are divided into three levels, i. e. , confident, less-confident, and unconfident and each group is trained separately in our collaborative optimization framework. Extensive experiments on various benchmarks demonstrate the superiority of our MATE compared to various state-of-the-art baselines.

IJCAI Conference 2025 Conference Paper

Physical Adversarial Camouflage Through Gradient Calibration and Regularization

  • Jiawei Liang
  • Siyuan Liang
  • Jianjie Huang
  • Chenxi Si
  • Ming Zhang
  • Xiaochun Cao

The advancement of deep object detectors has greatly affected safety-critical fields like autonomous driving. However, physical adversarial camouflage poses a significant security risk by altering object textures to deceive detectors. Existing techniques struggle with variable physical environments, facing two main challenges: 1) inconsistent sampling point densities across distances hinder the gradient optimization from ensuring local continuity, and 2) updating texture gradients from multiple angles causes conflicts, reducing optimization stability and attack effectiveness. To address these issues, we propose a novel adversarial camouflage framework based on gradient optimization. First, we introduce a gradient calibration strategy, which ensures consistent gradient updates across distances by propagating gradients from sparsely to unsampled texture points, thereby expanding the attack's effective range. Additionally, we develop a gradient decorrelation method, which prioritizes and orthogonalizes gradients based on loss values, enhancing stability and effectiveness in multi-angle optimization by eliminating redundant or conflicting updates. Extensive experimental results on various detection models, angles, and distances show that our method significantly surpasses the state-of-the-art, with an average attack success rate (ASR) increase of 13. 46\% across distances and 11. 03\% across angles. Furthermore, experiments in real-world settings confirm the method's threat potential, highlighting the urgent need for more robust autopilot systems less prone to spoofing.

EAAI Journal 2025 Journal Article

Removing visual occlusion of construction scaffolds via a two-step method combining semantic segmentation and image inpainting

  • Yuexiong Ding
  • Muyang Liu
  • Ming Zhang
  • Xiaowei Luo

With increasing computer vision (CV) applications in automated construction management, the visual occlusion issue caused by crisscrossing, wide-coverage, and immovable scaffolds has become one of the most challenging. This study proposes a novel deep learning-based two-step method combining pixel-level semantic segmentation and contextual image inpainting to remove scaffolds visually and restore the occluded visual information. A low-cost data synthesis method using only unlabeled data has also been developed to alleviate the shortage of labeled data for deep neural network (DNN) training. Experiments on the synthesized test data show that the proposed method achieves performances of 92% mean intersection over union (MIoU) for scaffold segmentation and over 82% structural similarity (SSIM) for scene restoration after removing scaffolds. This research set a precedent for addressing the visual occlusion issue of scaffolds, and the proposed method is verified in real-world cases that it helps existing CV models perform better in scaffolding scenarios.

NeurIPS Conference 2025 Conference Paper

SEGA: Shaping Semantic Geometry for Robust Hashing under Noisy Supervision

  • Yiyang Gu
  • Bohan Wu
  • Qinghua Ran
  • Rong-Cheng Tu
  • Xiao Luo
  • Zhiping Xiao
  • Wei Ju
  • Dacheng Tao

This paper studies the problem of learning hash codes from noisy supervision, which is a practical yet challenging task. This problem is important in extensive real-world applications such as image retrieval and cross-modal retrieval. However, most of the existing methods focus on label denoising to address this problem, but ignore the geometric structure of the hash space, which is critical for learning stable hash codes. Towards this end, this paper proposes a novel framework named Semantic Geometry Shaping (SEGA) that explicitly refines the semantic geometry of hash space. Specifically, we first learn dynamic class prototypes as semantic anchors and cluster hash embeddings around these prototypes to keep structural stability. We then leverage both the energy of predicted distributions and structure-based divergence to estimate the uncertainty of instances and calibrate the supervision in a soft manner. Moreover, we introduce structure-aware interpolation to improve the class boundaries. To verify the effectiveness of our design, we give the theoretical analysis for the proposed framework. Experiments on a range of widely-used retrieval datasets justify the superiority of our SEGA over extensive strong baselines under noisy supervision.

YNIMG Journal 2025 Journal Article

The association among individual gray matter volume of frontal-limbic circuitry, fatigue susceptibility, and comorbid neuropsychiatric symptoms following COVID-19

  • Xuan Niu
  • Wenrui Bao
  • Zhaoyao Luo
  • Pang Du
  • Heping Zhou
  • Haiyang Liu
  • Baoqi Wang
  • Huawen Zhang

BACKGROUND: Fatigue is often accompanied by comorbid sleep disturbance and psychiatric distress following the COVID-19 infection. However, identifying individuals at risk for developing post-COVID fatigue remains challenging. This study aimed to identify the neurobiological markers underlying fatigue susceptibility and further investigate their effect on COVID-19-related neuropsychiatric symptoms. METHODS: Individuals following a mild SARS-CoV-2 infection (COV+) underwent neuropsychiatric measurements (n = 335) and MRI scans (n = 271) within 1 month (baseline), and 191 (70.5 %) of the individuals were followed up 3 months after infection. Sixty-seven healthy controls (COV-) completed the same recruitment protocol. RESULTS: Whole-brain voxel-wise analysis showed that gray matter volume (GMV) during the acute phase did not differ between the COV+ and COV- groups. GMV in the right dorsolateral prefrontal cortex (DLPFC) and left dorsal anterior cingulate cortex (dACC) were associated with fatigue severity only in the COV+ group at baseline, which were assigned to the frontal system and limbic system, respectively. Furthermore, fatigue mediated the associations between volume differences in fatigue susceptibility and COVID-related sleep, post-traumatic stress disorder, anxiety and depression. Crucially, the initial GMV in the right DLPFC can predict fatigue symptoms 3 months after infection. CONCLUSIONS: We provide novel evidence on the neuroanatomical basis of fatigue vulnerability and emphasize that acute fatigue is an important link between early GMV in the frontal-limbic regions and comorbid neuropsychiatric symptoms at baseline and 3 months after infection. Our findings highlight the role of the frontal-limbic system in predisposing individuals to develop post-COVID fatigue.

AAAI Conference 2025 Conference Paper

TRACI: A Data-centric Approach for Multi-Domain Generalization on Graphs

  • Yusheng Zhao
  • Changhu Wang
  • Xiao Luo
  • Junyu Luo
  • Wei Ju
  • Zhiping Xiao
  • Ming Zhang

Graph neural networks (GNNs) have gained superior performance in graph-based prediction tasks with a variety of applications such as social analysis and drug discovery. Despite the remarkable progress, their performance often degrades on test graphs with distribution shifts. Existing domain adaptation methods rely on unlabeled test graphs during optimization, limiting their applicability to graphs in the wild. Towards this end, this paper studies the problem of multi-domain generalization on graphs, which utilizes multiple source graphs to learn a GNN with high performance on unseen target graphs. We propose a new approach named Topological Adversarial Learning with Prototypical Mixup (TRACI) to solve the problem. The fundamental principle behind our TRACI is to produce virtual adversarial and mixed graph samples from a data-centric view. In particular, TRACI enhances GNN generalization by employing a gradient-ascent strategy that considers both label prediction entropy and graph topology to craft challenging adversarial samples. Additionally, it generates domain-agnostic node representations by characterizing class-graph pair prototypes through latent distributions and applying multi-sample prototypical Mixup for distribution alignment across graphs. We further provide theoretical analysis showing that TRACI reduces the model's excess risk. Extensive experiments on various benchmark datasets demonstrate that TRACI outperforms state-of-the-art baselines, validating its effectiveness.

EAAI Journal 2024 Journal Article

A hierarchical deep model integrating economic facts for stock movement prediction

  • Jiahao Yang
  • Ming Zhang
  • Shuo Feng
  • Xuejun Zhang
  • Xing Bai

Accurate stock movement prediction is essential to profit from the stock market. However, this task is challenging due to the complexity and non-stationary nature of the market. Deep learning methods have obtained more attention and success in mining price movement patterns. However, some limitations affect their performances. In general, the stock market is ever-changing, and many factors affect stock movement, so capturing the stock movement patterns is hard without enough prior information. To tackle it, we consider employing economic facts to help improve the deep learning method. In this paper, we propose a novel Hierarchical Deep learning Model that fuses Economic Facts (HDMEF) to predict stock movement from the micro to the macro tiers: the individual, industry, and whole market tiers. Specifically, we present three well-designed modules to separately model them based on the Capital Asset Pricing Model (CAPM), the herding effects, and the holiday effects in the stock market. Experiments on the A-share CSI300 and CSI500 indexes demonstrate that our proposed method performs best on all test phases compared with previous competitive baselines, even an absolute improvement of 2%–3% on some test phases where all the baselines act poor, proving our method is more efficient and robust in different market conditions. In addition, we do an ablation study to analyze the role of various economic effects used in our model, and the results prove that each module is helpful for prediction.

EAAI Journal 2024 Journal Article

A novel health indicator by dominant invariant subspace on Grassmann manifold for state of health assessment of lithium-ion battery

  • Ying Zhang
  • Yan-Fu Li
  • Ming Zhang
  • Huan Wang

The precise estimation of the state of health (SoH) in Lithium-ion batteries (LiBs) relies heavily on a reliable health indicator (HI). Conventional indicators are often constructed by directly concatenating features from multiple sources. It overlooks significant non-linear and correlative information inherent in raw signals. To address this limitation, this paper introduces an innovative approach for SoH estimation in LiBs. Deep features extracted from signals of various sensors are obtained using denoising auto-encoders (DAEs). Then the dominant invariant subspaces (DIS) are calculated through the non-linear transformation of multi-source features on the Grassmann manifold. It can preserve essential and robust characteristics. The health indicator quantifies the geodesic distance of DIS using a projection metric. It provides a more comprehensive inclusion of nonlinear and correlation information. Consequently, this indicator offers heightened precision in discerning differences in health states. Validation of the proposed method is conducted using the NASA dataset. The result demonstrates its effectiveness on the SoH assessment and superiority to the state-of-the-art method.

IJCAI Conference 2024 Conference Paper

A Survey of Data-Efficient Graph Learning

  • Wei Ju
  • Siyu Yi
  • Yifan Wang
  • Qingqing Long
  • Junyu Luo
  • Zhiping Xiao
  • Ming Zhang

Graph-structured data, prevalent in domains ranging from social networks to biochemical analysis, serve as the foundation for diverse real-world systems. While graph neural networks demonstrate proficiency in modeling this type of data, their success is often reliant on significant amounts of labeled data, posing a challenge in practical scenarios with limited annotation resources. To tackle this problem, tremendous efforts have been devoted to enhancing graph machine learning performance under low-resource settings by exploring various approaches to minimal supervision. In this paper, we introduce a novel concept of Data-Efficient Graph Learning (DEGL) as a research frontier, and present the first survey that summarizes the current progress of DEGL. We initiate by highlighting the challenges inherent in training models with large labeled data, paving the way for our exploration into DEGL. Next, we systematically review recent advances on this topic from several key aspects, including self-supervised graph learning, semi-supervised graph learning, and few-shot graph learning. Also, we state promising directions for future research, contributing to the evolution of graph machine learning.

NeurIPS Conference 2024 Conference Paper

EGODE: An Event-attended Graph ODE Framework for Modeling Rigid Dynamics

  • Jingyang Yuan
  • Gongbo Sun
  • Zhiping Xiao
  • Hang Zhou
  • Xiao Luo
  • Junyu Luo
  • Yusheng Zhao
  • Wei Ju

This paper studies the problem of rigid dynamics modeling, which has a wide range of applications in robotics, graphics, and mechanical design. The problem is partly solved by graph neural network (GNN) simulators. However, these approaches cannot effectively handle the relationship between intrinsic continuity and instantaneous changes in rigid dynamics. Moreover, they usually neglect hierarchical structures across mesh nodes and objects in systems. In this paper, we propose a novel approach named Event-attend Graph ODE (EGODE) for effective rigid dynamics modeling. In particular, we describe the rigid system using both mesh node representations and object representations. To model continuous dynamics across hierarchical structures, we use a coupled graph ODE framework for the evolution of both types of representations over a long period. In addition, to capture instantaneous changes during the collision, we introduce an event module, which can effectively estimate the occurrence of the collision and update the states of both mesh node and object representations during evolution. Extensive experiments on a range of benchmark datasets validate the superiority of the proposed EGODE compared to various state-of-the-art baselines. The source code can be found at https: //github. com/yuanjypku/EGODE.

YNIMG Journal 2024 Journal Article

Individual differences of white matter characteristic along the anterior insula-based fiber tract circuit for pain empathy in healthy women and women with primary dysmenorrhea

  • Junya Mu
  • Leiming Wu
  • Chenxi Wang
  • Wanghuan Dun
  • Zilong Hong
  • Xinyue Feng
  • Ming Zhang
  • Jixin Liu

Pain empathy, defined as the ability of one person to understand another person's pain, shows large individual variations. The anterior insula is the core region of the pain empathy network. However, the relationship between white matter (WM) properties of the fiber tracts connecting the anterior insula with other cortical regions and an individual's ability to modulate pain empathy remains largely unclear. In this study, we outline an automatic seed-based fiber streamline (sFS) analysis method and multivariate pattern analysis (MVPA) to predict the levels of pain empathy in healthy women and women with primary dysmenorrhoea (PDM). Using the sFS method, the anterior insula-based fiber tract network was divided into five fiber cluster groups. In healthy women, interindividual differences in pain empathy were predicted only by the WM properties of the five fiber cluster groups, suggesting that interindividual differences in pain empathy may rely on the connectivity of the anterior insula-based fiber tract network. In women with PDM, pain empathy could be predicted by a single cluster group. The mean WM properties along the anterior insular-rostroventral area of the inferior parietal lobule further mediated the effect of pain on empathy in patients with PDM. Our results suggest that chronic periodic pain may lead to maladaptive plastic changes, which could further impair empathy by making women with PDM feel more pain when they see other people experiencing pain. Our study also addresses an important gap in the analysis of the microstructural characteristics of seed-based fiber tract network.

AAAI Conference 2024 Conference Paper

LLMEval: A Preliminary Study on How to Evaluate Large Language Models

  • Yue Zhang
  • Ming Zhang
  • Haipeng Yuan
  • Shichun Liu
  • Yongyao Shi
  • Tao Gui
  • Qi Zhang
  • Xuanjing Huang

Recently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first two questions, which are basically what tasks to give the LLM during testing and what kind of knowledge it should deal with. As for the third question, which is about what standards to use, the types of evaluators, how to score, and how to rank, there hasn't been much discussion. In this paper, we analyze evaluation methods by comparing various criteria with both manual and automatic evaluation, utilizing onsite, crowd-sourcing, public annotators and GPT-4, with different scoring methods and ranking systems. We propose a new dataset, LLMEval and conduct evaluations on 20 LLMs. A total of 2,186 individuals participated, leading to the generation of 243,337 manual annotations and 57,511 automatic evaluation results. We perform comparisons and analyses of different settings and conduct 10 conclusions that can provide some insights for evaluating LLM in the future. The dataset and the results are publicly available at https://github.com/llmeval. The version with the appendix are publicly available at https://arxiv.org/abs/2312.07398.

AAAI Conference 2024 Conference Paper

Preparing Lessons for Progressive Training on Language Models

  • Yu Pan
  • Ye Yuan
  • Yichun Yin
  • Jiaxin Shi
  • Zenglin Xu
  • Ming Zhang
  • Lifeng Shang
  • Xin Jiang

The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.

IJCAI Conference 2024 Conference Paper

Rank and Align: Towards Effective Source-free Graph Domain Adaptation

  • Junyu Luo
  • Zhiping Xiao
  • Yifan Wang
  • Xiao Luo
  • Jingyang Yuan
  • Wei Ju
  • Langechuan Liu
  • Ming Zhang

Graph neural networks (GNNs) have achieved impressive performance in graph domain adaptation. However, extensive source graphs could be unavailable in real-world scenarios due to privacy and storage concerns. To this end, we investigate an underexplored yet practical problem of source-free graph domain adaptation, which transfers knowledge from source models instead of source graphs to a target domain. To solve this problem, we introduce a novel GNN-based approach called Rank and Align (RNA), which ranks graph similarities with spectral seriation for robust semantics learning, and aligns inharmonic graphs with harmonic graphs which close to the source domain for subgraph extraction. In particular, to overcome label scarcity, we employ the spectral seriation algorithm to infer the robust pairwise rankings, which can guide semantic learning using a similarity learning objective. To depict distribution shifts, we utilize spectral clustering and the silhouette coefficient to detect harmonic graphs, which the source model can easily classify. To reduce potential domain discrepancy, we extract domain-invariant subgraphs from inharmonic graphs by an adversarial edge sampling process, which guides the invariant learning of GNNs. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our proposed RNA.

NeurIPS Conference 2023 Conference Paper

A*Net: A Scalable Path-based Reasoning Approach for Knowledge Graphs

  • Zhaocheng Zhu
  • Xinyu Yuan
  • Michael Galkin
  • Louis-Pascal Xhonneux
  • Ming Zhang
  • Maxime Gazeau
  • Jian Tang

Reasoning on large-scale knowledge graphs has been long dominated by embedding methods. While path-based methods possess the inductive capacity that embeddings lack, their scalability is limited by the exponential number of paths. Here we present A*Net, a scalable path-based method for knowledge graph reasoning. Inspired by the A* algorithm for shortest path problems, our A*Net learns a priority function to select important nodes and edges at each iteration, to reduce time and memory footprint for both training and inference. The ratio of selected nodes and edges can be specified to trade off between performance and efficiency. Experiments on both transductive and inductive knowledge graph reasoning benchmarks show that A*Net achieves competitive performance with existing state-of-the-art path-based methods, while merely visiting 10% nodes and 10% edges at each iteration. On a million-scale dataset ogbl-wikikg2, A*Net not only achieves a new state-of-the-art result, but also converges faster than embedding methods. A*Net is the first path-based method for knowledge graph reasoning at such scale.

YNIMG Journal 2023 Journal Article

APART-QSM: An improved sub-voxel quantitative susceptibility mapping for susceptibility source separation using an iterative data fitting method

  • Zhenghao Li
  • Ruimin Feng
  • Qiangqiang Liu
  • Jie Feng
  • Guoyan Lao
  • Ming Zhang
  • Jun Li
  • Yuyao Zhang

The brain tissue phase contrast in MRI sequences reflects the spatial distributions of multiple substances, such as iron, myelin, calcium, and proteins. These substances with paramagnetic and diamagnetic susceptibilities often colocalize in one voxel in brain regions. Both opposing susceptibilities play vital roles in brain development and neurodegenerative diseases. Conventional QSM methods only provide voxel-averaged susceptibility value and cannot disentangle intravoxel susceptibilities with opposite signs. Advanced susceptibility imaging methods have been recently developed to distinguish the contributions of opposing susceptibility sources for QSM. The basic concept of separating paramagnetic and diamagnetic susceptibility proportions is to include the relaxation rate R 2 * with R 2 ′ in QSM. The magnitude decay kernel, describing the proportionality coefficient between R 2 ′ and susceptibility, is an essential reconstruction coefficient for QSM separation methods. In this study, we proposed a more comprehensive complex signal model that describes the relationship between 3D GRE signal and the contributions of paramagnetic and diamagnetic susceptibility to the frequency shift and R 2 * relaxation. The algorithm is implemented as a constrained minimization problem in which the voxel-wise magnitude decay kernel and sub-voxel susceptibilities are determined alternately in each iteration until convergence. The calculated voxel-wise magnitude decay kernel could realistically model the relationship between the R 2 ′ relaxation and the volume susceptibility. Thus, the proposed method effectively prevents the errors of the magnitude decay kernel from propagating to the final susceptibility separation reconstruction. Phantom studies, ex vivo macaque brain experiments, and in vivo human brain imaging studies were conducted to evaluate the ability of the proposed method to distinguish paramagnetic and diamagnetic susceptibility sources. The results demonstrate that the proposed method provides state-of-the-art performances for quantifying brain iron and myelin compared to previous QSM separation methods. Our results show that the proposed method has the potential to simultaneously quantify whole brain iron and myelin during brain development and aging. The proposed model was also deployed with multiple-orientation complex GRE data input measurements, resulting in high-quality QSM separation maps with more faithful tissue delineation between brain structures compared to those reconstructed by single-orientation QSM separation methods.

AAAI Conference 2023 Conference Paper

GLCC: A General Framework for Graph-Level Clustering

  • Wei Ju
  • Yiyang Gu
  • Binqi Chen
  • Gongbo Sun
  • Yifang Qin
  • Xingyuming Liu
  • Xiao Luo
  • Ming Zhang

This paper studies the problem of graph-level clustering, which is a novel yet challenging task. This problem is critical in a variety of real-world applications such as protein clustering and genome analysis in bioinformatics. Recent years have witnessed the success of deep clustering coupled with graph neural networks (GNNs). However, existing methods focus on clustering among nodes given a single graph, while exploring clustering on multiple graphs is still under-explored. In this paper, we propose a general graph-level clustering framework named Graph-Level Contrastive Clustering (GLCC) given multiple graphs. Specifically, GLCC first constructs an adaptive affinity graph to explore instance- and cluster-level contrastive learning (CL). Instance-level CL leverages graph Laplacian based contrastive loss to learn clustering-friendly representations while cluster-level CL captures discriminative cluster representations incorporating neighbor information of each sample. Moreover, we utilize neighbor-aware pseudo-labels to reward the optimization of representation learning. The two steps can be alternatively trained to collaborate and benefit each other. Experiments on a range of well-known datasets demonstrate the superiority of our proposed GLCC over competitive baselines.

ICLR Conference 2023 Conference Paper

GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive Simulation

  • Ming Zhang
  • Shenghan Zhang
  • Zhenjie Yang
  • Lekai Chen
  • Jinliang Zheng
  • Chao Yang 0026
  • Chuming Li
  • Hang Zhou 0009

The emergence of various multi-agent environments has motivated powerful algorithms to explore agents' cooperation or competition. Even though this has greatly promoted the development of multi-agent reinforcement learning (MARL), it is still not enough to support further exploration on the behavior of swarm intelligence between multiple teams, and cooperation between multiple agents due to their limited scalability. To alleviate this, we introduce GoBigger, a scalable platform for cooperative-competition multi-agent interactive simulation. GoBigger is an enhanced environment for the Agar-like game, enabling the simulation of multiple scales of agent intra-team cooperation and inter-team competition. Compared with existing multi-agent simulation environments, our platform supports multi-team games with more than two teams simultaneously, which dramatically expands the diversity of agent cooperation and competition, and can more effectively simulate the swarm intelligent agent behavior. Besides, in GoBigger, the cooperation between the agents in a team can lead to much higher performance. We offer a diverse set of challenging scenarios, built-in bots, and visualization tools for best practices in benchmarking. We evaluate several state-of-the-art algorithms on GoBigger and demonstrate the potential of the environment. We believe this platform can inspire various emerging research directions in MARL, swarm intelligence, and large-scale agent interactive learning. Both GoBigger and its related benchmark are open-sourced. More information could be found at https://github.com/opendilab/GoBigger.

TMLR Journal 2023 Journal Article

RIGNN: A Rationale Perspective for Semi-supervised Open-world Graph Classification

  • Xiao Luo
  • Yusheng Zhao
  • Zhengyang Mao
  • Yifang Qin
  • Wei Ju
  • Ming Zhang
  • Yizhou Sun

Graph classification has gained growing attention in the graph machine learning community and a variety of semi-supervised methods have been developed to reduce the high cost of annotation. They usually combine graph neural networks (GNNs) and extensive semi-supervised techniques such as knowledge distillation. However, they adhere to the close-set assumption that unlabeled graphs all belong to known classes, limiting their applications in the real world. This paper goes further, investigating a practical problem of semi-supervised open-world graph classification where these unlabeled graph data could come from unseen classes. A novel approach named Rationale-Informed GNN (RIGNN) is proposed, which takes a rationale view to detect components containing the most information related to the label space and classify unlabeled graphs into a known class or an unseen class. In particular, RIGNN contains a relational detector and a feature extractor to produce effective rationale features, which maximize the mutual information with label information and exhibit sufficient disentanglement with non-rationale elements. Furthermore, we construct a graph-of-graph based on geometrical relationships, which gives instructions on enhancing rationale representations. In virtue of effective rationale representations, we can provide accurate and balanced predictions for unlabeled graphs. An extension is also made to accomplish effective open-set graph classification. We verify our proposed methods on four benchmark datasets in various settings and experimental results reveal the effectiveness of our proposed RIGNN compared with state-of-the-art methods.

TMLR Journal 2023 Journal Article

Zero-shot Node Classification with Graph Contrastive Embedding Network

  • Wei Ju
  • Yifang Qin
  • Siyu Yi
  • Zhengyang Mao
  • Kangjie Zheng
  • Luchen Liu
  • Xiao Luo
  • Ming Zhang

This paper studies zero-shot node classification, which aims to predict new classes (i.e., unseen classes) of nodes in a graph. This problem is challenging yet promising in a variety of real-world applications such as social analysis and bioinformatics. The key of zero-shot node classification is to enable the knowledge transfer of nodes from training classes to unseen classes. However, existing methods typically ignore the dependencies between nodes and classes, and fail to be organically integrated in a united way. In this paper, we present a novel framework called the Graph Contrastive Embedding Network (GraphCEN) for zero-shot node classification. Specifically, GraphCEN first constructs an affinity graph to model the relations between the classes. Then the node- and class-level contrastive learning (CL) are proposed to jointly learn node embeddings and class assignments in an end-to-end manner. The two-level CL can be optimized to mutually enhance each other. Extensive experiments indicate that our GraphCEN significantly outperforms the state-of-the-art approaches on multiple challenging benchmark datasets.

AAAI Conference 2022 Conference Paper

DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation Generation

  • Yifan Wang
  • Yiping Song
  • Shuai Li
  • Chaoran Cheng
  • Wei Ju
  • Ming Zhang
  • Sheng Wang

Citing and describing related literature are crucial to scientific writing. Many existing approaches show encouraging performance in citation recommendation, but are unable to accomplish the more challenging and onerous task of citation text generation. In this paper, we propose a novel disentangled representation based model DisenCite to automatically generate the citation text through integrating paper text and citation graph. A key novelty of our method compared with existing approaches is to generate context-specific citation text, empowering the generation of different types of citations for the same paper. In particular, we first build and make available a graph enhanced contextual citation dataset (GCite) with 25K edges in different types characterized by citation contained sections over 4. 8K research papers. Based on this dataset, we encode each paper according to both textual contexts and structure information in the heterogeneous citation graph. The resulted paper representations are then disentangled by the mutual information regularization between this paper and its neighbors in graph. Extensive experiments demonstrate the superior performance of our method comparing to state-of-the-art approaches. We further conduct ablation and case studies to reassure that the improvement of our method comes from generating the context-specific citation through incorporating the citation graph.

IJCAI Conference 2022 Conference Paper

Improving Transferability of Adversarial Examples with Virtual Step and Auxiliary Gradients

  • Ming Zhang
  • Xiaohui Kuang
  • Hu Li
  • Zhendong Wu
  • Yuanping Nie
  • Gang Zhao

Deep neural networks have been demonstrated to be vulnerable to adversarial examples, which fool networks by adding human-imperceptible perturbations to benign examples. At present, the practical transfer-based black-box attacks are attracting significant attention. However, most existing transfer-based attacks achieve only relatively limited success rates. We propose to improve the transferability of adversarial examples through the use of a virtual step and auxiliary gradients. Here, the “virtual step” refers to using an unusual step size and clipping adversarial perturbations only in the last iteration, while the “auxiliary gradients” refer to using not only gradients corresponding to the ground-truth label (for untargeted attacks), but also gradients corresponding to some other labels to generate adversarial perturbations. Our proposed virtual step and auxiliary gradients can be easily integrated into existing gradient-based attacks. Extensive experiments on ImageNet show that the adversarial examples crafted by our method can effectively transfer to different networks. For single-model attacks, our method outperforms the state-of-the-art baselines, improving the success rates by a large margin of 12%~28%. Our code is publicly available at https: //github. com/mingcheung/Virtual-Step-and-Auxiliary-Gradients.

IJCAI Conference 2022 Conference Paper

TGNN: A Joint Semi-supervised Framework for Graph-level Classification

  • Wei Ju
  • Xiao Luo
  • Meng Qu
  • Yifan Wang
  • Chong Chen
  • Minghua Deng
  • Xian-Sheng Hua
  • Ming Zhang

This paper studies semi-supervised graph classification, a crucial task with a wide range of applications in social network analysis and bioinformatics. Recent works typically adopt graph neural networks to learn graph-level representations for classification, failing to explicitly leverage features derived from graph topology (e. g. , paths). Moreover, when labeled data is scarce, these methods are far from satisfactory due to their insufficient topology exploration of unlabeled data. We address the challenge by proposing a novel semi-supervised framework called Twin Graph Neural Network (TGNN). To explore graph structural information from complementary views, our TGNN has a message passing module and a graph kernel module. To fully utilize unlabeled data, for each module, we calculate the similarity of each unlabeled graph to other labeled graphs in the memory bank and our consistency loss encourages consistency between two similarity distributions in different embedding spaces. The two twin modules collaborate with each other by exchanging instance similarity knowledge to fully explore the structure information of both labeled and unlabeled data. We evaluate our TGNN on various public datasets and show that it achieves strong performance.

YNIMG Journal 2022 Journal Article

The role of the interaction between the inferior parietal lobule and superior temporal gyrus in the multisensory Go/No-go task

  • Jiaying Sun
  • Jie Huang
  • Aijun Wang
  • Ming Zhang
  • Xiaoyu Tang

Information from multiple sensory modalities interacts. Using functional magnetic resonance imaging (fMRI), we aimed to identify the neural structures correlated with how cooccurring sound modulates the visual motor response execution. The reaction time (RT) to audiovisual stimuli was significantly faster than the RT to visual stimuli. Signal detection analyses showed no significant difference in the perceptual sensitivity (d') between audiovisual and visual stimuli, while the response criteria (β or c) of the audiovisual stimuli was decreased compared to the visual stimuli. The functional connectivity between the left inferior parietal lobule (IPL) and bilateral superior temporal gyrus (STG) was enhanced in Go processing compared with No-go processing of audiovisual stimuli. Furthermore, the left precentral gyrus (PreCG) showed enhanced functional connectivity with the bilateral STG and other areas of the ventral stream in Go processing compared with No-go processing of audiovisual stimuli. These results revealed that the neuronal network correlated with modulations of the motor response execution after the presentation of both visual stimuli along with cooccurring sound in a multisensory Go/Nogo task, including the left IPL, left PreCG, bilateral STG and some areas of the ventral stream. The role of the interaction between the IPL and STG in transforming audiovisual information into motor behavior is discussed. The current study provides a new perspective for exploring potential brain mechanisms underlying how humans execute appropriate behaviors on the basis of multisensory information.

YNIMG Journal 2021 Journal Article

Gender discrimination facilitates fMRI responses and connectivity to thermal pain

  • Ming Zhang
  • Yuqi Zhang
  • Yan Mu
  • Zhaoxing Wei
  • Yazhuo Kong

Gender discrimination is a serious social issue that has been shown to increase negative consequences, especially in females when accompanied by acute or chronic pain. Experiencing social pain through discrimination can increase an individual's evaluation of evoked physical pain. However, few studies have explored the mechanism underlying how gender discrimination modulates brain responses when individuals experience physical pain evoked by noxious stimuli. In this study, we addressed this issue using a gender discrimination fMRI paradigm with thermal pain stimulation. We found that discrimination indeed affected participants' own behavioral self-evaluation of noxious stimuli. Discrimination-encoded brain activations were identified in the temporopolar cortex, while brain activations to thermal stimuli after viewing pictures of discrimination were found in the dorsal anterior cingulate cortex (dACC). Brain activations in the temporopolar cortex and the dACC were correlated. Furthermore, pain perception-specific functional connectivity of the dACC-SII in the cue stage and the dACC-frontal in the pain stage were identified, suggesting a facilitative effect of gender discrimination on females' experience of physical pain. Our results indicate that the dACC may play a central role in mediating the affective aspect of physical pain after experiencing discrimination. These findings provide novel insights into the underlying mechanism of how gender discrimination facilitates females' experience of physical pain.

ICRA Conference 2021 Conference Paper

IMU Data Processing For Inertial Aided Navigation: A Recurrent Neural Network Based Approach

  • Ming Zhang
  • Mingming Zhang 0008
  • Yiming Chen 0001
  • Mingyang Li 0001

In this work, we propose a novel method for performing inertial aided navigation, by using deep neural net-works (DNNs). To date, most DNN inertial navigation methods focus on the task of inertial odometry, by taking gyroscope and accelerometer readings as input and regressing for integrated IMU poses (i. e. , position and orientation). While this design has been successfully applied on a number of applications, it is not of theoretical performance guarantee unless patterned motion is involved. This inevitably leads to significantly reduced accuracy and robustness in certain use cases. To solve this problem, we design a framework to compute observable IMU integration terms using DNNs, followed by the numerical pose integration and sensor fusion to achieve the performance gain. Specifically, we perform detailed analysis on the motion terms in IMU kinematic equations, propose a dedicated network design, loss functions, and training strategies for the IMU data processing, and conduct extensive experiments. The results show that our method is generally applicable and outperforms both traditional and DNN methods by wide margins.

YNIMG Journal 2021 Journal Article

MoDL-QSM: Model-based deep learning for quantitative susceptibility mapping

  • Ruimin Feng
  • Jiayi Zhao
  • He Wang
  • Baofeng Yang
  • Jie Feng
  • Yuting Shi
  • Ming Zhang
  • Chunlei Liu

Quantitative susceptibility mapping (QSM) has demonstrated great potential in quantifying tissue susceptibility in various brain diseases. However, the intrinsic ill-posed inverse problem relating the tissue phase to the underlying susceptibility distribution affects the accuracy for quantifying tissue susceptibility. Recently, deep learning has shown promising results to improve accuracy by reducing the streaking artifacts. However, there exists a mismatch between the observed phase and the theoretical forward phase estimated by the susceptibility label. In this study, we proposed a model-based deep learning architecture that followed the STI (susceptibility tensor imaging) physical model, referred to as MoDL-QSM. Specifically, MoDL-QSM accounts for the relationship between STI-derived phase contrast induced by the susceptibility tensor terms ( χ 13, χ 23 and χ 33 ) and the acquired single-orientation phase. The convolutional neural networks are embedded into the physical model to learn a regularization term containing prior information. χ 33 and phase induced by χ 13 and χ 23 terms were used as the labels for network training. Quantitative evaluation metrics were compared with recently developed deep learning QSM methods. The results showed that MoDL-QSM achieved superior performance, demonstrating its potential for future applications.

YNICL Journal 2021 Journal Article

Neurological effects of hemodialysis on white matter microstructure in end-stage renal disease

  • Junya Mu
  • Liang Ma
  • Shaohui Ma
  • Dun Ding
  • Peng Li
  • Xueying Ma
  • Ming Zhang
  • Jixin Liu

OBJECTIVES: To detect the effects of hemodialysis (HD) on the central nervous system (CNS), the present study forces the memory storage capacity and the difference in white matter (WM) microstructure characteristics among end-stage renal disease (ESRD) participants before HD initiation (ESRD-BHD), ESRD participants with maintenance HD (ESRD-MHD), and healthy participants (HCs). METHODS: Between 2016 and 2018, 56 ESRD-BHD, 39 ESRD-MHD, and 56 HCs were recruited for this study. The fractional anisotropy (FA) of tractography streamlines within the working memory network was investigated using a novel along-tracts analysis method. The relationship between WM microstructure and working memory scores, measured from an n-back task, were detected by multiple correlation analysis. RESULTS: As compared with HCs, a significantly lower FA was found along part of the WM in the working memory network in ESRD-BHD. In the group-difference location of ESRD-BHD and HCs, the FA of ESRD-MHD was reversed to normal levels in HCs. However, the FA in a new location was differentially reduced across groups: highest in HCs, intermediate in ESRD-BHD, and lowest in ESRD-MHD. Correlation analysis showed that a longer reaction time correlated to a lower FA, according to the following pattern: ESRD-BHD > ESRD-MHD > HCs. CONCLUSION: Despite the persisting abnormal brain structure, our findings suggest HD has a neuroprotective effect in ESRD patients.

NeurIPS Conference 2020 Conference Paper

Multi-agent Trajectory Prediction with Fuzzy Query Attention

  • Nitin Kamra
  • Hao Zhu
  • Dweep Kumarbhai Trivedi
  • Ming Zhang
  • Yan Liu

Trajectory prediction for scenes with multiple agents and entities is a challenging problem in numerous domains such as traffic prediction, pedestrian tracking and path planning. We present a general architecture to address this challenge which models the crucial inductive biases of motion, namely, inertia, relative motion, intents and interactions. Specifically, we propose a relational model to flexibly model interactions between agents in diverse environments. Since it is well-known that human decision making is fuzzy by nature, at the core of our model lies a novel attention mechanism which models interactions by making continuous-valued (fuzzy) decisions and learning the corresponding responses. Our architecture demonstrates significant performance gains over existing state-of-the-art predictive models in diverse domains such as human crowd trajectories, US freeway traffic, NBA sports data and physics datasets. We also present ablations and augmentations to understand the decision-making process and the source of gains in our model.

TIST Journal 2019 Journal Article

Combating Fake News

  • Karishma Sharma
  • Feng Qian
  • He Jiang
  • Natali Ruchansky
  • Ming Zhang
  • Yan Liu

The proliferation of fake news on social media has opened up new directions of research for timely identification and containment of fake news and mitigation of its widespread impact on public opinion. While much of the earlier research was focused on identification of fake news based on its contents or by exploiting users’ engagements with the news on social media, there has been a rising interest in proactive intervention strategies to counter the spread of misinformation and its impact on society. In this survey, we describe the modern-day problem of fake news and, in particular, highlight the technical challenges associated with it. We discuss existing methods and techniques applicable to both identification and mitigation, with a focus on the significant advances in each method and their advantages and limitations. In addition, research has often been limited by the quality of existing datasets and their specific application contexts. To alleviate this problem, we comprehensively compile and summarize characteristic features of available datasets. Furthermore, we outline new directions of research to facilitate future development of effective and interdisciplinary solutions.

IJCAI Conference 2018 Conference Paper

An Ensemble of Retrieval-Based and Generation-Based Human-Computer Conversation Systems

  • Yiping Song
  • Cheng-Te Li
  • Jian-Yun Nie
  • Ming Zhang
  • Dongyan Zhao
  • Rui Yan

Human-computer conversation systems have attracted much attention in Natural Language Processing. Conversation systems can be roughly divided into two categories: retrieval-based and generation-based systems. Retrieval systems search a user-issued utterance (namely a query ) in a large conversational repository and return a reply that best matches the query. Generative approaches synthesize new replies. Both ways have certain advantages but suffer from their own disadvantages. We propose a novel ensemble of retrieval-based and generation-based conversation system. The retrieved candidates, in addition to the original query, are fed to a reply generator via a neural network, so that the model is aware of more information. The generated reply together with the retrieved ones then participates in a re-ranking process to find the final reply to output. Experimental results show that such an ensemble system outperforms each single module by a large margin.

AAAI Conference 2018 Conference Paper

Learning the Joint Representation of Heterogeneous Temporal Events for Clinical Endpoint Prediction

  • Luchen Liu
  • Jianhao Shen
  • Ming Zhang
  • Zichang Wang
  • Jian Tang

The availability of a large amount of electronic health records (EHR) provides huge opportunities to improve health care service by mining these data. One important application is clinical endpoint prediction, which aims to predict whether a disease, a symptom or an abnormal lab test will happen in the future according to patients’ history records. This paper develops deep learning techniques for clinical endpoint prediction, which are effective in many practical applications. However, the problem is very challenging since patients’ history records contain multiple heterogeneous temporal events such as lab tests, diagnosis, and drug administrations. The visiting patterns of different types of events vary significantly, and there exist complex nonlinear relationships between different events. In this paper, we propose a novel model for learning the joint representation of heterogeneous temporal events. The model adds a new gate to control the visiting rates of different events which effectively models the irregular patterns of different events and their nonlinear correlations. Experiment results with real-world clinical data on the tasks of predicting death and abnormal lab tests prove the effectiveness of our proposed approach over competitive baselines.

AAAI Conference 2018 Conference Paper

Towards a Neural Conversation Model With Diversity Net Using Determinantal Point Processes

  • Yiping Song
  • Rui Yan
  • Yansong Feng
  • Yaoyuan Zhang
  • Dongyan Zhao
  • Ming Zhang

Typically, neural conversation systems generate replies based on the sequence-to-sequence (seq2seq) model. seq2seq tends to produce safe and universal replies, which suffers from the lack of diversity and information. Determinantal Point Processes (DPPs) is a probabilistic model defined on item sets, which can select the items with good diversity and quality. In this paper, we investigate the diversity issue in two different aspects, namely query-level and system-level diversity. We propose a novel framework which organically combines seq2seq model with Determinantal Point Processes (DPPs). The new framework achieves high quality in generated reply and significantly improves the diversity among them. Experiments show that our model achieves the best performance among various baselines in terms of both quality and diversity.

IJCAI Conference 2017 Conference Paper

Semi-supervised Learning over Heterogeneous Information Networks by Ensemble of Meta-graph Guided Random Walks

  • He Jiang
  • Yangqiu Song
  • Chenguang Wang
  • Ming Zhang
  • Yizhou Sun

Heterogeneous information networks (HINs) is a general representation of many real world applications. The difference between HIN and traditional homogeneous graphs is that the nodes and edges in HIN are with types. Then in the many applications, we need to consider the types to make the approach more semantically meaningful. For the applications that annotation is expensive, on natural way is to consider semi-supervised learning over HIN. In this paper, we present a semi-supervised learning algorithm constrained by the types of HINs. We first decompose the original HIN into several semantically meaningful sub-graphs based the meta-graphs composed of entity and relation types. Then we perform random walk over the sub-graphs to propagate the labels from labeled data to unlabeled data. After we obtain all the labels propagated by different trials of random walk guided by meta-graphs, we use an ensemble algorithm to vote for the final labeling results. We use two public available datasets, 20-newsgroups and RCV1 datasets to test our algorithm. Experimental results show that our algorithm is better than the traditional semi-supervised learning algorithms for HINs. One particular by-product of this work is that we show that previous random walk approach guided by meta-paths can be non-stationary, which is the major reason we propose a meta-graph guide random walk for semi-supervised learning over HINs.

IJCAI Conference 2017 Conference Paper

Socialized Word Embeddings

  • Ziqian Zeng
  • Yichun Yin
  • Yangqiu Song
  • Ming Zhang

Word embeddings have attracted a lot of attention. On social media, each user’s language use can be significantly affected by the user’s friends. In this paper, we propose a socialized word embedding algorithm which can consider both user’s personal characteristics of language use and the user’s social relationship on social media. To incorporate personal characteristics, we propose to use a user vector to represent each user. Then for each user, the word embeddings are trained based on each user’s corpus by combining the global word vectors and local user vector. To incorporate social relationship, we add a regularization term to impose similarity between two friends. In this way, we can train the global word vectors and user vectors jointly. To demonstrate the effectiveness, we used the latest large-scale Yelp data to train our vectors, and designed several experiments to show how user vectors affect the results.

IJCAI Conference 2016 Conference Paper

StalemateBreaker: A Proactive Content-Introducing Approach to Automatic Human-Computer Conversation

  • Xiang Li
  • Lili Mou
  • Rui Yan
  • Ming Zhang

Existing open-domain human-computer conversation systems are typically passive: they either synthesize or retrieve a reply provided with a human-issued utterance. It is generally presumed that humans should take the role to lead the conversation and introduce new content when a stalemate occurs, and that computers only need to "respond. " In this paper, we propose STALEMATEBREAKER, a conversation system that can proactively introduce new content when appropriate. We design a pipeline to determine when, what, and how to introduce new content during human-computer conversation. We further propose a novel reranking algorithm Bi-PageRank-HITS to enable rich interaction between conversation context and candidate replies. Experiments show that both the content-introducing approach and the reranking algorithm are effective. Our full STALEMATEBREAKER model outperforms a state-of-the-practice conversation system by +14. 4% p@1 when a stalemate occurs.

AAAI Conference 2016 Conference Paper

Text Classification with Heterogeneous Information Network Kernels

  • Chenguang Wang
  • Yangqiu Song
  • Haoran Li
  • Ming Zhang
  • Jiawei Han

Text classification is an important problem with many applications. Traditional approaches represent text as a bagof-words and build classifiers based on this representation. Rather than words, entity phrases, the relations between the entities, as well as the types of the entities and relations carry much more information to represent the texts. This paper presents a novel text as network classification framework, which introduces 1) a structured and typed heterogeneous information networks (HINs) representation of texts, and 2) a meta-path based approach to link texts. We show that with the new representation and links of texts, the structured and typed information of entities and relations can be incorporated into kernels. Particularly, we develop both simple linear kernel and indefinite kernel based on metapaths in the HIN representation of texts, where we call them HIN-kernels. Using Freebase, a well-known world knowledge base, to construct HIN for texts, our experiments on two benchmark datasets show that the indefinite HIN-kernel based on weighted meta-paths outperforms the state-of-theart methods and other HIN-kernels.

IJCAI Conference 2016 Conference Paper

Unsupervised Word and Dependency Path Embeddings for Aspect Term Extraction

  • Yichun Yin
  • Furu Wei
  • Li Dong
  • Kaimeng Xu
  • Ming Zhang
  • Ming Zhou

In this paper, we develop a novel approach to aspect term extraction based on unsupervised learning of distributed representations of words and dependency paths. The basic idea is to connect two words (w1 and w2) with the dependency path (r) between them in the embedding space. Specifically, our method optimizes the objective w1 + r ≈ w2 in the low-dimensional space, where the multi-hop dependency paths are treated as a sequence of grammatical relations and modeled by a recurrent neural network. Then, we design the embedding features that consider linear context and dependency context information, for the conditional random field (CRF) based aspect term extraction. Experimental results on the SemEval datasets show that, (1) with only embedding features, we can achieve state-of-the-art results; (2) our embedding method which incorporates the syntactic information among words yields better performance than other representative ones in aspect term extraction.

IJCAI Conference 2015 Conference Paper

Constrained Information-Theoretic Tripartite Graph Clustering to Identify Semantically Similar Relations

  • Chenguang Wang
  • Yangqiu Song
  • Dan Roth
  • Chi Wang
  • Jiawei Han
  • Heng Ji
  • Ming Zhang

In knowledge bases or information extraction results, differently expressed relations can be semantically similar (e. g. , (X, wrote, Y) and (X, ’s written work, Y)). Therefore, grouping semantically similar relations into clusters would facilitate and improve many applications, including knowledge base completion, information extraction, information retrieval, and more. This paper formulates relation clustering as a constrained tripartite graph clustering problem, presents an efficient clustering algorithm and exhibits the advantage of the constrained framework. We introduce several ways that provide side information via must-link and cannotlink constraints to improve the clustering results. Different from traditional semi-supervised learning approaches, we propose to use the similarity of relation expressions and the knowledge of entity types to automatically construct the constraints for the algorithm. We show improved relation clustering results on two datasets extracted from human annotated knowledge base (i. e. , Freebase) and open information extraction results (i. e. , ReVerb data).

IJCAI Conference 2015 Conference Paper

Opportunities or Risks to Reduce Labor in Crowdsourcing Translation? Characterizing Cost versus Quality via a PageRank-HITS Hybrid Model

  • Rui Yan
  • Yiping Song
  • Cheng-Te Li
  • Ming Zhang
  • Xiaohua Hu

Crowdsourcing machine translation shows advantages of lower expense in money to collect the translated data. Yet, when compared with translation by trained professionals, results collected from non-professional translators might yield lowquality outputs. A general solution for crowdsourcing practitioners is to employ a large amount of labor force to gather enough redundant data and then solicit from it. Actually we can further save money by avoid collecting bad translations. We propose to score Turkers by their authorities during observation, and then stop hiring the unqualified Turkers. In this way, we bring both opportunities and risks in crowdsourced translation: we can make it cheaper than cheaper while we might suffer from quality loss. In this paper, we propose a graphbased PageRank-HITS Hybrid model to distinguish authoritative workers from unreliable ones. The algorithm captures the intuition that good translation and good workers are mutually reinforced iteratively in the proposed frame. We demonstrate the algorithm will keep the performance while reduce work force and hence cut cost. We run experiments on the NIST 2009 Urdu-to-English evaluation set with Mechanical Turk, and quantitatively evaluate the performance in terms of BLEU score, Pearson correlation and real money.

AAAI Conference 2015 Conference Paper

Spectral Label Refinement for Noisy and Missing Text Labels

  • Yangqiu Song
  • Chenguang Wang
  • Ming Zhang
  • Hailong Sun
  • Qiang Yang

With the recent growth of online content on the Web, there have been more user generated data with noisy and missing labels, e. g. , social tags and voted labels from Amazon’s Mechanical Turks. Most of machine learning methods, which require accurate label sets, could not be trusted when the label sets were yet unreliable. In this paper, we provide a text label refinement algorithm to adjust the labels for such noisy and missing labeled datasets. We assume that the labeled sets can be refined based on the labels with certain confidence, and the similarity between data being consistent with the labels. We propose a label smoothness ratio criterion to measure the smoothness of the labels and the consistency between labels and data. We demonstrate the effectiveness of the label refining algorithm on eight labeled document datasets, and validate that the results are useful for generating better labels.

AAAI Conference 2014 Conference Paper

SUIT: A Supervised User-Item Based Topic Model for Sentiment Analysis

  • Fangtao Li
  • Sheng Wang
  • Shenghua Liu
  • Ming Zhang

Probabilistic topic models have been widely used for sentiment analysis. However, most of existing topic methods only model the sentiment text, but do not consider the user, who expresses the sentiment, and the item, which the sentiment is expressed on. Since different users may use different sentiment expressions for different items, we argue that it is better to incorporate the user and item information into the topic model for sentiment analysis. In this paper, we propose a new Supervised User-Item based Topic model, called SUIT model, for sentiment analysis. It can simultaneously utilize the textual topic and latent user-item factors. Our proposed method uses the tensor outer product of text topic proportion vector, user latent factor and item latent factor to model the sentiment label generalization. Extensive experiments are conducted on two datasets: review dataset and microblog dataset. The results demonstrate the advantages of our model. It shows significant improvement compared with supervised topic models and collaborative filtering methods.

AAAI Conference 2011 Conference Paper

Collaborative Users’ Brand Preference Mining across Multiple Domains from Implicit Feedbacks

  • Jian Tang
  • Jun Yan
  • Lei Ji
  • Ming Zhang
  • Shaodan Guo
  • Ning Liu
  • Xianfang Wang
  • Zheng Chen

Advanced e-applications require comprehensive knowledge about their users’ preferences in order to provide accurate personalized services. In this paper, we propose to learn users’ preferences to product brands from their implicit feedbacks such as their searching and browsing behaviors in user Web browsing log data. The user brand preference learning problem is challenge since (1) the users’ implicit feedbacks are extremely sparse in various product domains; and (2) we can only observe positive feedbacks from users’ behaviors. In this paper, we propose a latent factor model to collaboratively mine users’ brand preferences across multiple domains simultaneously. By collective learning, the learning processes in all the domains are mutually enhanced and hence the problem of data scarcity in each single domain can be effectively addressed. On the other hand, we learn our model with an adaption of the Bayesian personalized ranking (BPR) optimization criterion which is a general learning framework for collaborative filtering from implicit feedbacks. Experiments with both synthetic and real world datasets show that our proposed model significantly outperforms the baselines.

ICRA Conference 2008 Conference Paper

A microrobotic adherent cell injection system for investigating intracellular behavior of quantum dots

  • Wenhui Wang 0001
  • Yu Sun 0001
  • Ming Zhang
  • Robin Anderson
  • Lowell Langille
  • Warren Chan

This paper presents a semi-automated microrobotic system for adherent cell injection. Different from embryos/oocytes that have a spherical shape and regular morphology, adherent cells are flat with a thickness of a few micrometers and are highly irregular in morphology. Based on computer vision microscopy and motion control, the system coordinately controls a three-degrees-of-freedom microrobot and a precision XY stage. The microrobotic system demonstrates an injection speed of 25 endothelial cells per minute with a survival rate of 96% and a success rate of 82% (n=1012). The system has a high degree of performance consistency. It is immune to operator proficiency variations and from human fatigue, requiring a human operator to select injection destinations through computer mouse clicking as the only operator intervention. The microrobotic adherent cell injection system makes the injection of thousands of adherent cells practical and will enable our testing of intracellular behavior of semiconductive quantum dots (QDs).

v2026.09.13