Arrow Research search

Author name cluster

Dong Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

EAAI Journal 2026 Journal Article

A deep reinforcement learning approach for portfolio rebalancing with Dragon Pullback multi-stage candlestick pattern embedding

  • Yuyang Bai
  • Changsheng Zhang
  • Longhaoze Liu
  • Baiqing Sun
  • Haoxuan Sun
  • Shijia Wang
  • Dong Zhang

With the rapid advancement of artificial intelligence, deep reinforcement learning has emerged as a promising method for portfolio rebalancing. Existing deep reinforcement learning (DRL) methods for portfolio rebalancing typically rely on prices or price-based technical indicators as the state representation. However, such representation is sensitive to noise and tends to emphasize short-term and unstable price fluctuations, which often lead DRL agents to learn aggressive strategies that perform poorly in real markets with liquidity constraints, transaction costs and lot-size constraints. To address this issue, this paper proposes a deep reinforcement learning framework for portfolio rebalancing with Dragon Pullback multi-stage candlestick pattern embedding (DRL-DPMSC). The proposed approach consists of three modules. First, in the Dragon Pullback pattern capture module, Dragon Pullback patterns are efficiently captured online through a temporal-segment-based method. Then, captured patterns are modeled as a set of pattern-related features for price-trend characterization in the causal-discovery-guided feature selection module. Finally, these features are incorporated into the state representation and employed to generate portfolio rebalancing strategies via the Proximal Policy Optimization Clip model integrated with an asset-wise architecture. By introducing noise-robust Dragon Pullback multi-stage candlestick patterns that emphasize persistent price trends, DRL-DPMSC is able to make more reasonable rebalancing decisions, especially in constrained markets. To verify the effectiveness of the proposed DRL-DPMSC, an experiment is conducted on 5 test windows to compare DRL-DPMSC with six comparative methods under five different environment settings. The proposed method outperforms competitors in most cases on both the profit and risk-return metrics.

AAAI Conference 2026 Conference Paper

MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation

  • Shengwei Zhao
  • Jingwen Yao
  • Sitong Wei
  • Linhai Xu
  • Yuying Liu
  • Dong Zhang
  • Zhiqiang Tian
  • Shaoyi Du

Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments.

AAAI Conference 2026 Conference Paper

Modeling Item-Level Dynamic Variability with Residual Diffusion for Bundle Recommendation

  • Dong Zhang
  • Lin Li
  • Ming Li
  • Amran Bhuiyan
  • Meng Sun
  • Xiaohui Tao
  • Jimmy Huang

Existing solutions for bundle recommendation (BR) have achieved remarkable effectiveness for predicting the user’s preference for prebuilt bundles. However, bundle-item (B-I) affiliation will vary dynamically in real scenarios. For ex ample, a bundle themed as ‘casual outfit’ may add ‘hat’ or remove ‘watch’ due to factors such as seasonal variations, changes in user preferences or inventory adjustments. Our empirical study demonstrates that the performance of main stream BR models may fluctuate or decline under item-level variability. This paper makes the first attempt to address the above problem and proposes Residual Diffusion for Bundle Recommendation (RDiffBR) as a model-agnostic generative framework which can assist a BR model in adapting this sce nario. During the initial training of the BR model, RDiffBR employs a residual diffusion model to process the item-level bundle embeddings which are generated by the BR model to represent bundle theme via a forward-reverse process. In the inference stage, RDiffBR reverses item-level bundle em beddings obtained by the well-trained bundle model under B-I variability scenarios to generate the effective item-level bundle embeddings. In particular, the residual connection in our residual approximator significantly enhances BR mod els’ ability to generate high-quality item-level bundle embed dings. Experiments on six BRmodelsandfourpublicdatasets from different domains show that RDiffBR improves the per formance of Recall and NDCG of backbone BR models by up to 23%, while only increases training time about 4%.

TMLR Journal 2026 Journal Article

Towards Customized Knowledge Distillation for Efficient Dense Image Predictions

  • Dong Zhang
  • Pingcheng Dong
  • Long Chen
  • Kwang-Ting Cheng

It has been revealed that efficient dense image prediction (EDIP) models designed for AI chips, trained using the knowledge distillation (KD) framework, encounter two key challenges, including maintaining boundary region completeness and ensuring target region connectivity, despite their favorable real-time capacity to recognize the main object regions. In this work, we propose a customized boundary and context knowledge distillation (BCKD) method for EDIPs, which facilitates the targeted KD from large accurate teacher models to compact small student models. Specifically, the boundary distillation focuses on extracting explicit object-level boundaries from the hierarchical feature maps to enhance the student model's mask quality in boundary regions. Meanwhile, the context distillation leverages self-relations as a bridge to transfer implicit pixel-level contexts from the teacher model to the student model, ensuring strong connectivity in target regions. Our method is specifically designed for the EDIP tasks and is characterized by its simplicity and efficiency. Theoretical analysis and extensive experimental results across semantic segmentation, object detection, and instance segmentation on five representative datasets demonstrate the effectiveness of BCKD, resulting in well-defined object boundaries and smooth connecting regions.

YNIMG Journal 2026 Journal Article

Uniformity in happiness and uniqueness in sadness: Naturalistic emotional representation in major depression

  • Qingjin Liu
  • Xi Zhang
  • Jinpeng Niu
  • Kangjia Chen
  • Jie Xia
  • Yaohui He
  • Shuo Xu
  • Wei Li

Humans develop shared concepts of others' emotions to support adaptive social functioning, yet how these concepts are dynamically represented in major depressive disorder (MDD) during naturalistic movie viewing is not yet fully established. Using functional MRI, we examined patients with MDD (n = 55) and healthy controls (HCs; n = 62) as they freely viewed movie clips depicting happy and sad emotions. Neural similarity was quantified with inter-subject correlation at whole-brain, network, and regional levels, and its association with emotional traits was assessed using inter-subject representational similarity analysis. Compared with HCs, patients with MDD showed significantly reduced whole-brain similarity, particularly during sad contexts. Network analyses revealed that HCs exhibited increased similarity in the limbic network during sadness, reflecting a shared "sadness resonance," whereas patients with higher depressive severity showed widespread disruptions across visual, limbic, dorsal attention, and default mode networks. At the regional level, similarity in the inferior temporal gyrus and lateral occipital cortex was closely linked to individual differences in emotional awareness, with pronounced context- and region-specificity. These findings highlight neural decoupling and heterogeneity as core features of MDD and provide new evidence for potential biomarkers to inform risk assessment and personalized interventions.

JBHI Journal 2026 Journal Article

WN-Sleep: Modeling Whole-Night Data for Improved Sleep Staging Classification

  • Fang Zhou
  • Zhi Lu
  • Zhi Wu
  • Gaohan Ye
  • Lingjie Shu
  • Yu Pu
  • Beilei Wang
  • Dong Zhang

Sleep staging, crucial for diagnosing sleep disorders, requires precise recognition of physiological signals within 30-second epochs, a task fundamentally different from managing long-term semantic dependencies in natural language processing (NLP). Our model aims to refine the integration of local and global features for more accurate sleep stage classification. Following the American Academy of Sleep Medicine (AASM) guidelines, it focuses on rigorous intra-epoch feature extraction to ensure reliable identification of sleep stages. Moreover, our approach incorporates a global perspective by analyzing whole-night data, which is essential for handling transitional periods and ambiguities. Existing sequential modeling techniques often overlook the unique requirements of sleep staging, leading to performance declines when epochs extend beyond approximately 200. Our model addresses this by structurally processing local and global information and carefully balancing detailed intra-epoch analysis with an overarching view of sleep cycles through a gating mechanism. This gate mechanism selectively integrates long-term dependencies, optimizing the balance between local accuracy and global context. This approach represents a significant advancement over existing models, offering more accurate, reliable, and clinically relevant sleep staging. Extensive experiments on the SHHS, SleepEDF-20, and SleepEDF-78 datasets demonstrate that our method outperforms state-of-the-art approaches.

ICLR Conference 2025 Conference Paper

BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments

  • Xinghao Wang
  • Pengyu Wang 0006
  • Bo Wang 0084
  • Dong Zhang
  • Yunhua Zhou
  • Xipeng Qiu

Large language models (LLMs) have revolutionized numerous applications, yet their deployment remains challenged by memory constraints on local devices. While scaling laws have enhanced LLM capabilities, the primary bottleneck has shifted from $\textit{capability}$ to $\textit{availability}$, emphasizing the need for efficient memory management. Traditional compression methods, such as quantization, often require predefined compression ratios and separate compression processes for each setting, complicating deployment in variable memory environments. In this paper, we introduce $\textbf{BitStack}$, a novel, training-free weight compression approach that enables megabyte-level trade-offs between memory usage and model performance. By leveraging weight decomposition, BitStack can dynamically adjust the model size with minimal transmission between running memory and storage devices. Our approach iteratively decomposes weight matrices while considering the significance of each parameter, resulting in an approximately 1-bit per parameter residual block in each decomposition iteration. These blocks are sorted and stacked in storage as basic transmission units, with different quantities loaded based on current memory availability. Extensive experiments across a wide range of tasks demonstrate that, despite offering fine-grained size control, BitStack consistently matches or surpasses strong quantization baselines, particularly at extreme compression ratios. To the best of our knowledge, this is the first decomposition-based method that effectively bridges the gap to practical compression techniques like quantization. Code is available at https://github.com/xinghaow99/BitStack.

ICLR Conference 2025 Conference Paper

Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection

  • Chuhan Zhang
  • Chaoyang Zhu
  • Pingcheng Dong
  • Long Chen 0016
  • Dong Zhang

In pursuit of detecting unstinted objects that extend beyond predefined categories, prior arts of open-vocabulary object detection (OVD) typically resort to pretrained vision-language models (VLMs) for base-to-novel category generalization. However, to mitigate the misalignment between upstream image-text pretraining and downstream region-level perception, additional supervisions are indispensable, e.g., image-text pairs or pseudo annotations generated via self-training strategies. In this work, we propose CCKT-Det trained without any extra supervision. The proposed framework constructs a cyclic and dynamic knowledge transfer from language queries and visual region features extracted from VLMs, which forces the detector to closely align with the visual-semantic space of VLMs. Specifically, 1) we prefilter and inject semantic priors to guide the learning of queries, and 2) introduce a regional contrastive loss to improve the awareness of queries on novel objects. CCKT-Det can consistently improve performance as the scale of VLMs increases, all while requiring the detector at a moderate level of computation overhead. Comprehensive experimental results demonstrate that our method achieves performance gain of +2.9% and +10.2% AP_{50} over previous state-of-the-arts on the challenging COCO benchmark, both without and with a stronger teacher model.

IROS Conference 2025 Conference Paper

ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training

  • Dong Zhang
  • Jingwei Peng
  • Yuyang Jiao
  • Jiayuan Gu
  • Jingyi Yu 0001
  • Jiahao Chen

This paper presents a novel Expressive Facial Control (ExFace) method based on Diffusion Transformers, which achieves precise mapping from human facial blendshapes to bionic robot motor control. By incorporating an innovative model bootstrap training strategy, our approach not only generates high-quality facial expressions but also significantly improves accuracy and smoothness. Experimental results demonstrate that the proposed method outperforms previous methods in terms of accuracy, frames per second (FPS), and response time. Furthermore, we develop the ExFace dataset driven by human facial data. ExFace shows excellent real-time performance and natural expression rendering in applications such as robot performances and human-robot interactions, offering a new solution for bionic robot interaction.

NeurIPS Conference 2025 Conference Paper

Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph Generation

  • Lin Li
  • Chuhan ZHANG
  • Dong Zhang
  • Chong Sun
  • Chen Li
  • Long Chen

Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Existing OVSGG methods always adopt a two-stage pipeline: 1) Infusing knowledge into large-scale models via pre-training on large datasets; 2) Transferring knowledge from pre-trained models with fully annotated scene graphs during supervised fine-tuning. However, due to a lack of explicit interaction modeling, these methods struggle to distinguish between interacting and non-interacting instances of the same object category. This limitation induces critical issues in both stages of OVSGG: it generates noisy pseudo-supervision from mismatched objects during knowledge infusion, and causes ambiguous query matching during knowledge transfer. To this end, in this paper, we propose an interACtion-Centric end-to-end OVSGG framework (ACC) in an interaction-driven paradigm to minimize these mismatches. For interaction-centric knowledge infusion, ACC employs a bidirectional interaction prompt for robust pseudo-supervision generation to enhance the model's interaction knowledge. For interaction-centric knowledge transfer, ACC first adopts interaction-guided query selection that prioritizes pairing interacting objects to reduce interference from non-interacting ones. Then, it integrates interaction-consistent knowledge distillation to bolster robustness by pushing relational foreground away from the background while retaining general knowledge. Extensive experimental results on three benchmarks show that ACC achieves state-of-the-art performance, demonstrating the potential of interaction-centric paradigms for real-world applications.

ICLR Conference 2025 Conference Paper

Memory Efficient Transformer Adapter for Dense Predictions

  • Dong Zhang
  • Rui Yan
  • Pingcheng Dong
  • Kwang-Ting (Tim) Cheng

While current Vision Transformer (ViT) adapter methods have shown promising accuracy, their inference speed is implicitly hindered by inefficient memory access operations, e.g., standard normalization and frequent reshaping. In this work, we propose META, a simple and fast ViT adapter that can improve the model's memory efficiency and decrease memory time consumption by reducing the inefficient memory access operations. Our method features a memory-efficient adapter block that enables the common sharing of layer normalization between the self-attention and feed-forward network layers, thereby reducing the model's reliance on normalization operations. Within the proposed block, the cross-shaped self-attention is employed to reduce the model's frequent reshaping operations. Moreover, we augment the adapter block with a lightweight convolutional branch that can enhance local inductive biases, particularly beneficial for the dense prediction tasks, e.g., object detection, instance segmentation, and semantic segmentation. The adapter block is finally formulated in a cascaded manner to compute diverse head features, thereby enriching the variety of feature representations. Empirically, extensive evaluations on multiple representative datasets validate that META substantially enhances the predicted quality, while achieving a new state-of-the-art accuracy-efficiency trade-off. Theoretically, we demonstrate that META exhibits superior generalization capability and stronger adaptability.

IJCAI Conference 2025 Conference Paper

TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification

  • Rui Yan
  • Jin Wang
  • Hongyu Qu
  • Xiaoyu Du
  • Dong Zhang
  • Jinhui Tang
  • Tieniu Tan

Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap between modalities while the support-set cannot be tuned. To this end, we draw on each other's strengths and propose a novel framework, namely TEst-time Support-set Tuning for zero-shot Video Classification (TEST-V). It first dilates the support-set with multiple prompts (Multi-prompting Support-set Dilation, MSD) and then erodes the support-set via learnable weights to mine key cues dynamically (Temporal-aware Support-set Erosion, TSE). Specifically, i) MSD expands the support samples for each class based on multiple prompts inquired from LLMs to enrich the diversity of the support-set. ii) TSE tunes the support-set with factorized learnable weights according to the temporal prediction consistency in a self-supervised manner to dig pivotal supporting cues for each class. TEST-V achieves state-of-the-art results across four benchmarks and shows good interpretability.

NeurIPS Conference 2024 Conference Paper

SpeechAlign: Aligning Speech Generation to Human Preferences

  • Dong Zhang
  • Zhaowei Li
  • Shimin Li
  • Xin Zhang
  • Pengyu Wang
  • Yaqian zhou
  • Xipeng Qiu

Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of preference optimization to align speech outputs to human preferences is often neglected. This paper addresses this gap by first analyzing the distribution gap in codec language models, highlighting how it leads to discrepancies between the training and inference phases, which negatively affects performance. Then we explore leveraging preference optimization to bridge the distribution gap. We introduce SpeechAlign, an iterative self-improvement strategy that aligns speech language models to human preferences. SpeechAlign involves constructing a preference codec dataset contrasting golden codec tokens against synthetic tokens, followed by preference optimization to improve the codec language model. This cycle of improvement is carried out iteratively to steadily convert weak models to strong ones. Through both subjective and objective evaluations, we show that SpeechAlign can bridge the distribution gap and facilitating continuous self-improvement of the speech language model. Moreover, SpeechAlign exhibits robust generalization capabilities and works for smaller models. Demos are available at https: //0nutation. github. io/SpeechAlign. github. io/.

ICLR Conference 2024 Conference Paper

SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models

  • Xin Zhang
  • Dong Zhang
  • Shimin Li
  • Yaqian Zhou 0001
  • Xipeng Qiu

Current speech large language models build upon discrete speech representations, which can be categorized into semantic tokens and acoustic tokens. However, existing speech tokens are not specifically designed for speech language modeling. To assess the suitability of speech tokens for building speech language models, we established the first benchmark, SLMTokBench. Our results indicate that neither semantic nor acoustic tokens are ideal for this purpose. Therefore, we propose SpeechTokenizer, a unified speech tokenizer for speech large language models. SpeechTokenizer adopts the Encoder-Decoder architecture with residual vector quantization (RVQ). Unifying semantic and acoustic tokens, SpeechTokenizer disentangles different aspects of speech information hierarchically across different RVQ layers. Furthermore, We construct a Unified Speech Language Model (USLM) leveraging SpeechTokenizer. Experiments show that SpeechTokenizer performs comparably to EnCodec in speech reconstruction and demonstrates strong performance on the SLMTokBench benchmark. Also, USLM outperforms VALL-E in zero-shot Text-to-Speech tasks. Code and models are available at https://github.com/ZhangXInFD/SpeechTokenizer/.

ICLR Conference 2023 Conference Paper

A GNN-Guided Predict-and-Search Framework for Mixed-Integer Linear Programming

  • Qingyu Han
  • Linxin Yang
  • Qian Chen
  • Xiang Zhou
  • Dong Zhang
  • Akang Wang
  • Ruoyu Sun 0001
  • Xiaodong Luo

Mixed-integer linear programming (MILP) is widely employed for modeling combinatorial optimization problems. In practice, similar MILP instances with only coefficient variations are routinely solved, and machine learning (ML) algorithms are capable of capturing common patterns across these MILP instances. In this work, we combine ML with optimization and propose a novel predict-and-search framework for efficiently identifying high-quality feasible solutions. Specifically, we first utilize graph neural networks to predict the marginal probability of each variable, and then search for the best feasible solution within a properly defined ball around the predicted solution. We conduct extensive experiments on public datasets, and computational results demonstrate that our proposed framework achieves 51.1% and 9.9% performance improvements to MILP solvers SCIP and Gurobi on primal gaps, respectively.

JBHI Journal 2023 Journal Article

An Uncertainty-Aware and Sex-Prior Guided Biological Age Estimation From Orthopantomogram Images

  • Dong Zhang
  • Jing Yang
  • Shaoyi Du
  • Wenqing Bu
  • Yu-cheng Guo

Bone age, as a measure of biological age (BA), plays an important role in a variety of fields, including forensics, orthodontics, sports, and immigration. Despite its significance, accurate estimation of BA remains a challenge due to the uncertainty error between BA and chronological age (CA) caused by individual diversity and the difficult integration of multiple factors, such as sex, and identified or measured anatomical structures, into the estimation process. To address problems, we propose an uncertainty-aware and sex-prior guided biological age estimation from orthopantomogram images (OPGs), named UASP-BAE, which models uncertainty errors while setting sex dimorphism as tractive features to enhance age-related specific features, aiming to improve the accuracy of BA estimation. Furthermore, considering the global relevance of the anatomic structure, such as the mandible, teeth, maxillary sinus, etc. , a cross-attention module based on CNN and self-attention is proposed to mine the local texture and global semantic features of OPGs. Moreover, we design a novel age composition loss by cross-entropy, probability bias, and regression functions, aiming at evaluating BA's uncertainty errors and results to obtain an accurate and robust model. On 10703 OPGs from 5. 00 to 25. 00 years of age, our model had a best MAE value of 0. 8005 years and higher than the comparison popular algorithms, which also demonstrates the method's potential for improved accuracy in BA estimation.

JBHI Journal 2023 Journal Article

BMAnet: Boundary Mining With Adversarial Learning for Semi-Supervised 2D Myocardial Infarction Segmentation

  • Chenchu Xu
  • Yifei Wang
  • Dong Zhang
  • Longfei Han
  • Yanping Zhang
  • Jie Chen
  • Shuo Li

Automatic segmentation of myocardial infarction (MI) regions in late gadolinium-enhanced cardiac magnetic resonance images is an essential step in the computed diagnosis of myocardial infarction. Most of the current myocardial infarction region segmentation methods are based on fully supervised deep learning. However, cardiologists' annotation of myocardial infarction regions in cardiac magnetic resonance images during the diagnosis process is time-consuming and expensive. This paper proposes a semi-supervised myocardial infarction segmentation. It consists of two models: 1) a boundary mining model and 2) an adversarial learning model. The boundary mining model can solve the boundary ambiguity problem by enlarging the gap between the foreground and background features, thus segmenting the myocardial infarction region accurately. The adversarial learning model can make the boundary mining model learn from additional unlabeled data by evaluating the segmentation performance and providing pseudo supervision, which significantly increases the robustness of the boundary mining model. We conduct extensive experiments on an in-house myocardial magnetic resonance dataset. The experimental results on six evaluation metrics demonstrate that our method achieves excellent results in myocardial infarction segmentation and outperforms the state-of-the-art semi-supervised methods.

YNIMG Journal 2023 Journal Article

Deep learning-assisted identification and quantification of aneurysmal subarachnoid hemorrhage in non-contrast CT scans: Development and external validation of Hybrid 2D/3D UNet

  • Ping Hu
  • Haizhu Zhou
  • Tengfeng Yan
  • Hongping Miu
  • Feng Xiao
  • Xinyi Zhu
  • Lei Shu
  • Shuang Yang

Accurate stroke assessment and consequent favorable clinical outcomes rely on the early identification and quantification of aneurysmal subarachnoid hemorrhage (aSAH) in non-contrast computed tomography (NCCT) images. However, hemorrhagic lesions can be complex and difficult to distinguish manually. To solve these problems, here we propose a novel Hybrid 2D/3D UNet deep-learning framework for automatic aSAH identification and quantification in NCCT images. We evaluated 1824 consecutive patients admitted with aSAH to four hospitals in China between June 2018 and May 2022. Accuracy and precision, Dice scores and intersection over union (IoU), and interclass correlation coefficients (ICC) were calculated to assess model performance, segmentation performance, and correlations between automatic and manual segmentation, respectively. A total of 1355 patients with aSAH were enrolled: 931, 101, 179, and 144 in four datasets, of whom 326 were scanned with Siemens, 640 with Philips, and 389 with GE Medical Systems scanners. Our proposed deep-learning method accurately identified (accuracies 0.993-0.999) and segmented (Dice scores 0.550-0.897) hemorrhage in both the internal and external datasets, even combinations of hemorrhage subtypes. We further developed a convenient AI-assisted platform based on our algorithm to assist clinical workflows, whose performance was comparable to manual measurements by experienced neurosurgeons (ICCs 0.815-0.957) but with greater efficiency and reduced cost. While this tool has not yet been prospectively tested in clinical practice, our innovative hybrid network algorithm and platform can accurately identify and quantify aSAH, paving the way for fast and cheap NCCT interpretation and a reliable AI-based approach to expedite clinical decision-making for aSAH patients.

IJCAI Conference 2023 Conference Paper

Discrepancy-Guided Reconstruction Learning for Image Forgery Detection

  • Zenan Shi
  • Haipeng Chen
  • Long Chen
  • Dong Zhang

In this paper, we propose a novel image forgery detection paradigm for boosting the model learning capacity on both forgery-sensitive and genuine compact visual patterns. Compared to the existing methods that only focus on the discrepant-specific patterns (\eg, noises, textures, and frequencies), our method has a greater generalization. Specifically, we first propose a Discrepancy-Guided Encoder (DisGE) to extract forgery-sensitive visual patterns. DisGE consists of two branches, where the mainstream backbone branch is used to extract general semantic features, and the accessorial discrepant external attention branch is used to extract explicit forgery cues. Besides, a Double-Head Reconstruction (DouHR) module is proposed to enhance genuine compact visual patterns in different granular spaces. Under DouHR, we further introduce a Discrepancy-Aggregation Detector (DisAD) to aggregate these genuine compact visual patterns, such that the forgery detection capability on unknown patterns can be improved. Extensive experimental results on four challenging datasets validate the effectiveness of our proposed method against state-of-the-art competitors.

YNICL Journal 2021 Journal Article

Hippocampal subfield and anterior-posterior segment volumes in patients with sporadic amyotrophic lateral sclerosis

  • Shuangwu Liu
  • Qingguo Ren
  • Gaolang Gong
  • Yuan Sun
  • Bing Zhao
  • Xiaotian Ma
  • Na Zhang
  • Suyu Zhong

Neuroimaging studies of hippocampal volumes in patients with amyotrophic lateral sclerosis (ALS) have reported inconsistent results. Our aims were to demonstrate that such discrepancies are largely due to atrophy of different regions of the hippocampus that emerge in different disease stages of ALS and to explore the existence of co-pathology in ALS patients. We used the well-validated King's clinical staging system for ALS to classify patients into different disease stages. We investigated in vivo hippocampal atrophy patterns across subfields and anterior-posterior segments in different King's stages using structural MRI in 76 ALS patients and 94 health controls (HCs). The thalamus, corticostriatal tract and perforant path were used as structural controls to compare the sequence of alterations between these structures and the hippocampal subfields. Compared with HCs, ALS patients at King's stage 1 had lower volumes in the bilateral posterior subiculum and presubiculum; ALS patients at King's stage 2 exhibited lower volumes in the bilateral posterior subiculum, left anterior presubiculum and left global hippocampus; ALS patients at King's stage 3 showed significantly lower volumes in the bilateral posterior subiculum, dentate gyrus and global hippocampus. Thalamic atrophy emerged at King's stage 3. White matter tracts remained normal in a subset of ALS patients. Our study demonstrated that the pattern of hippocampal atrophy in ALS patients varies greatly across King's stages. Future studies in ALS patients that focus on the hippocampus may help to further clarify possible co-pathologies in ALS.

AAAI Conference 2021 Conference Paper

Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual Guidance

  • Dong Zhang
  • Suzhong Wei
  • Shoushan Li
  • Hanqian Wu
  • Qiaoming Zhu
  • Guodong Zhou

Multi-modal named entity recognition (MNER) aims to discover named entities in free text and classify them into predefined types with images. However, dominant MNER models do not fully exploit fine-grained semantic correspondences between semantic units of different modalities, which have the potential to refine multi-modal representation learning. To deal with this issue, we propose a unified multi-modal graph fusion (UMGF) approach for MNER. Specifically, we first represent the input sentence and image using a unified multi-modal graph, which captures various semantic relationships between multi-modal semantic units (words and visual objects). Then, we stack multiple graph-based multi-modal fusion layers that iteratively perform semantic interactions to learn node representations. Finally, we achieve an attentionbased multi-modal representation for each word and perform entity labeling with a CRF decoder. Experimentation on the two benchmark datasets demonstrates the superiority of our MNER model.

AAAI Conference 2021 Conference Paper

Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing

  • Dong Zhang
  • Xincheng Ju
  • Wei Zhang
  • Junhui Li
  • Shoushan Li
  • Qiaoming Zhu
  • Guodong Zhou

As an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus on multi-modal emotion recognition in a multilabel scenario. In this scenario, we consider not only the label-to-label dependency, but also the feature-to-label and modality-to-label dependencies. Particularly, we propose a heterogeneous hierarchical message passing network to effectively model above dependencies. Furthermore, we propose a new multi-modal multi-label emotion dataset based on partial time-series content to show predominant generalization of our model. Detailed evaluation demonstrates the effectiveness of our approach.

NeurIPS Conference 2020 Conference Paper

Causal Intervention for Weakly-Supervised Semantic Segmentation

  • Dong Zhang
  • Hanwang Zhang
  • Jinhui Tang
  • Xian-Sheng Hua
  • Qianru Sun

We present a causal inference framework to improve Weakly-Supervised Semantic Segmentation (WSSS). Specifically, we aim to generate better pixel-level pseudo-masks by using only image-level labels -- the most crucial step in WSSS. We attribute the cause of the ambiguous boundaries of pseudo-masks to the confounding context, e. g. , the correct image-level classification of "horse" and "person" may be not only due to the recognition of each instance, but also their co-occurrence context, making the model inspection (e. g. , CAM) hard to distinguish between the boundaries. Inspired by this, we propose a structural causal model to analyze the causalities among images, contexts, and class labels. Based on it, we develop a new method: Context Adjustment (CONTA), to remove the confounding bias in image-level classification and thus provide better pseudo-masks as ground-truth for the subsequent segmentation model. On PASCAL VOC 2012 and MS-COCO, we show that CONTA boosts various popular WSSS methods to new state-of-the-arts.

IROS Conference 2020 Conference Paper

Dual-SLAM: A framework for robust single camera navigation

  • Huajian Huang
  • Wen-Yan Lin
  • Siying Liu
  • Dong Zhang
  • Sai-Kit Yeung

SLAM (Simultaneous Localization And Mapping) seeks to provide a moving agent with real-time self-localization. To achieve real-time speed, SLAM incrementally propagates position estimates. This makes SLAM fast but also makes it vulnerable to local pose estimation failures. As local pose estimation is ill-conditioned, local pose estimation failures happen regularly, making the overall SLAM system brittle. This paper attempts to correct this problem. We note that while local pose estimation is ill-conditioned, pose estimation over longer sequences is well-conditioned. Thus, local pose estimation errors eventually manifest themselves as mapping inconsistencies. When this occurs, we save the current map and activate two new SLAM threads. One processes incoming frames to create a new map and the other, recovery thread, backtracks to link new and old maps together. This creates a Dual-SLAM framework that maintains real-time performance while being robust to local pose estimation failures. Evaluation on benchmark datasets shows Dual-SLAM can reduce failures by a dramatic 88%.

AAAI Conference 2020 Conference Paper

Rethinking the Bottom-Up Framework for Query-Based Video Localization

  • Long Chen
  • Chujie Lu
  • Siliang Tang
  • Jun Xiao
  • Dong Zhang
  • Chilie Tan
  • Xiaolin Li

In this paper, we focus on the task query-based video localization, i. e. , localizing a query in a long and untrimmed video. The prevailing solutions for this problem can be grouped into two categories: i) Top-down approach: It pre-cuts the video into a set of moment candidates, then it does classification and regression for each candidate; ii) Bottom-up approach: It injects the whole query content into each video frame, then it predicts the probabilities of each frame as a ground truth segment boundary (i. e. , start or end). Both two frameworks have respective shortcomings: the top-down models suffer from heavy computations and they are sensitive to the heuristic rules, while the performance of bottom-up models is behind the performance of top-down counterpart thus far. However, we argue that the performance of bottom-up framework is severely underestimated by current unreasonable designs, including both the backbone and head network. To this end, we design a novel bottom-up model: Graph-FPN with Dense Predictions (GDP). For the backbone, GDP firstly generates a frame feature pyramid to capture multi-level semantics, then it utilizes graph convolution to encode the plentiful scene relationships, which incidentally mitigates the semantic gaps in the multi-scale feature pyramid. For the head network, GDP regards all frames falling in the ground truth segment as the foreground, and each foreground frame regresses the unique distances from its location to bi-directional boundaries. Extensive experiments on two challenging query-based video localization tasks (natural language video localization and video relocalization), involving four challenging benchmarks (TACoS, Charades-STA, ActivityNet Captions, and Activity- VRL), have shown that GDP surpasses the state-of-the-art top-down models.

IJCAI Conference 2019 Conference Paper

Modeling both Context- and Speaker-Sensitive Dependence for Emotion Detection in Multi-speaker Conversations

  • Dong Zhang
  • Liangqing Wu
  • Changlong Sun
  • Shoushan Li
  • Qiaoming Zhu
  • Guodong Zhou

Recently, emotion detection in conversations becomes a hot research topic in the Natural Language Processing community. In this paper, we focus on emotion detection in multi-speaker conversations instead of traditional two-speaker conversations in existing studies. Different from non-conversation text, emotion detection in conversation text has one specific challenge in modeling the context-sensitive dependence. Besides, emotion detection in multi-speaker conversations endorses another specific challenge in modeling the speaker-sensitive dependence. To address above two challenges, we propose a conversational graph-based convolutional neural network. On the one hand, our approach represents each utterance and each speaker as a node. On the other hand, the context-sensitive dependence is represented by an undirected edge between two utterances nodes from the same conversation and the speaker-sensitive dependence is represented by an undirected edge between an utterance node and its speaker node. In this way, the entire conversational corpus can be symbolized as a large heterogeneous graph and the emotion detection task can be recast as a classification problem of the utterance nodes in the graph. The experimental results on a multi-modal and multi-speaker conversation corpus demonstrate the great effectiveness of the proposed approach.

NeurIPS Conference 2005 Conference Paper

Learning Influence among Interacting Markov Chains

  • Dong Zhang
  • Daniel Gatica-Perez
  • Samy Bengio
  • Deb Roy

We present a model that learns the influence of interacting Markov chains within a team. The proposed model is a dynamic Bayesian network (DBN) with a two-level structure: individual-level and group-level. Individual level models actions of each player, and the group-level models actions of the team as a whole. Experiments on synthetic multi-player games and a multi-party meeting corpus show the effectiveness of the proposed model.

ICRA Conference 2001 Conference Paper

The Switched Reluctance Motor Drive for the Direct-Drive Joint of the Robot

  • Hao Chen
  • Dong Zhang

The paper presents the principle of decoupling control of the phase voltage in the switched reluctance motor drive for the direct-drive joint of the robot. The motor drive system elements, such as the structure of the three-phase 6/10 structure switched reluctance motor and the rotor position, the main circuit topology of the three-phase bifilar winding power converter and the pulse width modulation control strategy, are described. The mathematical models of the main circuit of the power converter are also presented. The optimum range of the turn-on and turn-off angles of the main switches in the power converter are given by the criterion of reducing the pulsation of the output torque with a 2D finite element electromagnetic field calculation of the motor and the nonlinear simulation of the main circuit of the power converter with the control strategy.

v2026.09.13