Arrow Research search

Author name cluster

Mingyang Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

JBHI Journal 2026 Journal Article

DAFF-SNN: Dual Attention-Driven and Feature Fusion-Based Spiking Neural Network for Epilepsy Detection Based on Electroencephalogram

  • Tao Zhang
  • Lanqi He
  • Dingguo Zhang
  • Mingyang Li
  • Zhiyong Chang

Electroencephalogram (EEG) signals provide rich spatiotemporal brain features crucial for epilepsy diagnosis. Though traditional deep neural networks combining attention mechanisms excel in feature extraction, these models do not fully emulate the neural processing efficiency of the brain and they fall short in fully harnessing the spatiotemporal dynamics of EEG data. This study introduces the dual attention-driven and feature fusion-based spiking neural network (DAFF-SNN), which synergizes the spatiotemporal attention mechanism with SNNs’ inherent temporal processing prowess to enhance the precision and efficiency for epilepsy detection. The DAFF-SNN utilizes an adaptive spiking fusion module (ASFM) for feature integration, optimizing the feature fusion process by exploiting the spatiotemporal complementarity of EEG data through a spike-driven strategy. Evaluations on the Bonn, Beriut, CHB-MIT, and Siena datasets demonstrate DAFF-SNN’s high accuracies (100%, 97. 6%, 98. 2%, 99. 6% ), achieving comparable or superior performance to state-of-the-art ANN methods, highlighting its efficient epilepsy detection capability.

YNIMG Journal 2026 Journal Article

Emergence of functional topography in the neonatal white matter

  • Yongxuan Xu
  • Yufeng Xu
  • Junrui Zhang
  • Wenjie Dou
  • Mingyang Li
  • Yucen Sheng
  • Weihao Zheng
  • Baoming Li

The functional architecture of the cerebral cortex is well characterized by a hierarchical organization. Recent studies suggest that white matter (WM) demonstrates intrinsic functional dynamics comparable to cortical regions. However, how these patterns of WM functional organization initially emerge during early human development remains poorly understood. Using multimodal MRI data from the Developing Human Connectome Project (dHCP) comprising 399 infants (348 term-born, 51 preterm-born), we investigate WM functional gradients in the neonatal brain and assess their prognostic relevance. We show that neonatal WM exhibits a sophisticated, gray matter-like functional topography, characterized by principal gradients along sensorimotor-to-visual and unimodal-to-transmodal axes. These gradients demonstrate rapid, tract-specific developmental changes during the perinatal period and show significant associations with underlying myelination, indexed by the T1w/T2w ratio. Preterm birth distinctly disrupts this nascent functional architecture, notably reducing gradient values in the right corticospinal tract and increasing gradient values in the bilateral cingulum cingulate and forceps major. Importantly, we demonstrate that neonatal WM gradient features significantly predict language outcomes at 18 months, specifically within the preterm-born cohort. These findings establish neonatal WM functional organization as a crucial aspect of early brain development and highlight its potential utility as a biomarker for identifying infants at heightened risk for adverse neurodevelopmental outcomes.

EAAI Journal 2026 Journal Article

Image-plane geometric decoding for view-invariant indoor scene reconstruction

  • Mingyang Li
  • Yimeng Fan
  • Changsong Liu
  • Lixue Xu
  • Xin Wang
  • Yanyan Liu
  • Wei Zhang

Volume-based indoor scene reconstruction offers superior generalization and real-time potential. However, existing frameworks rely on weak multi-view geometric constraints, leading to quality degradation as input views decrease. In sparse-view scenarios, these methods often exhibit geometric fragmentation due to the lack of robust priors. To address this, we propose Image-Plane Geometric Decoding Reconstruction (IPDRecon) pipeline, a framework integrating geometric optical principles as inductive bias to systematically exploit single-view spatial information for view-invariant reconstruction. Our approach establishes a structured geometric constraint mechanism through three synergistic modules: the Pixel-level Confidence Encoder (PCE) leverages state–space modeling with diffuse reflection principles to extract distance and position awareness; the Affine Compensation Module (ACM) enforces rigid geometric constraints via affine invariance, enabling accurate recovery of complex structures under sparse views; and the Image-Plane Spatial Decoder (IPSD) employs a multi-source geometric prior fusion strategy to transform traditional back-projection into geometry-aware spatial encoding. Extensive experiments on benchmark datasets (ScanNet V2) demonstrate exceptional stability, achieving 79. 7% Precision and a 0. 722 harmonic mean of precision and recall (F-score). In robustness evaluations averaged on a per-scene basis across the validation set, our method shows remarkable resilience when reducing views from 100 to 60. It maintains a 99. 7% mean performance retention rate, with a per-scene coefficient of variation of 0. 24% and a maximum performance drop of only 0. 42%. These results confirm that our physics-guided approach provides a robust solution for high-fidelity reconstruction in view-limited applications.

AAAI Conference 2026 Conference Paper

Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems

  • Haowei Wang
  • Rupeng Zhang
  • Junjie Wang
  • Mingyang Li
  • Yuekai Huang
  • Dandan Wang
  • Qing Wang

Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by retrieving relevant documents from external corpora before generating responses. This approach significantly expands LLM capabilities by leveraging vast, up-to-date external knowledge. However, this reliance on external knowledge makes RAG systems vulnerable to corpus poisoning attacks that manipulate generated outputs via poisoned document injection. Existing poisoning attack strategies typically treat the retrieval and generation stages as disjointed, limiting their effectiveness. We propose Joint-GCG, the first framework to unify gradient-based attacks across both retriever and generator models through three innovations: (1) Cross-Vocabulary Projection for aligning embedding spaces, (2) Gradient Tokenization Alignment for synchronizing token-level gradient signals, and (3) Adaptive Weighted Fusion for dynamically balancing attacking objectives. Evaluations demonstrate that Joint-GCG achieves at most 25% and an average of 5% higher attack success rate than previous methods across multiple retrievers and generators. While optimized under a white-box assumption, the generated poisons show unprecedented transferability to unseen models. Joint-GCG's innovative unification of gradient-based attacks across retrieval and generation stages fundamentally reshapes our understanding of vulnerabilities within RAG systems.

YNIMG Journal 2026 Journal Article

Maturation and reorganization of structural connectivity in infants within half a year

  • Tingting Liu
  • Mingyang Li
  • Yuqing You
  • Hongxi Zhang
  • Ying Lv
  • Chai Ji
  • Yuting Li
  • Dan Wu

Although brain networks have been extensively investigated, the structural network refinement in the early postnatal period remains under-researched, a period during which axons undergo overproduction, elimination, and myelination, which may lead to alterations in the topology of intercortical connections. A total of 104 preterm infants with few complications were enrolled in the current study. The whole-brain tractography weighted by fiber density was performed for each participant using the second-order integration over the fiber orientation distribution method, which was then sifted to 1 million streamlines. The successive changes in the structural connectivity of infants were examined. The findings indicated a notable improvement in integration and segregation, characterized by a rapid initial increase in small-worldness followed by a deceleration, which is associated with the differential maturation of short- and long-range white matter. Global clustering coefficients increased with age, while node degree exhibited variability: frontal regions showed an increase, whereas temporoparietal-occipital regions demonstrated a decrease, indicating earlier maturation of the latter. Hemispheric hub edges, through short-range white matter connections between adjacent cortices, exhibited increased regularity and symmetry, potentially attributable to the earlier maturation of short-range fibers. Age-related changes in modularity and number indicate an increase in structural module segregation, characterized by declining modularity and stable composition, which reflects a distinct phase of connectivity reorganization. This study improves our understanding of early brain network development by identifying key topological maturation trends in the structural brain networks of early infants.

EAAI Journal 2025 Journal Article

A hybrid architecture of sparse convolutional neural network-transformer for enhanced spatial-geometric feature learning in surface reconstruction

  • Mingyang Li
  • Wei Zhang
  • Yanyan Liu
  • Xiang Feng
  • Changsong Liu
  • Yimeng Fan
  • Lixue Xu

Learning-based methods have garnered significant attention in indoor scene reconstruction tasks. However, researchers have often overlooked the crucial role of the surface prediction stage. Our study specifically focuses on this phase. According to our experiments and analysis, this phase primarily addresses spatial voxel occupancy and geometric structure maintenance. Simple structural designs are insufficient to effectively solve these problems. To address these challenges, we propose a hybrid model, which combines the strengths of Convolution Neural Networks and Transformer architectures for fine reconstruction. Additionally, we introduce several new techniques, including the Sparse Positional Attention mechanism, Sparse Channel Decoding Block, and Mixed Feature Fusion mechanism. These techniques, leveraging the characteristics of sparse computation, enhance feature utilization in both spatial and channel dimensions. With limited training and testing resources, our network achieves optimal results on the ScanNet dataset, improving precision and F-score by 2. 1% and 1. 6%, respectively, and reducing the Chamfer distance to 0. 055 m. To our knowledge, our model is the first use of hybrid structures in the surface prediction phase of an indoor scene reconstruction task. Moreover, we hope that our design and analysis can provide a new paradigm for task network design in this phase.

AAMAS Conference 2025 Conference Paper

Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization

  • Qian Kou
  • Mingyang Li
  • Zeyang Liu
  • Long Qian
  • Zhuoran Chen
  • Lipeng Wan
  • Xingyu Chen
  • Xuguang Lan

Multi-agent Preference-Based Reinforcement Learning (MAPbRL) is promising in offline policy learning by leveraging human preferences to replace complex manual reward designing. Current MAPbRL methods use complicated structures to realize better reward modeling with off-the-shelf MARL algorithms and obtain the joint policy based on it. However, it faces a severe preference-behavior mismatch problem stemming from the instability of RL training and global-local preference inconsistency datasets in offline MARL, resulting in potential suboptimal policy convergence. To address this problem, we propose Agent-aware Multi-Agent Direct Preference Optimization (AMADPO) by utilizing a multi-agent preference predictor to guide agent-aware direct optimization from imbalanced preference labels, which can learn coordination policy from both positive and negative segments. Experimental results in SMAC environment show substantial improvements in global-local preference inconsistency datasets, demonstrating the effectiveness of AMADPO in solving the preference-behavior mismatch problem.

AAAI Conference 2025 Conference Paper

ProtCLIP: Function-Informed Protein Multi-Modal Learning

  • Hanjing Zhou
  • Mingze Yin
  • Wei Wu
  • Mingyang Li
  • Kun Fu
  • Jintai Chen
  • Jian Wu
  • Zheng Wang

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.

AAAI Conference 2025 Conference Paper

Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer

  • Mingze Yin
  • Hanjing Zhou
  • Yiheng Zhu
  • Jialu Wu
  • Wei Wu
  • Mingyang Li
  • Kun Fu
  • Zheng Wang

Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery.

JBHI Journal 2025 Journal Article

TDSFE-Net: A Temporal Dual-Stream Feature Extraction Network for Depression Detection From EEG

  • Mingyang Li
  • Zhiwei Wang
  • Xi Yang
  • Tao Zhang

Early detection and diagnosis are critical for effective depression management. Although electroence-phalography (EEG) can provide an objective basis for the auxiliary diagnosis of depression, decoding depression-related brain activity from EEG is a highly challenging task due to the inherent complexity, dynamism, and non-linearity. Therefore, this study introduces a novel temporal dual-stream feature extraction network (TDSFE-Net) that incorporates multiple attention mechanisms. Specially, we first develop a dynamic fusion weight based local-global attention mechanism into the hierarchiclal temporal-separable convolutional network (TSCN) to automatically capture the temporal dynamic characteristics of the EEG signal. Subsequently, a channel-wise module is designed to reveal the key temporal information in spatial dimensions. Finally, a softmax with full conected layer is used as classifier. The TDSFE-Net achieved impressive classification accuracies of 98. 72%, 96. 91%, and 99. 53% on the MODMA, HUSM, and Hospital datasets, respectively. In addition, this study also reveals the pattern of correlation between the activity of specific brain regions and depression, providing a new perspective and scientific basis for discovering biomarkers and studying the neural mechanisms of depression.

NeurIPS Conference 2025 Conference Paper

Universal Visuo-Tactile Video Understanding for Embodied Interaction

  • Yifan Xie
  • Mingyang Li
  • Shoujie Li
  • Xingting Li
  • Guangyu Chen
  • Fei Ma
  • Fei Yu
  • Wenbo Ding

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model that enables universal Visuo-Tactile Video (VTV) understanding, bridging the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150, 000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, and scenario-based decision-making. Extensive experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile reasoning tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.

AAAI Conference 2024 Conference Paper

A New Benchmark and Model for Challenging Image Manipulation Detection

  • Zhenfei Zhang
  • Mingyang Li
  • Ming-Ching Chang

The ability to detect manipulation in multimedia data is vital in digital forensics. Existing Image Manipulation Detection (IMD) methods are mainly based on detecting anomalous features arisen from image editing or double compression artifacts. All existing IMD techniques encounter challenges when it comes to detecting small tampered regions from a large image. Moreover, compression-based IMD approaches face difficulties in cases of double compression of identical quality factors. To investigate the State-of-The-Art (SoTA) IMD methods in those challenging conditions, we introduce a new Challenging Image Manipulation Detection (CIMD) benchmark dataset, which consists of two subsets, for evaluating editing-based and compression-based IMD methods, respectively. The dataset images were manually taken and tampered with high-quality annotations. In addition, we propose a new two-branch network model based on HRNet that can better detect both the image-editing and compression artifacts in those challenging conditions. Extensive experiments on the CIMD benchmark show that our model significantly outperforms SoTA IMD methods on CIMD. The dataset is available at: https://github.com/ZhenfeiZ/CIMD.

YNIMG Journal 2024 Journal Article

Age-dependent functional development pattern in neonatal brain: An fMRI-based brain entropy study

  • Zhiyong Zhao
  • Yifan Shuai
  • Yihan Wu
  • Xinyi Xu
  • Mingyang Li
  • Dan Wu

The relationship between brain entropy (BEN) and early brain development has been established through animal studies. However, it remains unclear whether the BEN can be used to identify age-dependent functional changes in human neonatal brains and the genetic underpinning of the new neuroimaging marker remains to be elucidated. In this study, we analyzed resting-state fMRI data from the Developing Human Connectome Project, including 280 infants who were scanned at 37.5-43.5 weeks postmenstrual age. The BEN maps were calculated for each subject, and a voxel-wise analysis was conducted using a general linear model to examine the effects of age, sex, and preterm birth on BEN. Additionally, we evaluated the correlation between regional BEN and gene expression levels. Our results demonstrated that the BEN in the sensorimotor-auditory and association cortices, along the 'S-A' axis, was significantly positively correlated with postnatal age (PNA), and negatively correlated with gestational age (GA), respectively. Meanwhile, the BEN in the right rolandic operculum correlated significantly with both GA and PNA. Preterm-born infants exhibited increased BEN values in widespread cortical areas, particularly in the visual-motor cortex, when compared to term-born infants. Moreover, we identified five BEN-related genes (DNAJC12, FIG4, STX12, CETN2, and IRF2BP2), which were involved in protein folding, synaptic vesicle transportation and cell division. These findings suggest that the fMRI-based BEN can serve as an indicator of age-dependent brain functional development in human neonates, which may be influenced by specific genes.

JBHI Journal 2024 Journal Article

An Improved Statistical Modeling Approach to Individual Anticholinergic Drug Use Trend Analysis

  • Zhouyang Lou
  • Mingyang Li
  • Nan Kong
  • Noll L. Campbell
  • Wanzhu Tu

Anticholinergic (AC) drugs are commonly prescribed to older adults for treating diseases and chronic conditions, such as chronic obstructive pulmonary disease, urinary incontinence, gastrointestinal disorder, or simply pain and allergy. The high prevalence of AC drug use can have a detrimental effect on the mental health of older adults. We aim to improve the prediction of future trends of AC drug use at the individual level, with pharmacy refill data. The individual drug use data presents challenges in the modeling, such as data being discrete-valued with excess zeros and having significant unobserved heterogeneity in the trend pattern. To address these challenges, we propose a statistical model of hierarchical structure and an EM scheme for the model parameter estimation. We evaluate the proposed modeling approach through a numerical study with synthetic data and a case study with real-world pharmacy refill data. The simulation study show that our analysis method outperforms the existing ones (e. g. , reducing MSE significantly), particularly in terms of accurately predicting the trend pattern. The real-world case study further verifies the out-performance and demonstrate the advantageous features of our method. We expect the prediction tool developed based on our study can assist pharmacists' decision on initiating or strengthening behavioral interventions with the hope of discontinuing AC drug misuse.

EAAI Journal 2024 Journal Article

An intelligent decision support framework for nursing home resource planning with enhanced heterogeneous service demand modeling

  • Xuxue Sun
  • Nan Kong
  • Weiping Ding
  • Ying Li
  • Nazmus Sakib
  • Hao Zeng
  • Hongdao Meng
  • Chris Masterson

Demand-based nursing home resource planning is of great importance to ensure adequate resources (e. g. , beds and staffs) available to provide care services with desired quality, yet challenging. The challenge mainly lies in modeling heterogeneous demand of nursing home residents, reflected by various individual characteristics, diverse dwelling duration with multiple competing discharge dispositions, and diverse daily service need. Existing studies often assumed a homogeneous population of patients and neglected the complexity of demand heterogeneity and uncertainty, leading to biased demand estimation and misguided decisions. The objective of this work is to improve nursing home resource planning decisions in response to the complex demand heterogeneity and uncertainty. To address the challenges, we propose a novel knowledge-guided and data-driven decision support framework. This is the first work of integrating domain knowledge with predictive and decision analytics to enhance modeling fidelity and decision performance for nursing home resource planning. Specifically, to effectively capture different aspects of heterogeneous demand, we develop a novel knowledge-guided demand modeling module with predictive models, including a length-of-stay model with competing risk for duration analysis, a tree-based system for learning daily service need variations, and a demand simulator for capturing uncertainty of fluctuating demand. Moreover, to determine optimal capacity and staffing decisions under demand heterogeneity and uncertainty, we develop a demand-based decision-making module with effective optimization models and solution algorithms, ensuring satisfactory quality of care at reduced costs. Furthermore, to demonstrate the improved prediction and decision performances of the proposed framework, we provide a proof-of-the-concept case study using real data from our industrial collaborator and investigate how demand heterogeneity and uncertainty will impact resource planning decisions. The proposed framework also demonstrates its appealing adaptability under changing resident census compositions.

NeurIPS Conference 2024 Conference Paper

Bridge-IF: Learning Inverse Protein Folding with Markov Bridges

  • Yiheng Zhu
  • Jialu Wu
  • Qiuyi Li
  • Jiahuan Yan
  • Mingze Yin
  • Wei Wu
  • Mingyang Li
  • Jieping Ye

Inverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which predominantly employ a discriminative formulation, frequently encounter the error accumulation issue and often fail to capture the extensive variety of plausible sequences. To fill these gaps, we propose Bridge-IF, a generative diffusion bridge model for inverse folding, which is designed to learn the probabilistic dependency between the distributions of backbone structures and protein sequences. Specifically, we harness an expressive structure encoder to propose a discrete, informative prior derived from structures, and establish a Markov bridge to connect this prior with native sequences. During the inference stage, Bridge-IF progressively refines the prior sequence, culminating in a more plausible design. Moreover, we introduce a reparameterization perspective on Markov bridge models, from which we derive a simplified loss function that facilitates more effective training. We also modulate protein language models (PLMs) with structural conditions to precisely approximate the Markov bridge process, thereby significantly enhancing generation performance while maintaining parameter-efficient training. Extensive experiments on well-established benchmarks demonstrate that Bridge-IF predominantly surpasses existing baselines in sequence recovery and excels in the design of plausible proteins with high foldability. The code is available at https: //github. com/violet-sto/Bridge-IF.

YNIMG Journal 2024 Journal Article

Mechanisms underlying category learning in the human ventral occipito-temporal cortex

  • Xiangqi Luo
  • Mingyang Li
  • Jiahong Zeng
  • Zhiyun Dai
  • Zhenjiang Cui
  • Minhong Zhu
  • Mengxin Tian
  • Jiahao Wu

The human ventral occipito-temporal cortex (VOTC) has evolved into specialized regions that process specific categories, such as words, tools, and animals. The formation of these areas is driven by bottom-up visual and top-down nonvisual experiences. However, the specific mechanisms through which top-down nonvisual experiences modulate category-specific regions in the VOTC are still unknown. To address this question, we conducted a study in which participants were trained for approximately 13 h to associate three sets of novel meaningless figures with different top-down nonvisual features: the wordlike category with word features, the non-wordlike category with nonword features, and the visual familiarity condition with no nonvisual features. Pre- and post-training functional MRI (fMRI) experiments were used to measure brain activity during stimulus presentation. Our results revealed that training induced a categorical preference for the two training categories within the VOTC. Moreover, the locations of two training category-specific regions exhibited a notable overlap. Remarkably, within the overlapping category-specific region, training resulted in a dissociation in activation intensity and pattern between the two training categories. These findings provide important insights into how different nonvisual categorical information is encoded in the human VOTC.

YNIMG Journal 2023 Journal Article

Developmental pattern of individual morphometric similarity network in the human fetal brain

  • Ruoke Zhao
  • Cong Sun
  • Xinyi Xu
  • Zhiyong Zhao
  • Mingyang Li
  • Ruike Chen
  • Yao Shen
  • Yibin Pan

The development of the cerebral cortex during the fetal period is a complex yet well-coordinated process. MRI-based morphological brain network provides a powerful tool for describing this process at a network level. Due to the challenges of in-utero MRI acquisition and image processing, the fetal morphological brain network has not been established. In this study, utilizing high-resolution in-utero MRI data, we constructed an individual morphometric similarity network for each fetus based on multiple cortical features. The spatiotemporal development of morphological connections was described at the level of edge, node, and lobe, respectively. Based on graph theoretical method, the topology structure of fetal morphological network was characterized. Edge analysis demonstrated an increase of morphological dissimilarity between hemispheres with gestational age, especially for the parietal cortex. The limbic and parieto-occipital regions exhibited the most drastic changes of morphological connections at both the edge and node levels. Between- and within-lobe analysis illustrated that the limbic lobe became more similar to other lobes, while the parietal and occipital lobes became more dissimilar to other lobes. Graph theoretical analysis indicated that the small-world structure of the fetal morphological network appeared as early as 22 weeks and that the network topology exhibited an enhanced integration and reduced segregation during prenatal development. The findings obtained from the preterm-born neonates agreed well with those of the fetuses. In summary, this study fills a gap in prenatal morphological brain network research and provides a piece of important evidence for understanding the normal development of fetal brain connectome during the second-third trimester.

ICLR Conference 2023 Conference Paper

Incremental Learning of Structured Memory via Closed-Loop Transcription

  • Shengbang Tong
  • Xili Dai
  • Ziyang Wu
  • Mingyang Li
  • Brent Yi
  • Yi Ma 0001

This work proposes a minimal computational model for learning structured memories of multiple object classes in an incremental setting. Our approach is based on establishing a {\em closed-loop transcription} between the classes and a corresponding set of subspaces, known as a linear discriminative representation, in a low-dimensional feature space. Our method is simpler than existing approaches for incremental learning, and more efficient in terms of model size, storage, and computation: it requires only a single, fixed-capacity autoencoding network with a feature space that is used for both discriminative and generative purposes. Network parameters are optimized simultaneously without architectural manipulations, by solving a constrained minimax game between the encoding and decoding maps over a single rate reduction-based objective. Experimental results show that our method can effectively alleviate catastrophic forgetting, achieving significantly better performance than prior work of generative replay on MNIST, CIFAR-10, and ImageNet-50, despite requiring fewer resources.

YNIMG Journal 2023 Journal Article

Multi-modal multi-resolution atlas of the human neonatal cerebral cortex based on microstructural similarity

  • Mingyang Li
  • Xinyi Xu
  • Zuozhen Cao
  • Ruike Chen
  • Ruoke Zhao
  • Zhiyong Zhao
  • Xixi Dang
  • Kenichi Oishi

The neonatal period is a critical window for the development of the human brain and may hold implications for the long-term development of cognition and disorders. Multi-modal connectome studies have revealed many important findings underlying the adult brain but related studies were rare in the early human brain. One potential challenge is the lack of an appropriate and unbiased parcellation that combines structural and functional information in this population. Using 348 multi-modal MRI datasets from the developing human connectome project, we found that the information fused from the structural, diffusion, and functional MRI was relatively stable across MRI features and showed high reproducibility at the group level. Therefore, we generated automated multi-resolution parcellations (300 - 500 parcels) based on the similarity across multi-modal features using a gradient-based parcellation algorithm. In addition, to acquire a parcellation with high interpretability, we provided a manually delineated parcellation (210 parcels), which was approximately symmetric, and the adjacent areas around each boundary were statistically different in terms of the integrated similarity metric and at least one kind of original features. Overall, the present study provided multi-resolution and neonate-specific parcellations of the cerebral cortex based on multi-modal MRI properties, which may facilitate future studies of the human connectome in the early development period.

YNIMG Journal 2022 Journal Article

Developmental pattern of association fibers and their interaction with associated cortical microstructures in 0–5-month-old infants

  • Tingting Liu
  • Jiani Wu
  • Zhiyong Zhao
  • Mingyang Li
  • Ying Lv
  • Mingyan Li
  • Fusheng Gao
  • Yuqing You

Association fibers connect the cortical regions and experience rapid development involving myelination and axonal growth during infancy. Yet, the spatiotemporal patterns of microstructural changes along these tracts, as well as the developmental interaction between the white matter (WM) tracts and the cortical gray matter (cGM) connected to them, are mostly unknown during infancy. In this study, we performed a diffusion MRI-based tractography and microstructure study in a cohort of 89 healthy preterm-born infants with gestational age at birth between 28.1∼36.4 weeks and postmenstrual age at scan between 39.9∼59.9 weeks. Results revealed that several C-shaped fibers, such as the arcuate fasciculus, cingulum, and uncinate fasciculus, demonstrated symmetrical along-tract profiles; and the horizontally oriented running fibers, including the inferior fronto-occipital fasciculus and the inferior longitudinal fasciculus, demonstrated an anterior-posterior developmental gradient. This study characterized the along-tract profiles using fixel-based analysis and revealed that the fiber cross-section (FC) of all five association fibers demonstrated a fluctuating increase with age, while the fiber density (FD) monotonically increase with age. NODDI was utilized to analyze the microstructural development of cGM and indicated cGM connected to the anterior end of the association fibers developed faster than that of the posterior end during 0-5 months. Notably, a mediation analysis was used to explore the relation between the development of WM and associated cGM, and demonstrated a partial mediation effect of FD in WM on the development of intracellular volume (ICV) in cGM and a full mediation effect of ICV on the growth of FD in most fibers, suggesting a predominant mediation of cGM on the WM development. Furthermore, for assessing whether those results were biased by prematurity, we compared preterm- and term-born neonates with matched scan age, gender, and multiple births from the developing human connectome project (dHCP) dataset to assess the effect of preterm-birth, and the results indicated a similar developmental pattern of the association fibers and their attached cGM. These findings presented a comprehensive picture of the major association fibers during early infancy and deciphered the developmental interaction between WM and cGM in this period.

AAAI Conference 2022 Conference Paper

DuMLP-Pin: A Dual-MLP-Dot-Product Permutation-Invariant Network for Set Feature Extraction

  • Jiajun Fei
  • Ziyu Zhu
  • Wenlei Liu
  • Zhidong Deng
  • Mingyang Li
  • Huanjun Deng
  • Shuo Zhang

Existing permutation-invariant methods can be divided into two categories according to the aggregation scope, i. e. global aggregation and local one. Although the global aggregation methods, e. g. , PointNet and Deep Sets, get involved in simpler structures, their performance is poorer than the local aggregation ones like PointNet++ and Point Transformer. It remains an open problem whether there exists a global aggregation method with a simple structure, competitive performance, and even much fewer parameters. In this paper, we propose a novel global aggregation permutation-invariant network based on dual MLP dot-product, called DuMLP- Pin, which is capable of being employed to extract features for set inputs, including unordered or unstructured pixel, attribute, and point cloud data sets. We strictly prove that any permutation-invariant function implemented by DuMLP-Pin can be decomposed into two or more permutation-equivariant ones in a dot-product way as the cardinality of the given input set is greater than a threshold. We also show that the DuMLP- Pin can be viewed as Deep Sets with strong constraints under certain conditions. The performance of DuMLP-Pin is evaluated on several different tasks with diverse data sets. The experimental results demonstrate that our DuMLP-Pin achieves the best results on the two classification problems for pixel sets and attribute sets. On both the point cloud classification and the part segmentation, the accuracy of DuMLP-Pin is very close to the so-far best-performing local aggregation method with only a 1-2% difference, while the number of required parameters is significantly reduced by more than 85% in classification and 69% in segmentation, respectively. The code is publicly available on https: //github. com/JaronTHU/ DuMLP-Pin.

NeurIPS Conference 2022 Conference Paper

Revisiting Sparse Convolutional Model for Visual Recognition

  • Xili Dai
  • Mingyang Li
  • Pengyuan Zhai
  • Shengbang Tong
  • Xingjian Gao
  • Shao-Lun Huang
  • Zhihui Zhu
  • Chong You

Despite strong empirical performance for image classification, deep neural networks are often regarded as ``black boxes'' and they are difficult to interpret. On the other hand, sparse convolutional models, which assume that a signal can be expressed by a linear combination of a few elements from a convolutional dictionary, are powerful tools for analyzing natural images with good theoretical interpretability and biological plausibility. However, such principled models have not demonstrated competitive performance when compared with empirically designed deep networks. This paper revisits the sparse convolutional modeling for image classification and bridges the gap between good empirical performance (of deep learning) and good interpretability (of sparse convolutional models). Our method uses differentiable optimization layers that are defined from convolutional sparse coding as drop-in replacements of standard convolutional layers in conventional deep neural networks. We show that such models have equally strong empirical performance on CIFAR-10, CIFAR-100 and ImageNet datasets when compared to conventional neural networks. By leveraging stable recovery property of sparse modeling, we further show that such models can be much more robust to input corruptions as well as adversarial perturbations in testing through a simple proper trade-off between sparse regularization and data reconstruction terms.

YNIMG Journal 2022 Journal Article

The influence of visual deprivation on the development of the thalamocortical network: Evidence from congenitally blind children and adults

  • Junfeng Lin
  • Linjun Zhang
  • Runhua Guo
  • Saiyi Jiao
  • Xiaomeng Song
  • Suting Feng
  • Ke Wang
  • Mingyang Li

The thalamus is heavily involved in relaying sensory signals to the cerebral cortex. A relevant issue is how the deprivation of congenital visual sensory information modulates the development of the thalamocortical network. The answer is unclear because previous studies on this topic did not investigate network development, structure-function combinations, and cognition-related behaviors in the same study. To overcome these limitations, we recruited 30 congenitally blind subjects (8 children, 22 adults) and 31 sighted subjects (10 children, 21 adults), and conducted multiple analyses [i.e., gray matter volume (GMV) analysis using the voxel-based morphometry (VBM) method, resting-state functional connectivity (FC), and brain-behavior correlation]. We found that congenital blindness elicited significant changes in the development of GMV in visual and somatosensory thalamic regions. Blindness also resulted in significant changes in the development of FC between somatosensory thalamic regions and visual cortical regions as well as advanced information processing regions. Moreover, the somatosensory thalamic regions and their FCs with visual cortical regions were reorganized to process high-level tactile language information in blind individuals. These findings provide a refined understanding of the neuroanatomical and functional plasticity of the thalamocortical network.

YNIMG Journal 2020 Journal Article

Linguistic experience acquisition for novel stimuli selectively activates the neural network of the visual word form area

  • Mingyang Li
  • Yangwen Xu
  • Xiangqi Luo
  • Jiahong Zeng
  • Zaizhu Han

The human ventral visual cortex is functionally organized into different domains that sensitively respond to different categories, such as words and objects. There is heated debate over what principle constrains the locations of those domains. Taking the visual word form area (VWFA) as an example, we tested whether the word preference in this area originates from the bottom-up processes related to word shape (the shape hypothesis) or top-down connectivity of higher-order language regions (the connectivity hypothesis). We trained subjects to associate identical, meaningless, non-word-like figures with high-level features of either words or objects. We found that the word-feature learning for the figures elicited the neural activation change in the VWFA, and learning performance effectively predicted the activation strength of this area after learning. Word-learning effects were also observed in other language areas (i. e. , the left posterior superior temporal gyrus, postcentral gyrus, and supplementary motor area), with increased functional connectivity between the VWFA and the language regions. In contrast, object-feature learning was not associated with obvious activation changes in the language regions. These results indicate that high-level language features of stimuli can modulate the activation of the VWFA, providing supportive evidence for the connectivity hypothesis of words processing in the ventral occipitotemporal cortex.

IJCAI Conference 2020 Conference Paper

Overflow Aware Quantization: Accelerating Neural Network Inference by Low-bit Multiply-Accumulate Operations

  • Hongwei Xie
  • Yafei Song
  • Ling Cai
  • Mingyang Li

The inherent heavy computation of deep neural networks prevents their widespread applications. A widely used method for accelerating model inference is quantization, by replacing the input operands of a network using fixed-point values. Then the majority of computation costs focus on the integer matrix multiplication accumulation. In fact, high-bit accumulator leads to partially wasted computation and low-bit one typically suffers from numerical overflow. To address this problem, we propose an overflow aware quantization method by designing trainable adaptive fixed-point representation, to optimize the number of bits for each input tensor while prohibiting numeric overflow during the computation. With the proposed method, we are able to fully utilize the computing power to minimize the quantization loss and obtain optimized inference performance. To verify the effectiveness of our method, we conduct image classification, object detection, and semantic segmentation tasks on ImageNet, Pascal VOC, and COCO datasets, respectively. Experimental results demonstrate that the proposed method can achieve comparable performance with state-of-the-art quantization methods while accelerating the inference process by about 2 times.

AAAI Conference 2014 Conference Paper

Cross-Lingual Knowledge Validation Based Taxonomy Derivation from Heterogeneous Online Wikis

  • Zhigang Wang
  • Juanzi Li
  • Shuangjie Li
  • Mingyang Li
  • Jie Tang
  • Kuo Zhang
  • Kun Zhang

Creating knowledge bases based on the crowd-sourced wikis, like Wikipedia, has attracted significant research interest in the field of intelligent Web. However, the derived taxonomies usually contain many mistakenly imported taxonomic relations due to the difference between the user-generated subsumption relations and the semantic taxonomic relations. Current approaches to solving the problem still suffer the following issues: (i) the heuristic-based methods strongly rely on specific language dependent rules. (ii) the corpus-based methods depend on a large-scale high-quality corpus, which is often unavailable. In this paper, we formulate the cross-lingual taxonomy derivation problem as the problem of cross-lingual taxonomic relation prediction. We investigate different linguistic heuristics and language independent features, and propose a cross-lingual knowledge validation based dynamic adaptive boosting model to iteratively reinforce the performance of taxonomic relation prediction. The proposed approach successfully overcome the above issues, and experiments show that our approach significantly outperforms the designed state-of-the-art comparison methods.

v2026.09.13