Arrow Research search

Author name cluster

Jian Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

17 papers
2 author rows

Possible papers

17

AAAI Conference 2026 Conference Paper

IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment Generation

  • Jian Yu
  • Fei Shen
  • Cong Wang
  • Yanpeng Sun
  • Hao Tang
  • Qin Guo
  • Xiaoyu Du

Diffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. To address these limitations, we propose IMAGGarment+, an efficient diffusion framework for controllable and high-quality garment synthesis. It comprises two key modules designed for efficient and attribute-aware conditioning. First, we introduce an attribute-wise feature extractor (AFE) that disentangles key garment attributes, silhouette, logo, position, and color, into parallel latent streams. Each stream is optimized independently via LoRA, ensuring minimal parameter overhead while retaining expressive capacity. Second, we develop an attribute-adaptive attention (AA) module to inject attribute-specific cues into the generative process through a selective, layer-wise injection strategy. Specifically, silhouette and color features are injected into early decoder layers to guide structural and appearance formation, while logo features are propagated across all layers to ensure cross-scale consistency. Extensive experiments on fine-grained garment benchmarks demonstrate that IMAGGarment+ outperforms state-of-the-art baselines with less than 20% additional parameters, validating its effectiveness and efficiency.

AAAI Conference 2026 Conference Paper

MetaGameBO: Hierarchical Game-Theoretic Driven Robust Meta-Learning for Bayesian Optimization

  • Hui Li
  • Huafeng Liu
  • Yiran Fu
  • Shuyang Lin
  • Baoxin Zhang
  • Deqiang Ouyang
  • Liping Jing
  • Jian Yu

Meta-learning for Bayesian optimization accelerates optimization by leveraging knowledge from previous tasks, but existing methods optimize for average performance and fail on challenging outlier tasks critical in practice. These limitations become particularly severe when target tasks exhibit distribution shifts or when optimization budgets are limited in real-world applications. We introduce MetaGameBO, a hierarchical game-theoretic framework that formulates meta-learning as robust optimization through CVaR-based task selection and diversity-aware sample learning. Our approach incorporates uncertainty-aware adaptation via probabilistic embeddings and Thompson sampling for robust generalization to out-of-distribution targets. We establish theoretical guarantees including convergence to game-theoretic equilibria and improved sample complexity, and demonstrate substantial improvements with 95.7% reduction in average loss and 88.6% lower tail risk compared to state-of-the-art methods on challenging tasks and distribution shifts.

AAAI Conference 2025 Conference Paper

3D Measurement of Complex Textured Objects Based on Bidirectional Fringe Projection

  • Yuchong Chen
  • Jian Yu
  • Shaoyan Gai
  • Zeyu Cai
  • Feipeng Da

In structured light systems, the accuracy of measurement notably diminishes when assessing complex texture objects, especially encountering boundaries between various colors. To address this challenge, this paper meticulously analyzes and establishes an error model, elaborating the correlation between phase errors and the gradients of phase and gray-scale. Based on this analysis, a novel high-precision method is proposed for measuring complex texture objects via bidirectional fringe projection. This approach firstly leverages horizontal and vertical fringe projections to derive bidirectional phase information and calculates the angles between the tangent of the texture edges and the phase gradient. Subsequently, a refined temporal phase correction algorithm is formulated based on the epipolar matching algorithm and the devised error model, effectively mitigating numerical instability issues within the algorithm and significantly reducing errors of bidirectional phases. Ultimately, corrected point clouds are calculated based on bidirectional phases, and the obtained point clouds are merged to further diminish phase errors. Comparison experiments indicate that this method can reduce Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) by 65.74% and 67.75%, respectively. Compared to existing methods, it improves performance by 27.29% and 33.74%, respectively, demonstrating superior performance.

IROS Conference 2025 Conference Paper

DGETP: Dynamic Graph Attention Network for Embodied Task Planning

  • Pengfei Sun
  • Guiling Wang
  • Xinli Zhang
  • Jian Yu

With the development of embodied intelligence, many studies have made progress by incorporating scene graphs and GNN into task planning. However, most methods still face challenges in fully capturing the sequential relationships between agent actions and the environment, making it difficult to handle dynamic changes and complexity inherent in embodied tasks. This paper proposes a Dynamic Graph Attention Network for Embodied Task Planning (DGETP) to process scene graph sequences and robot graphs for dynamic environment perception. In DGETP, we design a Hierarchical Dynamic Graph Attention network (H-DGAT) by employing both structural and temporal attention mechanisms to model the dynamic evolution feature of the scene. A Dual-branch Action-object Predictor (DAP) is proposed in DGETP through introducing sequences of previous actions and objects to efficiently aggregate historical information. DAP captures temporal dependencies between past and future actions through explicit sequence modeling, and reduces prediction complexity via a dual-branch architecture that separates action and object prediction while preserving their correlations through targeted feature fusion. Experiments show that DGETP improves task accuracy by over 30% in seen scenes and over 15% in unseen scenes compared to other baselines. In complex scenes, DGETP demonstrates strong generalization ability. Finally, the simulation environment indicates that DGETP achieves more goals than most of the advanced task planning method.

JBHI Journal 2025 Journal Article

NRAG: A Knowledge-Enhanced LLM Framework for Interpretable Neurosurgical Disease Diagnosis in Outpatient and Emergency Settings

  • Haoyu Tian
  • Yiming Liu
  • Xinyu Dai
  • Xin Dong
  • Jian Yu
  • Wei Wei
  • Boran Wang
  • Xuezhong Zhou

Large language models (LLMs) have achieved state-of-the-art performance in numerous domains, yet their clinical deployment faces critical barriers, particularly insufficient reasoning in complex scenarios and limited interpretability. These challenges are exacerbated in neurosurgical diagnosis for outpatient and emergency settings, where time-sensitive decision-making, fragmented data, and complex comorbidities render conventional free-text-based modeling approaches unreliable. To address the limitations of existing LLMs in medical auxiliary diagnosis, particularly in interpretability and predictive performance, this study proposed NRAG, an auxiliary diagnosis method that combines LLMs with knowledge graphs (KGs). It extracts symptom descriptions from clinical records and performs personalized retrieval of associated paths in KG, and supplements potential patient symptoms to optimize the diagnosis model. Comparative experiments involving multiple general-domain and medical-domain LLMs, along with case studies, were conducted to validate the NRAG's effectiveness. Experimental results demonstrate that integrating KG significantly improves diagnosis accuracy, achieving an F1-score of 0. 8150. It also substantially improves model interpretability and performs excellently in expert evaluations. Ablation studies and comparative experiments with other general-domain and medical-domain LLMs confirm the superior performance of the proposed NRAG. NRAG effectively supplements missing symptom information and provides knowledge-path-based evidence for diagnosis results, while improving the precision and interpretability of intelligent diagnosis. Furthermore, this approach sets the foundation for intelligent diagnoses in neurosurgery while providing a methodological framework for the integration of in-depth clinical data mining with medical knowledge base resources.

EAAI Journal 2025 Journal Article

Short-term wind speed prediction method based on prior wind direction knowledge and multi-period decoupling

  • Zewen Shang
  • Xuewei Li
  • Zhiqiang Liu
  • Yingzhou Sun
  • Jian Yu
  • Mei Yu
  • Tianyi Xu
  • Wei Xiong

Accurate wind speed prediction improves power system stability and efficiency. Wind direction contains airflow dynamic information, but existing methods fail to fully capture it. The multi-periodicity of meteorological data causes irregular wind speed fluctuations, complicating the capture of correlations and variances across periods. To address these issues, this paper proposes the Multi-Period Wind Direction Graph Network(MWGN). Gaussian Spatial Module(GSM) calculates the relationship between wind turbines using Gaussian distribution along the crosswind direction, thus enhancing airflow motion capture. Multi-Period Time Series Decoupling Module(MPTSDM) selects the primary periods by analyzing the frequency domain and extracts local correlation and periodic correlation, to better model the changing pattern of time series. Compared with the existing models, MWGN achieves the consistent state-of-the-art in three datasets.

AAAI Conference 2025 Conference Paper

SS-GEN: A Social Story Generation Framework with Large Language Models

  • Yi Feng
  • Mingyang Song
  • Jiaqi Wang
  • Zhuang Chen
  • Guanqun Bi
  • Minlie Huang
  • Liping Jing
  • Jian Yu

Children with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines. Social Stories™ are traditionally crafted by psychology experts under strict constraints to address these challenges but are costly and limited in diversity. As Large Language Models (LLMs) advance, there's an opportunity to develop more automated, affordable, and accessible methods to generate Social Stories in real-time with broad coverage. However, adapting LLMs to meet the unique and strict constraints of Social Stories is a challenging issue. To this end, we propose SS-GEN, a Social Story GENeration framework with LLMs. Firstly, we develop a constraint-driven sophisticated strategy named StarSow to hierarchically prompt LLMs to generate Social Stories at scale, followed by rigorous human filtering to build a high-quality dataset. Additionally, we introduce quality assessment criteria to evaluate the effectiveness of these generated stories. Considering that powerful closed-source large models require very complex instructions and expensive API fees, we finally fine-tune smaller language models with our curated high-quality dataset, achieving comparable results at lower costs and with simpler instruction and deployment. This work marks a significant step in leveraging AI to personalize Social Stories cost-effectively for autistic children at scale, which we hope can encourage future research on special groups.

AAAI Conference 2023 Conference Paper

ImageNet Pre-training Also Transfers Non-robustness

  • Jiaming Zhang
  • Jitao Sang
  • Qi Yi
  • Yunfan Yang
  • Huiwen Dong
  • Jian Yu

ImageNet pre-training has enabled state-of-the-art results on many tasks. In spite of its recognized contribution to generalization, we observed in this study that ImageNet pre-training also transfers adversarial non-robustness from pre-trained model into fine-tuned model in the downstream classification tasks. We first conducted experiments on various datasets and network backbones to uncover the adversarial non-robustness in fine-tuned model. Further analysis was conducted on examining the learned knowledge of fine-tuned model and standard model, and revealed that the reason leading to the non-robustness is the non-robust features transferred from ImageNet pre-trained model. Finally, we analyzed the preference for feature learning of the pre-trained model, explored the factors influencing robustness, and introduced a simple robust ImageNet pre-training solution. Our code is available at https://github.com/jiamingzhang94/ImageNet-Pretraining-transfers-non-robustness.

AAAI Conference 2023 Conference Paper

Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification

  • Pengyu Xu
  • Lin Xiao
  • Bing Liu
  • Sijin Lu
  • Liping Jing
  • Jian Yu

Multi-label text classification (MLTC) involves tagging a document with its most relevant subset of labels from a label set. In real applications, labels usually follow a long-tailed distribution, where most labels (called as tail-label) only contain a small number of documents and limit the performance of MLTC. To facilitate this low-resource problem, researchers introduced a simple but effective strategy, data augmentation (DA). However, most existing DA approaches struggle in multi-label settings. The main reason is that the augmented documents for one label may inevitably influence the other co-occurring labels and further exaggerate the long-tailed problem. To mitigate this issue, we propose a new pair-level augmentation framework for MLTC, called Label-Specific Feature Augmentation (LSFA), which merely augments positive feature-label pairs for the tail-labels. LSFA contains two main parts. The first is for label-specific document representation learning in the high-level latent space, the second is for augmenting tail-label features in latent space by transferring the documents second-order statistics (intra-class semantic variations) from head labels to tail labels. At last, we design a new loss function for adjusting classifiers based on augmented datasets. The whole learning procedure can be effectively trained. Comprehensive experiments on benchmark datasets have shown that the proposed LSFA outperforms the state-of-the-art counterparts.

NeurIPS Conference 2023 Conference Paper

Neural Processes with Stability

  • Huafeng Liu
  • Liping Jing
  • Jian Yu

Unlike traditional statistical models depending on hand-specified priors, neural processes (NPs) have recently emerged as a class of powerful neural statistical models that combine the strengths of neural networks and stochastic processes. NPs can define a flexible class of stochastic processes well suited for highly non-trivial functions by encoding contextual knowledge into the function space. However, noisy context points introduce challenges to the algorithmic stability that small changes in training data may significantly change the models and yield lower generalization performance. In this paper, we provide theoretical guidelines for deriving stable solutions with high generalization by introducing the notion of algorithmic stability into NPs, which can be flexible to work with various NPs and achieves less biased approximation with theoretical guarantees. To illustrate the superiority of the proposed model, we perform experiments on both synthetic and real-world data, and the results demonstrate that our approach not only helps to achieve more accurate performance but also improves model robustness.

JBHI Journal 2022 Journal Article

A Progressive Generative Adversarial Method for Structurally Inadequate Medical Image Data Augmentation

  • ruixuan zhang
  • Wenhuan Lu
  • Xi Wei
  • Jialin Zhu
  • Han Jiang
  • Zhiqiang Liu
  • Jie Gao
  • Xuewei Li

The generation-based data augmentation method can overcome the challenge caused by the imbalance of medical image data to a certain extent. However, most of the current research focus on images with unified structure which are easy to learn. What is different is that ultrasound images are structurally inadequate, making it difficult for the structure to be captured by the generative network, resulting in the generated image lacks structural legitimacy. Therefore, a Progressive Generative Adversarial Method for Structurally Inadequate Medical Image Data Augmentation is proposed in this paper, including a network and a strategy. Our Progressive Texture Generative Adversarial Network alleviates the adverse effect of completely truncating the reconstruction of structure and texture during the generation process and enhances the implicit association between structure and texture. The Image Data Augmentation Strategy based on Mask-Reconstruction overcomes data imbalance from a novel perspective, maintains the legitimacy of the structure in the generated data, as well as increases the diversity of disease data interpretably. The experiments prove the effectiveness of our method on data augmentation and image reconstruction on Structurally Inadequate Medical Image both qualitatively and quantitatively. Finally, the weakly supervised segmentation of the lesion is the additional contribution of our method.

JMLR Journal 2021 Journal Article

Interpretable Deep Generative Recommendation Models

  • Huafeng Liu
  • Liping Jing
  • Jingxuan Wen
  • Pengyu Xu
  • Jiaqi Wang
  • Jian Yu
  • Michael K. Ng

User preference modeling in recommendation system aims to improve customer experience through discovering users’ intrinsic preference based on prior user behavior data. This is a challenging issue because user preferences usually have complicated structure, such as inter-user preference similarity and intra-user preference diversity. Among them, inter-user similarity indicates different users may share similar preference, while intra-user diversity indicates one user may have several preferences. In literatures, deep generative models have been successfully applied in recommendation systems due to its flexibility on statistical distributions and strong ability for non-linear representation learning. However, they suffer from the simple generative process when handling complex user preferences. Meanwhile, the latent representations learned by deep generative models are usually entangled, and may range from observed-level ones that dominate the complex correlations between users, to latent-level ones that characterize a user’s preference, which makes the deep model hard to explain and unfriendly for recommendation. Thus, in this paper, we propose an Interpretable Deep Generative Recommendation Model (InDGRM) to characterize inter-user preference similarity and intra-user preference diversity, which will simultaneously disentangle the learned representation from observed-level and latent-level. In InDGRM, the observed-level disentanglement on users is achieved by modeling the user-cluster structure (i.e., inter-user preference similarity) in a rich multimodal space, so that users with similar preferences are assigned into the same cluster. The observed-level disentanglement on items is achieved by modeling the intra-user preference diversity in a prototype learning strategy, where different user intentions are captured by item groups (one group refers to one intention). To promote disentangled latent representations, InDGRM adopts structure and sparsity-inducing penalty and integrates them into the generative procedure, which has ability to enforce each latent factor focus on a limited subset of items (e.g., one item group) and benefit latent-level disentanglement. Meanwhile, it can be efficiently inferred by minimizing its penalized upper bound with the aid of local variational optimization technique. Theoretically, we analyze the generalization error bound of InDGRM to guarantee its performance. A series of experimental results on four widely-used benchmark datasets demonstrates the superiority of InDGRM on recommendation performance and interpretability. [abs] [ pdf ][ bib ] &copy JMLR 2021. ( edit, beta )

JBHI Journal 2020 Journal Article

Deep Learning for Smartphone-Based Malaria Parasite Detection in Thick Blood Smears

  • Feng Yang
  • Mahdieh Poostchi
  • Hang Yu
  • Zhou Zhou
  • Kamolrat Silamut
  • Jian Yu
  • Richard J. Maude
  • Stefan Jaeger

Objective: This work investigates the possibility of automated malaria parasite detection in thick blood smears with smartphones. Methods: We have developed the first deep learning method that can detect malaria parasites in thick blood smear images and can run on smartphones. Our method consists of two processing steps. First, we apply an intensity-based Iterative Global Minimum Screening (IGMS), which performs a fast screening of a thick smear image to find parasite candidates. Then, a customized Convolutional Neural Network (CNN) classifies each candidate as either parasite or background. Together with this paper, we make a dataset of 1819 thick smear images from 150 patients publicly available to the research community. We used this dataset to train and test our deep learning method, as described in this paper. Results: A patient-level five-fold cross-evaluation demonstrates the effectiveness of the customized CNN model in discriminating between positive (parasitic) and negative image patches in terms of the following performance indicators: accuracy (93. 46% ± 0. 32%), AUC (98. 39% ± 0. 18%), sensitivity (92. 59% ± 1. 27%), specificity (94. 33% ± 1. 25%), precision (94. 25% ± 1. 13%), and negative predictive value (92. 74% ± 1. 09%). High correlation coefficients (>0. 98) between automatically detected parasites and ground truth, on both image level and patient level, demonstrate the practicality of our method. Conclusion: Promising results are obtained for parasite detection in thick blood smears for a smartphone application using deep learning methods. Significance: Automated parasite detection running on smartphones is a promising alternative to manual parasite counting for malaria diagnosis, especially in areas lacking experienced parasitologists.

ECAI Conference 2020 Conference Paper

Learning Contextualized Sentence Representations for Document-Level Neural Machine Translation

  • Pei Zhang
  • Xu Zhang
  • Wei Chen 0071
  • Jian Yu
  • Yanfeng Wang
  • Deyi Xiong

Document-level machine translation incorporates intersentential dependencies into the translation of a source sentence. In this paper, we propose a new framework to model cross-sentence dependencies by training neural machine translation (NMT) to predict both the target translation and surrounding sentences of a source sentence. By enforcing the NMT model to predict source context, we want the model to learn “contextualized” source sentence representations that capture document-level dependencies on the source side. We further propose two different methods to learn and integrate such contextualized sentence embeddings into NMT: a joint training method that jointly trains an NMT model with the source context prediction model and a pre-training & fine-tuning method that pretrains the source context prediction model on a large-scale monolingual document corpus and then fine-tunes it with the NMT model. Experiments on Chinese-English and English-German translation show that both methods can substantially improve the translation quality over a strong document-level Transformer baseline.

JBHI Journal 2019 Journal Article

HerGePred: Heterogeneous Network Embedding Representation for Disease Gene Prediction

  • Kuo Yang
  • Ruyu Wang
  • Guangming Liu
  • Zixin Shu
  • Ning Wang
  • Runshun Zhang
  • Jian Yu
  • Jianxin Chen

The discovery of disease-causing genes is a critical step towards understanding the nature of a disease and determining a possible cure for it. In recent years, many computational methods to identify disease genes have been proposed. However, making full use of disease-related (e. g. , symptoms) and gene-related (e. g. , gene ontology and protein-protein interactions) information to improve the performance of disease gene prediction is still an issue. Here, we develop a heterogeneous disease-gene-related network (HDGN) embedding representation framework for disease gene prediction (called HerGePred). Based on this framework, a low-dimensional vector representation (LVR) of the nodes in the HDGN can be obtained. Then, we propose two specific algorithms, namely, an LVR-based similarity prediction and a random walk with restart on a reconstructed heterogeneous disease-gene network (RWRDGN), to predict disease genes with high performance. First, to validate the rationality of the framework, we analyze the similarity-based overlap distribution of disease pairs and design an experiment for disease-gene association recovery, the results of which revealed that the LVR of nodes performs well at preserving the local and global network structure of the HDGN. Then, we apply tenfold cross validation and external validation to compare our methods with other well-known disease gene prediction algorithms. The experimental results show that the RW-RDGN performs better than the state-of-the-art algorithm. The prediction results of disease candidate genes are essential for molecular mechanism investigation and experimental validation. The source codes of HerGePred and experimental data are available at https://github.com/yangkuoone/HerGePred.

ICRA Conference 2018 Conference Paper

Axially and Radially Expandable Modular Helical Soft Actuator for Robotic Implantables

  • Eduardo R. Perez-Guagnelli
  • Sarunas Nejus
  • Jian Yu
  • Shuhei Miyashita
  • YanQiang Liu
  • Dana D. Damian

Soft robotics has advanced the field of biomedical engineering by creating safer technologies for interfacing with the human body. One of the challenges in this field is the realization of modular soft basic constituents and accessible assembly methods to increase the versatility of soft robots. We present a soft pneumatic actuator composed of two elastomeric strands that provide interdependent axial and radial expansion due to the modularity of the components and their helical arrangement. The actuator reaches 35% of elongation with respect to its initial height and both chambers achieve forces of 1N at about 19kPa. We describe the design, fabrication, modeling and benchtop testing of the soft actuator towards realizing 3D functional structures with potential medical applications. An example of application for soft medical robots is tissue regenerative for the long-gap esophageal atresia condition.

YNIMG Journal 2018 Journal Article

Changes in dynamic functional connections with aging

  • Lixia Tian
  • Qizhuo Li
  • Chaomurilige Wang
  • Jian Yu

Despite numerous studies on age-related changes in static functional connections (FCs), the available literature on the changes in dynamic FCs with aging is lacking. This study investigated the changes in dynamic FCs with aging based on resting state fMRI data of 61 healthy adults aged 30–85 years. The time-resolved FCs among 160 pre-defined regions of interest (ROIs) were first estimated using sliding-window correlation. Based on the dynamic FC matrices, we then analyzed the dynamic switches between different FC states using k-means clustering, and correlated age with the dwell time of each FC state across subjects. The elderly were observed to spend more time in an FC state characterized by weak interactions throughout the brain and less time in an FC state characterized by strong interactions within the sensory-motor network and the cognitive control network. These results may reflect an overall weakening of connections in the elderly, which support less efficient information transfer in them. Based on the dynamic FC matrices, we also evaluated the variability and amplitude of FC time-series, which measure the relative (to mean) and absolute strength of FC fluctuations, respectively, and correlated age with the two measures across subjects. Relatively weak age-vs-variability correlations were observed, but we did observe significant negative age-vs-amplitude correlations at both the global and regional level. These results indicate that amplitude may be another effective metric for assessing FC fluctuations, in addition to the widely-used variability metric. Moreover, the observed declines in the amplitude of FC fluctuations in the elderly may support the assumption that it should be the weakening of absolute interactions between brain regions, rather than toggling between positive and negative correlations, that causes the repeatedly reported widespread (static) FC decreases with aging. Overall, the present results not only reflect an overall weakening of connections in the elderly, but indicate the potential of dynamic FC analyses in studies of age-related psychiatric and neurological disorders.

v2026.09.13