Arrow Research search

Author name cluster

Yong Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

49 papers
2 author rows

Possible papers

49

EAAI Journal 2026 Journal Article

A logistic matrix factorization recommendation algorithm based on polynomial coefficient perturbation

  • Zhiqiang Zhang
  • Bo Li
  • Jiangzhou Deng
  • Yong Wang
  • Jianmei Ye
  • Zhuo Liu

Most current privacy-preserving recommendation schemes designed for explicit ratings have made significant progress. However, the privacy concerns arising from implicit feedback data have not received sufficient attention. To this end, we propose a novel logistic matrix factorization recommendation algorithm based on polynomial coefficient perturbation. This algorithm adopts logistic matrix factorization to fit implicit feedback data, while introducing perturbation into the objective function to protect user privacy. To manage the privacy budget efficiently, Taylor expansion is leveraged to approximate the objective function as a polynomial. Noise is only added to the first-order term to satisfy the differential privacy constraint, thereby minimizing the potential error accumulation. Theoretical analyses rigorously prove the privacy level and data utility of the proposed method. Experimental results on multiple datasets further demonstrate that our scheme can effectively protect user privacy while delivering good recommendation performance.

AAAI Conference 2026 Conference Paper

AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting

  • Renda Li
  • Hailang Huang
  • Fei Wei
  • Feng Xiong
  • Yong Wang
  • Xiangxiang Chu

Reinforcement learning (RL) has demonstrated considerable potential for enhancing reasoning in large language models (LLMs). However, existing methods suffer from Gradient Starvation and Policy Degradation when training directly on samples with mixed difficulty. To mitigate this, prior approaches leverage Chain-of-Thought (CoT) data, but the construction of high-quality CoT annotations remains labor-intensive. Alternatively, curriculum learning strategies have been explored but frequently encounter challenges, such as difficulty mismatch, reliance on manual curriculum design, and catastrophic forgetting. To address these issues, we propose AdaCuRL, a Adaptive Curriculum Reinforcement Learning framework that integrates coarse-to-fine difficulty estimation with adaptive curriculum scheduling. This approach dynamically aligns data difficulty with model capability and incorporates a data revisitation mechanism to mitigate catastrophic forgetting. Furthermore, AdaCuRL employs adaptive reference and sparse KL strategies to prevent Policy Degradation. Extensive experiments across diverse reasoning benchmarks demonstrate that AdaCuRL consistently achieves significant performance improvements on both LLMs and MLLMs.

AIIM Journal 2026 Journal Article

Adaptive time-frequency decomposition informer for pathological rest tremor sequence prediction

  • Feiyun Xiao
  • Ruixue Gao
  • Cheng Huang
  • Jingsong Mu
  • Yong Wang

Pathological tremor is a common symptom of various neurological disorders. Pathological tremor signal prediction is important to the tremor suppression equipment. However, how to improve the accuracy of multi-step prediction of tremor motion is a difficult problem. In this paper, an adaptive time-frequency decomposition informer (ATFDI) method is proposed to predict the signal of pathological tremor. For this method, an adaptive oscillator is used to model the tremor signal, and the tremor signal is decomposed into multiple sub-signals with narrow band and single main frequency. After eliminating the redundant signal, the Informer method is used to predict multiple sub-signals, and the actual prediction signal is obtained. Corresponding to prediction length of 100 ms, 300 ms, 500 ms, 700 ms, and 1000 ms, the root mean square error (RMSE), correlation coefficient (r), and mean absolute error (MAE) between the predicted signal and the actual signal are 0. 0827 ± 0. 0424, 0. 9186 ± 0. 1192, and 0. 0675 ± 0. 0357, respectively. The average RMSE, r and MAE between the predicted signal and the actual signal for 100 ms prediction length are 0. 0521 ± 0. 0227, 0. 9636 ± 0. 0360, and 0. 0443 ± 0. 0190, respectively. Under balanced leave-one-subject-out (LOSO) evaluation across 20 subjects and five horizons from 100 to 1000 ms, forecasting performance degraded with longer horizons with mean MAE increasing from 0. 133 to 0. 200 and mean correlation decreasing from 0. 813 to 0. 545. The corresponding required execution time is 34. 6 ms, which meets the real-time requirement. The proposed method is inspired by oscillatory properties of pathological tremor and attention-related signal modulation, and combines these ideas in an engineering prediction framework.

AAAI Conference 2026 Conference Paper

CaT-Diff: Cascaded Text-enhanced Diffusion Model for Time-Series Imputation

  • Changjian Xu
  • Yong Wang
  • Ruizheng Huang
  • Zhicheng Zhang
  • Wen Yin
  • Kexin Li

Most state-of-the-art time series imputation methods can leverage textual information to improve imputation quality, but they often struggle because they fail to effectively filter noisy information from large language model (LLM) derived textual information. Some existing solutions only filter over the entire token set, which can introduce erroneous conditional constraints, extreme token frequency effects and increased computational complexity. To address this, we propose CaT-Diff, a novel cascaded text-enhanced diffusion model for probabilistic imputation of multivariate time series under Missing Not At Random (MNAR) scenarios. To suppress irrelevant semantics and focus on context most predictive of missing values, CaT-Diff introduces an innovative Hierarchical Semantic Filter (HSF) that collaborates with a Mixture-of-Experts (MoE) Network. The MoE projects heterogeneous text embeddings into the time series latent space, and the HSF cascade-filters text embeddings from the segment level to the token level, thereby avoiding the pitfalls of direct token-level filtering and reducing overhead. We also incorporate a lightweight Missing Mechanism Estimator, jointly optimized with the denoising network to explicitly capture MNAR missingness patterns. Extensive tests on nine domains show that CaT-Diff outperforms state-of-the-art baselines. Our work presents a new approach for selectively fusing LLM-derived textual information.

EAAI Journal 2026 Journal Article

G-LFFN: A Global-Local Feature Fusion Network Leveraging Transformer-Encoder and Contrastive Learning for Multimodal Sentiment Analysis

  • Cong Liu
  • Yong Wang
  • Jing Yang
  • Xiaohui Tao
  • Jiaqi Liu

Due to the varieties of sentiment expressions, multimodal sentiment analysis for social media requires a comprehensive fusion of image and textual information. However, most of the previous studies have only modeled the inter-modal local or global interactions, ignoring inter-modal global and local co-influences, resulting in insufficient fusion of sentiment information. In addition, the introduction of multiple features may generate more sentiment-irrelevant information, thus leading to a weaker sentiment association of the fusion features. To solve the above issues, we propose a global-local feature fusion network model leveraging transformer-encoder and contrastive learning. Firstly, considering inter-modal global and local co-influences, the model extracts global and local features in the image. Secondly, we propose a cross-modal synchronous fusion transformer-encoder and its simplified version to capture inter-modal global and local consistent features, and combine it with soft self-attention to further enhance inter-modal interaction. On this basis, we utilize multiple contrastive learning to enhance the interactions among multiple fusion features and improve the sentiment associations of multimodal fusion features to assist the final sentiment analysis. Extensive experiments on three public multimodal datasets show that our model can adequately capture inter-modal global-local information interactions and effectively improve sentiment associations, thus demonstrating its validity and superiority.

EAAI Journal 2026 Journal Article

Hard constraint learning approaches with trainable influence functions for evolutionary equations

  • Yushi Zhang
  • Shuai Su
  • Yong Wang
  • Yanzhong Yao

This paper develops a novel deep learning approach for solving evolutionary equations, which integrates sequential learning strategies with an enhanced hard constraint strategy featuring trainable parameters, addressing the low computational accuracy of standard Physics-informed neural networks (PINNs) in large temporal domains. Sequential learning strategies divide a large temporal domain into multiple subintervals and solve them one by one in a chronological order, which naturally respects the principle of causality and improves the stability of the PINN solution. The improved hard constraint strategy strictly ensures the continuity and smoothness of the PINN solution at time interval nodes, and at the same time passes the information from the previous interval to the next interval, which avoids the incorrect/trivial solution at the position far from the initial time. Furthermore, by investigating the requirements of different types of equations on hard constraints, we design a novel influence function with trainable parameters for hard constraints, which provides theoretical and technical support for the effective implementations of hard constraint strategies, and significantly improves the universality and computational accuracy of our method. In addition, an adaptive time-domain partitioning algorithm is proposed, which plays an important role in the application of the proposed method as well as in the improvement of computational efficiency and accuracy. Numerical experiments verify the performance of the method. The data and code accompanying this paper are available at https: //github. com/zhizhi4452/HCS.

EAAI Journal 2026 Journal Article

Improved noising training detection Transformer based drone image detector

  • Lu Ding
  • Jinghua Deng
  • Xun Huang
  • Yong Wang

Transformer-based methods such as Detection Transformer (DETR) are playing an important role in the field of object detection. However, in drone image object detection, DETR has difficulty leveraging its algorithmic advantages due to issues such as small object size, large-scale changes, and uneven distribution of objects in the image. In response to the above issues, we propose an improved noising training detection Transformer (INT-DETR) algorithm aiming at solving the problem of small and dense distributed objects in drone image. Firstly, a noising training module is designed to randomly add varying noise levels to ground truth. This can improve stability of the model in matching complex objects and accelerate the focus on key dense object areas. Secondly, a one-to-many label matching assignment training algorithm is used to reduce the uncertainty in object classification caused by high density or occlusion. This procedure increases the accuracy of noising training for small object detection. Finally, evaluations on the VisDrone2019-DET and SeaDronesSeeV2 datasets validate the effectiveness of the proposed method. Experimental results demonstrate that INT-DETR achieves superior performance in drone image object detection.

EAAI Journal 2026 Journal Article

Multi-agent trajectory prediction with Hierarchical Coordinate-Based Representation

  • Yuanchen Zhu
  • Shuaiqi Fu
  • Yong Wang
  • Yanan Zhao
  • Huachun Tan

Trajectory prediction plays a crucial role in autonomous driving systems. Existing methods typically adopt agent-centric or scene-centric approaches to model driving scenarios. However, these approaches often lead to excessive redundant computations or pose estimation errors, which degrade both prediction efficiency and accuracy. To address these issues, we propose a novel multi-agent trajectory prediction model called Hierarchical Coordinate-Based Representation. This model decomposes the driving environment into two distinct components: global and local. In the global component, the interaction information is established and shared between all predicted agents, helping to reduce redundant computations. In the local component, each vehicle is assigned an individual reference frame, which mitigates the impact of pose variations and facilitates the extraction of temporal features. Furthermore, we introduce an adaptive anchor point generation method to tackle the challenge of capturing future trajectories across different driving scenarios. This method dynamically generates anchor points that are tailored to each specific scenario, guiding the prediction of trajectories for various modalities. The performance of the proposed model is evaluated on the Argoverse 1 and Argoverse 2 datasets. Experimental results demonstrate that Hierarchical Coordinate-Based Representation achieves competitive performance in terms of both efficiency and accuracy, outperforming state-of-the-art methods.

AAAI Conference 2026 Conference Paper

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

  • Jiao Li
  • Jian Lang
  • Xikai Tang
  • Wenzheng Shu
  • Ting Zhong
  • Qiang Gao
  • Yong Wang
  • Leiting Chen

Hate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular and ambiguous forms to evade censorship, resulting in substantial semantic drift and rendering previously trained models ineffective. Test-Time Adaptation (TTA) offers a solution by adapting models during inference to narrow the cross-domain gap, while conventional TTA methods target mild distribution shifts and struggle with the severe semantic drift in HVD. To tackle these challenges, we propose SCANNER, the first TTA framework tailored for HVD. Motivated by the insight that, despite the evolving nature of hateful manifestations, their underlying cores remain largely invariant (i.e., targeting is still based on characteristics like gender, race, etc), we leverage these stable cores as a bridge to connect the source and target domains. Specifically, SCANNER initially reveals the stable cores from the ambiguous layout in evolving hateful content via a principled centroid-guided alignment mechanism. To alleviate the impact of outlier-like samples that are weakly correlated with centroids during the alignment process, SCANNER enhances the prior by incorporating a sample-level adaptive centroid alignment strategy, promoting more stable adaptation. Furthermore, to mitigate semantic collapse from overly uniform outputs within clusters, SCANNER introduces an intra-cluster diversity regularization that encourages the cluster-wise semantic richness. Experiments show that SCANNER outperforms all baselines, with an average gain of 4.69% in Macro-F1 over the best.

YNIMG Journal 2026 Journal Article

Spinal cord stimulation improves brain connectivity and consciousness level in patients with disorders of consciousness

  • Adilijiang Aihemaitiniyazi
  • Tiemin Li
  • Huawei Zhang
  • Da Wei
  • Pu Cai
  • Wei Wang
  • Guoming Luan
  • Yong Wang

OBJECTIVE: Spinal cord stimulation (SCS) is an advanced neuromodulation technology in disorders of consciousness (DOC) field. However, research on the modulation effects and mechanisms of SCS is limited. METHOD: We proposed a study design (SCS and sham) to study the short-term effects of 20 minutes' SCS, in which resting state EEG and Coma Recovery Scale-Revised (CRS-R) were used to measure the changes in neural and behavioral activity caused by SCS. We used the Genuine Permutation Cross Mutual Information(G_PCMI) to analyze EEG data and study changes in cortical connectivity during SCS. Finally, all patients' CRS-R results were obtained after 6 months' SCS treatment. RESULTS: Short-term SCS (20 min) did not alter the patient's CRS-R score, but long-term SCS (6 months) can improve the CRS-R scores of all patients. EEG results show G_PCMI of the frontal and central brain regions significantly change before and after short-term SCS (p < 0.01) and PCMI of the F-P, F-O regions have significant differences before and after short-term SCS (p < 0.05). Besides, the G_PCMI changes in frontal, parietal, F-P and F-O regions show a significant positive correlation with CRS-R changes (r = 0.80, 0.66, 0.68 and 0.72; p < 0.05). However, the sham group showed no significant G_PCMI changes. CONCLUSION: SCS can improve the awareness level of DOC patients. SCS improves cortical short- and long-distance connectivity of DOC patients may contribute the improvement of consciousness level.

AAAI Conference 2026 Conference Paper

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

  • Ziyu Ma
  • Chenhui Gou
  • Yiming Hu
  • Yong Wang
  • Bohan Zhuang
  • Jianfei Cai

Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and high inference cost. To address these challenges, task-vector-based methods have been explored by inserting compact representations of many-shot in-context demonstrations into model activations. However, existing task-vector-based methods either overlook the importance of where to insert task vectors or struggle to determine suitable values for each location. To this end, we propose a novel Sensitivity-aware Task Vector insertion framework (STV) to figure out where and what to insert. Our key insight is that activation deltas across query-context pairs exhibit consistent structural patterns, providing a reliable cue for insertion. Based on the identified sensitive-aware locations, we construct a pre-clustered activation bank for each location by clustering the activation values, and then apply reinforcement learning to choose the most suitable one to insert. We evaluate STV across a range of multimodal models (e.g., Qwen-VL, Idefics-2) and tasks (e.g., VizWiz, OK-VQA), demonstrating its effectiveness and showing consistent improvements over previous task-vector-based methods with strong generalization.

EAAI Journal 2025 Journal Article

Camouflaged Object Detection with boundary localization in complex backgrounds

  • Guangjian Zhang
  • Zhengming Yang
  • Yong Wang
  • Yuliang Chen
  • Duoqian Miao

The primary challenge of Camouflaged Object Detection (COD) lies in the high similarity between the target and the complex background, making it difficult for the human eye to distinguish them. Based on the phenomenon that human attention shifts between the target and the background when observing objects, we propose a network model named MENet. This model adopts a three-stage decoupled architecture of “localization-interaction-fusion. ” In the localization stage, we utilize an attention mechanism-based backbone network (Pyramid Vision Transformer V2, abbreviated as PVT-V2) to generate multi-level features, which can initially locate the target area. In the interaction stage, we design a Contour-Aware Edge Module (CAEM) and an Area Decoder (AD) to capture the target edges and background information, respectively, thereby achieving precise localization of the target boundary and reducing interference from background noise. Furthermore, we developed a Boundary Guidance Module (BGM) that effectively injects boundary cues and relevant background information separately into the multi-level features, enhancing the model’s ability to detect target edges in complex backgrounds. In the fusion stage, we design two Feature Fusion Modules (FFM and KFFM) to effectively merge multi-level features with precise boundaries and de-noised features, thereby enhancing the prediction performance of camouflaged objects. Extensive experiments on three challenging benchmark datasets demonstrate that our MENet outperforms many existing state-of-the-art methods. Our method leverages artificial intelligence (AI) techniques to improve the accuracy of camouflaged object and pest detection in complex visual environments. Our code is publicly available at: https: //github. com/yang19950966666/MENet.

ICRA Conference 2025 Conference Paper

METDrive: Multimodal End-to-End Autonomous Driving with Temporal Guidance

  • Ziang Guo
  • Xinhao Lin
  • Zakhar Yagudin
  • Artem Lykov
  • Yong Wang
  • Yanqiang Li
  • Dzmitry Tsetserukou

Multimodal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both static and dynamic aspects of the driving environment is enhanced, thereby improving the safety of autonomous driving. In this paper, we introduce METDrive, an end-to-end system that leverages temporal guidance from the embedded time series features of ego states, including rotation angles, steering, throttle signals, and waypoint vectors. The geometric features derived from the perception sensor data and the time series features of ego state data jointly guide the waypoint prediction with the proposed temporal guidance loss function. We evaluated METDrive on the CARLA leaderboard benchmarks, achieving a driving score of 70%, a route completion score of 94%, and an infraction score of 0. 78.

JBHI Journal 2025 Journal Article

MiRNA-disease Association Prediction via Cosine Annealing and Multi-Head Self-Attention in HyperGCN

  • Rong Zhu
  • Jie Zheng
  • Zheng-Hua Chang
  • Jin-Xing Liu
  • Jun-Liang Shang
  • Yong Wang

Biological studies have demonstrated that understanding the association between miRNAs and disease is critical for disease prevention, assessment, and therapy. However, traditional experimental methods for inferring these connections are not only costly but also inefficient. Hence, there is a pressing need to develop novel methods to improve the accuracy and efficiency of forecasting. Currently, graph convolutional networks (GCNs) techniques are one of the mainstream methods for predicting disease correlations. Nevertheless, traditional GCNs suffer from gradient vanishing and gradient explosion problems when dealing with long-range dependencies. To overcome these problems, we suggest a new approach called HGCMMDA, which relies on HyperGCN and combines a cosine annealing algorithm and a multi-head self-attention mechanism. In HGCMMDA, similarity networks for miRNAs and diseases are constructed, and GCN is used for feature extraction. A heterogeneity hypergraph is then built via HyperGCN for improved information propagation. Multi-head self-attention captures diverse node relations, while cosine annealing adjusts the learning rate. A combined BCE-Dice loss ensures accurate prediction. To evaluate the effectiveness of the proposed method, a comprehensive set of experiments was conducted using the Human microRNA Disease Database (HMDD v3. 2). The method achieved a peak area under the receiver operating characteristic curve (AUC) of 0. 9515, along with competitive performance in other evaluation metrics. The experimental findings indicate that HGCMMDA achieves notable enhancements over previously established approaches. These results strongly support the assertion that HGCMMDA serves as a dependable and effective framework for comprehensively exploring the intricate associations between microRNAs and human diseases.

EAAI Journal 2025 Journal Article

Multi-scale frequency attention fusion network for infrared and visible image fusion

  • Yong Wang
  • Xueyuan Zhao
  • Jianfei Pu
  • Lulu Zhang
  • Duoqian Miao

The goal of visible and infrared image fusion is to generate a composite image that not only preserves fine-grained textures from visible images but also highlights salient targets from infrared images. However, existing methods often struggle to capture detailed features and typically rely solely on spatial domain information, overlooking the complementary advantages offered by frequency domain features. To address these limitations, we propose a feature-level fusion approach based on a multi-scale frequency attention fusion network, which incorporates a spatial-frequency attention fusion module with cross-attention and a multi-scale depthwise separable convolution block equipped with coordinate attention. Moreover, a multi-scale compensation fusion is incorporated within the spatial frequency attention fusion module to reduce cross-modal domain discrepancies. This work applies deep learning-based artificial intelligence (AI) techniques to the field of multi-modal image fusion. Experiments on three public datasets show that our method achieves competitive performance across multiple evaluation metrics compared to state-of-the-art methods. Our method demonstrates superior performance in preserving fine details and enhancing target saliency. The code will be available at https: //github. com/ioschunsheng1230.

EAAI Journal 2025 Journal Article

Point cloud semantic segmentation network based on graph convolution and attention mechanism

  • Nan Yang
  • Yong Wang
  • Lei Zhang
  • Bin Jiang

Point cloud data provides rich three-dimensional spatial information. Accurate three-dimensional point cloud semantic segmentation algorithms enhance environmental understanding and perception, with wide-ranging applications in autonomous driving and scene analysis. However, Graph Neural Networks often struggle to retain semantic relationships among neighboring points during feature extraction, potentially leading to the omission of critical features during aggregation. To address these challenges, we propose a novel network, the Feature-Enhanced Residual Attention Network. This network includes an innovative graph convolution module, the Neighborhood-Enhanced Convolutional Aggregation Module, which utilizes K-Nearest Neighbor and Dilated K-Nearest Neighbor techniques to construct diverse dynamic graphs and aggregate features, thereby prioritizing essential information. This approach significantly enhances the expressiveness and generalization capabilities of the network. Additionally, we introduce a new spatial attention module designed to capture semantic relationships among points. Experimental results demonstrate that the Feature-Enhanced Residual Attention Network outperforms benchmark models, achieving an average intersection ratio of 61. 3% and an overall accuracy of 86. 7% on the Stanford Large-Scale Three-dimensional Indoor Spaces dataset, thereby significantly improving semantic segmentation performance.

EAAI Journal 2025 Journal Article

The multi-depot pickup and delivery vehicle routing problem with time windows and dynamic demands

  • Yong Wang
  • Mengyuan Gou
  • Siyu Luo
  • Jianxin Fan
  • Haizhong Wang

The rapid development of the urban logistics recycling industry, combined with the complexity of the pickup and delivery networks, has created a surge in dynamic customer demands and exacerbated the difficulty of logistics resource sharing. Accordingly, this work focuses on a multi-depot pickup and delivery vehicle routing problem with time windows and dynamic demands, which incorporates resource sharing. A bi-objective mathematical model is formulated to minimize the total operating cost and number of vehicles. A three-dimensional affinity propagation clustering and an adaptive nondominated sorting genetic algorithm-II are combined to find Pareto optimal solutions. A dynamic demand insertion strategy is proposed to determine the vehicle service sequences for dynamic situations. Combined with an elite iteration mechanism to prevent the proposed algorithm from falling into search stagnation and improve the convergence performance. The superiority of the proposed algorithm is verified by comparing with CPLEX solver (i. e. , ILOG CPLEX Optimization Studio 12. 10), multi-objective ant colony optimization, multi-objective particle swarm optimization, multi-objective evolutionary algorithm, multi-objective genetic algorithm, and decomposition-based multi-objective evolutionary algorithm with tabu search. Besides, the proposed model and algorithm are tested by a real-world case study in Chongqing city, China, and the further analysis indicates that significant improvement can be achieved. Furthermore, by incorporating the recognition and prediction techniques of artificial intelligence on dynamic demand data, the proposed approach can realize the self-optimization of multi-depot vehicle routes and the precise allocation of logistics resources in dynamic environments. This study is conducive to the construction of a digitally-intelligent urban logistics system.

EAAI Journal 2025 Journal Article

Transfer learning framework integrating attention mechanism and domain adaptation for Low Earth Orbit satellite network traffic prediction

  • Yan Zhang
  • Yong Wang
  • Qingsong Zhao
  • Yadi Zhai
  • Zhi Lin
  • Luda Zhao
  • Yihua Hu

Traffic prediction is a crucial prerequisite for planning and even network security in Low Earth Orbit (LEO) satellite networks (LSNs). This paper designed a transfer learning framework for LSN traffic prediction that leverages attention mechanism and domain adaptation. Firstly, by integrating five-dimensional data including global population distribution, local time coefficient, the Internet penetration rate of the country, daily data volume of a single Internet user, and global aeronautical traffic demand, a traffic model that could characterize the traffic situation of the area covered by LEO satellite within a specific time range was constructed. Considering the problem of insufficient online traffic data, knowledge was transferred from terrestrial network traffic (source domain) to satellite network traffic (target domain) by incorporating the Domain-Adversarial Neural Network (DANN) method to tackle the data distribution discrepancies between the source and target domains. Finally, by combining DANN with the attention mechanism, the domain-invariant features of the source domain and the target domain were extracted to predict satellite network traffic accurately. Experimental results show that compared to baseline models, the error of the proposed framework in terms of root mean square error measurement is reduced by 9. 57% to 33. 47% and 18. 85% to 38. 99% in the two simulated LSN traffic scenarios. Moreover, this framework has low computational complexity among other transfer learning models, which can lay a foundation for subsequent satellite traffic planning and network security.

JBHI Journal 2025 Journal Article

Transformer-Based Weakly Supervised Learning for Whole Slide Lung Cancer Image Classification

  • Jianpeng An
  • Yong Wang
  • Qing Cai
  • Gang Zhao
  • Stephan Dooper
  • Geert Litjens
  • Zhongke Gao

Image analysis can play an important role in supporting histopathological diagnoses of lung cancer, with deep learning methods already achieving remarkable results. However, due to the large scale of whole-slide images (WSIs), creating manual pixel-wise annotations from expert pathologists is expensive and time-consuming. In addition, the heterogeneity of tumors and similarities in the morphological phenotype of tumor subtypes have caused inter-observer variability in annotations, which limits optimal performance. Effective use of weak labels could potentially alleviate these issues. In this paper, we propose a two-stage transformer-based weakly supervised learning framework called Simple Shuffle-Remix Vision Transformer (SSRViT). Firstly, we introduce a Shuffle-Remix Vision Transformer (SRViT) to retrieve discriminative local tokens and extract effective representative features. Then, the token features are selected and aggregated to generate sparse representations of WSIs, which are fed into a simple transformer-based classifier (SViT) for slide-level prediction. Experimental results demonstrate that the performance of our proposed SSRViT is significantly improved compared with other state-of-the-art methods in discriminating between adenocarcinoma, pulmonary sclerosing pneumocytoma and normal lung tissue (accuracy of 96. 9 ${\%}$ and AUC of 99. 6 ${\%}$ ).

NeurIPS Conference 2024 Conference Paper

Classifier Clustering and Feature Alignment for Federated Learning under Distributed Concept Drift

  • Junbao Chen
  • Jingfeng Xue
  • Yong Wang
  • Zhenyan Liu
  • Lu Huang

Data heterogeneity is one of the key challenges in federated learning, and many efforts have been devoted to tackling this problem. However, distributed concept drift with data heterogeneity, where clients may additionally experience different concept drifts, is a largely unexplored area. In this work, we focus on real drift, where the conditional distribution $P(\mathcal{Y}|\mathcal{X})$ changes. We first study how distributed concept drift affects the model training and find that local classifier plays a critical role in drift adaptation. Moreover, to address data heterogeneity, we study the feature alignment under distributed concept drift, and find two factors that are crucial for feature alignment: the conditional distribution $P(\mathcal{Y}|\mathcal{X})$ and the degree of data heterogeneity. Motivated by the above findings, we propose FedCCFA, a federated learning framework with classifier clustering and feature alignment. To enhance collaboration under distributed concept drift, FedCCFA clusters local classifiers at class-level and generates clustered feature anchors according to the clustering results. Assisted by these anchors, FedCCFA adaptively aligns clients' feature spaces based on the entropy of label distribution $P(\mathcal{Y})$, alleviating the inconsistency in feature space. Our results demonstrate that FedCCFA significantly outperforms existing methods under various concept drift settings. Code is available at https: //github. com/Chen-Junbao/FedCCFA.

EAAI Journal 2024 Journal Article

Condition monitoring for nuclear turbines with improved dynamic partial least squares and local information increment

  • Yixiong Feng
  • Zetian Zhao
  • Bingtao Hu
  • Yong Wang
  • Hengyuan Si
  • Zhaoxi Hong
  • Jianrong Tan

Performing online condition monitoring for nuclear turbines in the rapidly changing environment is a challenging but imperative task to enhance the safety and reliability of nuclear power plants. Given the nonlinear and dynamic properties of nuclear turbine operation, this paper proposes an innovative method for condition monitoring. Specifically, the paper first redesigns time augmented matrices based on lagged data to reflect the process dynamics. Subsequently, a dynamic auto-regressive model, integrated with the variant of kernel partial least squares, is built between input and output variables, which represents auto-correlations and cross-correlations of operation data simultaneously. The prediction of the model serves as the baseline for the monitoring indicator. Additionally, since the operation process involves variable working excitation and random noise, making static control limits insufficient to satisfy the requirements of condition monitoring, the proposed method utilizes a novel monitoring indicator based on local information increment. The indicator comprehensively incorporates the prediction value and past measurement for monitoring statistics and control limits. Finally, the proposed method is applied to a real nuclear turbine operation process, and the results are compared with three other methods to demonstrate its superiority.

TCS Journal 2024 Journal Article

Finding the edges in optimal Hamiltonian cycles based on frequency quadrilaterals

  • Yong Wang

The sufficient condition is improved for identifying more edges in optimal Hamiltonian cycles based on frequency quadrilaterals. If the frequency 5s related to an edge occupy more than two-thirds with respect to all frequency quadrilaterals containing it and the others are frequency 3s, this edge is in an optimal Hamiltonian cycle. We also proved that if an edge is contained in an optimal Hamiltonian cycle of one graph, it will be contained in an optimal Hamiltonian cycle of a second graph with a probability bigger than two-thirds as the second graph is expanded from the first graph by adding one vertex. The experiments are conducted to approve the findings.

AAAI Conference 2024 Short Paper

Improving IP Geolocation With Target-Centric IP Graph (Student Abstract)

  • Kai Yang
  • Jiayang Li
  • Wenxin Tai
  • Zhenhui Li
  • Ting Zhong
  • Guangqiang Yin
  • Yong Wang

Accurate IP geolocation is indispensable for location-aware applications. While recent advances based on router-centric IP graphs are considered cutting-edge, one challenge remain: the prevalence of sparse IP graphs (14.24% with fewer than 10 nodes, 9.73% isolated) limits graph learning. To mitigate this issue, we designate the target host as the central node and aggregate multiple last-hop routers to construct the target-centric IP graph, instead of relying solely on the router with the smallest last-hop latency as in previous works. Experiments on three real-world datasets show that our method significantly improves the geolocation accuracy compared to existing baselines.

AAAI Conference 2024 Conference Paper

On the Concept Trustworthiness in Concept Bottleneck Models

  • Qihan Huang
  • Jie Song
  • Jingwen Hu
  • Haofei Zhang
  • Yong Wang
  • Mingli Song

Concept Bottleneck Models (CBMs), which break down the reasoning process into the input-to-concept mapping and the concept-to-label prediction, have garnered significant attention due to their remarkable interpretability achieved by the interpretable concept bottleneck. However, despite the transparency of the concept-to-label prediction, the mapping from the input to the intermediate concept remains a black box, giving rise to concerns about the trustworthiness of the learned concepts (i.e., these concepts may be predicted based on spurious cues). The issue of concept untrustworthiness greatly hampers the interpretability of CBMs, thereby hindering their further advancement. To conduct a comprehensive analysis on this issue, in this study we establish a benchmark to assess the trustworthiness of concepts in CBMs. A pioneering metric, referred to as concept trustworthiness score, is proposed to gauge whether the concepts are derived from relevant regions. Additionally, an enhanced CBM is introduced, enabling concept predictions to be made specifically from distinct parts of the feature map, thereby facilitating the exploration of their related regions. Besides, we introduce three modules, namely the cross-layer alignment (CLA) module, the cross-image alignment (CIA) module, and the prediction alignment (PA) module, to further enhance the concept trustworthiness within the elaborated CBM. The experiments on five datasets across ten architectures demonstrate that without using any concept localization annotations during training, our model improves the concept trustworthiness by a large margin, meanwhile achieving superior accuracy to the state-of-the-arts. Our code is available at https://github.com/hqhQAQ/ProtoCBM.

EAAI Journal 2024 Journal Article

SCGRFuse: An infrared and visible image fusion network based on spatial/channel attention mechanism and gradient aggregation residual dense blocks

  • Yong Wang
  • Jianfei Pu
  • Duoqian Miao
  • L. Zhang
  • Lulu Zhang
  • Xin Du

The goal of image fusion is to retain the strengths of different images in the fused result. However, existing fusion algorithms are often complex in design and overlook the influence of attention mechanisms on deep features. To address these issues, we propose an image fusion network based on spatial/channel attention mechanisms and gradient-aggregated residual dense blocks(SCGRFuse). Firstly, we design a novel gradient-aggregated residual dense block (GRXDB) that combines the advantages of ResNeXt and DenseNet, which integrating the Sobel and Laplacian operators to preserve both strong and weak texture features. Then, we introduce spatial and channel attention mechanisms to refine the channel and spatial information of feature maps, enhancing their information capturing capability. Additionally, we leverage a pooling fusion block to merge the refined spatial and channel feature maps, yielding high-quality fusion features. Compared to the existing state-of-the-art methods, experimental results on the MSRS, RoadScene and TNO datasets demonstrate the outstanding fusion performance of our proposed approach. In addition, in the task-driven experiments, SCGRFuse achieved an mIoU accuracy of 71. 37%.

JBHI Journal 2023 Journal Article

A Real-Time Bionic Method Inspired by Neural Oscillators for Estimation and Extraction of Pathological Tremor

  • Feiyun Xiao
  • Biqing Zhong
  • Yanming Wu
  • Yong Wang

The typical representative of pathological tremor is Parkinson's disease. One of the pathogenesis is that the synchronized neural oscillations within and between brain areas are affected. Inspired by this, this work proposes an algorithm based on neural oscillator to extract voluntary motion and estimate tremor motion in real time, which is named as RTBNO. This algorithm is composed of multiple adaptive modified Hopf oscillators linear combiner. The combiner is divided into two parts: one is used to estimate tremor motion and the other is applied to estimate voluntary motion. As it is updated iteratively in real time, this method has no phase delay. The performance of the proposed method was verified by the simulated action tremor and the actual experimental results of twenty Parkinson's disease patients. For the rest tremor signals of patients, the mean Root Mean Square Error (RMSE) values between the estimated signal and the actual signal was 0. 0272±0. 0077. The mean RMSE values between the estimated voluntary movement from action tremor and the actual voluntary movement were 0. 0360±0. 0097 (pick and put motion) and 0. 0380±0. 0083 (drawing motion). The execution time for the corresponding 10 seconds data was 0. 0478s. The comparison results between the proposed method and the existing methods demonstrated the effectiveness of the proposed method.

EAAI Journal 2023 Journal Article

A Semi-Supervised Network Framework for low-light image enhancement

  • Jin Chen
  • Yong Wang
  • Yujuan Han

Existing supervised learning-based low-light image enhancement algorithms treat all degradations as a whole, resulting in limited enhancement performance. It is difficult for fully unsupervised learning to recover more hidden details, making the augmented results unsatisfactory for visual needs. To overwhelm the limitations of supervised and unsupervised learning, we propose a Semi-Supervised Network Framework (SSNF) to enhance low-light images. Specifically, we decouple the low-light image enhancement task into two stages. In the first stage of the SSNF, we employ methods based on information entropy and Retinex to improve the visibility of images. It is worth mentioning that this stage is a lightweight self-supervised network, which only needs to input low-light images and undergo minute-level training to achieve brightness improvement. In the second stage of the SSNF, we utilize U-Net and residual networks to remove problems such as noise and degradation existing in the first-stage enhancement results, thereby improving the visual properties of the enhanced images. It overwhelms the challenge of dealing with low-light images directly. We conduct extensive experiments on datasets such as LOL, synthetic, DICM, etc. The experimental results show that SSNF exhibits better visual effects and outperforms other advanced methods in performance metrics.

EAAI Journal 2023 Journal Article

An adaptive matching control method of multiple turboshaft engines

  • Yong Wang
  • Chuang Ji
  • Zhihua Xi
  • Haibo Zhang
  • Qijun Zhao

In order to overcome the problem that the conventional multi-engine matching control method never balance the service lives of engines and transmission system synchronously while engines have individual differences and performance degradations, an adaptive matching control method of multiple turboshaft engines is proposed and designed. Firstly, based on the deep neural network (DNN), the onboard adaptive model is established to simulate the coupling dynamics of multiple engines. It can automatically trace engine output torques, discharge temperatures of gas turbine and rotational speeds of compressor. Then, the numerical optimization problem of multi-engine matching is built to obtain the optimal discharge temperatures of gas turbine online. It is featured with a multi-objective function that integrates the maximum deviations of engine output torques, discharge temperatures of gas turbine and the relative rotational speeds of compressor. Finally, the optimal discharge temperatures of gas turbine are input to the conventional matching strategy to accomplish the adaptive matching control. The results demonstrate that compared with the conventional matching strategy, when the number of turboshaft engine is three, the adaptive matching control method can dramatically decrease the maximum matching error of compressor speeds and engine output torques by more than 40% and 60% individually with the maximum deviation of the discharge temperatures of gas turbine no more than 3%. It proves to be conducive to balancing the service lives of multiple engines and transmission system.

EAAI Journal 2023 Journal Article

MIANet: Multi-level temporal information aggregation in mixed-periodicity time series forecasting tasks

  • Sheng Wang
  • Xi Chen
  • Dongliang Ma
  • Chen Wang
  • Yong Wang
  • Honggang Qi
  • Gongjian Zhou
  • Qingli Li

Regular human activities generate a large number of time series with mixed periodicity that can reflect human behavior patterns and the societal working mechanism. When forecasting these time series, nonlinear neural networks often encounter some limitations, such as utilizing mixed-periodic patterns, balancing multi-level information, incorporating future vision, forecasting delays and scale insensitivity, which affect the forecasting accuracy. To address these problems, we propose the Multi-level Information Aggregation Network (MIANet), a novel neural network with four key characteristics: (i) a novel folded recurrent structure that dynamically updates the local and mini-local information at a global range in a compact manner; (ii) a new recurrent unit called Folded Convolution Aggregation Temporal Memory (FCATM) that extracts and aggregates neighbor-trends in local and mini-local data; (iii) a fusing decoder structure that promotes the sharing of forward–backward future information and adaptively adjusts relationships among adjacent points; and (iv) a new Skip-Autoregressive (SAR) linear strategy that addresses scale sensitivity issues. The SAR can be embedded as a plug-and-play component into other deep learning (DL) models. Compared with other baseline methods, MIANet obtains statistically significant improvements on six real-world datasets, as demonstrated by conducting two-sample t-tests, indicating that the MIANet can be applied to various predictive scenarios, such as road occupancy, electricity consumption, pedestrian flow and urban noise.

NeurIPS Conference 2023 Conference Paper

Optimized Covariance Design for AB Test on Social Network under Interference

  • Qianyi Chen
  • Bo Li
  • Lu Deng
  • Yong Wang

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the Stable Unit Treatment Value Assumption (SUTVA) and poses great challenge to experimental design. Existing network experimental design research was mostly based on the unbiased Horvitz-Thompson (HT) estimator with substantial data trimming to ensure unbiasedness at the price of high resultant estimation variance. In this paper, we strive to balance the bias and variance in designing randomized network experiments. Under a potential outcome model with 1-hop interference, we derive the bias and variance of the standard HT estimator and reveal their relation to the network topological structure and the covariance of the treatment assignment vector. We then propose to formulate the experimental design problem as to optimize the covariance matrix of the treatment assignment vector to achieve the bias and variance balance by minimizing the mean squared error (MSE) of the estimator. An efficient projected gradient descent algorithm is presented to the implement of the desired randomization scheme. Finally, we carry out extensive simulation studies to demonstrate the advantages of our proposed method over other existing methods in many settings, with different levels of model misspecification.

EAAI Journal 2022 Journal Article

A novel fractional time-delayed grey Bernoulli forecasting model and its application for the energy production and consumption prediction

  • Yong Wang
  • Xinbo He
  • Lei Zhang
  • Xin Ma
  • Wenqing Wu
  • Rui Nie
  • Pei Chi
  • Yuyang Zhang

Energy affects the stable and sustainable development of social economy. Energy prediction plays an important role in the process of China’s energy market transformation. Scientific and reasonable energy predicting method can help government to make decisions effectively, and then adjust energy structure and industrial layout. The energy field is full of fractional order phenomenon and nonlinear disturbance. Aiming at the energy data sets with the characteristics of scarcity, complexity and nonlinear, a mathematical model including time delay term and Bernoulli equation can be used to fit this trend. A new fractional time-delayed grey Bernoulli model is proposed, and the new model has a wider application in the nonlinear field. The model is discretized by integral, and the least square estimation of the linear parameters and the approximate time response equation are obtained. The Grey Wolf Optimizer (GWO) is used to search the optimal parameters of the model. In addition, the energy prediction model is established from the perspective of renewable energy and fossil energy, and the effectiveness of the model is verified by three actual cases of renewable energy, crude oil and fossil fuel. Compared with the other seven grey models, the results show that the new model has higher prediction performance. Finally, the energy development trend in the next few years is predicted by using the proposed model, and relevant conclusions are drawn according to the prediction results.

EAAI Journal 2022 Journal Article

A novel self-adaptive fractional multivariable grey model and its application in forecasting energy production and conversion of China

  • Yong Wang
  • Li Wang
  • Lingling Ye
  • Xin Ma
  • Wenqing Wu
  • Zhongsen Yang
  • Xinbo He
  • Lei Zhang

Energy production and conversion have a significant impact on the economic development of all countries in the world. China’s energy production and conversion are large. Therefore, accurate mid-to-long term China’s energy production and conversion forecasting is becoming more and more important for integrating energy systems and energy strategic planning. For this purpose, a novel fractional grey sequence is proposed based on Grunwald–Letnikov fractional calculus. Furthermore, a novel self-adaptive fractional multivariable grey model is proposed based on the novel sequence. In this article, we compare several classical optimization algorithms and finally choose Particle Swarm Optimization (PSO) to compute the parameters. In addition, Monte-Carlo simulation and probability density analysis (PDA) are presented in this article to verify the model’s performance. Monte-Carlo simulation reduces the randomness of the results of the model runs to a certain extent. Probability density analysis visualizes this randomness through kernel density estimation (KDE). This paper compares the new model with the existing seven grey models and predicts the total energy consumption per capita, energy conversion efficiency and total renewable energy in China, respectively. The experimental results show that the new model is superior to the other seven models in terms of stability and prediction accuracy.

YNIMG Journal 2022 Journal Article

Accuracy and reliability of diffusion imaging models

  • Nicole A. Seider
  • Babatunde Adeyemo
  • Ryland Miller
  • Dillan J. Newbold
  • Jacqueline M. Hampton
  • Kristen M. Scheidter
  • Jerrel Rutlin
  • Timothy O. Laumann

Diffusion imaging aims to non-invasively characterize the anatomy and integrity of the brain's white matter fibers. We evaluated the accuracy and reliability of commonly used diffusion imaging methods as a function of data quantity and analysis method, using both simulations and highly sampled individual-specific data (927-1442 diffusion weighted images [DWIs] per individual). Diffusion imaging methods that allow for crossing fibers (FSL's BedpostX [BPX], DSI Studio's Constant Solid Angle Q-Ball Imaging [CSA-QBI], MRtrix3's Constrained Spherical Deconvolution [CSD]) estimated excess fibers when insufficient data were present and/or when the data did not match the model priors. To reduce such overfitting, we developed a novel Bayesian Multi-tensor Model-selection (BaMM) method and applied it to the popular ball-and-stick model used in BedpostX within the FSL software package. BaMM was robust to overfitting and showed high reliability and the relatively best crossing-fiber accuracy with increasing amounts of diffusion data. Thus, sufficient data and an overfitting resistant analysis method enhance precision diffusion imaging. For potential clinical applications of diffusion imaging, such as neurosurgical planning and deep brain stimulation (DBS), the quantities of data required to achieve diffusion imaging reliability are lower than those needed for functional MRI.

AAAI Conference 2022 Short Paper

Large-Scale IP Usage Identification via Deep Ensemble Learning (Student Abstract)

  • Zhiyuan Wang
  • Fan Zhou
  • Kunpeng Zhang
  • Yong Wang

Understanding users’ behavior via IP addresses is essential towards numerous practical IP-based applications such as online content delivery, fraud prevention, and many others. Among which profiling IP address has been extensively studied, such as IP geolocation and anomaly detection. However, less is known about the scenario of an IP address, e. g. , dedicated enterprise network or home broadband. In this work, we initiate the first attempt to address a large-scale IP scenario prediction problem. Specifically, we collect IP scenario data from four regions and propose a novel deep ensemble learning-based model to learn IP assignment rules and complex feature interactions. Extensive experiments support that our method can make accurate IP scenario identification and generalize from data in one region to another.

YNIMG Journal 2021 Journal Article

Spontaneous transient brain states in EEG source space in disorders of consciousness

  • Yang Bai
  • Jianghong He
  • Xiaoyu Xia
  • Yong Wang
  • Yi Yang
  • Haibo Di
  • Xiaoli Li
  • Ulf Ziemann

Spontaneous transient states were recently identified by functional magnetic resonance imaging and magnetoencephalography in healthy subjects. They organize and coordinate neural activity in brain networks. How spontaneous transient states are altered in abnormal brain conditions is unknown. Here, we conducted a transient state analysis on resting-state electroencephalography (EEG) source space and developed a state transfer analysis to patients with disorders of consciousness (DOC). They uncovered different neural coordination patterns, including spatial power patterns, temporal dynamics, spectral shifts, and connectivity construction varies at potentially very fast (millisecond) time scales, in groups with different consciousness levels: healthy subjects, patients in minimally conscious state (MCS), and patients with vegetative state/unresponsive wakefulness syndrome (VS/UWS). Machine learning based on transient state features reveal high classification accuracy between MCS and VS/UWS. This study developed methodology of transient states analysis on EEG source space and abnormal brain conditions. Findings correlate spontaneous transient states with human consciousness and suggest potential roles of transient states in brain disease assessment.

AAAI Conference 2020 Conference Paper

Go From the General to the Particular: Multi-Domain Translation with Domain Transformation Networks

  • Yong Wang
  • Longyue Wang
  • Shuming Shi
  • Victor O.K. Li
  • Zhaopeng Tu

The key challenge of multi-domain translation lies in simultaneously encoding both the general knowledge shared across domains and the particular knowledge distinctive to each domain in a unified model. Previous work shows that the standard neural machine translation (NMT) model, trained on mixed-domain data, generally captures the general knowledge, but misses the domain-specific knowledge. In response to this problem, we augment NMT model with additional domain transformation networks to transform the general representations to domain-specific representations, which are subsequently fed to the NMT decoder. To guarantee the knowledge transformation, we also propose two complementary supervision signals by leveraging the power of knowledge distillation and adversarial learning. Experimental results on several language pairs, covering both balanced and unbalanced multi-domain translation, demonstrate the effectiveness and universality of the proposed approach. Encouragingly, the proposed unified model achieves comparable results with the fine-tuning approach that requires multiple models to preserve the particular knowledge. Further analyses reveal that the domain transformation networks successfully capture the domain-specific knowledge as expected. 1

IJCAI Conference 2020 Conference Paper

Lexical-Constraint-Aware Neural Machine Translation via Data Augmentation

  • Guanhua Chen
  • Yun Chen
  • Yong Wang
  • Victor O. K. Li

Leveraging lexical constraint is extremely significant in domain-specific machine translation and interactive machine translation. Previous studies mainly focus on extending beam search algorithm or augmenting the training corpus by replacing source phrases with the corresponding target translation. These methods either suffer from the heavy computation cost during inference or depend on the quality of the bilingual dictionary pre-specified by user or constructed with statistical machine translation. In response to these problems, we present a conceptually simple and empirically effective data augmentation approach in lexical constrained neural machine translation. Specifically, we make constraint-aware training data by first randomly sampling the phrases of the reference as constraints, and then packing them together into the source sentence with a separation symbol. Extensive experiments on several language pairs demonstrate that our approach achieves superior translation results over the existing systems, improving translation of constrained sentences without hurting the unconstrained ones.

AAAI Conference 2020 Conference Paper

PSENet: Psoriasis Severity Evaluation Network

  • Yi Li
  • Zhe Wu
  • Shuang Zhao
  • Xian Wu
  • Yehong Kuang
  • YangTian Yan
  • Shen Ge
  • Kai Wang

Psoriasis is a chronic skin disease which affects hundreds of millions of people around the world. This disease cannot be fully cured and requires lifelong caring. If the deterioration of Psoriasis is not detected and properly treated in time, it could cause serious complications or even lead to a life threat. Therefore, a quantitative measurement that can track the Psoriasis severity is necessary. Currently, PASI (Psoriasis Area and Severity Index) is the most frequently used measurement in clinical practices. However, PASI has the following disadvantages: (1) Time consuming: calculating PASI usually takes more than 30 minutes which poses a heavy burden on dermatologists; and (2) Inconsistency: due to the complexity of PASI calculation, different or even the same dermatologist could give different scores for the same case. To overcome these drawbacks, we propose PSENet which applies deep neural networks to estimate Psoriasis severity based on skin lesion images. Different from typical deep learning frameworks for image processing, PSENet has the following characteristics: (1) PSENet introduces a score re- fine module which is able to capture the visual features of skin at both coarse and fine-grained granularities; (2) PSENet uses siamese structure in training and accepts pairwise inputs, which reduces the dependency on large amount of training data; and (3) PSENet can not only estimate the severity, but also locate the skin lesion regions from the input image. To train and evaluate PSENet, we work with professional dermatologists from a top hospital and spend years in building a golden dataset. The experimental results show that PSENet can achieve the mean absolute error of 2. 21 and the accuracy of 77. 87% in pair comparison, outperforming baseline methods. Overall, PSENet not only relieves dermatologists from the dull PASI calculation but also enables patients to track Psoriasis severity in a much more convenient manner.

EAAI Journal 2020 Journal Article

Robust RGB-D tracking via compact CNN features

  • Yong Wang
  • Xian Wei
  • Lingkun Luo
  • Wen Wen
  • Yang Wang

Feature representation is at the core of visual tracking. This paper presents a robust tracking method in RGB-D videos. Firstly, the RGB and depth images are separately encoded using a hierarchical convolutional neural network (CNN) features. Secondly, in order to reduce computation cost, we exploit random projection to compress the CNN features. The high dimensional CNN features are randomly projected into a low dimensional feature space. The correlation filter tracking framework is then independently carried out in RGB and depth images. And backward tracking scheme is adopted to evaluate the tracking results in these two images. The final position is determined according to the tracked location in the two image channels. In addition, model updating is implemented adaptively. Our tracker is evaluated on two RGB-D benchmark datasets and achieves comparable results to the other state-of-the-art RGB-D tracking methods.

IROS Conference 2019 Conference Paper

Local Pose optimization with an Attention-based Neural Network

  • Yiling Liu
  • Hesheng Wang 0001
  • Fan Xu 0004
  • Yong Wang
  • Weidong Chen 0001
  • Qirong Tang

In this paper, we propose a novel pose optimizer which can be inserted into either supervised or unsupervised end-to-end visual odometry for the purpose of local pose optimization. The pose optimizer is an analogue of the pose graph optimization used in traditional VSLAM algorithms. Local pose optimization is performed by an attention-based neural network which iteratively refines the predicted pose estimates of an image snippet. Instead of complicated graph convolutional network, the attention mechanism based on geometric consistency of trajectory constraint is utilized because pose features whose spatial distribution is not important can be flattened to vectors and then processed. The pose optimizer is aimed at improving pose estimation accuracy by redistributing errors of pose estimates. Quantitative and qualitative evaluation of the proposed approach on the KITTI Odometry dataset [1] is presented to demonstrate its effectiveness in improving pose estimation accuracy and minimizing pose drift.

YNICL Journal 2019 Journal Article

Quantification of white matter cellularity and damage in preclinical and early symptomatic Alzheimer's disease

  • Qing Wang
  • Yong Wang
  • Jingxia Liu
  • Courtney L. Sutphen
  • Carlos Cruchaga
  • Tyler Blazey
  • Brian A. Gordon
  • Yi Su

Interest in understanding the roles of white matter (WM) inflammation and damage in the pathophysiology of Alzheimer disease (AD) has been growing significantly in recent years. However, in vivo magnetic resonance imaging (MRI) techniques for imaging inflammation are still lacking. An advanced diffusion-based MRI method, neuro-inflammation imaging (NII), has been developed to clinically image and quantify WM inflammation and damage in AD. Here, we employed NII measures in conjunction with cerebrospinal fluid (CSF) biomarker classification (for β-amyloid (Aβ) and neurodegeneration) to evaluate 200 participants in an ongoing study of memory and aging. Elevated NII-derived cellular diffusivity was observed in both preclinical and early symptomatic phases of AD, while disruption of WM integrity, as detected by decreased fractional anisotropy (FA) and increased radial diffusivity (RD), was only observed in the symptomatic phase of AD. This may suggest that WM inflammation occurs earlier than WM damage following abnormal Aβ accumulation in AD. The negative correlation between NII-derived cellular diffusivity and CSF Aβ42 level (a marker of amyloidosis) may indicate that WM inflammation is associated with increasing Aβ burden. NII-derived FA also negatively correlated with CSF t-tau level (a marker of neurodegeneration), suggesting that disruption of WM integrity is associated with increasing neurodegeneration. Our findings demonstrated the capability of NII to simultaneously image and quantify WM cellularity changes and damage in preclinical and early symptomatic AD. NII may serve as a clinically feasible imaging tool to study the individual and composite roles of WM inflammation and damage in AD.

EAAI Journal 2019 Journal Article

Robust visual tracking based on response stability

  • Yong Wang
  • Xinbin Luo
  • Lu Ding
  • Shan Fu
  • Xian Wei

In this paper, a new approach of response stability based for visual object tracking is developed. This approach proposes a response stability criterion to measure the tracking quality and fuse tracking results of multiple layers of a convolutional neural network (CNN). Inspired by recent detection based methods for visual tracking, the detection capability of EdgeBoxes is investigated, and proposes to re-detect target when tracking failure occurs. In addition, 3D locally adaptive regression kernels (LARK) feature is employed in correlation filter based tracking framework. Extensive experimental results and performance compared with state-of-the-art tracking algorithms on challenging benchmark datasets show that our method is more accurate and robust.

AAAI Conference 2018 Conference Paper

Search Engine Guided Neural Machine Translation

  • Jiatao Gu
  • Yong Wang
  • Kyunghyun Cho
  • Victor O.K. Li

In this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. In the first stage– retrieval stage–, an off-the-shelf, black-box search engine is used to retrieve a small subset of sentence pairs from a training set given a source sentence. These pairs are further filtered based on a fuzzy matching score based on edit distance. In the second stage–translation stage–, a novel translation model, called search engine guided NMT (SEG-NMT), seamlessly uses both the source sentence and a set of retrieved sentence pairs to perform the translation. Empirical evaluation on three language pairs (En-Fr, En-De, and En-Es) shows that the proposed approach significantly outperforms the baseline approach and the improvement is more significant when more relevant sentence pairs were retrieved.

IS Journal 2014 Journal Article

An Energy-Efficient and Swarm Intelligence-Based Routing Protocol for Next-Generation Sensor Networks

  • Yong Wang
  • Changle Li
  • Yulong Duan
  • Jin Yang
  • Xiang Cheng

After providing a brief overview of routing protocols for next-generation sensor networks (NGSNs), the authors propose Bee-Sensor-C, an energy-efficient, swarm intelligence-based, and scalable multipath routing protocol that integrates dynamic clustering, multipath routing, and bee-inspired routing to meet the performance requirements of NGSNs. A performance evaluation is also provided.

YNIMG Journal 2014 Journal Article

Quantifying white matter tract diffusion parameters in the presence of increased extra-fiber cellularity and vasogenic edema

  • Chia-Wen Chiang
  • Yong Wang
  • Peng Sun
  • Tsen-Hsuan Lin
  • Kathryn Trinkaus
  • Anne H. Cross
  • Sheng-Kwei Song

The effect of extra-fiber structural and pathological components confounding diffusion tensor imaging (DTI) computation was quantitatively investigated using data generated by both Monte-Carlo simulations and tissue phantoms. Increased extent of vasogenic edema, by addition of various amount of gel to fixed normal mouse trigeminal nerves or by increasing non-restricted isotropic diffusion tensor components in Monte-Carlo simulations, significantly decreased fractional anisotropy (FA) and increased radial diffusivity, while less significantly increased axial diffusivity derived by DTI. Increased cellularity, mimicked by graded increase of the restricted isotropic diffusion tensor component in Monte-Carlo simulations, significantly decreased FA and axial diffusivity with limited impact on radial diffusivity derived by DTI. The MC simulation and tissue phantom data were also analyzed by the recently developed diffusion basis spectrum imaging (DBSI) to simultaneously distinguish and quantify the axon/myelin integrity and extra-fiber diffusion components. Results showed that increased cellularity or vasogenic edema did not affect the DBSI-derived fiber FA, axial or radial diffusivity. Importantly, the extent of extra-fiber cellularity and edema estimated by DBSI correlated with experimentally added gel and Monte-Carlo simulations. We also examined the feasibility of applying 25-direction diffusion encoding scheme for DBSI analysis on coherent white matter tracts. Results from both phantom experiments and simulations suggested that the 25-direction diffusion scheme provided comparable DBSI estimation of both fiber diffusion parameters and extra-fiber cellularity/edema extent as those by 99-direction scheme. An in vivo 25-direction DBSI analysis was performed on experimental autoimmune encephalomyelitis (EAE, an animal model of human multiple sclerosis) optic nerve as an example to examine the validity of derived DBSI parameters with post-imaging immunohistochemistry verification. Results support that in vivo DBSI using 25-direction diffusion scheme correctly reflect the underlying axonal injury, demyelination, and inflammation of optic nerves in EAE mice.

ICRA Conference 2011 Conference Paper

Hybrid map-based navigation for intelligent wheelchair

  • Yong Wang
  • Weidong Chen

A navigation system based on hybrid map for intelligent wheelchair is presented. The system is consisted of hybrid map building, localization, path planning and trajectory following. The hybrid map includes a series of small probabilistic grid maps (PGM) and a global topological map (GTM). They are built simultaneously and easily using the human-guided method. Then on the hybrid map, the localization and the real-time path planning algorithms are realized smartly and effectively. The experiments and applications results show that the human-guided method integrates both the computer's modeling ability and the human's sensory perception to the environment. The hybrid map is easy to solve the loop-closure and doorway problems that enhances the robustness against uncertainty of sensors. It also can improve the efficiency in large-scale SLAM. Further more, at an elderly home we did the activities of daily living (ADL) testing and at Shanghai Expo 2010 we demonstrated the wheelchair system by offering trial rides to visitors.

IJCAI Conference 2011 Conference Paper

Local and Structural Consistency for Multi-Manifold Clustering

  • Yong Wang
  • Yuan Jiang
  • Yi Wu
  • Zhi-Hua Zhou

Data sets containing multi-manifold structures are ubiquitous in real-world tasks, and effective grouping of such data is an important yet challenging problem. Though there were many studies on this problem, it is not clear on how to design principled methods for the grouping of multiple hybrid manifolds. In this paper, we show that spectral methods are potentially helpful for hybridmanifold clustering when the neighborhood graph is constructed to connect the neighboring samples from the same manifold. However, traditional algorithms which identify neighbors according to Euclidean distance will easily connect samples belonging to different manifolds. To handle this drawback, we propose a new criterion, i. e. , local and structural consistency criterion, which considers the neighboring information as well as the structural information implied by the samples. Based on this criterion, we develop a simple yet effective algorithm, named Local and Structural Consistency (LSC), for clustering with multiple hybrid manifolds. Experiments show that LSC achieves promising performance.

AAAI Conference 2011 Conference Paper

Localized K-Flats

  • Yong Wang
  • Yuan Jiang
  • Yi Wu
  • Zhi-Hua Zhou

K-flats is a model-based linear manifold clustering algorithm which has been successfully applied in many real-world scenarios. Though some previous works have shown that K-flats doesn’t always provide good performance, little effort has been devoted to analyze its inherent deficiency. In this paper, we address this challenge by showing that the deteriorative performance of K-flats can be attributed to the usual reconstruction error measure and the infinitely extending representations of linear models. Then we propose Localized K-flats algorithm (LKF), which introduces localized representations of linear models and a new distortion measure, to remove confusion among different clusters. Experiments on both synthetic and real-world data sets demonstrate the efficiency of the proposed algorithm. Moreover, preliminary experiments show that LKF has the potential to group manifolds with nonlinear structure.

v2026.09.13