Arrow Research search

Author name cluster

Tong Xia

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

JBHI Journal 2026 Journal Article

Continuous Mobile Audio Monitoring for Sleep Apnea Detection

  • Jing Han
  • Tong Xia
  • Cecilia Mascolo

Audio-based sleep apnea detection methods hold great potential to improve access to diagnosis, by providing unattended sleep apnea screening at home via sound collected from mobile sensors during sleep. Our research involved a thorough comparison and evaluation of tracheal and ambient microphone recordings for sleep apnea detection with different granularities. Utilising a variety of acoustic representations and sophisticated deep learning architectures, we performed an extensive analysis on the open PSG-Audio dataset, which encompasses over 850 hours of audio data from 194 subjects. For sleep apnea classification, the most effective model showed a 90. 8 % accuracy in detecting sleep apnea, 83. 3 % accuracy when hypopneic and apneic events were detected separately, and 75. 7 % accuracy when apneic events were further divided into three sub-categories. On overnight recordings, the model achieved a sensitivity of 0. 93 and a specificity of 1. 0 for moderate sleep apnea screening, and a sensitivity of 0. 84 and a specificity of 0. 97 for severe sleep apnea screening. This research also provided a unique study to compare and combine respiratory sounds from two different types of sensors for sleep apnea detection. The high performance of our model provides a promising avenue for enabling remote diagnosis and monitoring of sleep apnea.

NeurIPS Conference 2024 Conference Paper

Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking

  • Yuwei Zhang
  • Tong Xia
  • Jing Han
  • Yu Y. Wu
  • Georgios Rizos
  • Yang Liu
  • Mohammed Mosuily
  • Jagmohan Chauhan

Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets ($\sim$136K samples, over 400 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https: //github. com/evelyn0414/OPERA.

JBHI Journal 2024 Journal Article

Uncertainty-Aware Health Diagnostics via Class-Balanced Evidential Deep Learning

  • Tong Xia
  • Ting Dang
  • Jing Han
  • Lorena Qendro
  • Cecilia Mascolo

Uncertainty quantification is critical for ensuring the safety of deep learning-enabled health diagnostics, as it helps the model account for unknown factors and reduces the risk of misdiagnosis. However, existing uncertainty quantification studies often overlook the significant issue of class imbalance, which is common in medical data. In this paper, we propose a class-balanced evidential deep learning framework to achieve fair and reliable uncertainty estimates for health diagnostic models. This framework advances the state-of-the-art uncertainty quantification method of evidential deep learning with two novel mechanisms to address the challenges posed by class imbalance. Specifically, we introduce a pooling loss that enables the model to learn less biased evidence among classes and a learnable prior to regularize the posterior distribution that accounts for the quality of uncertainty estimates. Extensive experiments using benchmark data with varying degrees of imbalance and various naturally imbalanced health data demonstrate the effectiveness and superiority of our method. Our work pushes the envelope of uncertainty quantification from theoretical studies to realistic healthcare application scenarios. By enhancing uncertainty estimation for class-imbalanced data, we contribute to the development of more reliable and practical deep learning-enabled health diagnostic systems.

JBHI Journal 2023 Journal Article

MT-FiST: A Multi-Task Fine-Grained Spatial-Temporal Framework for Surgical Action Triplet Recognition

  • Yuchong Li
  • Tong Xia
  • Huoling Luo
  • Baochun He
  • Fucang Jia

Surgical action triplet recognition plays a significant role in helping surgeons facilitate scene analysis and decision-making in computer-assisted surgeries. Compared to traditional context-aware tasks such as phase recognition, surgical action triplets, comprising the instrument, verb, and target, can offer more comprehensive and detailed information. However, current triplet recognition methods fall short in distinguishing the fine-grained subclasses and disregard temporal correlation in action triplets. In this article, we propose a multi-task fine-grained spatial-temporal framework for surgical action triplet recognition named MT-FiST. The proposed method utilizes a multi-label mutual channel loss, which consists of diversity and discriminative components. This loss function decouples global task features into class-aligned features, enabling the learning of more local details from the surgical scene. The proposed framework utilizes partial shared-parameters LSTM units to capture temporal correlations between adjacent frames. We conducted experiments on the CholecT50 dataset proposed in the MICCAI 2021 Surgical Action Triplet Recognition Challenge. Our framework is evaluated on the private test set of the challenge to ensure fair comparisons. Our model apparently outperformed state-of-the-art models in instrument, verb, target, and action triplet recognition tasks, with mAPs of 82. 1% (+4. 6%), 51. 5% (+4. 0%), 45. 50% (+7. 8%), and 35. 8% (+3. 1%), respectively. The proposed MT-FiST boosts the recognition of surgical action triplets in a context-aware surgical assistant system, further solving multi-task recognition by effective temporal aggregation and fine-grained features.

AAAI Conference 2021 Conference Paper

AttnMove: History Enhanced Trajectory Recovery via Attentional Network

  • Tong Xia
  • Yunhan Qi
  • Jie Feng
  • Fengli Xu
  • Funing Sun
  • Diansheng Guo
  • Yong Li

A considerable amount of mobility data has been accumulated due to the proliferation of location-based service. Nevertheless, compared with mobility data from transportation systems like the GPS module in taxis, this kind of data is commonly sparse in terms of individual trajectories in the sense that users do not access mobile services and contribute their data all the time. Consequently, the sparsity inevitably weakens the practical value of the data even it has a high user penetration rate. To solve this problem, we propose a novel attentional neural network-based model, named AttnMove, to densify individual trajectories by recovering unobserved locations at a fine-grained spatial-temporal resolution. To tackle the challenges posed by sparsity, we design various intraand inter- trajectory attention mechanisms to better model the mobility regularity of users and fully exploit the periodical pattern from long-term history. We evaluate our model on two real-world datasets, and extensive results demonstrate the performance gain compared with the state-of-the-art methods. This also shows that, by providing high-quality mobility data, our model can benefit a variety of mobility-oriented down-stream applications.

NeurIPS Conference 2021 Conference Paper

COVID-19 Sounds: A Large-Scale Audio Dataset for Digital Respiratory Screening

  • Tong Xia
  • Dimitris Spathis
  • Chlo{\"e} Brown
  • J Ch
  • Andreas Grammenos
  • Jing Han
  • Apinan Hasthanasombat
  • Erika Bondareva

Audio signals are widely recognised as powerful indicators of overall health status, and there has been increasing interest in leveraging sound for affordable COVID-19 screening through machine learning. However, there has also been scepticism regarding the initial efforts, due to perhaps the lack of reproducibility, large datasets and transparency which unfortunately is often an issue with machine learning for health. To facilitate the advancement and openness of audio-based machine learning for respiratory health, we release a dataset consisting of 53, 449 audio samples (over 552 hours in total) crowd-sourced from 36, 116 participants through our COVID-19 Sounds app. Given its scale, this dataset is comprehensive in terms of demographics and spectrum of health conditions. It also provides participants' self-reported COVID-19 testing status with 2, 106 samples tested positive. To the best of our knowledge, COVID-19 Sounds is the largest multi-modal dataset of COVID-19 respiratory sounds: it consists of three modalities including breathing, cough, and voice recordings. Additionally, in this paper, we report on several benchmarks for two principal research tasks: respiratory symptoms prediction and COVID-19 prediction. For these tasks we demonstrate performance with a ROC-AUC of over 0. 7, confirming both the promise of machine learning approaches based on these types of datasets as well as the usability of our data for such tasks. We describe a realistic experimental setting that hopes to pave the way to a fair performance evaluation of future models. In addition, we reflect on how the released dataset can help to scale some existing studies and enable new research directions, which inspire and benefit a wide range of future works.

IJCAI Conference 2020 Conference Paper

A Sequential Convolution Network for Population Flow Prediction with Explicitly Correlation Modelling

  • Jie Feng
  • Ziqian Lin
  • Tong Xia
  • Funing Sun
  • Diansheng Guo
  • Yong Li

Population flow prediction is one of the most fundamental components in many applications from urban management to transportation schedule. It is challenging due to the complicated spatial-temporal correlation. While many studies have been done in recent years, they fail to simultaneously and effectively model the spatial correlation and temporal variations among population flows. In this paper, we propose Convolution based Sequential and Cross Network (CSCNet) to solve them. On the one hand, we design a CNN based sequential structure with progressively merging the flow features from different time in different CNN layers to model the spatial-temporal information simultaneously. On the other hand, we make use of the transition flow as the proxy to efficiently and explicitly capture the dynamic correlation between different types of population flows. Extensive experiments on 4 datasets demonstrate that CSCNet outperforms the state-of-the-art baselines by reducing the prediction error around 7. 7%∼10. 4%.

TIST Journal 2020 Journal Article

DeepApp

  • Tong Xia
  • Yong Li
  • Jie Feng
  • Depeng Jin
  • Qing Zhang
  • Hengliang Luo
  • Qingmin Liao

Smartphone mobile application (App) usage prediction, i.e., which Apps will be used next, is beneficial for user experience improvement. Through an in-depth analysis on a real-world dataset, we find that App usage is highly spatio-temporally correlated and personalized. Given the ability to model complex spatio-temporal contexts, we aim to apply deep learning to achieve high prediction accuracy. However, the personalization yields a problem: training one network for each individual suffers from data scarcity, yet training one deep neural network for all users often fails to uncover user preference. In this article, we propose a novel App usage prediction framework, named DeepApp, to achieve context-aware prediction via multi-task learning. To tackle the challenge of data scarcity, we train one general network for multiple users to share common patterns. To better utilize the spatio-temporal contexts, we supplement a location prediction task in the multi-task learning framework to learn spatio-temporal relations. As for the personalization, we add a user identification task to capture user preference. We evaluate DeepApp on the large-scale dataset by extensive experiments. Results demonstrate that DeepApp outperforms the start-of-the-art baseline by 6.44%.

v2026.09.13