Arrow Research search

Author name cluster

Weiwei Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

AAAI Conference 2025 Conference Paper

Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation

  • Jinda Xu
  • Yuhao Song
  • Daming Wang
  • Weiwei Zhao
  • Minghua Chen
  • Kangliang Chen
  • Qinya Li

In an era overwhelmed by vast amounts of data, the effective curation of web-crawl datasets is essential for optimizing model performance. This paper tackles the challenges associated with the unstructured and heterogeneous nature of such datasets. Traditional heuristic curation methods often inadequately capture complex features, resulting in biases and the exclusion of relevant data. We introduce an advanced, learning-driven approach, Ensemble Curation Of DAta ThroUgh Multimodal Operators, called EcoDatum, which employs a novel quality-guided deduplication method to balance feature distribution. EcoDatum strategically integrates various unimodal and multimodal data curation operators within a weak supervision ensemble framework, utilizing automated optimization to effectively score each data point. EcoDatum, which significantly improves the data curation quality and efficiency, outperforms existing state-of-the-art (SOTA) techniques, ranking 1st on the DataComp leaderboard with an average performance score of 0.182 across 38 diverse evaluation datasets. This represents a 28% improvement over the DataComp baseline method, demonstrating its effectiveness in improving dataset curation and model training efficiency.

YNIMG Journal 2025 Journal Article

RETRACTED: Test-retest reliability of coupling between cerebrospinal fluid flow and global brain activity after normal sleep and sleep deprivation

  • Weiwei Zhao
  • Joy Rao
  • Ruosi Wang
  • Ya Chai
  • Tianxin Mao
  • Peng Quan
  • Yao Deng
  • Wenwen Chen

The glymphatic system (GS) plays a key role in maintaining brain homeostasis by clearing metabolic waste during sleep, with the coupling between global blood-oxygen-level-dependent (gBOLD) and cerebrospinal fluid (CSF) signals serving as a potential marker for glymphatic clearance function. However, the test-retest reliability and spatial heterogeneity of gBOLD-CSF coupling after different sleep conditions remain unclear. In this study, we assessed the test-retest reliability of gBOLD-CSF coupling following either normal sleep or total sleep deprivation (TSD) in 64 healthy adults under controlled laboratory conditions. The reliability was high after normal sleep (ICC = 0.763) but decreased following TSD (ICC = 0.581). Moreover, spatial heterogeneity was evident in participants with normal sleep, with lower-order networks (visual, somatomotor, and attention) showing higher ICC values compared to higher-order networks (default-mode, limbic, and frontoparietal). This spatial variation was less distinct in the TSD group. These results demonstrate the robustness of the gBOLD-CSF coupling method and emphasize the significance of considering sleep history in glymphatic function research.

TIST Journal 2021 Journal Article

A GDPR-compliant Ecosystem for Speech Recognition with Transfer, Federated, and Evolutionary Learning

  • Di Jiang
  • Conghui Tan
  • Jinhua Peng
  • Chaotao Chen
  • Xueyang Wu
  • Weiwei Zhao
  • Yuanfeng Song
  • Yongxin Tong

Automatic Speech Recognition (ASR) is playing a vital role in a wide range of real-world applications. However, Commercial ASR solutions are typically “one-size-fits-all” products and clients are inevitably faced with the risk of severe performance degradation in field test. Meanwhile, with new data regulations such as the European Union’s General Data Protection Regulation (GDPR) coming into force, ASR vendors, which traditionally utilize the speech training data in a centralized approach, are becoming increasingly helpless to solve this problem, since accessing clients’ speech data is prohibited. Here, we show that by seamlessly integrating three machine learning paradigms (i.e., T ransfer learning, F ederated learning, and E volutionary learning (TFE)), we can successfully build a win-win ecosystem for ASR clients and vendors and solve all the aforementioned problems plaguing them. Through large-scale quantitative experiments, we show that with TFE, the clients can enjoy far better ASR solutions than the “one-size-fits-all” counterpart, and the vendors can exploit the abundance of clients’ data to effectively refine their own ASR products.

TIST Journal 2021 Journal Article

Industrial Federated Topic Modeling

  • Di Jiang
  • Yongxin Tong
  • Yuanfeng Song
  • Xueyang Wu
  • Weiwei Zhao
  • Jinhua Peng
  • Rongzhong Lian
  • Qian Xu

Probabilistic topic modeling has been applied in a variety of industrial applications. Training a high-quality model usually requires a massive amount of data to provide comprehensive co-occurrence information for the model to learn. However, industrial data such as medical or financial records are often proprietary or sensitive, which precludes uploading to data centers. Hence, training topic models in industrial scenarios using conventional approaches faces a dilemma: A party (i.e., a company or institute) has to either tolerate data scarcity or sacrifice data privacy. In this article, we propose a framework named Industrial Federated Topic Modeling (iFTM), in which multiple parties collaboratively train a high-quality topic model by simultaneously alleviating data scarcity and maintaining immunity to privacy adversaries. iFTM is inspired by federated learning, supports two representative topic models (i.e., Latent Dirichlet Allocation and SentenceLDA) in industrial applications, and consists of novel techniques such as private Metropolis-Hastings, topic-wise normalization, and heterogeneous model integration. We conduct quantitative evaluations to verify the effectiveness of iFTM and deploy iFTM in two real-life applications to demonstrate its utility. Experimental results verify iFTM’s superiority over conventional topic modeling.

v2026.09.13