Arrow Research search

Author name cluster

Huan Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

47 papers
2 author rows

Possible papers

47

TMLR Journal 2026 Journal Article

A Survey of Token Compression for Efficient Multimodal Large Language Models

  • Kele Shao
  • Keda TAO
  • Kejia Zhang
  • Sicheng Feng
  • Mu Cai
  • Yuzhang Shang
  • Haoxuan You
  • Can Qin

Multimodal large language models (MLLMs) have made remarkable strides, largely driven by their ability to process increasingly long and complex contexts, such as high-resolution images, extended video sequences, and lengthy audio input. While this ability significantly enhances MLLM capabilities, it introduces substantial computational challenges, primarily due to the quadratic complexity of self-attention mechanisms with numerous input tokens. To mitigate these bottlenecks, token compression has emerged as an auspicious and critical approach, efficiently reducing the number of tokens during both training and inference. In this paper, we present the first systematic survey and synthesis of the burgeoning field of multimodal long context token compression. Recognizing that effective compression strategies are deeply tied to the unique characteristics and redundancies of each modality, we categorize existing approaches by their primary data focus, enabling researchers to quickly access and learn methods tailored to their specific area of interest: (1) image-centric compression, which addresses spatial redundancy in visual data; (2) video-centric compression, which tackles spatio-temporal redundancy in dynamic sequences; and (3) audio-centric compression, which handles temporal and spectral redundancy in acoustic signals. Beyond this modality-driven categorization, we further dissect methods based on their underlying mechanisms, including transformation-based, similarity-based, attention-based, and query-based approaches. By providing a comprehensive and structured overview, this survey aims to consolidate current progress, identify key challenges, and inspire future research directions in this rapidly evolving domain.

AAAI Conference 2026 Conference Paper

OmniBench: A Comprehensive Benchmark Integrating Real-World, Time-sensitive, and Multi-Hop Questions with a Multi-Dimensional Hybrid Evaluation Framework

  • Wenjie Wang
  • Yufeng Jiang
  • Ge Sun
  • Chenghang Dong
  • Zheng Jun
  • Li Mengjie
  • Lixin Chen
  • Huan Wang

Recently, with the increasing capabilities of Large Language Models (LLMs), AI applications have gradually emerged to solve various problems in people's daily lives, so accurately measuring their performance and reliability is paramount. However, existing benchmarks predominantly rely on closed-ended, multiple-choice or short-answer question formats. While useful for assessment, these formats exhibit a significant gap compared to the diverse and open-ended nature of questions posed by real-world users. To bridge this gap, we produce OmniBench, a comprehensive open-domain benchmark. OmniBench is uniquely composed of authentic, user-generated questions harvested from real-world interactions on various websites and applications, covering 16 rigorously defined knowledge domains and 5 crucial user intents derived from a large-scale analysis of the mass corpus. Crucially, we propose three automated data construction pipelines that enable the continuous and periodic updating of the benchmark dataset. This approach not only ensures that the questions can keep up with current events, but also effectively mitigates the critical issue of data contamination prevalent in static benchmarks. Moreover, a multi-dimensional hybrid evaluation framework named OmniEval is proposed for evaluating the responses. This framework combines diverse metrics and evaluation methods to capture nuanced aspects of answer performance. Extensive validation demonstrates that this evaluation framework exhibits strong alignment with human judgments, ensuring the reliability of the benchmark results.

AAAI Conference 2026 Conference Paper

Toward Multimodal Fake News Detection by Multi-perspective Rationale Generation and Verification

  • Junyang Chen
  • Yueqian Li
  • Ka Chung Ng
  • Huan Wang
  • Liang-Jie Zhang

The rapid proliferation of social media platforms has led to a surge in multimodal fake news, where deceptive content often combines text and images to mislead audiences. Traditional unimodal detection methods struggle to address the complexity of such content, necessitating holistic multimodal approaches. While the latest advancements in Multimodal Large Language Models (MLLMs) offer new opportunities for enhancing detection performance by analyzing multi-dimensional features, including source credibility, cross-modal contradictions, emotional bias, and manipulative writing patterns, these methods suffer from a key flaw: a susceptibility to hallucinations or erroneous reasoning, which can lead to flawed conclusions and ultimately biased detection results. We propose the Multimodal Fake News Detection via Multi-perspective Rationale Generation and Verification (MMRGV) model to mitigate this challenge. Our method employs a cross-verification mechanism to screen and reconcile contradictions among different rationales, thereby preserving the LLM's analytical advantages while mitigating the impact of erroneous reasoning or hallucinations on the final detection. Subsequently, these optimized rationales are fused via an adaptive weighting strategy to output a robust final prediction. Extensive experiments on three benchmark datasets (Twitter, Weibo, and GossipCop) demonstrate the superiority of our method, achieving state-of-the-art accuracy of 0.9972, 0.9663, and 0.8772, respectively, and significantly outperforming existing baselines. These results validate the effectiveness of multi-perspective rationale generation and cross-verification in enhancing multimodal fake news detection, offering a resilient solution to combat misinformation in the era of generative AI.

AAAI Conference 2026 Conference Paper

Towards Zero-Shot Diabetic Retinopathy Grading: Learning Generalized Knowledge via Prompt-Driven Matching and Emulating

  • Huan Wang
  • Haoran Li
  • Yuxin Lin
  • Huaming Chen
  • Jun Yan
  • Lijuan Wang
  • Jiahua Shi
  • Qihao Xu

As one of the primary causes of visual impairment, Diabetic Retinopathy (DR) requires accurate and robust grading to facilitate timely diagnosis and intervention. Different from conventional DR grading methods that utilize single-view images, recent clinical studies have revealed that multi-view fundus images can significantly enhance DR grading performance by expanding the field of view (FOV). However, there is a long-tailed distribution problem in fundus image analysis, i.e., a high prevalence of mild DR grades and a low prevalence of rare ones (e.g., cases of high severity), which presents a significant challenge to developing a unified model capable of detecting rare or unseen DR grades not encountered during training. In this paper, we propose ProME-DR, a Prompt-driven zero-shot DR grading framework, which leverages prompt Matching and Emulating to recognize the unseen DR categories and views beyond the training set. ProME-DR disentangles the training process into two stages to learn generalized knowledge for novel DR disease grading. Initially, ProME-DR leverages two sets of prompt units to capture semantic and inter-view consistency knowledge via a split-and-mask manner, gathering instance-level DR visual clues. Subsequently, it constructs a concept-aware emulator to generate context prompt units, linking extensible knowledge learned from the previously seen DR attributes for zero-shot DR grading. Extensive experiments conducted on eight datasets and various scenarios confirm the superiority of ProME-DR.

AAAI Conference 2026 Conference Paper

Zero-shot Recommendation: Towards Class Semantic Relation Learning for Inferring Labels of Unseen Micro-videos

  • Junyang Chen
  • Huan Wang
  • Yirui Wu
  • Qiuzhen Lin
  • Yunfeng Diao
  • Junkai Ji

Micro-video label prediction plays a pivotal role on contemporary video-sharing platforms, such as Kwai and Tiktok. The emergence of video content lacking labels presents a formidable challenge for conventional user interest prediction methods. This paper addresses the challenge of micro-video label prediction, particularly for unseen videos, by proposing a zero-shot method called Class Semantic Relation Learning (CSRL). Unlike traditional user interest prediction models, CSRL leverages the pre-trained Large Language Model (LLM) to enhance prediction accuracy for unlabeled videos. The novelty of CSRL lies in its integration of three key components: a raw feature autoencoder, LLM-enhanced features, and a decomposed graph network. The decomposed graph network is specifically designed to disentangle the relationships between labeled and unlabeled videos, offering a significant improvement over previous methods. By fusing hidden topics with LLM-enhanced text, CSRL effectively handles sparse video features. Experiments on large-scale datasets from the Kwai platform show that CSRL achieves state-of-the-art results, with up to 44.64% improvement in Hit Ratio (HR), highlighting its superiority over existing zero-shot recommendation models in predicting user interests within the user-video network.

IJCAI Conference 2025 Conference Paper

ABNet: Mitigating Sample Imbalance in Anomaly Detection Within Dynamic Graphs

  • Yifan Hong
  • Muhammad Asif Ali
  • Huan Wang
  • Junyang Chen
  • Di Wang

In dynamic graphs, detecting anomalous nodes faces challenges due to sample imbalance, stemming from the scarcity of anomalous samples and feature representation bias. Existing methods often use unsupervised or semi-supervised learning to extract anomalous samples from unlabeled data, but struggle to obtain enough anomalous instances due to their low occurrence. Moreover, GNN-based approaches often prioritize normal samples, neglecting rare anomalies. To address these issues, we propose the Anomaly Balance Network (ABNet), designed to alleviate sample imbalance and enhance anomaly detection. ABNet includes three key components: a feature extractor that compares node features across time points to avoid bias, an anomaly augmenter that amplifies anomaly details and generates diverse anomalous samples, and an anomaly detector using meta-learning to adapt to graph evolution. Experimental results show that ABNet outperforms existing methods on three real-world datasets, effectively addressing sample imbalance.

ICLR Conference 2025 Conference Paper

Accessing Vision Foundation Models via ImageNet-1K

  • Yitian Zhang
  • Xu Ma 0005
  • Yue Bai
  • Huan Wang
  • Yun Fu 0001

Vision foundation models are renowned for the generalization ability due to massive training data. Nevertheless, they demand tremendous training resources, and the training data is often inaccessible, e.g., CLIP, DINOv2, posing great challenges to developing derivatives that could facilitate the research. In this work, we offer a very simple and general solution, named Proteus, to distill foundation models into smaller equivalents on ImageNet-1K without access to the original training data. Specifically, we remove the designs from conventional knowledge distillation settings that result in dataset bias and present three levels of training objectives, i.e., token, patch, and feature, to maximize the efficacy of knowledge transfer. In this manner, Proteus is trained at ImageNet-level costs with surprising ability, facilitating the accessibility of training foundation models for the broader research community. When leveraging DINOv2-g/14 as the teacher, Proteus-L/14 matches the performance of the Oracle method DINOv2-L/14 (142M training data) across 19 benchmarks and outperforms other vision foundation models including CLIP-L/14 (400M), OpenCLIP-L/14 (400M/2B) and SynCLR-L/14 (600M) with a significantly smaller training set of 1.2M images.

NeurIPS Conference 2025 Conference Paper

APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

  • Akshara Prabhakar
  • Zuxin Liu
  • Ming Zhu
  • Jianguo Zhang
  • Tulika Manoj Awalgaonkar
  • Shiyu Wang
  • Zhiwei Liu
  • Haolin Chen

Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that generates verifiable and diverse multi-turn agent data. In the first phase, our agentic pipeline produces detailed task blueprints with ground-truth actions, leveraging a committee of LLM reviewers and iterative feedback loops. These blueprints are then transformed into complete interaction trajectories through simulated human-agent interplay. We train a family of models---the xLAM-2-fc-r series with sizes ranging from 1B to 70B parameters. Our models outperform frontier models such as GPT-4o and Claude 3. 5 on $\tau$-bench and BFCL benchmarks, with the smaller models surpassing their larger counterparts, particularly in multi-turn settings, while maintaining superior consistency across multiple trials. Comprehensive experiments demonstrate that our verified blueprint-to-details approach yields high-quality training data, enabling the development of more reliable, efficient, and capable agents. We open-source both the synthetic data collected and the trained xLAM-2-fc-r models to advance research in AI agents. Dataset: https: //huggingface. co/datasets/Salesforce/APIGen-MT-5k & Models: https: //huggingface. co/collections/Salesforce/xlam-2-67ef5be12949d8dcdae354c4

UAI Conference 2025 Conference Paper

Collapsing Sequence-Level Data-Policy Coverage via Poisoning Attack in Offline Reinforcement Learning

  • Xue Zhou
  • Dapeng Man
  • Chen Xu 0008
  • Fanyi Zeng
  • Tao Liu 0038
  • Huan Wang
  • Shucheng He
  • Chaoyang Gao

Offline reinforcement learning (RL) heavily relies on the coverage of pre-collected data over the target policy’s distribution. Existing studies aim to improve data-policy coverage to mitigate distributional shifts, but overlook security risks from insufficient coverage, and the single-step analysis is not consistent with the multi-step decision-making nature of offline RL. To address this, we introduce the sequence-level concentrability coefficient to quantify coverage, and reveal its exponential amplification on the upper bound of estimation errors through theoretical analysis. Building on this, we propose the Collapsing Sequence-Level Data-Policy Coverage (CSDPC) poisoning attack. Considering the continuous nature of offline RL data, we convert state-action pairs into decision units, and extract representative decision patterns that capture multi-step behavior. We identify rare patterns likely to cause insufficient coverage, and poison them to reduce coverage and exacerbate distributional shifts. Experiments show that poisoning just 1% of the dataset can degrade agent performance by 90%. This finding provides new perspectives for analyzing and safeguarding the security of offline RL.

EAAI Journal 2025 Journal Article

Deep adaptive wavelet autoencoder with mutually independent empirical cumulative distribution for unsupervised motor anomaly detection

  • Pinze Ren
  • Ning Zhu
  • Dandan Peng
  • Liyuan Ren
  • Huan Wang

—Permanent magnet synchronous motors (PMSMs) are widely used in industrial applications but remain vulnerable to stator faults, such as inter-coil and inter-turn short circuits. Although recent deep learning-based fault detection methods have shown promise, they typically rely on large volumes of labelled fault data for training. To address this limitation, this paper proposes a novel unsupervised fault detection framework, termed Deep Adaptive Wavelet Autoencoder (DAWA) with Mutually Independent Empirical Cumulative Distribution (MIECD), specifically designed for PMSM fault detection. DAWA utilizes convolutional neural networks to learn adaptive wavelet filters through fast discrete wavelet transform, allowing for fully learnable, threshold-free extraction of fine-grained signal patterns. The resulting latent features are then mapped by MIECD into a mutually independent space via independent component analysis (ICA). Without assuming any prior data distribution, MIECD estimates empirical cumulative distributions (ECDs), computes tail probabilities across dimensions, and aggregates them into a unified anomaly score. Experimental results on motor vibration datasets demonstrate the effectiveness of the proposed method, showing average accuracy improvements of 15. 85 % for Interturn and 15. 16 % for Intercoil fault detection compared to conventional data-driven baselines across various operating conditions.

IJCAI Conference 2025 Conference Paper

Diffuse& Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object Detection

  • Jianzhou Wang
  • Yirui Wu
  • Lixin Yuan
  • Wenxiao Zhang
  • Jun Liu
  • Junyang Chen
  • Huan Wang
  • Wenhai Wang

Incremental Object Detection(IOD) targets at progressively extending capability of object detectors to recognize new classes. However, representation confusion between old and new classes leads to catastrophic forgetting. To alleviate this problem, we propose DiffKA, with intrinsic knowledge generated and aggregated by forward and backward diffusion, gradually establishing rigid class boundary. With incremental streaming data, forward diffusion spreads information to generate potential inter-class associations among new- and old-class prototypes within a hierarchical tree, named as Intrinsic Correlation Tree(ICTree), to store intrinsic knowledge. Afterwards, backward diffusion refines and aggregates the generated knowledge in ICTree, explicitly establishing rigid class boundary to mitigate representation confusion. To keep semantic consistency with extreme IOD settings, we reorganize semantic relevance of old- and new-class prototypes in paradigms to adaptively and effectively update DiffKA. Experiments on MS COCO dataset show DiffKA achieves state-of-the-art performance on IOD tasks with significant advantages.

NeurIPS Conference 2025 Conference Paper

FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware Guidance

  • Ying Li
  • Chengfei Lyu
  • Huan Wang

Visual AutoRegressive (VAR) modeling employs a next-scale decoding paradigm that progresses from coarse structures to fine details. While enhancing fidelity and scalability, this approach challenges two fundamental assumptions of conventional dynamic inference: semantic stability (intermediate outputs approximating final results) and monotonic locality (smooth representation evolution across layers), which renders existing dynamic inference methods ineffective for VAR models. To address this challenge, we propose FreqExit, a unified training framework that enables dynamic inference in VAR without altering its architecture or compromising output quality. FreqExit is based on a key insight: high-frequency details are crucial for perceptual quality and tend to emerge only in later decoding stages. Leveraging this insight, we design targeted mechanisms that guide the model to learn more effectively through frequency-aware supervision. The proposed framework consists of three components: (1) a curriculum-based supervision strategy with progressive layer dropout and early exit loss; (2) a wavelet-domain high-frequency consistency loss that aligns spectral content across different generation steps; and (3) a lightweight self-supervised frequency-gated module that guides adaptive learning of both structural and detailed spectral components. On ImageNet 256×256, FreqExit achieves up to 2× speedup with only minor degradation, and delivers 1. 3× acceleration without perceptible quality loss. This enables runtime-adaptive acceleration within a unified model, offering a favorable trade-off between efficiency and fidelity for for practical and flexible deployment.

NeurIPS Conference 2025 Conference Paper

HoliTom: Holistic Token Merging for Fast Video Large Language Models

  • Kele Shao
  • Keda TAO
  • Can Qin
  • Haoxuan You
  • Yang Sui
  • Huan Wang

Video large language models (video LLMs) excel at video comprehension but face significant computational inefficiency due to redundant video tokens. Existing token pruning methods offer solutions. However, approaches operating within the LLM (inner-LLM pruning), such as FastV, incur intrinsic computational overhead in shallow layers. In contrast, methods performing token pruning before the LLM (outer-LLM pruning) primarily address spatial redundancy within individual frames or limited temporal windows, neglecting the crucial global temporal dynamics and correlations across longer video sequences. This leads to sub-optimal spatio-temporal reduction and does not leverage video compressibility fully. Crucially, the synergistic potential and mutual influence of combining these strategies remain unexplored. To further reduce redundancy, we introduce HoliTom, a novel training-free holistic token merging framework. HoliTom employs outer-LLM pruning through global redundancy-aware temporal segmentation, followed by spatial-temporal merging to reduce visual tokens by over 90%, significantly alleviating the LLM's computational burden. Complementing this, we introduce a robust inner-LLM token similarity-based merging approach, designed for superior performance and compatibility with outer-LLM pruning. Evaluations demonstrate our method's promising efficiency-performance trade-off on LLaVA-OneVision-7B, reducing computational costs to 6. 9% of FLOPs while maintaining 99. 1% of the original performance. Furthermore, we achieve a 2. 28× reduction in Time-To-First-Token (TTFT) and a 1. 32× acceleration in decoding throughput, highlighting the practical benefits of our integrated pruning approach for efficient video LLMs inference.

NeurIPS Conference 2025 Conference Paper

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs

  • Kejia Zhang
  • Keda TAO
  • Jiasheng Tang
  • Huan Wang

Large vision-language models (LVMs) extend large language models (LLMs) with visual perception capabilities, enabling them to process and interpret visual information. A major challenge compromising their reliability is object hallucination that LVMs may generate plausible but factually inaccurate information. We propose a novel \textit{visual adversarial perturbation (VAP)} method to mitigate this hallucination issue. VAP alleviates LVM hallucination by applying strategically optimized visual noise without altering the base model. Our approach formulates hallucination suppression as an optimization problem, leveraging adversarial strategies to generate beneficial visual perturbations that enhance the model's factual grounding and reduce parametric knowledge bias. Extensive experimental results demonstrate that our method consistently reduces object hallucinations across 8 state-of-the-art LVMs, validating its efficacy across diverse evaluations.

NeurIPS Conference 2025 Conference Paper

Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting

  • Yiren Lu
  • Yunlai Zhou
  • Yiran Qiao
  • Chaoda Song
  • Tuo Liang
  • Jing Ma
  • Huan Wang
  • Yu Yin

Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with dynamic scenes due to the complexities of motion modeling. In this paper, we propose Segment then Splat, a 3D-aware open vocabulary segmentation approach for both static and dynamic scenes based on Gaussian Splatting. Segment then Splat reverses the long established approach of "segmentation after reconstruction'' by dividing Gaussians into distinct object sets before reconstruction. Once reconstruction is complete, the scene is naturally segmented into individual objects, achieving true 3D segmentation. This design eliminates both geometric and semantic ambiguities, as well as Gaussian–object misalignment issues in dynamic scenes. It also accelerates the optimization process, as it eliminates the need for learning a separate language field. After optimization, a CLIP embedding is assigned to each object to enable open-vocabulary querying. Extensive experiments on various datasets demonstrate the effectiveness of our proposed method in both static and dynamic scenarios.

AAAI Conference 2025 Conference Paper

Text2Data: Low-Resource Data Generation with Textual Control

  • Shiyu Wang
  • Yihao Feng
  • Tian Lan
  • Ning Yu
  • Yu Bai
  • Ran Xu
  • Huan Wang
  • Caiming Xiong

Natural language serves as a common and straightforward control signal for humans to interact seamlessly with machines. Recognizing the importance of this interface, the machine learning community is investing considerable effort in generating data that is semantically coherent with textual instructions. While strides have been made in text-to-data generation spanning image editing, audio synthesis, video creation, and beyond, low-resource areas characterized by expensive annotations or complex data structures, such as molecules, motion dynamics, and time series, often lack textual labels. This deficiency impedes supervised learning, thereby constraining the application of advanced generative models for text-to-data tasks. In response to these challenges in the low-resource scenario, we propose Text2Data, a novel approach that utilizes unlabeled data to understand the underlying data distribution through an unsupervised diffusion model. Subsequently, it undergoes controllable finetuning via a novel constraint optimization-based learning objective that ensures controllability and effectively counteracts catastrophic forgetting. Comprehensive experiments demonstrate that Text2Data is able to achieve enhanced performance regarding controllability across various modalities, including molecules, motions and time series, when compared to existing baselines.

EAAI Journal 2024 Journal Article

A novel health indicator by dominant invariant subspace on Grassmann manifold for state of health assessment of lithium-ion battery

  • Ying Zhang
  • Yan-Fu Li
  • Ming Zhang
  • Huan Wang

The precise estimation of the state of health (SoH) in Lithium-ion batteries (LiBs) relies heavily on a reliable health indicator (HI). Conventional indicators are often constructed by directly concatenating features from multiple sources. It overlooks significant non-linear and correlative information inherent in raw signals. To address this limitation, this paper introduces an innovative approach for SoH estimation in LiBs. Deep features extracted from signals of various sensors are obtained using denoising auto-encoders (DAEs). Then the dominant invariant subspaces (DIS) are calculated through the non-linear transformation of multi-source features on the Grassmann manifold. It can preserve essential and robust characteristics. The health indicator quantifies the geodesic distance of DIS using a projection metric. It provides a more comprehensive inclusion of nonlinear and correlation information. Consequently, this indicator offers heightened precision in discerning differences in health states. Validation of the proposed method is conducted using the NASA dataset. The result demonstrates its effectiveness on the SoH assessment and superiority to the state-of-the-art method.

NeurIPS Conference 2024 Conference Paper

APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets

  • Zuxin Liu
  • Thai Hoang
  • Jianguo Zhang
  • Ming Zhu
  • Tian Lan
  • Shirley Kokane
  • Juntao Tan
  • Weiran Yao

The advancement of function-calling agent models requires diverse, reliable, and high-quality datasets. This paper presents APIGen, an automated data generation pipeline designed to synthesize high-quality datasets for function-calling applications. We leverage APIGen and collect 3, 673 executable APIs across 21 different categories to generate diverse function-calling datasets in a scalable and structured manner. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, improving its reliability and correctness. We demonstrate that models trained with our curated datasets, even with only 7B parameters, can achieve state-of-the-art performance on the Berkeley Function-Calling Benchmark, outperforming multiple GPT-4 models. Moreover, our 1B model achieves exceptional performance, surpassing GPT-3. 5-Turbo and Claude-3 Haiku. We release a dataset containing 60, 000 high-quality entries, aiming to advance the field of function-calling agent domains. The dataset and models are available on the project homepage \url{https: //apigen-pipeline. github. io/}.

NeurIPS Conference 2024 Conference Paper

Slicing Vision Transformer for Flexible Inference

  • Yitian Zhang
  • Huseyin Coskun
  • Xu Ma
  • Huan Wang
  • Ke Ma
  • Xi Chen
  • Derek H. Hu
  • Yun Fu

Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe that smaller ViTs are intrinsically the sub-networks of a larger ViT with different widths. Thus, we propose a general framework, named Scala, to enable a single network to represent multiple smaller ViTs with flexible inference capability, which aligns with the inherent design of ViT to vary from widths. Concretely, Scala activates several subnets during training, introduces Isolated Activation to disentangle the smallest sub-network from other subnets, and leverages Scale Coordination to ensure each sub-network receives simplified, steady, and accurate learning objectives. Comprehensive empirical validations on different tasks demonstrate that with only one-shot training, Scala learns slimmable representation without modifying the original ViT structure and matches the performance of Separate Training. Compared with the prior art, Scala achieves an average improvement of 1. 6% on ImageNet-1K with fewer parameters.

AAAI Conference 2024 Conference Paper

Sparse Enhanced Network: An Adversarial Generation Method for Robust Augmentation in Sequential Recommendation

  • Junyang Chen
  • Guoxuan Zou
  • Pan Zhou
  • Wu Yirui
  • Zhenghan Chen
  • Houcheng Su
  • Huan Wang
  • Zhiguo Gong

Sequential Recommendation plays a significant role in daily recommendation systems, such as e-commerce platforms like Amazon and Taobao. However, even with the advent of large models, these platforms often face sparse issues in the historical browsing records of individual users due to new users joining or the introduction of new products. As a result, existing sequence recommendation algorithms may not perform well. To address this, sequence-based data augmentation methods have garnered attention. Existing sequence enhancement methods typically rely on augmenting existing data, employing techniques like cropping, masking prediction, random reordering, and random replacement of the original sequence. While these methods have shown improvements, they often overlook the exploration of the deep embedding space of the sequence. To tackle these challenges, we propose a Sparse Enhanced Network (SparseEnNet), which is a robust adversarial generation method. SparseEnNet aims to fully explore the hidden space in sequence recommendation, generating more robust enhanced items. Additionally, we adopt an adversarial generation method, allowing the model to differentiate between data augmentation categories and achieve better prediction performance for the next item in the sequence. Experiments have demonstrated that our method achieves a remarkable 4-14% improvement over existing methods when evaluated on the real-world datasets. (https://github.com/junyachen/SparseEnNet)

EAAI Journal 2024 Journal Article

Wavelet-powered hierarchical frequency filtering framework for autonomous vehicle sensors fault diagnosis and correction under open environments

  • Huan Wang
  • Yan-Fu Li

The reliable operation of autonomous vehicles heavily depends on the functionality of numerous sensors that facilitate environmental perception and vehicle control. However, sensor malfunctions can significantly undermine the dependability and safety of autonomous vehicles. Despite the significant advancements made in sensor fault diagnosis through deep learning, the technology's susceptibility to noise and limited frequency analysis capacity renders it less effective in addressing frequency aliasing and noise interference issues. In addition, repairing incorrect sensor data holds greater significance for autonomous vehicles than merely identifying sensor failures. To this end, this paper proposes a wavelet-powered hierarchical frequency filtering framework (WavePHF) for autonomous vehicle sensors' fault diagnosis and correction. First, this study designs a multi-sensor fault identification and correction joint architecture to detect sensor faults and repair fault signals in time. Second, this study integrates the discrete wavelet transform (DWT) algorithm into the deep model to significantly enhance the model's frequency analysis capabilities and suppress frequency aliasing and noise interference. In particular, DWT is deeply integrated into the convolutional modules' parameter updating and optimization process, thereby facilitating mutual learning and promotion. Third, this study proposes an adaptive frequency weighting block for low-frequency features and an adaptive frequency discarding block for high-frequency features to achieve hierarchical frequency filtering in a coarse-to-fine manner and efficiently identify fault categories and restore pure signals. The WavePHF is evaluated on a real autonomous driving dataset. Experiments show that WavePHF has excellent autonomous vehicle sensors fault diagnosis and correction capabilities and competitive performance in different urban scenarios.

NeurIPS Conference 2023 Conference Paper

Latent Graph Inference with Limited Supervision

  • Jianglin Lu
  • Yi Xu
  • Huan Wang
  • Yue Bai
  • Yun Fu

Latent graph inference (LGI) aims to jointly learn the underlying graph structure and node representations from data features. However, existing LGI methods commonly suffer from the issue of supervision starvation, where massive edge weights are learned without semantic supervision and do not contribute to the training loss. Consequently, these supervision-starved weights, which determine the predictions of testing samples, cannot be semantically optimal, resulting in poor generalization. In this paper, we observe that this issue is actually caused by the graph sparsification operation, which severely destroys the important connections established between pivotal nodes and labeled ones. To address this, we propose to restore the corrupted affinities and replenish the missed supervision for better LGI. The key challenge then lies in identifying the critical nodes and recovering the corrupted affinities. We begin by defining the pivotal nodes as k-hop starved nodes, which can be identified based on a given adjacency matrix. Considering the high computational burden, we further present a more efficient alternative inspired by CUR matrix decomposition. Subsequently, we eliminate the starved nodes by reconstructing the destroyed connections. Extensive experiments on representative benchmarks demonstrate that reducing the starved nodes consistently improves the performance of state-of-the-art LGI methods, especially under extremely limited supervision (6. 12% improvement on Pubmed with a labeling rate of only 0. 3%).

JMLR Journal 2023 Journal Article

Merlion: End-to-End Machine Learning for Time Series

  • Aadyot Bhatnagar
  • Paul Kassianik
  • Chenghao Liu
  • Tian Lan
  • Wenzhuo Yang
  • Rowan Cassius
  • Doyen Sahoo
  • Devansh Arpit

We introduce Merlion, an open-source machine learning library for time series. It features a unified interface for many commonly used models and datasets for forecasting and anomaly detection on both univariate and multivariate time series, along with standard pre/post-processing layers. It has several modules to improve ease-of-use, including a no-code visual dashboard, anomaly score calibration to improve interpetability, AutoML for hyperparameter tuning and model selection, and model ensembling. Merlion also provides an evaluation framework that simulates the live deployment of a model in production, and a distributed computing backend to run time series models at industrial scale. This library aims to provide engineers and researchers a one-stop solution to rapidly develop models for their specific time series needs and benchmark them across multiple datasets. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

NeurIPS Conference 2023 Conference Paper

SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds

  • Yanyu Li
  • Huan Wang
  • Qing Jin
  • Ju Hu
  • Pavlo Chemerys
  • Yun Fu
  • Yanzhi Wang
  • Sergey Tulyakov

Text-to-image diffusion models can create stunning images from natural language descriptions that rival the work of professional artists and photographers. However, these models are large, with complex network architectures and tens of denoising iterations, making them computationally expensive and slow to run. As a result, high-end GPUs and cloud-based inference are required to run diffusion models at scale. This is costly and has privacy implications, especially when user data is sent to a third party. To overcome these challenges, we present a generic approach that, for the first time, unlocks running text-to-image diffusion models on mobile devices in **less than 2 seconds**. We achieve so by introducing efficient network architecture and improving step distillation. Specifically, we propose an efficient UNet by identifying the redundancy of the original model and reducing the computation of the image decoder via data distillation. Further, we enhance the step distillation by exploring training strategies and introducing regularization from classifier-free guidance. Our extensive experiments on MS-COCO show that our model with $8$ denoising steps achieves better FID and CLIP scores than Stable Diffusion v$1. 5$ with $50$ steps. Our work democratizes content creation by bringing powerful text-to-image diffusion models to the hands of users.

YNICL Journal 2023 Journal Article

The relationship between disrupted anhedonia-related circuitry and suicidal ideation in major depressive disorder: A network-based analysis

  • Xiaoqin Wang
  • Yi Xia
  • Rui Yan
  • Huan Wang
  • Hao Sun
  • Yinghong Huang
  • Lingling Hua
  • Hao Tang

BACKGROUND: Several epidemiological studies and psychological models have suggested that major depressive disorder (MDD) with anhedonia is associated with suicidal ideation (SI). However, little is known about whether the functional network pattern and intrinsic topologically disrupted in patients with anhedonia are related to SI. METHODS: The resting-fMRI by applying network-based statistic (NBS) and graph-theory analyses was estimated in 273 patients with MDD (144 high anhedonia [HA], 129 low anhedonia [LA]) and 150 healthy controls. In addition, we quantified the SI scores of each patient. Finally, the mediation analysis assessed whether anhedonia symptoms could mediate the relationship between anhedonia-related network metrics and SI. RESULT: The NBS analysis demonstrated that individuals with HA have a single abnormally increased functional connectivity component in a frontal-limbic circuit (termed the "anhedonia-related network", including the frontal cortex, striatum, anterior cingulate cortex and amygdala). The graph-theory analysis demonstrated that the anhedonia-related network showed a significantly disrupted topological organization (lower gamma and lambda), which the small-world property trend randomized. Furthermore, the anhedonia symptoms could mediate the relationship between the anhedonia-related network metrics (the mean functional connectivity values, the area under the curves values of gamma and nodal local efficiency in nucleus accumbens) and SI. CONCLUSIONS: We found that disruption of the reward-related network in MDD leads to SI through anhedonia symptoms. These findings show the abnormal topological construction of functional brain network organization in anhedonia, shedding light on the neurological processes underlying SI in MDD patients with anhedonia symptoms.

NeurIPS Conference 2023 Conference Paper

Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection

  • Yu Bai
  • Fan Chen
  • Huan Wang
  • Caiming Xiong
  • Song Mei

Neural sequence models based on the transformer architecture have demonstrated remarkable \emph{in-context learning} (ICL) abilities, where they can perform new tasks when prompted with training and test examples, without any parameter update to the model. This work first provides a comprehensive statistical theory for transformers to perform ICL. Concretely, we show that transformers can implement a broad class of standard machine learning algorithms in context, such as least squares, ridge regression, Lasso, learning generalized linear models, and gradient descent on two-layer neural networks, with near-optimal predictive power on various in-context data distributions. Using an efficient implementation of in-context gradient descent as the underlying mechanism, our transformer constructions admit mild size bounds, and can be learned with polynomially many pretraining sequences. Building on these ``base'' ICL algorithms, intriguingly, we show that transformers can implement more complex ICL procedures involving \emph{in-context algorithm selection}, akin to what a statistician can do in real life---A \emph{single} transformer can adaptively select different base ICL algorithms---or even perform qualitatively different tasks---on different input sequences, without any explicit prompting of the right algorithm or task. We both establish this in theory by explicit constructions, and also observe this phenomenon experimentally. In theory, we construct two general mechanisms for algorithm selection with concrete examples: pre-ICL testing, and post-ICL validation. As an example, we use the post-ICL validation mechanism to construct a transformer that can perform nearly Bayes-optimal ICL on a challenging task---noisy linear models with mixed noise levels. Experimentally, we demonstrate the strong in-context algorithm selection capabilities of standard transformer architectures.

NeurIPS Conference 2023 Conference Paper

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

  • Can Qin
  • Shu Zhang
  • Ning Yu
  • Yihao Feng
  • Xinyi Yang
  • Yingbo Zhou
  • Huan Wang
  • Juan Carlos Niebles

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl, a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.

EAAI Journal 2023 Journal Article

Wavelet integrated attention network with multi-resolution frequency learning for mixed-type wafer defect recognition

  • Yuxiang Wei
  • Huan Wang

Wafer defect recognition has been an important measure in developing the manufacturing process. In real-life manufacturing, however, defects can be complicated, and their wafer maps are often accompanied by noise. This encourages us to build a noise-robust framework with outstanding performance and sound interpretability on defect recognition. Therefore, this paper chooses discrete wavelet transform for its clear physical meaning and ability of frequency learning. Moreover, it outperforms traditional down-sampling operations in the preservation of fringe information. Based on this, we propose a multiresolution wavelet integrated attention network (MRWA-Net). Specifically, we design a learnable discrete wavelet transform layer (DWT-Layer), which expands convolutional neural network’s (CNN’s) feature learning space to the wavelet domain. This helps the framework procure hidden information from different frequency components and their location information. Furthermore, we utilize different levels of wavelet transform to interpret the images with different resolutions, thus learning features from different perspectives. Additionally, we insert a frequency-location attention module (FLA) to select the useful frequency-location information captured by DWT-Layer. The proposed approach is evaluated on a dataset with 38015 subjects and 38 types of defects and reaches 98. 84% accuracy. To demonstrate the noise-robustness of our framework, we further compare it with other state-of-the-art methods on wafer maps with different ratios of additional noise. The results show that our framework excels other methods under all noise ratios and exhibits more notable excellence on data accompanied by a higher ratio of noise. Finally, we present visualizing analysis to demonstrate that the proposed DWT-Layer can learn from different frequency bands and retrieve information with multiple resolutions.

YNIMG Journal 2023 Journal Article

White matter BOLD signals at 7 Tesla reveal visual field maps in optic radiation and vertical occipital fasciculus

  • Huan Wang
  • Xiaoxiao Wang
  • Yanming Wang
  • Du Zhang
  • Yan Yang
  • Yifeng Zhou
  • Bensheng Qiu
  • Peng Zhang

There is growing evidence that blood-oxygen-level-dependent (BOLD) activity in the white matter (WM) can be detected by functional magnetic resonance imaging (fMRI). However, the functional relevance and significance of WM BOLD signals remain controversial. Here we investigated whether 7T BOLD fMRI can reveal fine-scale functional organizations of a WM bundle. Population receptive field (pRF) analyses of the 7T retinotopy dataset from the Human Connectome Project revealed clear contralateral retinotopic organizations of two visual WM bundles: the optic radiation (OR) and the vertical occipital fasciculus (VOF). The retinotopic maps of OR are highly consistent with post-mortem dissections and diffusion tractographies, while the VOF maps are compatible with the dorsal and ventral visual areas connected by the WM. Similar to the grey matter (GM) visual areas, both WM bundles show over-representations of the central visual field and increasing pRF size with eccentricity. Hemodynamic response functions of visual WM were slower and wider compared with those of GM areas. These findings clearly demonstrate that WM BOLD at 7 Tesla is closely coupled with neural activity related to axons, encoding highly specific information that can be used to characterize fine-scale functional organizations of a WM bundle.

NeurIPS Conference 2022 Conference Paper

Ensemble of Averages: Improving Model Selection and Boosting Performance in Domain Generalization

  • Devansh Arpit
  • Huan Wang
  • Yingbo Zhou
  • Caiming Xiong

In Domain Generalization (DG) settings, models trained independently on a given set of training domains have notoriously chaotic performance on distribution shifted test domains, and stochasticity in optimization (e. g. seed) plays a big role. This makes deep learning models unreliable in real world settings. We first show that this chaotic behavior exists even along the training optimization trajectory of a single model, and propose a simple model averaging protocol that both significantly boosts domain generalization and diminishes the impact of stochasticity by improving the rank correlation between the in-domain validation accuracy and out-domain test accuracy, which is crucial for reliable early stopping. Taking advantage of our observation, we show that instead of ensembling unaveraged models (that is typical in practice), ensembling moving average models (EoA) from independent runs further boosts performance. We theoretically explain the boost in performance of ensembling and model averaging by adapting the well known Bias-Variance trade-off to the domain generalization setting. On the DomainBed benchmark, when using a pre-trained ResNet-50, this ensemble of averages achieves an average of $68. 0\%$, beating vanilla ERM (w/o averaging/ensembling) by $\sim 4\%$, and when using a pre-trained RegNetY-16GF, achieves an average of $76. 6\%$, beating vanilla ERM by $\sim 6\%$.

NeurIPS Conference 2022 Conference Paper

Look More but Care Less in Video Recognition

  • Yitian Zhang
  • Yue Bai
  • Huan Wang
  • Yi Xu
  • Yun Fu

Existing action recognition methods typically sample a few frames to represent each video to avoid the enormous computation, which often limits the recognition performance. To tackle this problem, we propose Ample and Focal Network (AFNet), which is composed of two branches to utilize more frames but with less computation. Specifically, the Ample Branch takes all input frames to obtain abundant information with condensed computation and provides the guidance for Focal Branch by the proposed Navigation Module; the Focal Branch squeezes the temporal size to only focus on the salient frames at each convolution block; in the end, the results of two branches are adaptively fused to prevent the loss of information. With this design, we can introduce more frames to the network but cost less computation. Besides, we demonstrate AFNet can utilize less frames while achieving higher accuracy as the dynamic selection in intermediate features enforces implicit temporal modeling. Further, we show that our method can be extended to reduce spatial redundancy with even less cost. Extensive experiments on five datasets demonstrate the effectiveness and efficiency of our method.

TCS Journal 2022 Journal Article

Multi-attribute based influence maximization in social networks: Algorithms and analysis

  • Qiufen Ni
  • Jianxiong Guo
  • Hongmin W. Du
  • Huan Wang

The most valuable feature of social networks is that they can generate contents for users and spread them quickly on the network, which is a very important platform for viral marketing. Most of the related work on viral marketing focuses on the spread of single information, while a product may associate with multiple attributes in real life. Information about multiple attributes of a product propagates in the social networks simultaneously and independently. The attribute information that a user receives will determine whether he would purchase the product or not. We extend the traditional single information influence maximization problem to the multi-attribute based influence maximization problem. We also present the Multi-dimensional IC model (MIC model) for the proposed problem, then formulate the problem as the Multi-attribute based Influence Maximization Problem (MIMP). The objective function for MIMP is proved to be non-submodular, then we solve the problem with two different algorithms: the Sandwich Algorithm and the Supermodular Algorithm, whose solutions can get a m a x { f ( S U ) f ‾ ( S U ), f _ ( S L ⁎ ) f ( S o ⁎ ) } ( 1 − 1 / e ) approximation ratio and an 1 / ( d + 2 ) approximation ratio to the optimal solution, respectively. Experiments based on the real world social network datasets verify the effectiveness and correctness of our proposed solutions.

NeurIPS Conference 2022 Conference Paper

Parameter-Efficient Masking Networks

  • Yue Bai
  • Huan Wang
  • Xu Ma
  • Yitian Zhang
  • Zhiqiang Tao
  • Yun Fu

A deeper network structure generally handles more complicated non-linearity and performs more competitively. Nowadays, advanced network designs often contain a large number of repetitive structures (e. g. , Transformer). They empower the network capacity to a new level but also increase the model size inevitably, which is unfriendly to either model restoring or transferring. In this study, we are the first to investigate the representative potential of fixed random weights with limited unique values by learning diverse masks and introduce the Parameter-Efficient Masking Networks (PEMN). It also naturally leads to a new paradigm for model compression to diminish the model size. Concretely, motivated by the repetitive structures in modern neural networks, we utilize one random initialized layer, accompanied with different masks, to convey different feature mappings and represent repetitive network modules. Therefore, the model can be expressed as \textit{one-layer} with a bunch of masks, which significantly reduce the model storage cost. Furthermore, we enhance our strategy by learning masks for a model filled by padding a given random weights vector. In this way, our method can further lower the space complexity, especially for models without many repetitive architectures. We validate the potential of PEMN learning masks on random weights with limited unique values and test its effectiveness for a new compression paradigm based on different network architectures. Code is available at \href{https: //github. com/yueb17/PEMN}{\textcolor{magenta}{https: //github. com/yueb17/PEMN}}.

NeurIPS Conference 2022 Conference Paper

Policy Optimization for Markov Games: Unified Framework and Faster Convergence

  • Runyu Zhang
  • Qinghua Liu
  • Huan Wang
  • Caiming Xiong
  • Na Li
  • Yu Bai

This paper studies policy optimization algorithms for multi-agent reinforcement learning. We begin by proposing an algorithm framework for two-player zero-sum Markov Games in the full-information setting, where each iteration consists of a policy update step at each state using a certain matrix game algorithm, and a value update step with a certain learning rate. This framework unifies many existing and new policy optimization algorithms. We show that the \emph{state-wise average policy} of this algorithm converges to an approximate Nash equilibrium (NE) of the game, as long as the matrix game algorithms achieve low weighted regret at each state, with respect to weights determined by the speed of the value updates. Next, we show that this framework instantiated with the Optimistic Follow-The-Regularized-Leader (OFTRL) algorithm at each state (and smooth value updates) can find an $\mathcal{\widetilde{O}}(T^{-5/6})$ approximate NE in $T$ iterations, and a similar algorithm with slightly modified value update rule achieves a faster $\mathcal{\widetilde{O}}(T^{-1})$ convergence rate. These improve over the current best $\mathcal{\widetilde{O}}(T^{-1/2})$ rate of symmetric policy optimization type algorithms. We also extend this algorithm to multi-player general-sum Markov Games and show an $\mathcal{\widetilde{O}}(T^{-3/4})$ convergence rate to Coarse Correlated Equilibria (CCE). Finally, we provide a numerical example to verify our theory and investigate the importance of smooth value updates, and find that using ''eager'' value updates instead (equivalent to the independent natural policy gradient algorithm) may significantly slow down the convergence, even on a simple game with $H=2$ layers.

IJCAI Conference 2022 Conference Paper

Recent Advances on Neural Network Pruning at Initialization

  • Huan Wang
  • Can Qin
  • Yue Bai
  • Yulun Zhang
  • Yun Fu

Neural network pruning typically removes connections or neurons from a pretrained converged model; while a new pruning paradigm, pruning at initialization (PaI), attempts to prune a randomly initialized network. This paper offers the first survey concentrated on this emerging pruning fashion. We first introduce a generic formulation of neural network pruning, followed by the major classic pruning topics. Then, as the main body of this paper, a thorough and structured literature review of PaI methods is presented, consisting of two major tracks (sparse training and sparse selection). Finally, we summarize the surge of PaI compared to PaT and discuss the open problems. Apart from the dedicated literature review, this paper also offers a code base for easy sanity-checking and benchmarking of different PaI methods.

JMLR Journal 2022 Journal Article

WarpDrive: Fast End-to-End Deep Multi-Agent Reinforcement Learning on a GPU

  • Tian Lan
  • Sunil Srinivasa
  • Huan Wang
  • Stephan Zheng

WarpDrive is a flexible, lightweight, and easy-to-use open-source framework for end-to-end deep multi-agent reinforcement learning (MARL) on a Graphics Processing Unit (GPU), available at https://github.com/salesforce/warp-drive. It addresses key system bottlenecks when applying MARL to complex environments with high-dimensional state, observation, or action spaces. For example, WarpDrive eliminates data copying between the CPU and GPU and runs thousands of simulations and agents in parallel. It also enables distributed training on multiple GPUs and scales to millions of agents. In all, WarpDrive enables orders-of-magnitude faster MARL compared to common CPU-GPU implementations. For example, WarpDrive yields 2.9 million environment steps/second with 2000 environments and 1000 agents (at least 100× faster than a CPU version) in a 2d-Tag simulation. It is user-friendly: e.g., it provides a lightweight, extendable Python interface and flexible environment wrappers. It is also compatible with PyTorch. In all, WarpDrive offers a platform to significantly accelerate reinforcement learning research and development. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2022. ( edit, beta )

NeurIPS Conference 2022 Conference Paper

What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical Perspective

  • Huan Wang
  • Suhas Lohit
  • Michael N. Jones
  • Yun Fu

Knowledge distillation (KD) is a general neural network training approach that uses a teacher model to guide the student model. Existing works mainly study KD from the network output side (e. g. , trying to design a better KD loss function), while few have attempted to understand it from the input side. Especially, its interplay with data augmentation (DA) has not been well understood. In this paper, we ask: Why do some DA schemes (e. g. , CutMix) inherently perform much better than others in KD? What makes a "good" DA in KD? Our investigation from a statistical perspective suggests that a good DA scheme should reduce the covariance of the teacher-student cross-entropy. A practical metric, the stddev of teacher’s mean probability (T. stddev), is further presented and well justified empirically. Besides the theoretical understanding, we also introduce a new entropy-based data-mixing DA scheme, CutMixPick, to further enhance CutMix. Extensive empirical studies support our claims and demonstrate how we can harvest considerable performance gains simply by using a better DA scheme in knowledge distillation. Code: https: //github. com/MingSun-Tse/Good-DA-in-KD.

NeurIPS Conference 2021 Conference Paper

Aligned Structured Sparsity Learning for Efficient Image Super-Resolution

  • Yulun Zhang
  • Huan Wang
  • Can Qin
  • Yun Fu

Lightweight image super-resolution (SR) networks have obtained promising results with moderate model size. Many SR methods have focused on designing lightweight architectures, which neglect to further reduce the redundancy of network parameters. On the other hand, model compression techniques, like neural architecture search and knowledge distillation, typically consume considerable memory and computation resources. In contrast, network pruning is a cheap and effective model compression technique. However, it is hard to be applied to SR networks directly, because filter pruning for residual blocks is well-known tricky. To address the above issues, we propose aligned structured sparsity learning (ASSL), which introduces a weight normalization layer and applies $L_2$ regularization to the scale parameters for sparsity. To align the pruned locations across different layers, we propose a \emph{sparsity structure alignment} penalty term, which minimizes the norm of soft mask gram matrix. We apply aligned structured sparsity learning strategy to train efficient image SR network, named as ASSLN, with smaller model size and lower computation than state-of-the-art methods. We conduct extensive comparisons with lightweight SR networks. Our ASSLN achieves superior performance gains over recent methods quantitatively and visually.

NeurIPS Conference 2021 Conference Paper

Evaluating State-of-the-Art Classification Models Against Bayes Optimality

  • Ryan Theisen
  • Huan Wang
  • Lav R. Varshney
  • Caiming Xiong
  • Richard Socher

Evaluating the inherent difficulty of a given data-driven classification problem is important for establishing absolute benchmarks and evaluating progress in the field. To this end, a natural quantity to consider is the \emph{Bayes error}, which measures the optimal classification error theoretically achievable for a given data distribution. While generally an intractable quantity, we show that we can compute the exact Bayes error of generative models learned using normalizing flows. Our technique relies on a fundamental result, which states that the Bayes error is invariant under invertible transformation. Therefore, we can compute the exact Bayes error of the learned flow models by computing it for Gaussian base distributions, which can be done efficiently using Holmes-Diaconis-Ross integration. Moreover, we show that by varying the temperature of the learned flow models, we can generate synthetic datasets that closely resemble standard benchmark datasets, but with almost any desired Bayes error. We use our approach to conduct a thorough investigation of state-of-the-art classification models, and find that in some --- but not all --- cases, these models are capable of obtaining accuracy very near optimal. Finally, we use our method to evaluate the intrinsic "hardness" of standard benchmark datasets.

NeurIPS Conference 2021 Conference Paper

Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement Learning

  • Tengyang Xie
  • Nan Jiang
  • Huan Wang
  • Caiming Xiong
  • Yu Bai

Recent theoretical work studies sample-efficient reinforcement learning (RL) extensively in two settings: learning interactively in the environment (online RL), or learning from an offline dataset (offline RL). However, existing algorithms and theories for learning near-optimal policies in these two settings are rather different and disconnected. Towards bridging this gap, this paper initiates the theoretical study of *policy finetuning*, that is, online RL where the learner has additional access to a "reference policy" $\mu$ close to the optimal policy $\pi_\star$ in a certain sense. We consider the policy finetuning problem in episodic Markov Decision Processes (MDPs) with $S$ states, $A$ actions, and horizon length $H$. We first design a sharp *offline reduction* algorithm---which simply executes $\mu$ and runs offline policy optimization on the collected dataset---that finds an $\varepsilon$ near-optimal policy within $\widetilde{O}(H^3SC^\star/\varepsilon^2)$ episodes, where $C^\star$ is the single-policy concentrability coefficient between $\mu$ and $\pi_\star$. This offline result is the first that matches the sample complexity lower bound in this setting, and resolves a recent open question in offline RL. We then establish an $\Omega(H^3S\min\{C^\star, A\}/\varepsilon^2)$ sample complexity lower bound for *any* policy finetuning algorithm, including those that can adaptively explore the environment. This implies that---perhaps surprisingly---the optimal policy finetuning algorithm is either offline reduction or a purely online RL algorithm that does not use $\mu$. Finally, we design a new hybrid offline/online algorithm for policy finetuning that achieves better sample complexity than both vanilla offline reduction and purely online RL algorithms, in a relaxed setting where $\mu$ only satisfies concentrability partially up to a certain time step. Overall, our results offer a quantitative understanding on the benefit of a good reference policy, and make a step towards bridging offline and online RL.

NeurIPS Conference 2021 Conference Paper

Sample-Efficient Learning of Stackelberg Equilibria in General-Sum Games

  • Yu Bai
  • Chi Jin
  • Huan Wang
  • Caiming Xiong

Real world applications such as economics and policy making often involve solving multi-agent games with two unique features: (1) The agents are inherently asymmetric and partitioned into leaders and followers; (2) The agents have different reward functions, thus the game is general-sum. The majority of existing results in this field focuses on either symmetric solution concepts (e. g. Nash equilibrium) or zero-sum games. It remains open how to learn the Stackelberg equilibrium ---an asymmetric analog of the Nash equilibrium---in general-sum games efficiently from noisy samples. This paper initiates the theoretical study of sample-efficient learning of the Stackelberg equilibrium, in the bandit feedback setting where we only observe noisy samples of the reward. We consider three representative two-player general-sum games: bandit games, bandit-reinforcement learning (bandit-RL) games, and linear bandit games. In all these games, we identify a fundamental gap between the exact value of the Stackelberg equilibrium and its estimated version using finitely many noisy samples, which can not be closed information-theoretically regardless of the algorithm. We then establish sharp positive results on sample-efficient learning of Stackelberg equilibrium with value optimal up to the gap identified above, with matching lower bounds in the dependency on the gap, error tolerance, and the size of the action spaces. Overall, our results unveil unique challenges in learning Stackelberg equilibria under noisy bandit feedback, which we hope could shed light on future research on this topic.

NeurIPS Conference 2021 Conference Paper

Slow Learning and Fast Inference: Efficient Graph Similarity Computation via Knowledge Distillation

  • Can Qin
  • Handong Zhao
  • Lichen Wang
  • Huan Wang
  • Yulun Zhang
  • Yun Fu

Graph Similarity Computation (GSC) is essential to wide-ranging graph applications such as retrieval, plagiarism/anomaly detection, etc. The exact computation of graph similarity, e. g. , Graph Edit Distance (GED), is an NP-hard problem that cannot be exactly solved within an adequate time given large graphs. Thanks to the strong representation power of graph neural network (GNN), a variety of GNN-based inexact methods emerged. To capture the subtle difference across graphs, the key success is designing the dense interaction with features fusion at the early stage, which, however, is a trade-off between speed and accuracy. For slow learning of graph similarity, this paper proposes a novel early-fusion approach by designing a co-attention-based feature fusion network on multilevel GNN features. To further improve the speed without much accuracy drop, we introduce an efficient GSC solution by distilling the knowledge from the slow early-fusion model to the student one for fast inference. Such a student model also enables the offline collection of individual graph embeddings, speeding up the inference time in orders. To address the instability through knowledge transfer, we decompose the dynamic joint embedding into the static pseudo individual ones for precise teacher-student alignment. The experimental analysis on the real-world datasets demonstrates the superiority of our approach over the state-of-the-art methods on both accuracy and efficiency. Particularly, we speed up the prior art by more than 10x on the benchmark AIDS data.

NeurIPS Conference 2021 Conference Paper

Understanding the Under-Coverage Bias in Uncertainty Estimation

  • Yu Bai
  • Song Mei
  • Huan Wang
  • Caiming Xiong

Estimating the data uncertainty in regression tasks is often done by learning a quantile function or a prediction interval of the true label conditioned on the input. It is frequently observed that quantile regression---a vanilla algorithm for learning quantiles with asymptotic guarantees---tends to *under-cover* than the desired coverage level in reality. While various fixes have been proposed, a more fundamental understanding of why this under-coverage bias happens in the first place remains elusive. In this paper, we present a rigorous theoretical study on the coverage of uncertainty estimation algorithms in learning quantiles. We prove that quantile regression suffers from an inherent under-coverage bias, in a vanilla setting where we learn a realizable linear quantile function and there is more data than parameters. More quantitatively, for $\alpha>0. 5$ and small $d/n$, the $\alpha$-quantile learned by quantile regression roughly achieves coverage $\alpha - (\alpha-1/2)\cdot d/n$ regardless of the noise distribution, where $d$ is the input dimension and $n$ is the number of training data. Our theory reveals that this under-coverage bias stems from a certain high-dimensional parameter estimation error that is not implied by existing theories on quantile regression. Experiments on simulated and real data verify our theory and further illustrate the effect of various factors such as sample size and model capacity on the under-coverage bias in more practical setups.

NeurIPS Conference 2020 Conference Paper

Towards Understanding Hierarchical Learning: Benefits of Neural Representations

  • Minshuo Chen
  • Yu Bai
  • Jason D. Lee
  • Tuo Zhao
  • Huan Wang
  • Caiming Xiong
  • Richard Socher

Deep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data. However, how they make use of the intermediate representations are not explained by recent theories that relate them to ``shallow learners'' such as kernels. In this work, we demonstrate that intermediate \emph{neural representations} add more flexibility to neural networks and can be advantageous over raw inputs. We consider a fixed, randomly initialized neural network as a representation function fed into another trainable network. When the trainable network is the quadratic Taylor model of a wide two-layer network, we show that neural representation can achieve improved sample complexities compared with the raw input: For learning a low-rank degree-$p$ polynomial ($p \geq 4$) in $d$ dimension, neural representation requires only $\widetilde{O}(d^{\ceil{p/2}})$ samples, while the best-known sample complexity upper bound for the raw input is $\widetilde{O}(d^{p-1})$. We contrast our result with a lower bound showing that neural representations do not improve over the raw input (in the infinite width limit), when the trainable network is instead a neural tangent kernel. Our results characterize when neural representations are beneficial, and may provide a new perspective on why depth is important in deep learning.

IJCAI Conference 2013 Conference Paper

Exact Recovery of Sparsely-Used Dictionaries

  • Daniel A. Spielman
  • Huan Wang
  • John Wright

We consider the problem of learning sparsely used dictionaries with an arbitrary square dictionary and a random, sparse coefficient matrix. We prove that O(n log n) samples are sufficient to uniquely determine the coefficient matrix. Based on this proof, we design a polynomial-time algorithm, called Exact Recovery of Sparsely-Used Dictionaries (ER- SpUD), and prove that it probably recovers the dictionary and coefficient matrix when the coefficient matrix is sufficiently sparse. Simulation results show that ER-SpUD reveals the true dictionary as well as the coefficients with probability higher than many state-of-the-art algorithms.

YNIMG Journal 2013 Journal Article

Local mechanical properties of white matter structures in the human brain

  • Curtis L. Johnson
  • Matthew D.J. McGarry
  • Armen A. Gharibans
  • John B. Weaver
  • Keith D. Paulsen
  • Huan Wang
  • William C. Olivero
  • Bradley P. Sutton

The noninvasive measurement of the mechanical properties of brain tissue using magnetic resonance elastography (MRE) has emerged as a promising method for investigating neurological disorders. To date, brain MRE investigations have been limited to reporting global mechanical properties, though quantification of the stiffness of specific structures in the white matter architecture may be valuable in assessing the localized effects of disease. This paper reports the mechanical properties of the corpus callosum and corona radiata measured in healthy volunteers using MRE and atlas-based segmentation. Both structures were found to be significantly stiffer than overall white matter, with the corpus callosum exhibiting greater stiffness and less viscous damping than the corona radiata. Reliability of both local and global measures was assessed through repeated experiments, and the coefficient of variation for each measure was less than 10%. Mechanical properties within the corpus callosum and corona radiata demonstrated correlations with measures from diffusion tensor imaging pertaining to axonal microstructure.

IJCAI Conference 2007 Conference Paper

  • Huan Wang
  • Shuicheng Yan
  • Thomas Huang
  • Xiaoou Tang

Recently, substantial efforts have been devoted to the subspace learning techniques based on tensor representation, such as 2DLDA, DATER and Tensor Subspace Analysis (TSA). In this context, a vital yet unsolved problem is that the computational convergency of these iterative algorithms is not guaranteed. In this work, we present a novel solution procedure for general tensor-based subspace learning, followed by a detailed convergency proof of the solution projection matrices and the objective function value. Extensive experiments on real-world databases verify the high convergence speed of the proposed procedure, as well as its superiority in classification capability over traditional solution procedures.

v2026.09.13