Arrow Research search
Back to ECAI

ECAI 2025

GCQ-ViT: Group-Aware Collaborative Post-Training Quantization for Vision Transformers

Conference Paper Accepted Paper Artificial Intelligence

Abstract

Post-training quantization (PTQ) is widely utilized in Vision Transformers (ViTs) for its computational efficiency and retraining elimination. However, the unique architecture of ViTs introduces significant quantization challenges. Dynamic fluctuations in channel activations, particularly post-LayerNorm, result in distributional mismatches. Additionally, the heavy-tailed nature of post-Softmax activations compromises the accurate representation of critical attention regions, vital for ViT performance. Moreover, weight quantization at low bit-widths leads to a loss of structural information, degrading global feature representation. To address these challenges, we introduce the Group-aware Collaborative Quantization framework (GCQ-ViT), which significantly improves both the accuracy and efficiency of ViT quantization. The GCQ-ViT framework integrates a novel dynamic perception grouping quantization mechanism to ensure distributional consistency within groups, thus reducing hardware expense. It also utilizes a self-adaptive displaced uniform log2 quantizer, optimizing shift factors and nonlinear intervals to enhance representation in high-density regions of post-Softmax activations. Additionally, we propose a dynamic dimension-aware error compensation method to correct quantization errors across channel dimensions using a residual mean compensation skill, ensuring robust feature preservation. Extensive experiments on image classification, object detection, and instance segmentation tasks demonstrate that GCQ-ViT outperforms the current leading PTQ methods, setting a new benchmark for ViT quantization.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
European Conference on Artificial Intelligence
Archive span
1982-2025
Indexed papers
5223
Paper id
961621712160416043
v2026.09.13