Arrow Research search

Author name cluster

Zining Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

EAAI Journal 2026 Journal Article

Beyond network community: Authority hierarchy reveals more

  • Zining Wang
  • Ziyu Zhang
  • Jun Tang
  • Xian Wu
  • Qingtao Pan
  • Zhaolin Lv
  • Haosen Wang
  • Xing Wang

Most complex networks, if not all, inherently possess community and hierarchichal structure. The hierarchy between nodes within these communities provides a more refined perspective for network analysis and optimization compared to the mesoscale community structure. To this end, we introduce a novel method for Local Community Division based on Authority Hierarchy (LCDAH). Our method advances network data mining by constructing an authority hierarchy graph -- a directed structure that explicitly models pairwise authority relationships. Within this graph, densely connected core nodes are efficiently identified at its apex and serve as co-leaders for community formation; communities are subsequently assigned to each node by traversing downward from these cores through the graph. The method not only detects community boundaries with high accuracy, outperforming benchmarks on six real-world networks, but also reveals the internal hierarchical structure, offering insights beyond mere partitioning. We demonstrate its utility in two data mining applications: image clustering via network transformation and analysis of an international trade network, validating its effectiveness in modeling complex systems.

AAAI Conference 2026 Conference Paper

I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle

  • Zhaoquan Yuan
  • Zining Wang
  • Yuankang Pan
  • Ao Luo
  • Wei Li
  • Xiao Wu
  • Changsheng Xu

Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks.

ICML Conference 2025 Conference Paper

HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration

  • Yushi Huang
  • Zining Wang
  • Ruihao Gong
  • Jing Liu 0048
  • Xinjie Zhang
  • Jinyang Guo
  • Xianglong Liu 0001
  • Jun Zhang 0004

Diffusion Transformers (DiTs) excel in generative tasks but face practical deployment challenges due to high inference costs. Feature caching, which stores and retrieves redundant computations, offers the potential for acceleration. Existing learning-based caching, though adaptive, overlooks the impact of the prior timestep. It also suffers from misaligned objectives- aligned predicted noise vs. high-quality images -between training and inference. These two discrepancies compromise both performance and efficiency. To this end, we harmonize training and inference with a novel learning-based caching framework dubbed HarmoniCa. It first incorporates Step-Wise Denoising Training (SDT) to ensure the continuity of the denoising process, where prior steps can be leveraged. In addition, an Image Error Proxy-Guided Objective (IEPO) is applied to balance image quality against cache utilization through an efficient proxy to approximate the image error. Extensive experiments across $8$ models, $4$ samplers, and resolutions from $256\times256$ to $2K$ demonstrate superior performance and speedup of our framework. For instance, it achieves over $40\%$ latency reduction (*i. e. *, $2. 07\times$ theoretical speedup) and improved performance on PixArt-$\alpha$. Remarkably, our *image-free* approach reduces training time by $25\%$ compared with the previous method. Our code is available at https: //github. com/ModelTC/HarmoniCa.

AAAI Conference 2025 Conference Paper

InstructOCR: Instruction Boosting Scene Text Spotting

  • Chen Duan
  • Qianyi Jiang
  • Pei Fu
  • Jiamin Chen
  • Shengxi Li
  • Zining Wang
  • Shan Guo
  • Junfeng Luo

In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.

ICML Conference 2024 Conference Paper

Compressing Large Language Models by Joint Sparsification and Quantization

  • Jinyang Guo
  • Jianyu Wu
  • Zining Wang
  • Jiaheng Liu
  • Ge Yang
  • Yifu Ding
  • Ruihao Gong
  • Haotong Qin

In this paper, we introduce a novel model compression technique named Joint Sparsification and Quantization (JSQ), explicitly tailored for large language models (LLMs). Traditional methods employ either sparsification or quantization individually to compress LLMs, leading to performance degradation at high compression ratios. In contrast, our JSQ approach integrates sparsification and quantization cohesively. As sparsification tend to preserve outliers that is harmful to quantization, we introduce a novel sparsity metric to serves as a bridge between the sparsification and quantization. Moreover, it is proven outliers in LLMs have significant impact but harmful to compression. Current solutions are highly coupled with quantization process, which is not helpful to sparsification. To this end, we also introduce a search-based activation editor to automatically eliminate relatively useless outliers. Comprehensive experiments across various datasets and architectures affirm the efficacy of our JSQ framework. Notably, our JSQ achieves 7. 96$\times$ computation reduction without crashing for the representative model LLaMA. This accomplishment stands in stark contrast to the limitations of most state-of-the-art LLM compression methods, which typically fail under such extreme compression ratios. Our code is released at https: //github. com/uanu2002/JSQ.

AAAI Conference 2024 Conference Paper

Fast and Controllable Post-training Sparsity: Learning Optimal Sparsity Allocation with Global Constraint in Minutes

  • Ruihao Gong
  • Yang Yong
  • Zining Wang
  • Jinyang Guo
  • Xiuying Wei
  • Yuqing Ma
  • Xianglong Liu

Neural network sparsity has attracted many research interests due to its similarity to biological schemes and high energy efficiency. However, existing methods depend on long-time training or fine-tuning, which prevents large-scale applications. Recently, some works focusing on post-training sparsity (PTS) have emerged. They get rid of the high training cost but usually suffer from distinct accuracy degradation due to neglect of the reasonable sparsity rate at each layer. Previous methods for finding sparsity rates mainly focus on the training-aware scenario, which usually fails to converge stably under the PTS setting with limited data and much less training cost. In this paper, we propose a fast and controllable post-training sparsity (FCPTS) framework. By incorporating a differentiable bridge function and a controllable optimization objective, our method allows for rapid and accurate sparsity allocation learning in minutes, with the added assurance of convergence to a predetermined global sparsity rate. Equipped with these techniques, we can surpass the state-of-the-art methods by a large margin, e.g., over 30\% improvement for ResNet-50 on ImageNet under the sparsity rate of 80\%. Our plug-and-play code and supplementary materials are open-sourced at https://github.com/ModelTC/FCPTS.

IROS Conference 2020 Conference Paper

Inferring Spatial Uncertainty in Object Detection

  • Zining Wang
  • Di Feng
  • Yiyang Zhou
  • Lars Rosenbaum
  • Fabian Timm
  • Klaus Dietmayer
  • Masayoshi Tomizuka
  • Wei Zhan

The availability of real-world datasets is the prerequisite for developing object detection methods for autonomous driving. While ambiguity exists in object labels due to error-prone annotation process or sensor observation noises, current object detection datasets only provide deterministic annotations without considering their uncertainty. This precludes an in-depth evaluation among different object detection methods, especially for those that explicitly model predictive probability. In this work, we propose a generative model to estimate bounding box label uncertainties from LiDAR point clouds, and define a new representation of the probabilistic bounding box through spatial distribution. Comprehensive experiments show that the proposed model represents uncertainties commonly seen in driving scenarios. Based on the spatial distribution, we further propose an extension of IoU, called the Jaccard IoU (JIoU), as a new evaluation metric that incorporates label uncertainty. Experiments on the KITTI and the Waymo Open Datasets show that JIoU is superior to IoU when evaluating probabilistic object detectors.

ICRA Conference 2020 Conference Paper

Precise 3D Calibration of Wafer Handling Robot by Visual Detection and Tracking of Elliptic-shape Wafers

  • Zining Wang
  • Masayoshi Tomizuka

This work provides a framework for the 3D calibration of wafers and a wafer handling robot by monocular vision. The proposed method precisely reconstructs the 3D poses of wafers from a set of images captured by the camera mounted on the robot. In addition, it calibrates the robot kinematics simultaneously. A robust ellipse detection and tracking algorithm based on the edge arcs is developed to recognize wafers among images. Then a joint optimization is constructed from a multi-object pose graph to solve the 3D poses of wafers and other calibration parameters of the robot-camera system. The proposed tracking method is able to associate multiple incomplete elliptic segments using a Gaussian Mixture Model-based registration algorithm. The algorithm is point-based where no feature descriptor is required. The proposed 3D pose optimization incorporates shape constraints, and is more accurate than the point-wise reconstruction produced by classic bundle adjustment methods.

ICRA Conference 2016 Conference Paper

Robust two-degree-of-freedom iterative learning control for flexibility compensation of industrial robot manipulators

  • Cong Wang 0015
  • Minghui Zheng
  • Zining Wang
  • Masayoshi Tomizuka

Most industrial robots are actuated using geared motors with no direct load side measurement. The flexibility introduced by the gear reducer causes transmission errors and vibrations, which limits the adoption of robot manipulators in many demanding applications. This paper presents a lean and efficient scheme of iterative learning control (ILC) to compensate for the joint flexibility of industrial robot manipulators. A two-degree-of-freedom ILC method is introduced. Compared with the dual-stage ILC that has been previously proposed for servo flexibility compensation, the method is more effective and also enables a leaner implementation. In addition, in order to handle system variation, a robust synthesis method is developed by using H∞ and μ techniques in an innovative way. The proposed method is analyzed using simulation studies as well as tested on an actual industrial robot manipulator.

v2026.09.13