Arrow Research search

Author name cluster

Jiawei Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

EAAI Journal 2026 Journal Article

Structure-aware coarse-to-fine upsampling network for arbitrary-scale super-resolution of remote sensing images

  • Shiyan Wang
  • Jiawei Zhao
  • Jiajia Tang

Arbitrary-scale super-resolution (ASSR) methods aim to reconstruct high-resolution images with arbitrary scaling factors by learning the mapping between the latent codes in the coordinate domain and the Red-Green-Blue (RGB) values in the spatial domain. However, most ASSR methods based on implicit neural representation rely solely on local feature aggregation for latent code generation, which inherently limits receptive fields and fails to capture global structural patterns. This limitation significantly degrades performance, especially in large-scale remote sensing super-resolution tasks. To address these issues, we propose a Structure-aware Coarse-to-Fine arbitrary-scale Super-Resolution (SC2FSR) framework. SC2FSR employs a triple-branch architecture to jointly extract structural priors and multi-level features, constructing enhanced contextual representations through the proposed Structure-Guided Feature Interaction (SGFI). The SGFI module utilizes cascaded High-Order Channel Attention (HOCA) to facilitate feature integration across shallow textures, semantic information, and geometric structures, simultaneously generating higher-order statistics and global contextual cues for latent code enrichment. Furthermore, a Coarse-to-Fine Upsampling (C2FUP) pipeline is established, where latent codes are first realigned with structural priors via multi-layer perceptron for coordinate-wise dense prediction, then refined through structure-aware weighted filters. These multi-scale filters effectively integrate structural patterns from local to global, thereby enlarging the receptive field and refining the detailed performance of high-resolution images. Extensive experiments on benchmark datasets demonstrate that SC2FSR not only outperforms state-of-the-art methods but also achieves advanced performance at non-integer and large scales while preserving fine structural details.

IROS Conference 2025 Conference Paper

A Study on the Generation of Single Cell Droplets via the Combination of Lateral-Field Optoelectronic Tweezers and Electrowetting-on-Dielectric

  • Jiawei Zhao
  • Shunxiao Huang
  • Hongyi Xiong
  • Chunyuan Gan
  • Jingwen Ye
  • Wenyan Niu
  • Lin Feng 0002

Microfluidic technology is currently a popular approach in the field of single-cell research, which is used to reveal the heterogeneity among cells. However, most of the existing microfluidic technologies for single-cell research lack the ability to control the microenvironment of single cells after isolating them. In this work, a technology that combines lateral-field optoelectronic tweezers (LOET) with electrowetting-on-dielectric (EWOD) is used to separate cells into single cells and then encapsulate each single cell within an individual droplet, generating single-cell droplets. More importantly, it also enables the control of the microenvironment of the separated single cells. The driving control of the single - cell droplets is achieved through the EWOD, which has good application prospects in the field of single- cell research.

NeurIPS Conference 2025 Conference Paper

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

  • Haizhong Zheng
  • Yang Zhou
  • Brian Bartoldson
  • Bhavya Kailkhura
  • Fan Lai
  • Jiawei Zhao
  • Beidi Chen

Reinforcement learning, such as PPO and GRPO, has powered recent breakthroughs in LLM reasoning. Scaling rollout to sample more prompts enables models to selectively use higher-quality data for training, which can stabilize RL training and improve model performance, but at the cost of significant computational overhead. In this paper, we first show that a substantial portion of this overhead can be avoided by skipping uninformative prompts before rollout. Our analysis of reward dynamics reveals a strong temporal consistency in prompt value: prompts that are uninformative in one epoch of training are likely to remain uninformative in near future epochs. Based on these insights, we propose GRESO (GRPO with Efficient Selective Rollout), an online, lightweight pre-rollout filtering algorithm that predicts and skips uninformative prompts using reward training dynamics. By evaluating GRESO on a broad range of math reasoning benchmarks and models, like Qwen2. 5-Math-1. 5B, DeepSeek-R1-Distill-Qwen-1. 5B, Qwen2. 5-Math-7B, Qwen2. 5-14B, and Qwen2. 5-32B, we show that GRESO achieves up to 2. 4x wall-clock time speedup in rollout and up to 2. 0x speedup in total training time without accuracy degradation. We make our code publicly available at https: //github. com/Infini-AI-Lab/GRESO/.

ICML Conference 2025 Conference Paper

From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

  • Ajay Kumar Jaiswal
  • Yifan Wang
  • Lu Yin 0006
  • Shiwei Liu 0003
  • Runjin Chen
  • Jiawei Zhao
  • Ananth Grama
  • Yuandong Tian

Large Language Models (LLMs) matrices can often be expressed in low-rank format with potential to relax memory and compute resource requirements. Unlike previous works which pivot around developing novel matrix decomposition algorithms, in this work we focus to study the emerging non-uniform low-rank properties across weight matrices in LLMs through the lens of stabilizing gradient subspace. Firstly, we provide a theoretical framework to understand the stabilization of gradient subspaces through Hessian analysis. Secondly, we empirically establish a consequential relationship between the gradient dynamics and low-rank expressiveness of weight matrices. Our findings reveal that different LLM components exhibit varying levels of converged low-rank structure, necessitating a non-uniform rank reduction across them to minimize performance drop due to compression. In view of that, we present Weight Low-Rank Projection (WeLore) that unifies weight compression and memory-efficient fine-tuning as ONE, in a data-agnostic and one-shot way. Going beyond only as a compression technique, WeLore categorizes weight matrices into Low-rank Components (LRCs) and Non-Low-rank Components (N-LRCs) based on their ability to express themselves as low-rank. Our gradient dynamics perspective illustrate that LRCs tend to have better finetuning capabilities and their standalone finetuning can closely mimic (sometimes outperform) the training loss trajectory and performance of full-finetuning with notable memory and compute footprint reduction. All codes and checkpoints will be released.

IROS Conference 2025 Conference Paper

High-Precision Parallel Manipulation of Multi-Particle System Using Optoelectronic Tweezers

  • Shunxiao Huang
  • Jiawei Zhao
  • Chunyuan Gan
  • Zijin Zeng
  • Hongyi Xiong
  • Jingwen Ye
  • Wenyan Niu
  • Ao Wang

This paper presents a multi-particle parallel manipulation optoelectronic tweezers system integrated with computer vision technology, enabling the parallel and precise manipulation of dozens of particles. This system significantly enhances manipulation efficiency while maintaining high precision. By real-time monitoring of particle motion and light patterns, the system can rapidly adjust and optimize its manipulation strategy, thereby improving the stability and reliability of multi-particle synchronization in complex environments. Extensive experimental results demonstrate the system’s outstanding performance. For instance, it can quickly arrange complex patterns and letter sequences, facilitate the coordinated assembly of organoids from particle groups, and efficiently perform the precise separation and arrangement of mixed particles. The core advantage of this system lies in its high parallelism and flexibility, enabling it to handle large-scale synchronous manipulation tasks with exceptional operating accuracy. With continuous technological advancements and the broadening of application scenarios, this system is expected to have a profound impact in fields such as cell sorting, micro-device assembly, and organoid construction, providing robust support for research and technological development in these areas.

NeurIPS Conference 2025 Conference Paper

ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization

  • Zechun Liu
  • Changsheng Zhao
  • Hanxian Huang
  • Sijia Chen
  • Jing Zhang
  • Jiawei Zhao
  • Scott Roy
  • Lisa Jin

The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, others propose that 1. 58-bit offers superior results. However, the lack of a cohesive framework for different bits has left such conclusions relatively tenuous. We present ParetoQ, the first unified framework that facilitates rigorous comparisons across 1-bit, 1. 58-bit, 2-bit, 3-bit, and 4-bit quantization settings. Our findings reveal a notable learning transition between 2 and 3 bits: For 3-bits and above, the fine-tuned models stay close to their original pre-trained distributions, whereas for learning 2-bit networks or below, the representations change drastically. By optimizing training schemes and refining quantization functions, ParetoQ surpasses all previous methods tailored to specific bit widths. Remarkably, our ParetoQ ternary 600M-parameter model even outperforms the previous SoTA ternary 3B-parameter model in accuracy, using only one-fifth of the parameters. Extensive experimentation shows that ternary, 2-bit, and 3-bit quantization maintains comparable performance in the size-accuracy trade-off and generally exceeds 4-bit and binary quantization. Considering hardware constraints, 2-bit quantization offers promising potential for memory reduction and speedup.

ICRA Conference 2024 Conference Paper

Dynamic Adaptive Imaging System on Optoelectronic Tweezers Platform

  • Ao Wang
  • Chunyuan Gan
  • Haocheng Han
  • Hongyi Xiong
  • Jiawei Zhao
  • Chutian Wang
  • Lin Feng 0002

Optoelectronic tweezers (OET) has shown great promise in various applications, especially in the precise manipulation of microparticles and microorganisms on a micron and nanometer scale. This technology significantly enhances the efficiency of single-cell sorting and the development of antibody-based drugs. However, conventional OET platforms are limited by issues such as low autofocusing accuracy, restricted imaging field of view, and uneven illumination. To overcome these limitations, we have innovatively developed a dynamic adaptive imaging system. By incorporating peak-finding and in situ Gaussian blur compensation algorithms, we achieved rapid automatic focusing and illumination shadow compensation across an expanded field of view. At the same time, the system can also dynamically adjust compensation parameters under different lighting conditions. Our system has successfully completed comprehensive scanning of the optoelectronic tweezers chip, achieving a 60% reduction in autofocus time and a 15. 8% improvement in lighting uniformity. Moreover, this imaging system demonstrates robust versatility and can serve as a reference for other optical systems.

ICML Conference 2024 Conference Paper

GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

  • Jiawei Zhao
  • Zhenyu Zhang 0015
  • Beidi Chen
  • Zhangyang Wang
  • Anima Anandkumar
  • Yuandong Tian

Training Large Language Models (LLMs) presents significant memory challenges, predominantly due to the growing size of weights and optimizer states. Common memory-reduction approaches, such as low-rank adaptation (LoRA), add a trainable low-rank matrix to the frozen pre-trained weight in each layer, reducing trainable parameters and optimizer states. However, such approaches typically underperform training with full-rank weights in both pre-training and fine-tuning stages since they limit the parameter search to a low-rank subspace and alter the training dynamics, and further, may require full-rank warm start. In this work, we propose Gradient Low-Rank Projection (GaLore), a training strategy that allows full-parameter learning but is more memory-efficient than common low-rank adaptation methods such as LoRA. Our approach reduces memory usage by up to 65. 5% in optimizer states while maintaining both efficiency and performance for pre-training on LLaMA 1B and 7B architectures with C4 dataset with up to 19. 7B tokens, and on fine-tuning RoBERTa on GLUE tasks. Our 8-bit GaLore further reduces optimizer memory by up to 82. 5% and total training memory by 63. 3%, compared to a BF16 baseline. Notably, we demonstrate, for the first time, the feasibility of pre-training a 7B model on consumer GPUs with 24GB memory (e. g. , NVIDIA RTX 4090) without model parallel, checkpointing, or offloading strategies.

TMLR Journal 2024 Journal Article

Incremental Spatial and Spectral Learning of Neural Operators for Solving Large-Scale PDEs

  • Robert Joseph George
  • Jiawei Zhao
  • Jean Kossaifi
  • Zongyi Li
  • Anima Anandkumar

Fourier Neural Operators (FNO) offer a principled approach to solving challenging partial differential equations (PDE) such as turbulent flows. At the core of FNO is a spectral layer that leverages a discretization-convergent representation in the Fourier domain, and learns weights over a fixed set of frequencies. However, training FNO presents two significant challenges, particularly in large-scale, high-resolution applications: (i) Computing Fourier transform on high-resolution inputs is computationally intensive but necessary since fine-scale details are needed for solving many PDEs, such as fluid flows, (ii) selecting the relevant set of frequencies in the spectral layers is challenging, and too many modes can lead to overfitting, while too few can lead to underfitting. To address these issues, we introduce the Incremental Fourier Neural Operator (iFNO), which progressively increases both the number of frequency modes used by the model as well as the resolution of the training data. We empirically show that iFNO reduces total training time while maintaining or improving generalization performance across various datasets. Our method demonstrates a 38% lower testing error, using 20% fewer frequency modes compared to the existing FNO, while also achieving up to 46% faster training and a 2.8x reduction in model size.

NeurIPS Conference 2024 Conference Paper

Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences Training

  • Cheng Luo
  • Jiawei Zhao
  • Zhuoming Chen
  • Beidi Chen
  • Anima Anandkumar

We introduce Mini-Sequence Transformer (MsT), a simple and effective methodology for highly efficient and accurate LLM training with extremely long sequences. MsT partitions input sequences and iteratively processes mini-sequences to reduce intermediate memory usage. Integrated with activation recomputation, it enables significant memory savings in both forward and backward passes. In experiments with the Llama3-8B model, with MsT, we measure no degradation in throughput or convergence even with 12x longer sequences than standard implementations. MsT is fully general, implementation-agnostic, and requires minimal code changes to integrate with existing LLM training frameworks. Integrated with the huggingface library, MsT successfully extends the maximum context length of Qwen, Mistral, and Gemma-2 by 12-24x.

NeurIPS Conference 2024 Conference Paper

S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity

  • Xinyu Yang
  • Jixuan Leng
  • Geyang Guo
  • Jiawei Zhao
  • Ryumei Nakada
  • Linjun Zhang
  • Huaxiu Yao
  • Beidi Chen

Current PEFT methods for LLMs can achieve high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability. Utilizing this key insight, we propose a family of Structured Sparse Fine-Tuning (S${^2}$FT) methods for LLMs, which concurrently achieve state-of-the-art fine-tuning performance, training efficiency, and inference scalability. S${^2}$FT accomplishes this by "selecting sparsely and computing densely". Based on the coupled structures in LLMs, \model selects a few attention heads and channels in the MHA and FFN modules for each Transformer block, respectively. Next, it co-permutes the weight matrices on both sides of all coupled structures to connect the selected subsets in each layer into a dense submatrix. Finally, S${^2}$FT performs in-place gradient updates on all selected submatrices. Through theoretical analyses and empirical results, our method prevents forgetting while simplifying optimization, delivers SOTA performance on both commonsense and arithmetic reasoning with 4. 6% and 1. 3% average improvements compared to LoRA, and surpasses full FT by 11. 5% when generalizing to various domains after instruction tuning. Using our partial back-propagation algorithm, S${^2}$FT saves training memory up to 3$\times$ and improves latency by 1. 5-2. 7$\times$ compared to full FT, while achieving an average 10\% improvement over LoRA on both metrics. We further demonstrate that the weight updates in S${^2}$FT can be decoupled into adapters, enabling effective fusion, fast switch, and efficient parallelism when serving multiple fine-tuned models.

IROS Conference 2023 Conference Paper

Microrobot Control Method Based on Movement of Field Free Point in Gradient Magnetic Field

  • Chutian Wang
  • Yiming Ji
  • Xinyun Luo
  • Chunyuan Gan
  • Jiapeng Yang
  • Jiawei Zhao
  • Luyao Wang
  • Lin Feng 0002

The untethered microrobots driven by multiple external physics fields have promising ability in minimally invasive disease treatments. One common type of the driving fields is gradient magnetic field, which can provide microrobots with adequate driving force in complicated environment. In this study, a control method of microrobot through gradient magnetic field system is presented, which is realized by moving the field free point (FFP) to produce an alterable magnetic driving force. A confirmatory experiment of the robot reciprocating motion control is undertaken in a 1D gradient magnetic robot system. The control method could be applied to further studies on in vivo applications of targeted microrobot drug delivery system.

IROS Conference 2023 Conference Paper

Parallel Cell Array Patterning and Target Cell Lysis on an Optoelectronic Micro-Well Device

  • Chunyuan Gan
  • Hongyi Xiong
  • Jiawei Zhao
  • Ao Wang
  • Chutian Wang
  • Shuzhang Liang
  • Jiaying Zhang
  • Lin Feng 0002

This work presents a novel electrical method, implemented in the form of a microfluidic device, for cell arraying and target cell lysis. The microfluidic device contains a micro-well array on the photoconductive layer based on the optoelectronic tweezers (OET) method, where parallel cell manipulation is performed. As cell suspension flows over the micro-wells, cells can be actively captured in the micro-wells by light-induced dielectrophoresis (DEP) forces, form the designed pattern array in less than 120 s. The single-cell capture rate is over 83 % in the patterned cell array, and about 94% of micro-wells are occupied by cells. Then, the target cell in the specific micro-well is illuminated and lysed by electroporation in 5 seconds. The micro-well barriers and DEP forces block the influence of the flow, and a relatively closed space is critical to preserve the cell lysates. Through experiments, light-induced DEP force cell capture and target cell electroporation can be modulated by changing the light patterns and the applied signal. This device, based on the OET and dynamic electroporation, allows the rapidity in the cell capture and target lysis at the single-cell level and can enable single-cell-based studies, such as molecular diagnostics and disease detection.

IROS Conference 2023 Conference Paper

UMIRobot: An Open-{Software, Hardware} Low-Cost Robotic Manipulator for Education

  • Murilo M. Marinho
  • Hung-Ching Lin
  • Jiawei Zhao

Robot teleoperation has been studied for the past 70 years and is relevant in many contexts, such as in the handling of hazardous materials and telesurgery. The COVID19 pandemic has rekindled interest in this topic, but the existing robotic education kits fall short of being suitable for teleoperated robotic manipulator learning. In addition, the global restrictions of motion motivated large investments in online/hybrid education. In this work, a newly developed robotics education kit and its ecosystem are presented which is used as the backbone of an online/hybrid course in teleoperated robots. The students are divided into teams. Each team designs, fabricates (3D printing and assembling), and implements a control strategy for a master device and gripper. Coupling those with the UMIRobot, provided as a kit, the students compete in a teleoperation challenge. The kit is low cost (<100USD), which allows higher-learning institutions to provide one kit per student and they can learn in a risk-free environment. As of now, 73 such kits have been assembled and sent to course participants in eight countries. As major success stories, we show an example of gripper and master designed for the proposed course. In addition, we show a teleoperated task between Japan and Bangladesh executed by course participants. Design files, videos, source code, and more information are available at https://mmmarinho.github.io/UMIRobot/

TMLR Journal 2022 Journal Article

ZerO Initialization: Initializing Neural Networks with only Zeros and Ones

  • Jiawei Zhao
  • Florian Tobias Schaefer
  • Anima Anandkumar

Deep neural networks are usually initialized with random weights, with adequately selected initial variance to ensure stable signal propagation during training. However, selecting the appropriate variance becomes challenging especially as the number of layers grows. In this work, we replace random weight initialization with a fully deterministic initialization scheme, viz., ZerO, which initializes the weights of networks with only zeros and ones (up to a normalization factor), based on identity and Hadamard transforms. Through both theoretical and empirical studies, we demonstrate that ZerO is able to train networks without damaging their expressivity. Applying ZerO on ResNet achieves state-of-the-art performance on various datasets, including ImageNet, which suggests random weights may be unnecessary for network initialization. In addition, ZerO has many benefits, such as training ultra deep networks (without batch-normalization), exhibiting low-rank learning trajectories that result in low-rank and sparse solutions, and improving training reproducibility.

NeurIPS Conference 2020 Conference Paper

Learning compositional functions via multiplicative weight updates

  • Jeremy Bernstein
  • Jiawei Zhao
  • Markus Meister
  • Ming-Yu Liu
  • Anima Anandkumar
  • Yisong Yue

Compositionality is a basic structural feature of both biological and artificial neural networks. Learning compositional functions via gradient descent incurs well known problems like vanishing and exploding gradients, making careful learning rate tuning essential for real-world applications. This paper proves that multiplicative weight updates satisfy a descent lemma tailored to compositional functions. Based on this lemma, we derive Madam---a multiplicative version of the Adam optimiser---and show that it can train state of the art neural network architectures without learning rate tuning. We further show that Madam is easily adapted to train natively compressed neural networks by representing their weights in a logarithmic number system. We conclude by drawing connections between multiplicative weight updates and recent findings about synapses in biology.

v2026.09.13