Arrow Research search

Author name cluster

Jin Zhou

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

AAAI Conference 2026 Conference Paper

SCORE: Semantic Collage by Optimizing Rendered Elements

  • Zefan Shao
  • Jin Zhou
  • Hongliang Yang
  • Pengfei Xu

Collage is a powerful medium for visual expression, traditionally demanding significant artistic expertise and manual effort. Existing methods often struggle with a trade-off between semantic expression and the visual fidelity of the constituent images. To address this, we introduce SCORE (Semantic Collage by Optimizing Rendered Elements), a novel text-driven framework that automates the creation of semantically rich and structurally sound collages. Our key innovation is to shift the optimization process entirely into the image space. By employing a differentiable renderer, we can backpropagate gradients from a powerful, pre-trained text-to-image model directly to the spatial parameters, including position, rotation, and scale, of each image element. We leverage Variational Score Distillation (VSD) to provide robust semantic guidance from a text prompt, ensuring the final layout aligns with the desired concept. Crucially, our ''minimal editing'' principle preserves the integrity of the original elements by forgoing any content-level modifications. The layout is refined by a joint loss function that combines the VSD-based semantic loss with structural regularizers that penalize overlap and enforce boundary constraints. The output of SCORE is a parametric, structured representation that allows further editing and downstream use. Our work reduces the barrier to creative expression and provides a new, powerful paradigm for organizing visual contents.

AAAI Conference 2026 Conference Paper

StrokeFusion: Vector Sketch Generation via Joint Stroke-UDF Encoding and Latent Sequence Diffusion

  • Jin Zhou
  • Yi Zhou
  • Hongliang Yang
  • Pengfei Xu
  • Hui Huang

In the field of sketch generation, raster-format trained models often produce non-stroke artifacts, while vector-format trained models typically lack a holistic understanding of sketches, leading to compromised recognizability. Moreover, existing methods struggle to extract common features from similar elements (e.g., eyes of animals) appearing at varying positions across sketches. To address these challenges, we propose StrokeFusion, a two-stage framework for vector sketch generation. It contains a dual-modal sketch feature learning network that maps strokes into a high-quality latent space. This network decomposes sketches into normalized strokes and jointly encodes stroke sequences with Unsigned Distance Function (UDF) maps, representing sketches as sets of stroke feature vectors. Building upon this representation, our framework exploits a stroke-level latent diffusion model that simultaneously adjusts stroke position, scale, and trajectory during generation. This enables high-fidelity stroke generation while supporting stroke interpolation editing. Extensive experiments across multiple sketch datasets, demonstrate that our framework outperforms state-of-the-art techniques, validating its effectiveness in preserving structural integrity and semantic features.

NeurIPS Conference 2025 Conference Paper

$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training

  • Jin Zhou
  • Kaiwen Wang
  • Jonathan Chang
  • Zhaolin Gao
  • Nathan Kallus
  • Kilian Weinberger
  • Kianté Brantley
  • Wen Sun

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inherited from pre-training. In this work, we introduce $Q\sharp$, a value-based algorithm for KL-regularized RL that guides the reference policy using the optimal regularized $Q$ function. We propose to learn the optimal $Q$ function using distributional RL on an aggregated online dataset. Unlike prior value-based baselines that guide the model using unregularized $Q$-values, our method is theoretically principled and provably learns the optimal policy for the KL-regularized RL problem. Empirically, $Q\sharp$ outperforms prior baselines in math reasoning benchmarks while maintaining a smaller KL divergence to the reference policy. Theoretically, we establish a reduction from KL-regularized RL to no-regret online learning, providing the first bounds for deterministic MDPs under only realizability. Thanks to distributional RL, our bounds are also variance-dependent and converge faster when the reference policy has small variance. In sum, our results highlight $Q\sharp$ as an effective approach for post-training LLMs, offering both improved performance and theoretical guarantees. The code can be found at \url{https: //github. com/jinpz/q_sharp}.

ICRA Conference 2025 Conference Paper

Dashing for the Golden Snitch: Multi-Drone Time-Optimal Motion Planning with Multi-Agent Reinforcement Learning

  • Xian Wang
  • Jin Zhou
  • Yuanli Feng
  • Jiahao Mei
  • Jiming Chen 0001
  • Shuo Li

Recent innovations in autonomous drones have facilitated time-optimal flight in single-drone configurations, and enhanced maneuverability in multi-drone systems by applying optimal control and learning-based methods. However, few studies have achieved time-optimal motion planning for multi-drone systems, particularly during highly agile maneuvers or in dynamic scenarios. This paper presents a decentralized policy network using multi-agent reinforcement learning for time-optimal multi-drone flight. To strike a balance between flight efficiency and collision avoidance, we introduce a soft collision-free mechanism inspired by optimization-based methods. By customizing PPO in a centralized training, decentralized execution (CTDE) fashion, we unlock higher efficiency and stability in training while ensuring lightweight implementation. Extensive simulations show that, despite slight performance tradeoffs compared to single-drone systems, our multi-drone approach maintains near-time-optimal performance with a low collision rate. Real-world experiments validate our method, with two quadrotors using the same network as in simulation achieving a maximum speed of 13. 65 m/s and a maximum body rate of 13. 4 rad/s in a 5. 5 m × 5. 5 m × 2. 0 m space across various tracks, relying entirely on onboard computation [video 3 3 https://youtu.be/KACuFMtGGpo][code 4 4 https://github.com/KafuuChikai/Dashing-for-the-Golden-Snitch-Multi-Drone-RL].

ICRA Conference 2025 Conference Paper

Gate-Aware Online Planning for Two-Player Autonomous Drone Racing

  • Fangguo Zhao
  • Jiahao Mei
  • Jin Zhou
  • Yuanyi Chen
  • Jiming Chen
  • Shuo Li

The flying speed of autonomous quadrotors has increased significantly in the field of autonomous drone racing. However, most research primarily focuses on the aggressive flight of a single quadrotor, simplifying the racing gate traversal problem to a waypoint passing problem that neglects the orientations of the racing gates or implicitly considers the waypoint direction during path planning. In this paper, we propose a systematic method called Pairwise Model Predictive Control (PMPC) that can guide two quadrotors online to navigate racing gates with minimal time and without collisions. The flight task is initially simplified as a point-mass model waypoint passing problem to provide time optimal reference through an efficient two-step velocity search method. Subsequently, we utilize the spatial configuration of the racing track to compute the optimal heading at each gate, maximizing the visibility of subsequent gates for the quadrotors. To address varying gate orientations, we introduce a novel Magnetic Induction Line-based spatial curve to guide the quadrotors through racing gates of different orientations. Furthermore, we formulate a nonlinear optimization problem that uses the point-mass trajectory as initial values and references to enhance solving efficiency. The feasibility of the proposed method is validated through both simulation and real-world experiments. In real-world tests, the two quadrotors achieved a top speed of $6. 1m/s$ on a 7-waypoint racing track within a compact flying arena of $5m\times 4m\times 2m$.

ICLR Conference 2025 Conference Paper

MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science

  • Erle Zhu
  • Yadi Liu
  • Zhe Zhang
  • Xujun Li
  • Jin Zhou
  • Xinjie Yu
  • Minlie Huang
  • Hongning Wang

Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. However, their performance is still lacking in physical domains that require understanding diagrams with complex physical structures and quantitative analysis based on multi-modal information. To address this, we develop a new framework, named **M**ulti-Modal Scientific Re**A**soning with **P**hysics Perception and **S**imulation (**MAPS**) based on an MLLM. MAPS decomposes expert-level multi-modal reasoning task into physical diagram understanding via a Physical Perception Model (PPM) and reasoning with physical knowledge via a simulator. The PPM module is obtained by fine-tuning a visual language model using carefully designed synthetic data with paired physical diagrams and corresponding simulation language descriptions. At the inference stage, MAPS integrates the simulation language description of the input diagram provided by PPM and results obtained through a Chain-of-Simulation process with MLLM to derive the underlying rationale and the final answer. Validated using our collected college-level circuit analysis problems, MAPS significantly improves reasoning accuracy of MLLM and outperforms all existing models. The results confirm MAPS offers a promising direction for enhancing multi-modal scientific reasoning ability of MLLMs. We will release our code, model and dataset used for our experiments upon publishing of this paper.

EAAI Journal 2025 Journal Article

Multi-agent reinforcement learning for vibration control of regenerative active suspension

  • Xiaotian Gao
  • Yu Du
  • Shiyuan Han
  • Wenxiu Zhao
  • Jin Zhou
  • Tong Zhang
  • C.L. Philip Chen

The comfort and responsiveness of active suspension systems have surpassed those of traditional suspension systems, but their development and implementation have been significantly constrained by high power requirements. In response to the current drive for energy conservation and emission reduction, regenerative active suspension systems effectively address the issues of high power consumption and energy loss while demonstrating considerable market potential. The main contribution of the investigation lies in designing a novel regenerative active suspension system specifically for electric vehicles powered by electric motors. The system integrates an electric motor and a generator to ensure precise power delivery and energy recovery from suspension dynamics. To fulfill the dual requirements of energy regeneration and ride smoothness, the Double Proximal Policy Optimization-Multi Dimensional Output (DPPO-MDO) algorithm has been devised based on multi-agent reinforcement learning. The algorithm features two agents equipped with the Proximal Policy Optimization (PPO) algorithm, each managing the operations of the motor and generator respectively. Furthermore, to accommodate diverse scenarios and balance comfort with energy efficiency, economic and comfort modes have been developed for the regenerative active suspension through the DPPO-MDO algorithm. Comprehensive experimental analysis demonstrates that both redesigned reward function and innovative reinforcement learning framework significantly enhance the application of multi-agent reinforcement learning in regenerative active suspension systems.

IROS Conference 2025 Conference Paper

Online Motion Planning for Quadrotor Multi-Point Navigation Using Efficient Imitation Learning-Based Strategy

  • Jin Zhou
  • Jiahao Mei
  • Fangguo Zhao
  • Jiming Chen 0001
  • Shuo Li

Over the past decade, there has been a remarkable surge in utilizing quadrotors for various purposes due to their simple structure and aggressive maneuverability. One of the key challenges is online time-optimal trajectory generation and control technique. This paper proposes an imitation learning-based online solution to efficiently navigate the quadrotor through multiple waypoints with near-time-optimal performance. The neural networks (WN&CNets) are trained to learn the control law from the dataset generated by the time-consuming CPC algorithm and then deployed to generate the optimal control commands online to guide the quadrotors. To address the challenge of limited training data and the hover maneuver at the final waypoint, we propose a transition phase strategy that utilizes MINCO trajectories to help the quadrotor ‘jump over’ the stop-and-go maneuver when switching waypoints. Our method is demonstrated in both simulation and real-world experiments, achieving a maximum speed of 5. 6m/s while navigating through 7 waypoints in a confined space of 5. 5m × 5. 5m × 2. 0m [video 3 ]. The results show that with a slight loss in optimality, the WN&CNets significantly reduce the processing time and enable online control for multi-point flight tasks.

NeurIPS Conference 2025 Conference Paper

Value-Guided Search for Efficient Chain-of-Thought Reasoning

  • Kaiwen Wang
  • Jin Zhou
  • Jonathan Chang
  • Zhaolin Gao
  • Nathan Kallus
  • Kianté Brantley
  • Wen Sun

In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method does not require a fine-grained notion of ``step, '' which is difficult to define for long-context reasoning models. By collecting a dataset of 2. 5 million reasoning traces, we train a 1. 5B token-level value model and apply it to DeepSeek models for improved performance with test-time compute scaling. We find that block-wise value-guided search (\texttt{VGS}) with a final weighted majority vote achieves better test-time scaling than standard methods such as majority voting or best-of-$n$. Moreover, \texttt{VGS} significantly reduces the inference FLOPs required to achieve the same performance of majority voting. Our dataset, model and codebase are open-sourced at \codeurl.

IROS Conference 2024 Conference Paper

An Observability Constrained Downward-Facing Optical-Flow-Aided Visual-Inertial Odometry

  • Dandi Liu
  • Jiahao Mei
  • Jin Zhou
  • Shuo Li

Visual-Inertial Odometry (VIO) has been widely used by autonomous drones as an onboard navigation method. However, it suffers from drifts especially in scenarios where the environments have few texture features such as an empty room with solid color walls. Optical flow sensors are another type of onboard sensor used by drones that face downward and measure the velocity by detecting changes in pixels between consecutive images, which don’t introduce accumulative error. In this work, we present an efficient tight-coupled estimator to improve the accuracy of VIO by fusing the measurements of a downward-facing optical flow sensor into the VIO framework consistently. We further analyze the observability of the estimators and prove that there are four unobservable directions in the ideal case and then we utilize OC-EKF to maintain the consistency of the estimator. Furthermore, we extend an adaptive weighting algorithm to the proposed method, which can better adapt to the scenes where feature tracking is less accurate. Finally, both simulation and real-world experiments demonstrate the feasibility of the proposed method.

IROS Conference 2023 Conference Paper

Aggressive Trajectory Generation for a Swarm of Autonomous Racing Drones

  • Yuyang Shen
  • Jin Zhou
  • Danzhe Xu
  • Fangguo Zhao
  • Jinming Xu 0002
  • Jiming Chen 0001
  • Shuo Li

Autonomous drone racing is becoming an excellent platform to challenge quadrotors' autonomy techniques including planning, navigation and control technologies. However, most research on this topic mainly focuses on single drone scenarios. In this paper, we describe a novel time-optimal trajectory generation method for generating time-optimal trajectories for a swarm of quadrotors to fly through pre-defined waypoints with their maximum maneuverability without collision. We verify the method in the Gazebo simulations where a swarm of 5 quadrotors can fly through a complex 6-waypoint racing track in a $35m\times 35m$ space with a top speed of 14m/s. Flight tests are performed on two quadrotors passing through 3 waypoints in a $4m\times 2m$ flight arena to demonstrate the feasibility of the proposed method in the real world. Both simulations and real-world flight tests show that the proposed method can generate the optimal aggressive trajectories for a swarm of autonomous racing drones. The method can also be easily transferred to other types of robot swarms.

JBHI Journal 2022 Journal Article

Breast Tumor Classification Based on MRI-US Images by Disentangling Modality Features

  • Mengyun Qiao
  • Chencheng Liu
  • Zeju Li
  • Jin Zhou
  • Qin Xiao
  • Shichong Zhou
  • Cai Chang
  • Yajia Gu

Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) and ultrasound (US), which are two common modalities for clinical breast tumor diagnosis besides Mammograms, can provide different and complementary information for the same tumor regions. Although many machine learning methods have been proposed for breast tumor classification based on either single modality, it remains unclear how to further boost the classification performance by utilizing paired multi-modality information with different dimensions. In this paper, we propose MRI-US multi-modality network (MUM-Net) to classify breast tumor into different subtypes based on 3D MR and 2D US images. The key insight of MUM-Net is that we explicitly distill modality-agnostic features for tumor classification. Specifically, we first adopt a discrimination-adaption module to decompose features into modality-agnostic and modality-specific ones with min-max training strategies. Then, we propose a feature fusion module to increase the compactness of the modality-agnostic features by utilizing an affinity matrix with nearest neighbour selection. We build a paired MRI-US breast tumor classification dataset containing 502 cases with three clinical indicators to validate the proposed method. In three tasks including lymph node metastasis, histological grade and Ki-67 level, MUM-Net achieves AUC scores of 0. 8581, 0. 8965 and 0. 8577, outperforming other counterparts which are based on single task or single modality by a wide margin. In addition, we find that the extracted modality-agnostic features can help the network focus on the tumor regions in both modalities.

v2026.09.13