Arrow Research search

Author name cluster

Qingyi Gu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

AAAI Conference 2026 Conference Paper

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

  • Fuqiang Gu
  • Yuanke Li
  • Xianlei Long
  • Kangping Ji
  • Chao Chen
  • Qingyi Gu
  • Zhenliang Ni

Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic range conditions due to limitations of frame cameras. Event cameras offer complementary advantages such as high temporal resolution and low latency, yet lack color and texture, making them insufficient on their own. To address this, recent research has explored multimodal fusion of RGB and event data; however, many existing approaches are computationally expensive and focus primarily on spatial fusion, neglecting the temporal dynamics inherent in event streams. In this work, we propose MambaSeg, a novel dual-branch semantic segmentation framework that employs parallel Mamba encoders to efficiently model RGB images and event streams. To reduce cross-modal ambiguity, we introduce the Dual-Dimensional Interaction Module (DDIM), comprising a Cross-Spatial Interaction Module (CSIM) and a Cross-Temporal Interaction Module (CTIM), which jointly perform fine-grained fusion along both spatial and temporal dimensions. This design improves cross-modal alignment, reduces ambiguity, and leverages the complementary properties of each modality. Extensive experiments on the DDD17 and DSEC datasets demonstrate that MambaSeg achieves state-of-the-art segmentation performance while significantly reducing computational cost, showcasing its promise for efficient, scalable, and robust multimodal perception.

AAAI Conference 2026 Conference Paper

SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model

  • Jing Zhang
  • Zhikai Li
  • Chengzhi Hu
  • Xuewen Liu
  • Qingyi Gu

Segment Anything Model (SAM) exhibits remarkable zero-shot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when applied to SAM, owing to its specialized model components and promptable workflow: (i) The mask decoder's attention exhibits extreme activation outliers, and we find that aggressive clipping (even 100x), without smoothing or isolation, is effective in suppressing outliers while maintaining performance. Unfortunately, traditional distribution-based metrics (e.g., MSE) fail to provide such large-scale clipping. (ii) Existing quantization reconstruction methods neglect semantic interactivity of SAM, leading to misalignment between image feature and prompt intention. To address the above issues, we propose SAQ-SAM in this paper, which boosts PTQ for SAM from the perspective of semantic alignment. Specifically, we propose Perceptual-Consistency Clipping, which exploits attention focus overlap to promote aggressive clipping while preserving semantic capabilities. Furthermore, we propose Prompt-Aware Reconstruction, which incorporates image-prompt interactions by leveraging cross-attention in mask decoder, thus facilitating alignment in both distribution and semantic. Moreover, to ensure the interaction efficiency, we design a layer-skipping strategy for image tokens in encoder. Extensive experiments are conducted on various SAM sizes and tasks, including instance segmentation, oriented object detection, and semantic segmentation, and the results show that our method consistently exhibits advantages. For example, when quantizing SAM-B to 4-bit, SAQ-SAM achieves 11.7% higher mAP than the baseline in instance segmentation task.

IROS Conference 2025 Conference Paper

SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

  • Xianlei Long
  • Xiaxin Zhu
  • Fangming Guo
  • Wanyi Zhang
  • Qingyi Gu
  • Chao Chen 0004
  • Fuqiang Gu

Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a Spike-driven Lightweight Transformer-based Network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model’s parameters. Then, to enhance the long-range contextual feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9. 06% and 9. 39% mIoU, respectively, with extremely 4. 58× lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0.

ICRA Conference 2024 Conference Paper

A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching

  • Xianlei Long
  • Hui Zhao
  • Chao Chen 0004
  • Fuqiang Gu
  • Qingyi Gu

In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficient object searching, and accurate localization. To address these challenges, this paper presents a hybrid system that incorporates a wide-angle camera, a high-speed search camera, and a galvano-mirror. In this system, the wide-angle camera offers panoramic images as prior information, which helps the search camera capture detailed images of the targeted objects. This integrated approach enhances the overall efficiency and effectiveness of wide-area visual detection systems. Specifically, in this study, we introduce a wide-angle camera-based method to generate a panoramic probability map (PPM) for estimating high-probability regions of target object presence. Then, we propose a probability searching module that uses the PPM-generated prior information to dynamically adjust the sampling range and refine target coordinates based on uncertainty variance computed by the object detector. Finally, the integration of PPM and the probability searching module yields an efficient hybrid vision system capable of achieving 120 fps multi-object search and detection. Extensive experiments are conducted to verify the system’s effectiveness and robustness.

IROS Conference 2022 Conference Paper

A Flexible Calibration Algorithm for High-speed Bionic Vision System based on Galvanometer

  • Qing Li 0046
  • Mengjuan Chen
  • Qingyi Gu
  • Idaku Ishii

Traditional gimbal-based bionic eye systems usually use a multi-degree-of-freedom mechanical platform to move the camera freely, which makes the structure complex and bulky. The galvanometer-based reflective bionic eye system uses a galvanometer to replace the traditional mechanical rotation structure, which separates the camera from the gimbal system, greatly simplifying the structure. However, there are currently few methods for calibrating such systems, mostly for object detection and tracking. In this paper, a flexible method for high-precision calibration of a galvanometer-based reflective bionic eye system is proposed. In this method, a planar target is used for the calibration of the bionic eye system. The effectiveness and accuracy of the method are evaluated by the reprojection error of the control voltage and the spatial localization of the binocular system. Experiments show that the error of the control voltage after calibration is less than 0. 2%. At an indoor distance of about 7 m, the RMSE of spatial visual localization is less than 0. 3 cm.

ICRA Conference 2020 Conference Paper

Natural Scene Facial Expression Recognition with Dimension Reduction Network

  • Shenhua Hu
  • Yiming Hu
  • Jianquan Li
  • Xianlei Long
  • Mengjuan Chen
  • Qingyi Gu

As an external manifestation of human emotions, expression recognition plays an important role in human-computer interaction. Although existing expression recognition methods performs perfectly on constrained frontal faces, there are still many challenges in expression recognition in natural scenes due to different unrestricted conditions. Expression classification belongs to a pattern recognition problem where intra-class distance is greater than the inter-class distance, which leads to severe over-fitting when using neural networks for expression recognition. This paper proposes a novel net-work structure called Dimension Reduction Network which can effectively reduce generalization error. By adding a data dimension reduction module before the general classification network, a lot of redundant information is filtered, and only useful information is left. This can reduce the interference by irrelevant information when performing classification tasks and reduce generalization error. The proposed method does not require any modification to the classification network, only a small dimension reduction module needs to be added in front of the classification network. However, it can effectively reduce generalization error. We designed big and tiny versions of Dimension Reduction Network, both exceeds our baseline on AffectNet data set. The big version of our proposed method surpassed the state-of-the-art methods by more than 1. 2% on AffectNet data set. Our code will open source 3 when the paper is accepted.

IROS Conference 2017 Conference Paper

12, 000-fps Multi-object detection using HOG descriptor and SVM classifier

  • Jianquan Li
  • Yingjie Yin
  • Xilong Liu
  • De Xu
  • Qingyi Gu

This paper describes a high-frame-rate (HFR) vision system that can detect multiple objects in an image of 512 × 512 pixels at 12, 000 frames per seconds (fps). An optimized algorithm is proposed based on conventional Histograms of Oriented Gradient (HOG) descriptor and Support Vector Machine (SVM) classifier algorithms for hardware implementation. By implementing the proposed algorithm on a field-programmable gate array (FPGA) of a high-speed vision platform, multi-object in an image can be detected at 12, 000 fps under complex background. In hardware implementation, 64 pixels were processed in parallel with 80 MHz camera clock. Source image and detection results can be transferred to personal computer (PC) in real-time for recording or post-processing. Our developed HFR multi-object detection system was verified by performing several evaluations.

ICRA Conference 2016 Conference Paper

Control scheme of nongrasping manipulation based on virtual connecting constraint

  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Qingyi Gu
  • Idaku Ishii

The research field of nongrasping manipulation is a maturing area in robotic motion control. However, the common principles of motion planning for nongrasping manipulation systems have not yet been established. This paper proposes the concept of virtual connecting manipulation as a generalized motion planning framework for nongrasping manipulation systems. As a preliminary result, we had previously realized a flower-stick juggling task called “propeller motion” using an actual experimental system. In this paper, we apply the virtual connecting manipulation concept to a flower-stick juggling task and analyze the generated motion from the view point of analytical methodology. We conduct a stability analysis of the generated cyclic motion of the flower stick by using a Poincaré map, and the analytical results show that the generated cyclic motion is asymptotically stable.

ICRA Conference 2015 Conference Paper

A scheme for manipulating a passive object using an active plate

  • Tadayoshi Aoyama
  • Yuji Harada
  • Qingyi Gu
  • Takeshi Takaki
  • Idaku Ishii

We propose a novel scheme for manipulating a passive object using an active plate. The objective of this study is to control an object's orientation with respect to the gravitational force direction by using an active plate for realizing hitherto unrealized object motion. In this context, motions of the object and active plate are designed to be cyclic. A state vector composed of the object's angle and angular velocity is defined, and the cyclic motion is expressed as a nonlinear discrete system. Fixed points of the state vector are searched for in the designed cyclic motion. A stability analysis around the fixed points is conducted using a Poincaré map. As a result, the fixed points are shown to be asymptotically stable. Finally, experimental results are used to verify that the object's angle can be manipulated with the designed cyclic motion using the plate.

IROS Conference 2015 Conference Paper

Realization of flower stick rotation using robotic arm

  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Takumi Miura
  • Qingyi Gu
  • Idaku Ishii

Flower stick juggling is a dexterous task done by skillful jugglers. We aim to realize dexterous tasks done by humans using robotic systems. This work focuses on flower stick juggling and proposes a feedback control strategy for a flower stick juggling task called “propeller” as one of the robotic dexterous manipulations. The propeller motion is modeled by considering combined flower stick and a robotic manipulator. We developed a control strategy that allows stable cyclic rotation of the flower stick in the air. The control parametars in the control strategy are searched through numerical simulation. Finally, the flower stick propeller motion is realized using an actual robotic system.

ICRA Conference 2014 Conference Paper

Rapid vision-based shape and motion analysis system for fast-flowing cells in a microchannel

  • Qingyi Gu
  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Idaku Ishii

This paper proposes a novel method for simultaneous cell shape and motion analysis in rapid microchannel flows based on a multi-object feature extraction algorithm with a frame-straddling high-speed vision platform. This system can synchronize two camera inputs that share the same view with only a very small sub-microsecond time delay. Real-time video processing is performed using the hardware logic by extracting the moment features of multiple cells at 2000 fps or more, which are obtained from the two camera inputs, and their frame-straddling time can be adjusted from 0 to 0. 5 ms in 9. 9 ns steps. After setting the frame-straddling time within a certain range to avoid large image displacements between the two camera inputs, the frame-straddling high-speed vision platform can perform simultaneous shape and motion analysis of cells in fast microchannel flows of 1 m/s or greater. The results of real-time experiments conducted to analyze the deformabilities, velocities, and shapes of fast-flowing sea urchin egg cells in straight and L-type microchannels verified the efficacy of our vision-based cell analysis system.

IROS Conference 2014 Conference Paper

Real-time LOC-based morphological cell analysis system using high-speed vision

  • Qingyi Gu
  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Idaku Ishii
  • Ayumi Takemoto
  • Naoaki Sakamoto

In this paper, a high-speed vision-based morphological analysis system for fast-flowing cells in a microchannel implementing a multi-object feature extraction algorithm on a high-speed vision platform is proposed. Real-time video processing is performed in hardware logic by extracting the moment features and bounding boxes of multiple cells in 512×256-pixel images at 2000 fps. The extracted cell regions are pushed into a first-in-first-out (FIFO) buffer for real-time image-based morphological analysis after being shrunk proportionally to a certain size. By extracting the bounding boxes of the cell regions using hardware logic and shrinking the cell region to a certain size to reduce processing time, our high-speed vision system can perform fast morphological analysis of cells at 2 ms/cell in fast microchannel flows. The results of real-time experiments conducted to analyze the size, eccentricity, and transparency of fertilized sea urchin eggs fast flowing in microchannels verify the efficacy of our vision-based cell analysis system.

IROS Conference 2013 Conference Paper

A fast multi-camera tracking system with heterogeneous lenses

  • Xiaorong Zhao
  • Qingyi Gu
  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Idaku Ishii

We have developed a fast target tracking system that utilizes four cameras with lenses of different focal lengths to track an object without blurring images, even when the object moves in the depth direction away from the cameras in three-dimensional (3-D) space. This system can maintain a well-focused camera view by switching among the four input images, instead of using lens motor control. The multi-camera system was mounted on a two-axis mechanical active vision platform. The active vision control and camera-view switching are executed by processing color 512 × 512 images from the four camera inputs at 500 fps in real time on a high-speed vision platform. The performance of our system was verified by its tracking results for objects moving rapidly in 3-D space.

IROS Conference 2013 Conference Paper

Fast 3-D shape measurement using blink-dot projection

  • Jun Chen
  • Qingyi Gu
  • Hao Gao
  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Idaku Ishii

We propose a novel dot-pattern-projection three-dimensional (3-D) shape measurement method that can measure 3-D displacements of blink dots projected onto a measured object accurately even when it moves rapidly or is observed from a camera as moving rapidly. In our method, blinking dot patterns, in which each dot changes its size at different timings corresponding to its identification (ID) number, are projected from a projector at a high frame rate. 3-D shapes can be obtained without any miscorrespondence of the projected dots between frames by simultaneous tracking and identification of multiple dots projected onto a measured 3-D object in a camera view. Our method is implemented on a field-programmable gate array (FPGA)-based high-frame-rate (HFR) vision platform that can track and recognize as much as 15×15 blink-dot pattern in a 512×512 image in real time at 1000 fps, synchronized with an HFR projector. We demonstrate the performance of our system by showing real-time 3-D measurement results when our system is mounted on a parallel link manipulator as a sensing head.

IROS Conference 2013 Conference Paper

Real-time feature-based video mosaicing at 500 fps

  • Ken-ichi Okumura
  • Sushil Raut
  • Qingyi Gu
  • Tadayoshi Aoyama
  • Takeshi Takaki
  • Idaku Ishii

We conducted high-frame-rate (HFR) video mo-saicing for real-time synthesis of a panoramic image by implementing an improved feature-based video mosaicing algorithm on a field-programmable gate array (FPGA)-based high-speed vision platform. In the implementation of the mosaicing algorithm, feature point extraction was accelerated by implementing a parallel processing circuit module for Harris corner detection in the FPGA on the high-speed vision platform. Feature point correspondence matching can be executed for hundreds of selected feature points in the current frame by searching those in the previous frame in their neighbor ranges, assuming that frame-to-frame image displacement becomes considerably smaller in HFR vision. The system we developed can mosaic 512×512 images at 500 fps as a single synthesized image in real time by stitching the images based on their estimated frame-to-frame changes in displacement and orientation. The results of an experiment conducted, in which an outdoor scene was captured using a hand-held camera-head that was quickly moved by hand, verify the performance of our system.

ICRA Conference 2010 Conference Paper

2000 fps real-time vision system with high-frame-rate video recording

  • Idaku Ishii
  • Tetsuro Tatebe
  • Qingyi Gu
  • Yuta Moriue
  • Takeshi Takaki
  • Kenji Tajima

This paper introduces a high-speed vision system called IDP Express, which can execute real-time image processing and high frame rate video recording simultaneously. In IDP Express, a dedicated FPGA (Field Programmable Gate Array) board processes 512 × 512 pixel images from two camera heads by implementing image processing algorithms as hardware logic; the input images and processed results are transferred to standard PC memory at a rate of 2000 fps or more. Owing to the simultaneous high-frame-rate video processing and recording, IDP Express can be used as an intelligent video logger for long-term high-speed phenomenon analysis even when the measured objects move quickly in a wide area. We applied IDP Express to a mechanical target tracking system to record a high-frame-rate video at high resolution for a crucial moment, which is magnified by tracking the measured objects with visual feedback control. Several experiments on moving objects that undergo sudden shape deformation were performed. The results of the experiments involving the explosion of a rotating balloon and the crash of falling custard pudding have been provided to verify the effectiveness of IDP Express.

v2026.09.13