Arrow Research search

Author name cluster

Jia Guo

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
2 author rows

Possible papers

13

IROS Conference 2025 Conference Paper

Action Recognition for Underwater Gesture Communication in Human Diver and Robot Teaming

  • Zi-Hao Zhang
  • E. Baker Herrin
  • Jia Guo
  • Aditya Penumarti
  • Zilong He
  • Andres Pulido
  • Jane Shin

This paper presents a Spatio-Temporal Transformer-based algorithm for underwater diver hand gesture recognition, forming a key component of diver-robot teaming. Existing computer vision-based approaches primarily rely on frame-wise gesture detection, which often fails to capture motion continuity and suffers under degraded underwater visibility. The presented method integrates temporal modeling to (i) improve recognition accuracy by capturing spatio-temporal patterns in hand motion, and (ii) increase robustness in challenging underwater environments by leveraging sequential image data, thereby mitigating the impact of intermittent misclassifications. The system is evaluated using real-world underwater footage, demonstrating high recognition accuracy and robustness to lighting fluctuations and partial occlusions. The results highlight the effectiveness and practicality of the presented method for real-world diver-robot collaboration, establishing a foundation for more reliable and intelligent underwater human-robot collaboration.

TMLR Journal 2024 Journal Article

Large Language Models Synergize with Automated Machine Learning

  • Jinglue Xu
  • Jialong Li
  • Zhen Liu
  • NAV Suryanarayanan
  • Guoyuan Zhou
  • Jia Guo
  • Hitoshi Iba
  • Kenji Tei

Recently, program synthesis driven by large language models (LLMs) has become increasingly popular. However, program synthesis for machine learning (ML) tasks still poses significant challenges. This paper explores a novel form of program synthesis, targeting ML programs, by combining LLMs and automated machine learning (autoML). Specifically, our goal is to fully automate the generation and optimization of the code of the entire ML workflow, from data preparation to modeling and post-processing, utilizing only textual descriptions of the ML tasks. To manage the length and diversity of ML programs, we propose to break each ML program into smaller, manageable parts. Each part is generated separately by the LLM, with careful consideration of their compatibilities. To ensure compatibilities, we design a testing technique for ML programs. Unlike traditional program synthesis, which typically relies on binary evaluations (i.e., correct or incorrect), evaluating ML programs necessitates more than just binary judgments. Our approach automates the numerical evaluation and optimization of these programs, selecting the best candidates through autoML techniques. In experiments across various ML tasks, our method outperforms existing methods in 10 out of 12 tasks for generating ML programs. In addition, autoML significantly improves the performance of the generated ML programs. In experiments, given the textual task description, our method, Text-to-ML, generates the complete and optimized ML program in a fully autonomous process. The implementation of our method is available at https://github.com/JLX0/llm-automl.

IJCAI Conference 2023 Conference Paper

RaMLP: Vision MLP via Region-aware Mixing

  • Shenqi Lai
  • Xi Du
  • Jia Guo
  • Kaipeng Zhang

Recently, MLP-based architectures achieved impressive results in image classification against CNNs and ViTs. However, there is an obvious limitation in that their parameters are related to image sizes, allowing them to process only fixed image sizes. Therefore, they cannot directly adapt dense prediction tasks (e. g. , object detection and semantic segmentation) where images are of various sizes. Recent methods tried to address it but brought two new problems, long-range dependencies or important visual cues are ignored. This paper presents a new MLP-based architecture, Region-aware MLP (RaMLP), to satisfy various vision tasks and address the above three problems. In particular, we propose a well-designed module, Region-aware Mixing (RaM). RaM captures important local information and further aggregates these important visual clues. Based on RaM, RaMLP achieves a global receptive field even in one block. It is worth noting that, unlike most existing MLP-based architectures that adopt the same spatial weights to all samples, RaM is region-aware and adaptively determines weights to extract region-level features better. Impressively, our RaMLP outperforms state-of-the-art ViTs, CNNs, and MLPs on both ImageNet-1K image classification and downstream dense prediction tasks, including MS-COCO object detection, MS-COCO instance segmentation, and ADE20K semantic segmentation. In particular, RaMLP outperforms MLPs by a large margin (around 1. 5% Apb or 1. 0% mIoU) on dense prediction tasks. The training code could be found at https: //github. com/xiaolai-sqlai/RaMLP.

NeurIPS Conference 2023 Conference Paper

ReContrast: Domain-Specific Anomaly Detection via Contrastive Reconstruction

  • Jia Guo
  • Shuai Lu
  • Lize Jia
  • Weihang Zhang
  • Huiqi Li

Most advanced unsupervised anomaly detection (UAD) methods rely on modeling feature representations of frozen encoder networks pre-trained on large-scale datasets, e. g. ImageNet. However, the features extracted from the encoders that are borrowed from natural image domains coincide little with the features required in the target UAD domain, such as industrial inspection and medical imaging. In this paper, we propose a novel epistemic UAD method, namely ReContrast, which optimizes the entire network to reduce biases towards the pre-trained image domain and orients the network in the target domain. We start with a feature reconstruction approach that detects anomalies from errors. Essentially, the elements of contrastive learning are elegantly embedded in feature reconstruction to prevent the network from training instability, pattern collapse, and identical shortcut, while simultaneously optimizing both the encoder and decoder on the target domain. To demonstrate our transfer ability on various image domains, we conduct extensive experiments across two popular industrial defect detection benchmarks and three medical image UAD tasks, which shows our superiority over current state-of-the-art methods.

NeurIPS Conference 2022 Conference Paper

WT-MVSNet: Window-based Transformers for Multi-view Stereo

  • Jinli Liao
  • Yikang Ding
  • Yoli Shavit
  • Dihe Huang
  • Shihao Ren
  • Jia Guo
  • Wensen Feng
  • Kai Zhang

Recently, Transformers have been shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global feature aggregation in multi-view stereo. We introduce a Window-based Epipolar Transformer (WET) which reduces matching redundancy by using epipolar constraints. Since point-to-line matching is sensitive to erroneous camera pose and calibration, we match windows near the epipolar lines. A second Shifted WT is employed for aggregating global information within cost volume. We present a novel Cost Transformer (CT) to replace 3D convolutions for cost volume regularization. In order to better constrain the estimated depth maps from multiple views, we further design a novel geometric consistency loss (Geo Loss) which punishes unreliable areas where multi-view consistency is not satisfied. Our WT multi-view stereo method (WT-MVSNet) achieves state-of-the-art performance across multiple datasets and ranks $1^{st}$ on Tanks and Temples benchmark. Code will be available upon acceptance.

YNIMG Journal 2021 Journal Article

Cerebrovascular reactivity measurements using simultaneous 15O-water PET and ASL MRI: Impacts of arterial transit time, labeling efficiency, and hematocrit

  • Moss Y Zhao
  • Audrey P Fan
  • David Yen-Ting Chen
  • Magdalena J. Sokolska
  • Jia Guo
  • Yosuke Ishii
  • David D Shin
  • Mohammad Mehdi Khalighi

O-water PET and ASL MRI data from 19 healthy subjects. CVR and CBF measured by the ASL techniques were compared using PET as the reference technique. The impacts of blood T1 and labeling efficiency on ASL were assessed using individual measurements of hematocrit and flow velocity data of the carotid and vertebral arteries measured using phase-contrast MRI. We found that multi-PLD PCASL is the ASL technique most consistent with PET for CVR quantification (group mean CVR of the whole brain = 42±19% and 40±18% respectively). Single-PLD ASL underestimated the CVR of the whole brain significantly by 15±10% compared with PET (p<0.01, paired t-test). Changes in ATT pre- and post-acetazolamide was the principal factor affecting ASL-based CVR quantification. Variations in labeling efficiency and blood T1 had negligible effects.

TCS Journal 2021 Journal Article

Vertex-pancyclicity of the (n,k)-bubble-sort networks

  • Xin Wang
  • Chaoqun Ma
  • Jia Guo

The cycle embedding is an important problem of networks, which can determine the fault tolerance of the networks. A network can be viewed as a graph. Let m be an integer with m ≥ 4, G be a graph and w ∈ V ( G ) be an arbitrary vertex. The graph G is vertex-pancyclic if G has a cycle C l of length l with w ∈ V ( C l ) for every l ∈ { 3, 4, ⋯, | V ( G ) | } and G is m-weak-vertex-pancyclic if G has a cycle C l of length l with w ∈ V ( C l ) for every l ∈ { m, m + 1, ⋯, | V ( G ) | }. Let G ′ be a bipartite graph and w ′ ∈ V ( G ′ ) be an arbitrary vertex. The graph G ′ is vertex-bipancyclic if G ′ has a cycle C h of length h with w ′ ∈ V ( C h ) for any even integer h with 4 ≤ h ≤ | V ( G ′ ) |. In this paper, we study the cycle embedding in the ( n, k ) -bubble-sort network B n, k. We obtain that (1) B n, 1 is vertex-pancyclic for n ≥ 3. (2) B n, n − 1 is vertex-bipancyclic for n ≥ 4. (3) B 4, 2 and B 5, 2 are 6-weak-vertex-pancyclic and B 5, 3 is vertex-pancyclic. (4) B n, k is vertex-pancyclic for n ≥ 6 with 2 ≤ k ≤ n − 2 and every constructed cycle of B n, k contains a residual edge for n ≥ 4 with 2 ≤ k ≤ n − 2.

IS Journal 2020 Journal Article

General Learning Modeling for AUV Position Tracking

  • Jia Guo
  • Rui Jiang
  • Bo He
  • Tianhong Yan
  • Shuzhi Sam Ge

In this article, we propose a nonlinear model based on hidden-layer neural networks and local Gaussian process regression. The hidden-layer neural networks have short training time, whereas the local Gaussian process regression is more suitable for nonlinearity. According to the abovementioned advantages of hidden-layer neural network and local Gaussian process regression, we get the measurement learning model and global process learning model offline and local process learning model online. Then, the proposed learning-based models have been applied to extended Kalman filter-simultaneous localization and mapping (EKF-SLAM), Rao-Blackwellised particle filter formulation of simultaneous localisation and mapping (FastSLAM), and unscented Kalman filter--simultaneous localization and mapping (UKF-SLAM) for autonomous underwater vehicles position tracking with field experiments. Compared with the conventional process and measurement models, improved position tracking performance can be achieved without cumbersome system analysis nor identification, benefiting from the proposed model. Experiments in more than 3883-m run show that the root-mean-square error has been improved by 36. 95%, 37. 14%, and 27. 65% in EKF-SLAM, FastSLAM, and UKF-SLAM frameworks, respectively.

YNIMG Journal 2020 Journal Article

The potential for gas-free measurements of absolute oxygen metabolism during both baseline and activation states in the human brain

  • Eulanca Y. Liu
  • Jia Guo
  • Aaron B. Simon
  • Frank Haist
  • David J. Dubowitz
  • Richard B. Buxton

Quantitative functional magnetic resonance imaging methods make it possible to measure cerebral oxygen metabolism (CMRO2) in the human brain. Current methods require the subject to breathe special gas mixtures (hypercapnia and hyperoxia). We tested a noninvasive suite of methods to measure absolute CMRO2 in both baseline and dynamic activation states without the use of special gases: arterial spin labeling (ASL) to measure baseline and activation cerebral blood flow (CBF), with concurrent measurement of the blood oxygenation level dependent (BOLD) signal as a dynamic change in tissue R2*; VSEAN to estimate baseline O2 extraction fraction (OEF) from a measurement of venous blood R2, which in combination with the baseline CBF measurement yields an estimate of baseline CMRO2; and FLAIR-GESSE to measure tissue R2′ to estimate the scaling parameter needed for calculating the change in CMRO2 in response to a stimulus with the calibrated BOLD method. Here we describe results for a study sample of 17 subjects (8 female, mean age = 25. 3 years, range 21–31 years). The primary findings were that OEF values measured with the VSEAN method were in good agreement with previous PET findings, while estimates of the dynamic change in CMRO2 in response to a visual stimulus were in good agreement between the traditional hypercapnia calibration and calibration based on R2′. These results support the potential of gas-free methods for quantitative physiological measurements.

TCS Journal 2019 Journal Article

The g-good-neighbor conditional diagnosability of the crossed cubes under the PMC and MM* model

  • Jia Guo
  • Desai Li
  • Mei Lu

The significant increase in the number of processors of the multiprocessor system increases its vulnerability to component failures. Diagnosability is an important indicator in measuring the reliability of interconnection networks. The g-good-neighbor conditional faulty set is a faulty set that each fault-free vertex is adjacent to at least g fault-free vertices. The g-good-neighbor conditional diagnosability gives the maximum cardinality of g-good-neighbor conditional faulty set that the system is guaranteed to identify. This paper we establish the g-good-neighbor conditional diagnosability of the crossed cube C Q n under the PMC and MM* model.

YNICL Journal 2018 Journal Article

Temporal lobe epilepsy lateralization using retrospective cerebral blood volume MRI

  • Xinyang Feng
  • Marla J. Hamberger
  • Hannah C. Sigmon
  • Jia Guo
  • Scott A. Small
  • Frank A. Provenzano

Steady-state cerebral blood volume (CBV) is tightly coupled to regional cerebral metabolism, and CBV imaging is a variant of MRI that has proven useful in mapping brain dysfunction. CBV derived from exogenous contrast-enhanced MRI can generate sub-millimeter functional maps. Higher resolution helps to more accurately interrogate smaller cortical regions, such as functionally distinct regions of the hippocampus. Many MRIs have fortuitously adequate sequences required for CBV mapping. However, these scans vary substantially in acquisition parameters. Here, we determined whether previously acquired contrast-enhanced MRI scans ordered in patients with unilateral temporal lobe epilepsy can be used to generate hippocampal CBV. We used intrinsic reference regions to correct for intensity scaling on a research CBV dataset to identify white matter as a robust marker for scaling correction. Next, we tested the technique on a sample of unilateral focal epilepsy patients using clinical MRI scans. We find evidence suggestive of significant hypometabolism in the ipsilateral-hippocampus of unilateral TLE subjects. We also highlight the subiculum as a potential driver of this effect. This study introduces a technique that allows CBV maps to be generated retrospectively from clinical scans, potentially with broad application for mapping dysfunction throughout the brain.

TCS Journal 2017 Journal Article

Conditional diagnosability of the round matching composition networks

  • Jia Guo
  • Mei Lu

The conditional diagnosability of many well-known networks has been explored. In this paper, we analyze the conditional diagnosability of a family of networks, called the round matching composition networks, which are a class of networks composed of r ( r ≥ 4 ) components of the same order linked by r perfect matchings. Applying the result, we determine the conditional diagnosability of the k-ary n-cubes and the recursive circulant graphs under the PMC model and the MM model, respectively.

TCS Journal 2016 Journal Article

The extra connectivity of bubble-sort star graphs

  • Jia Guo
  • Mei Lu

Connectivity plays an important role in measuring the fault tolerance of a multiprocessor system in the case of vertices failures. Extra connectivity is an important indicator of a network's ability for diagnosis and fault tolerance. In this paper, we analyze the fault tolerance ability of the bubble-sort star network, denoted by B S n, a well-known interconnection network proposed for multiprocessor systems, and establish the g-extra connectivity for 1 ≤ g ≤ 3.

v2026.09.13