Arrow Research search

Author name cluster

Jie Guo

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
1 author row

Possible papers

14

EAAI Journal 2026 Journal Article

Robust adaptive-neighbor-induced optimization to nonnegative matrix factorization with regularized strategies in the framework of semi-supervised learning

  • Jie Guo
  • Ting Li
  • Jialu Liu
  • Zhong Wan
  • Fang Zhang

As an efficient tool in artificial intelligence, nonnegative matrix factorization (NMF) is widely used for data clustering and feature discovery, yet existing models are often sensitive to noise and outliers and lack effective mechanisms to exploit limited supervisory information in semi-supervised settings. To address these limitations, this paper proposes a novel robust NMF optimization model within a semi-supervised learning framework, introducing a reconstruction-error-based loss function to bolster robustness and an adaptive neighbor induced strategy to propagate pairwise constraints via dynamic similarity graphs, along with a dataset-adaptive mechanism to refine sample similarity weighting. For this model, we develop an efficient optimization algorithm with convergence guarantees. Extensive experiments on twelve public image and text datasets demonstrate that the proposed method outperforms state-of-the-art alternatives across multiple clustering metrics, confirming its effectiveness in noisy environments and demonstrating its capacity to leverage supervisory information for improved clustering performance.

NeurIPS Conference 2025 Conference Paper

Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models

  • Xiaoyu Zhan
  • Wenxuan Huang
  • Hao Sun
  • Xinyu Fu
  • Changfeng Ma
  • Shaosheng Cao
  • Bohan Jia
  • Shaohui Lin

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.

JBHI Journal 2025 Journal Article

Development of a tongue image-based machine learning tool for the diagnosis of colorectal cancer: a prospective multicentre clinical cohort study

  • Xiaohe Sun
  • Letian Huang
  • Libo Qu
  • Cheng Chen
  • Xing Zeng
  • Zuojian Zhou
  • Hongyan Li
  • Jin Sun

Colorectal cancer (CRC) remains a persistent major global health burden, with traditional diagnostic methods like colonoscopy suffering from suboptimal patient compliance rates. This study develops an intelligent diagnostic model based on tongue images to assist in CRC diagnosis, leveraging the integrative potential of traditional tongue diagnosis and modern machine learning. Between June 2023 and July 2024, we collected and processed 1, 389 tongue images from CRC patients and 1, 543 from non-colorectal cancer (NCRC) participants. Our methodology combines innovative image segmentation using the Segment Anything Model (SAM) with Grounding DINO, extracts both hand-crafted features (color, texture, shape) and deep learning features via Swin-Transformer, and employs feature fusion and selection techniques. The diagnostic model achieves an accuracy of 87. 93% (F1-score: 0. 9072) in internal validation. In an independent external cohort of 119 CRC patients and 221 NCRC participants, it demonstrates 85. 18% precision (recall: 85%, F1-score: 0. 8507). This noninvasive, cost-effective approach demonstrates significant potential as a complementary screening tool for CRC, particularly in regions with limited access to conventional diagnostic resources.

EAAI Journal 2025 Journal Article

Hypergraph induced semi-supervised orthogonal nonnegative matrix factorization with label and constraint propagation

  • Jie Guo
  • Ting Li
  • Jialu Liu
  • Zhong Wan

As a popular technology of artificial intelligence, nonnegative matrix factorization (NMF) aims at clustering and finding the differentially expressed features of each cluster. However, for complex high-dimensional sample data, it is still a challenge to design more appropriate NMF optimization models and develop more efficient algorithms to solve this model in view of enhanced theoretical properties and numerical performance. In this paper, a novel NMF optimization model with regularization is proposed such that the NMF is performed by a semi-supervised approach, as well as incorporating the strategies of hypergraph induced label propagation and constraint propagation. Specifically, different from existing NMF methods, the hypergraph structure underlying the data, together with the simple graph information, is employed to guide the pairwise constraint propagation in our built model. In recognition of sample similarity, a dataset-adaptive strategy is proposed to update the weight matrix of the graphs. By adding dual orthogonality on the factor matrices in the objective function, interpretability and feature independence of the built model are enhanced. Then, an algorithm is developed to efficiently solve this complicated model. Theoretically, it is proved that the developed algorithms are well defined and convergent. Numerically, extensive tests on the proposed model and algorithm are performed, which validate that they outperform the state-of-the-art ones in terms of different metrics of evaluating clustering performance when they are applied into solution of the problems from eight public datasets.

AAAI Conference 2025 Conference Paper

IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target Detection

  • Mingjin Zhang
  • Xiaolong Li
  • Fei Gao
  • Jie Guo

Infrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learning methods struggle to suppress noise and background interference while preserving fine details, leading to missed detections and false alarms. To address these issues, we propose IRMamba, an encoder-decoder architecture featuring Pixel Difference Mamba (PDMamba) and a Layer Restoration Module (LRM). Specifically, PDMamba integrates the intensity and directional information of pixel differences between scanning positions and their central neighborhoods into the state equation of the state space model (SSM). This enhances target detail representation and suppresses background interference by capturing local 2D dependencies from a global perspective. In addition, LRM incorporates the double-depth image prior into the iterative convergence algorithm, and utilizes the inter-layer interrelationships to gradually reverse the separation of the target layer, achieving noise suppression and refined reconstruction of the image mask. Experiments conducted on multiple public datasets, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1K, demonstrate the significant advantages of IRMamba over SOTA methods.

AAAI Conference 2025 Conference Paper

MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection

  • Mingjin Zhang
  • Yuanjun Ouyang
  • Fei Gao
  • Jie Guo
  • Qiming Zhang
  • Jing Zhang

In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.

IJCAI Conference 2025 Conference Paper

Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging

  • Mingjin Zhang
  • Longyi Li
  • Fei Gao
  • Qiming Zhang
  • Jie Guo

The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method.

AAAI Conference 2025 Conference Paper

Real-Time Neural Denoising with Render-Aware Knowledge Distillation

  • Mengxun Kong
  • Jie Guo
  • Chen Wang
  • Ye Yuan
  • Yanwen Guo

Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we present a render-aware knowledge distillation (RAKD) framework, specifically designed for Monte Carlo denoising. We meticulously delineate the Knowledge Distillation (KD) process within RAKD, emphasizing three pivotal techniques: the strategic incorporation of an auxiliary unlabeled dataset, the integration of adversarial learning through generative adversarial network (GAN), and the application of parameter transfer for robust model initialization. These approaches are harmoniously combined to distill knowledge effectively, enabling our student model to adeptly strike a balance between preserving high-frequency details and reducing low-frequency noise. Finally, our results demonstrate that RAKD achieves state-of-the-art quality while upholding real-time performance, successfully tackling the computational constraints faced by resource-limited devices.

EAAI Journal 2024 Journal Article

A dual-level consensus model for large-scale group decision-making driven by trust relationships in social networks

  • Jie Guo
  • Zilong Wang
  • Zhiwen Zhang

In large-scale group decision-making, obtaining reasonable and reliable decision results is the research focus. To this end, a dual-level consensus model for large-scale group decision-making is constructed. In the model, we first presented a method for measuring the two-dimensional similarity between decision-makers. The method considered the opinion similarity at evaluation scales and ranking results level, which provided a more comprehensive measure of the opinion differences between decision-makers. A hierarchical clustering algorithm based on the two-dimensional similarity-trust matrix is proposed. By considering the two-dimensional similarity and trust relationships among decision-makers, the method contributed to the formation of subgroups with excellent cooperative and communicative atmospheres. Furthermore, we introduced a dual-level consensus measures method to comprehensively and accurately assess the group consensus level. To increase the efficiency of decision-making and the satisfaction of decision-makers, we proposed a personalized feedback adjustment mechanism for decision-makers. Additionally, for the subgroups that have not reached consensus, a social networks DeGroot model is constructed to adjust the opinion of subgroups. Finally, a case study of low-carbon supplier selection verified the feasibility and applicability of the model. The effectiveness and superiority of the model are demonstrated by comparative and simulation analysis. The analysis results show that the model is conducive to promoting cooperation and coordination among decision-makers and obtaining highly acceptable and quality decision results.

AAAI Conference 2024 Conference Paper

IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel Pruning

  • Mingjin Zhang
  • Handi Yang
  • Jie Guo
  • Yunsong Li
  • Xinbo Gao
  • Jing Zhang

Infrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and computation inefficiencies. In this pioneering study, we introduce the concept of utilizing network pruning to enhance the efficiency of IRSTD. Due to the challenge posed by low signal-to-noise ratios and the absence of detailed semantic information in infrared images, directly applying existing pruning techniques yields suboptimal performance. To address this, we propose a novel wavelet structure-regularized soft channel pruning method, giving rise to the efficient IRPruneDet model. Our approach involves representing the weight matrix in the wavelet domain and formulating a wavelet channel pruning strategy. We incorporate wavelet regularization to induce structural sparsity without incurring extra memory usage. Moreover, we design a soft channel reconstruction method that preserves important target information against premature pruning, thereby ensuring an optimal sparse structure while maintaining overall sparsity. Through extensive experiments on two widely-used benchmarks, our IRPruneDet method surpasses established techniques in both model complexity and accuracy. Specifically, when employing U-net as the baseline network, IRPruneDet achieves a 64.13% reduction in parameters and a 51.19% decrease in FLOPS, while improving IoU from 73.31% to 75.12% and nIoU from 70.92% to 74.30%. The code is available at https://github.com/hd0013/IRPruneDet.

EAAI Journal 2024 Journal Article

Temporal Contrastive and Spatial Enhancement Coarse Grained Network for Weakly Supervised Group Activity Recognition

  • Jie Guo
  • Yongxin Ge

Group activity recognition (GAR) is an increasingly popular topic in the field of computer vision. Numerous researchers have proposed a range of methods to achieve outstanding recognition performance. However, these methods invariably require fine-grained personal feature extraction and a large network architecture to aggregate individual features or reason person relationships. To mitigate the need for a bloated portfolio of annotations and high training costs, weak supervision has emerged as a promising approach. Under the weak supervision paradigm, only coarse-grained labels are used during network training. Nevertheless, this method poses two key challenges. Firstly, it is limited in its ability to model temporal relationships among individual persons, and secondly, it tends to focus on less relevant information, thereby leading to suboptimal network parameter optimization. Both of these challenges result in erroneous temporal information judgment and training inefficiencies. To address these challenges within the weak supervision paradigm, we propose a novel Temporal Contrastive and Spatial Enhancement Coarse-Grained Network (TCSE-CGN) to solve the GAR problem. TCSE-CGN comprises two simple yet effective streams, namely the Spatial Enhancement Stream and the Temporal Contrastive Stream. After extracting features using only several RGB frames, half of the extracted feature is sent to the Spatial Enhancement Stream for enhancement using an attention mechanism. Consequently, the network automatically learns more representative information. The remaining feature is sent to the Temporal Contrastive Stream, which uses contrastive learning to model temporal relationships among all RGB frame-level features. Specifically, the network is guided to learn the hidden semantic temporal information about inter-frame sequences. Network parameters are optimized using a combination of universe cross-entropy loss and a novel temporal contrastive loss. Comprehensive experiments are conducted on two widely used datasets, namely the Volleyball dataset and the Collective dataset, to demonstrate the effectiveness of TCSE-CGN. Results show that TCSE-CGN performs competitively with other works that require more supervision and a larger architecture.

JBHI Journal 2022 Journal Article

Learning Binary Semantic Embedding for Large-Scale Breast Histology Image Analysis

  • Xingbo Liu
  • Xiao Kang
  • Xiushan Nie
  • Jie Guo
  • Shaohua Wang
  • Yilong Yin

With the progress of clinical imaging innovation and machine learning, the computer-assisted diagnosis of breast histology images has attracted broad attention. Nonetheless, the use of computer-assisted diagnoses has been blocked due to the incomprehensibility of customary classification models. In view of this question, we propose a novel method for L earning B inary S emantic E mbedding (LBSE). In this study, bit balance and uncorrelation constraints, double supervision, discrete optimization and asymmetric pairwise similarity are seamlessly integrated for learning binary semantic-preserving embedding. Moreover, a fusion-based strategy is carefully designed to handle the intractable problem of parameter setting, saving huge amounts of time for boundary tuning. Based on the above-mentioned proficient and effective embedding, classification and retrieval are simultaneously performed to give interpretable image-based deduction and model helped conclusions for breast histology images. Extensive experiments are conducted on three benchmark datasets to approve the predominance of LBSE in different situations.

IJCAI Conference 2022 Conference Paper

SAR-to-Optical Image Translation via Neural Partial Differential Equations

  • Mingjin Zhang
  • Chengyu He
  • Jing Zhang
  • Yuxiang Yang
  • Xiaoqi Peng
  • Jie Guo

Synthetic Aperture Radar (SAR) becomes prevailing in remote sensing while SAR images are challenging to interpret by human visual perception due to the active imaging mechanism and speckle noise. Recent researches on SAR-to-optical image translation provide a promising solution and have attracted increasing attentions, though still suffering from low optical image quality with geometric distortion due to the large domain gap. In this paper, we mitigate this issue from a novel perspective, i. e. , neural partial differential equations (PDE). First, based on the efficient numerical scheme for solving PDE, i. e. , Taylor Central Difference (TCD), we devise a basic TCD residual block to build the backbone network, which promotes the extraction of useful information in SAR images by aggregating and enhancing features from different levels. Furthermore, inspired by the Perona-Malik Diffusion (PMD), we devise a PMD neural module to implement feature diffusion through layers, aiming at removing the noises in smooth regions while preserving the geometric structures. Assembling them together, we propose a novel SAR-to-Optical image translation network named S2O-NPDE, which delivers optical images with finer structures and less noise while enjoying an explainability advantage from explicit mathematical derivation. Experiments on the popular SEN1-2 dataset show that our model outperforms state-of-the-art methods in terms of both objective metrics and visual quality.

NeurIPS Conference 2022 Conference Paper

Unsupervised Point Cloud Completion and Segmentation by Generative Adversarial Autoencoding Network

  • Changfeng Ma
  • Yang Yang
  • Jie Guo
  • Fei Pan
  • Chongjun Wang
  • Yanwen Guo

Most existing point cloud completion methods assume the input partial point cloud is clean, which is not practical in practice, and are Most existing point cloud completion methods assume the input partial point cloud is clean, which is not the case in practice, and are generally based on supervised learning. In this paper, we present an unsupervised generative adversarial autoencoding network, named UGAAN, which completes the partial point cloud contaminated by surroundings from real scenes and cutouts the object simultaneously, only using artificial CAD models as assistance. The generator of UGAAN learns to predict the complete point clouds on real data from both the discriminator and the autoencoding process of artificial data. The latent codes from generator are also fed to discriminator which makes encoder only extract object features rather than noises. We also devise a refiner for generating better complete cloud with a segmentation module to separate the object from background. We train our UGAAN with one real scene dataset and evaluate it with the other two. Extensive experiments and visualization demonstrate our superiority, generalization and robustness. Comparisons against the previous method show that our method achieves the state-of-the-art performance on unsupervised point cloud completion and segmentation on real data.

v2026.09.13