Arrow Research search

Author name cluster

Guohui Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

AAAI Conference 2026 Conference Paper

SGP4SR: Separated-Modality Guided User Preference Learning for Multimodal Sequential Recommendation

  • Changhong Li
  • Zhiqiang Guo
  • Guohui Li
  • Zhong Yang
  • Chuhang Hong

With the booming development of multimodal data (e.g., image, text) on internet platforms, multimodal sequential recommendation methods continue to emerge. Most existing methods incorporate item modal features as auxiliary information, typically concatenating them to learn unified user representations. However, these methods directly use modal features for representation learning, neglecting the impact of inherent modal noise. We argue that internal-modal noise and cross-modal noise hinder the acquisition of more accurate user representations. To address this problem, we propose SGP4SR - Separated-modality Guided user Preference learning for multimodal Sequential Recommendation. Globally, the user preference modeling is carried out from a separated-modality perspective to alleviate cross-modal noise. Locally, for each individual modality, we use item relationship graphs and user interest centers, aggregated with ID embeddings, to replace direct modal features, thereby mitigating internal-modal noise. Finally, user representations from both separated-modality and multimodal perspectives participate in prediction independently. In experiments conducted on four real-world datasets, our method outperforms state-of-the-art approaches, achieving an average performance improvement of up to 8.84% over the best baseline. The comprehensive experiments further validate the superior noise tolerance and robustness of our method.

AAAI Conference 2026 Conference Paper

Unsupervised Combinatorial Probabilistic Reasoning: Probabilistic Coin Change Problem

  • Zhongdi Qu
  • Yingheng Wang
  • Utku Umur Acikalin
  • Aaron M. Ferber
  • Goncalo J. Gouveia
  • Brandon Bills
  • Guohui Li
  • Joshua Kline

We introduce the Probabilistic Coin Change Problem (PCCP), a novel variant of the classical Combination Coin Change Problem (CCCP), motivated by a real-world scientific inverse task. The goal of CCCP is to enumerate all unordered combinations of coin denominations that sum to a given target. In PCCP, each coin type’s value follows a discrete probability distribution, and the aggregate value of a combination of coins is thus stochastic. Given a set of such coin types and noisy observations of total sums, the task is to infer the most likely latent coin combination. To address the combinatorial and probabilistic complexity of PCCP, we propose DeepProReasoner (Deep Combinatorial Probabilistic Reasoning with Embedded Representations), an unsupervised, end-to-end, deep-learning framework that integrates combinatorial reasoning, latent-space modeling, and differentiable probabilistic reasoning. The model is trained using a reconstruction loss between the observed empirical distribution and a decoded probability mass function (PMF), enabling efficient gradient-based search over a continuous relaxation of the combinatorial space. We evaluate DeepProReasoner on two instances of PCCP: (1) a synthetic Candy Mix problem for ablation studies, and (2) a real-world task of molecular formula inference from ultrahigh resolution mass spectrometry (MS) data. Besides the two given instances, PCCP captures a wide range of inverse settings in biology, chemistry, environmental sciences, and medicine, where latent combinatorial structures give rise to noisy aggregate observations through stochastic processes. Our results show that DeepProReasoner achieves high accuracy and robustness, outperforming state-of-the-art methods.

EAAI Journal 2025 Journal Article

Underwater acoustic signal recognition system with multi-scale hybrid cepstral feature strategy and joint deep network

  • Hong Yang
  • Jinmei Li
  • Guohui Li
  • Chao Wang

In this paper, we propose a new underwater acoustic signal recognition system to address the recognition difficulties caused by the susceptibility of signals to complex noise interference in underwater acoustic environments. Specifically, the proposed system includes two stages: feature extraction and recognition. Feature extraction: a multi-scale hybrid cepstral feature strategy is proposed. It uses new singular spectrum decomposition to obtain multi-scale components and then extracts the Mel-frequency cepstral coefficients, inverse Mel-frequency cepstral coefficients, Gammatone frequency cepstral coefficients, and linear prediction cepstral coefficients of each component. After feature enhancement and selection, a novel multi-scale hybrid cepstral feature set is constructed. This feature set realizes the complementarity and enhancement of different cepstral features and effectively solves the problems of single feature expression and data redundancy. Recognition: a new joint deep network model is proposed. It adopts the unique design of one-dimensional convolutional neural network (1DCNN) and bidirectional gated recursive unit (BiGRU), which realizes the mutual complement of spatial information extracted by 1DCNN and dependent information captured by BiGRU and effectively improves the processing ability of the model for complex feature sets. In addition, the Kepler optimization algorithm and self-concern mechanism are introduced into the network, which solves the problem of selecting network parameters and improves the focus ability of the model on key features. By setting up multiple groups of comparison and ablation experiments, the recognition results of underwater acoustic data, including ship-radiated noise signals and marine biological signals, show that the recognition accuracy of the proposed system reaches 96. 11 % and 98. 67 %, respectively, which is better than all comparison methods. In addition, we further verified that the system still has high robustness under a low signal-to-noise ratio, which provides new ideas for research in the field of underwater acoustic signal recognition.

AAAI Conference 2024 Conference Paper

LGMRec: Local and Global Graph Learning for Multimodal Recommendation

  • Zhiqiang Guo
  • Jianjun Li
  • Guohui Li
  • Chaoyang Wang
  • Si Shi
  • Bin Ruan

The multimodal recommendation has gradually become the infrastructure of online media platforms, enabling them to provide personalized service to users through a joint modeling of user historical behaviors (e.g., purchases, clicks) and item various modalities (e.g., visual and textual). The majority of existing studies typically focus on utilizing modal features or modal-related graph structure to learn user local interests. Nevertheless, these approaches encounter two limitations: (1) Shared updates of user ID embeddings result in the consequential coupling between collaboration and multimodal signals; (2) Lack of exploration into robust global user interests to alleviate the sparse interaction problems faced by local interest modeling. To address these issues, we propose a novel Local and Global Graph Learning-guided Multimodal Recommender (LGMRec), which jointly models local and global user interests. Specifically, we present a local graph embedding module to independently learn collaborative-related and modality-related embeddings of users and items with local topological relations. Moreover, a global hypergraph embedding module is designed to capture global user and item embeddings by modeling insightful global dependency relations. The global embeddings acquired within the hypergraph embedding space can then be combined with two decoupled local embeddings to improve the accuracy and robustness of recommendations. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of our LGMRec over various state-of-the-art recommendation baselines, showcasing its effectiveness in modeling both local and global user interests.

NeurIPS Conference 2024 Conference Paper

Sim2Real-Fire: A Multi-modal Simulation Dataset for Forecast and Backtracking of Real-world Forest Fire

  • Yanzhi Li
  • Keqiu Li
  • Guohui Li
  • Zumin Wang
  • Changqing Ji
  • Lubo Wang
  • Die Zuo
  • Qing Guo

The latest research on wildfire forecast and backtracking has adopted AI models, which require a large amount of data from wildfire scenarios to capture fire spread patterns. This paper explores using cost-effective simulated wildfire scenarios to train AI models and apply them to the analysis of real-world wildfire. This solution requires AI models to minimize the Sim2Real gap, a brand-new topic in the fire spread analysis research community. To investigate the possibility of minimizing the Sim2Real gap, we collect the Sim2Real-Fire dataset that contains 1M simulated scenarios with multi-modal environmental information for training AI models. We prepare 1K real-world wildfire scenarios for testing the AI models. We also propose a deep transformer, S2R-FireTr, which excels in considering the multi-modal environmental information for forecasting and backtracking the wildfire. S2R-FireTr surpasses state-of-the-art methods in real-world wildfire scenarios.

EAAI Journal 2023 Journal Article

A new hybrid prediction model of COVID-19 daily new case data

  • Guohui Li
  • Jin Lu
  • Kang Chen
  • Hong Yang

With the emergence of new mutant corona virus disease 2019 (COVID-19) strains such as Delta and Omicron, the number of infected people in various countries has reached a new high. Accurate prediction of the number of infected people is of far-reaching sig Nificance to epidemiological prevention in all countries of the world. In order to improve the prediction accuracy of COVID-19 daily new case data, a new hybrid prediction model of COVID-19 is proposed, which consists of four modules: decomposition, complexity judgment, prediction and error correction. Firstly, singular spectrum decomposition is used to decompose the COVID-19 data into singular spectrum components (SSC). Secondly, the complexity judgment is innovatively divided into high-complexity SSC and low-complexity SSC by neural network estimation time entropy. Thirdly, an improved LSSVM by GODLIKE optimization algorithm, named GLSSVM, is proposed to improve its prediction accuracy. Then, each low-complexity SSC is predicted by ARIMA, and each high-complexity SSC is predicted by GLSSVM, and the prediction error of each high-complexity SSC is predicted by GLSSVM. Finally, the predicted results are combined and reconstructed. Simulation experiments in Japan, Germany and Russia show that the proposed model has the highest prediction accuracy and the lowest prediction error. Diebold Mariano (DM) test is introduced to evaluate the model comprehensively. Taking Japan as an example, compared with ARIMA prediction model, the RMSE, average error and MAPE of the proposed model are reduced by 93. 17%, 91. 42% and 81. 20% respectively.

AAAI Conference 2023 Conference Paper

HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question Answering

  • Zhiyuan Ma
  • Zhihuan Yu
  • Jianjun Li
  • Guohui Li

Visual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answer quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pre-training tasks are generally incompatible with downstream question answering tasks, which makes the knowledge from the language model not well transferable to downstream tasks, and greatly limits their performance in few-shot scenarios; (2) Under-fitting. They generally do not integrate human priors to compensate for universal knowledge from language models, so as to fit the challenging VQA problem and generate reliable answers. To address these issues, we propose HybridPrompt, a cloze- and verify-style hybrid prompt framework with bridging language models and human priors in prompt tuning for VQA. Specifically, we first modify the input questions into the cloze-style prompts to narrow the gap between upstream pre-training tasks and downstream VQA task, which ensures that the universal knowledge in the language model can be better transferred to subsequent human prior-guided prompt tuning. Then, we imitate the cognitive process of human brain to introduce topic and sample related priors to construct a dynamic learnable prompt template for human prior-guided prompt learning. Finally, we add fixed-length learnable free-parameters to further enhance the generalizability and scalability of prompt learning in the VQA model. Experimental results verify the effectiveness of HybridPrompt, showing that it achieves competitive performance against previous methods on widely-used VQAv2 dataset and obtains new state-of-the-art results. Our code is released at: https://github.com/zhizhi111/hybrid.

EAAI Journal 2022 Journal Article

A new traffic flow prediction model based on cosine similarity variational mode decomposition, extreme learning machine and iterative error compensation strategy

  • Hong Yang
  • Yuanxun Cheng
  • Guohui Li

Traffic flow data (TFD) prediction is a hot research area in intelligent transportation system. TFD is non-stationary and nonlinear, so it has become a challenge to predict it accurately. In order to improve TFD prediction accuracy, a new TFD prediction model based on cosine similarity variational mode decomposition (CSVMD), extreme learning machine (ELM), and iterative error compensation strategy, named CSVMD-ELM-error, is proposed. To solve mode number K value selection of variational mode decomposition, CSVMD is proposed, which realizes the self-adaptive determination of K value. The idea of CSVMD-ELM-error is roughly as follows. Firstly, CSVMD decomposes TFD into a series of intrinsic mode functions (IMFs), and ELM is established for each IMF component. Then, in order to further improve the prediction accuracy, ELM is used to correct the prediction error of each IMF. Finally, the revised IMF error results and the IMF prediction results are reconstructed to complete the prediction. Four TFDs and nine comparison models are used for simulation experiment, the experimental result shows that CSVMD-ELM-error has the best prediction accuracy and has an effective application in TFD prediction.

AAAI Conference 2011 Conference Paper

Integrating Community Question and Answer Archives

  • Wei Wei
  • Gao Cong
  • Xiaoli Li
  • See-Kiong Ng
  • Guohui Li

Question and answer pairs in Community Question Answering (CQA) services are organized into hierarchical structures or taxonomies to facilitate users to find the answers for their questions conveniently. We observe that different CQA services have their own knowledge focus and used different taxonomies to organize their question and answer pairs in their archives. As there are no simple semantic mappings between the taxonomies of the CQA services, the integration of CQA services is a challenging task. The existing approaches on integrating taxonomies ignore the hierarchical structures of the source taxonomy. In this paper, we propose a novel approach that is capable of incorporating the parent-child and sibling information in the hierarchical structures of the source taxonomy for accurate taxonomy integration. Our experimental results with real world CQA data demonstrate that the proposed method significantly outperforms state-of-the-art methods.

v2026.09.13