Arrow Research search

Author name cluster

Yu Gu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

17 papers
2 author rows

Possible papers

17

AAAI Conference 2026 Conference Paper

JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error Correction

  • Yuhao Zhan
  • Yuqing Zhang
  • Jing Yuan
  • Qixiang Ma
  • Zhiqi Yang
  • Yu Gu
  • Zemin Liu
  • Fei Wu

Existing Grammatical Error Correction (GEC) systems suffer from limited reference diversity, leading to underestimated evaluation and restricted model generalization. To address this issue, we introduce the Judge of Edit-Level Validity (JELV), an automated framework to validate correction edits from grammaticality, faithfulness, and fluency. Using our proposed human-annotated Pair-wise Edit-level Validity Dataset (PEVData) as benchmark, JELV offers two implementations: a multi-turn LLM-as-Judges pipeline achieving 90% agreement with human annotators, and a distilled DeBERTa classifier with 85% precision on valid edits. We then apply JELV to reclassify misjudged false positives in evaluation and derive a comprehensive evaluation metric by integrating false positive decoupling and fluency scoring, resulting in state-of-the-art correlation with human judgments. We also apply JELV to filter LLM-generated correction candidates, expanding the BEA19's single-reference dataset containing 38,692 source sentences. Retraining top GEC systems on this expanded dataset yields measurable performance gains. JELV provides a scalable solution for enhancing reference diversity and strengthening both evaluation and model generalization.

EAAI Journal 2026 Journal Article

MiHi-Unet: A Mix-Transformer and hierarchical feature fusion U-shaped framework for optic disc and cup segmentation

  • Zhaocan Yang
  • Shuang Liang
  • Huixiang Liu
  • Yu Gu

Glaucoma is a major cause of irreversible vision loss. Timely screening of glaucoma, along with appropriate treatment, can effectively mitigate the progression of vision loss in patients. The glaucoma diagnosis method involves measuring the cup-to-disc ratio through the segmentation of the optic disc (OD) and optic cup (OC). In color fundus photograph, vessels overlap with OD and OC, and the edges of them are blurred, which impact the segmentation accuracy. In this paper, a deep learning framework is proposed for the joint segmentation of the OD and OC. It consists of Mix-Transformer (MiT) encoder, dilated criss-cross convolution module (D3CM), hierarchical feature fusion and refinement module (HFFRM), and local and global attention (LGA) decoder. The MiT encoder is employed to extract global features. The D3CM is used in the bottleneck of the network to fuse semantic features, which ensures the connectivity of the OD and OC. The hierarchical features are fused by the HFFRM and incorporated into the LGA decoder to increase the segmentation accuracy. The obtained results show that the proposed model outperforms existing methods using three indicators. Specifically, on three public datasets, the OD segmentation yields the highest Dice values of 97. 91, 97. 13, and 96. 79%, and the highest Intersection over Union (IoU) values of 96. 00, 94. 42, and 93. 85%, respectively. The OC segmentation yields the highest Dice values of 92. 63, 88. 42, and 89. 46%, and the highest IoU values of 87. 79, 80. 25, and 81. 62%, respectively. The excellent performance indicates its potential to assist clinicians in objective decision-making and support precision medicine.

EAAI Journal 2026 Journal Article

Multi-grained detail-enhanced and patch-aware network based on bird sound recognition

  • Lin Duan
  • Lidong Yang
  • Dawei Niu
  • Yong Guo
  • Yu Gu

Combining deep learning and bird sound recognition strongly supports monitoring bird species and maintaining ecological balance. However, in outdoor environments, the extraction of bird sound features is often hindered by environmental noise, making it challenging for models to learn the fine-grained features of bird sounds fully. And single-scale feature extraction is harrowing to cover the time–frequency domain feature information of bird sounds in multiple dimensions. To address these issues, this paper proposes a multi-grained detail-enhanced and patch-aware network. The model utilizes densely connected time delay neural network as the backbone network and introduces the multi-grained detail-enhanced convolution, which combines vanilla convolutions with differential convolutions in the horizontal, vertical, angular, and central levels, and incorporates multi-grained pooling strategies to learn fine-grained acoustic features at different levels. To further overcome the limitations of single-scale feature extraction, the branch patch-aware attention module is proposed. This module collaboratively captures local details and global contextual information through a multi-branch structure and patch partitioning of different sizes. On the three datasets, the method achieved accuracies of 96. 29%, 86. 51%, and 97. 40%, respectively. This achievement demonstrates the precise capture and parsing ability of the method for audio feature information.

NeurIPS Conference 2025 Conference Paper

AlignedGen: Aligning Style Across Generated Images

  • Jiexuan Zhang
  • Yiheng Du
  • Qian Wang
  • Weiqi Li
  • Yu Gu
  • Jian Zhang

Diffusion-based generative models struggle to maintain high style consistency across generated images via text description. Although several style-aligned image generation methods have been proposed to address this issue, they exhibit suboptimal performance and are primarily built upon the U-Net architecture, limiting their compatibility with DiT diffusion models like Flux that has emerged as a predominant model in the field of image generation. To address these limitations, we propose AlignedGen, a novel training-free style-aligned image generation method for DiT models to significantly enhance style consistency across generated images. Specifically, AlignedGen incorporates two key components to achieve this: Shifted Position Embedding (ShiftPE) and Advanced Attention Sharing (AAS). ShiftPE alleviates the text controllability degradation observed in prior methods when applied to DiT models through its non-overlapping position indices design, while AAS comprises three specialized techniques to unleash the full potential of DiT for style-aligned generation. Furthermore, to broaden the applicability of our method, we present an efficient query, key, and value feature extraction algorithm, enabling our method to seamlessly incorporate external images as style references. Extensive experimental results validate that our method effectively enhances style consistency across generated images while maintaining favorable text controllability. Code: https: //github. com/Jiexuanz/AlignedGen.

AAAI Conference 2025 Conference Paper

CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder

  • Jianwei Cui
  • Yu Gu
  • Shihao Chen
  • Jie Zhang
  • Liping Chen
  • Lirong Dai

Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-end modeling is effective in the fields of SVS and Text to Speech (TTS). In this work, we thus present a fully end-to-end SVS method together with a chunkwise streaming inference to address the latency issue for practical usages. Note that this is the first attempt to fully implement end-to-end streaming audio synthesis using latent representations in VAE. We have made specific improvements to enhance the performance of streaming SVS using latent representations. Experimental results demonstrate that the proposed method achieves synthesized audio with high expressiveness and pitch accuracy in both streaming SVS and TTS tasks.

AAAI Conference 2025 Conference Paper

Hierarchical Gradient-Based Genetic Sampling for Accurate Prediction of Biological Oscillations

  • Heng Rao
  • Yu Gu
  • Jason Zipeng Zhang
  • Ge Yu
  • Yang Cao
  • Minghan Chen

Biological oscillations are periodic changes in various signaling processes crucial for the proper functioning of living organisms. These oscillations are modeled by ordinary differential equations, with coefficient variations leading to diverse periodic behaviors, typically measured by oscillatory frequencies. This paper explores sampling techniques for neural networks to model the relationship between system coefficients and oscillatory frequency. However, the scarcity of oscillations in the vast coefficient space results in many samples exhibiting non-periodic behaviors, and small coefficient changes near oscillation boundaries can significantly alter oscillatory properties. This leads to non-oscillatory bias and boundary sensitivity, making accurate predictions difficult. While existing importance and uncertainty sampling approaches partially mitigate these challenges, they either fail to resolve the sensitivity problem or result in redundant sampling. To address these limitations, we propose the Hierarchical Gradient-based Genetic Sampling (HGGS) framework, which improves the accuracy of neural network predictions for biological oscillations. The first layer, Gradient-based Filtering, extracts sensitive oscillation boundaries and removes redundant non-oscillatory samples, creating a balanced coarse dataset. The second layer, Multi-grid Genetic Sampling, utilizes residual information to refine these boundaries and explore new high-residual regions, increasing data diversity for model training. Experimental results demonstrate that HGGS outperforms seven comparative sampling methods across four biological systems, highlighting its effectiveness in enhancing sampling and prediction accuracy.

TMLR Journal 2025 Journal Article

Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

  • Yu Gu
  • Kai Zhang
  • Yuting Ning
  • Boyuan Zheng
  • Boyu Gou
  • Tianci Xue
  • Cheng Chang
  • Sanjari Srivastava

Language agents based on large language models (LLMs) have demonstrated great promise in automating web-based tasks. Recent work has shown that incorporating advanced planning algorithms, e.g., tree search, is advantageous over reactive planning for web agents. However, unlike simulated sandbox environments, real-world environments such as the web are rife with irreversible actions. This undermines the feasibility of backtracking, a cornerstone of (tree) search. Overly relying on test-time search also hurts efficiency. We advocate model-based planning for web agents that employs a world model to simulate and deliberate over the outcome of each candidate action before committing to one. We systematically explore this paradigm by: (1) Proposing a model-based planning framework, WebDreamer, which employs LLMs to serve as both world models and value functions; (2) Training specialized LLMs as world models with a scalable data synthesis pipeline. Empirical results demonstrate that WebDreamers achieves substantial performance improvements over reactive baselines. It is competitive, while being - times more efficient, with tree search in sandbox environments (VisualWebArena) and also works effectively on real-world websites (Online-Mind2Web and Mind2Web-Live). Furthermore, our trained world model, Dreamer-7B, performs comparable to GPT-4o, highlighting the potential of specialized world models for efficient and effective planning in complex web environments. All code, models, and data are publicly available at https://github.com/OSU-NLP-Group/WebDreamer

NeurIPS Conference 2025 Conference Paper

LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph Completion

  • Yuan Guo
  • Qian Ma
  • Hui Li
  • Qiao Ning
  • Furui Zhan
  • Yu Gu
  • Ge Yu
  • Shikai Guo

Multi-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e. g. , text, images, videos) associated with entities. This approach facilitates the automatic discovery of previously unobserved factual knowledge. However, existing MMKGC methods encounter several critical challenges: (i) the imbalance of inter-entity information across different modalities; (ii) the heterogeneity of intra-entity multimodal information; and (iii) for a given entity, the informational contributions of different modalities are inconsistent across contexts. In this paper, we propose a novel L arge model-driven B alanced M ultimodal K nowledge G raph C ompletion framework, termed LBMKGC. Subsequently, to bridge the semantic gap between heterogeneous modalities, LBMKGC aligns the multimodal embeddings of entities semantically by using the CLIP (Contrastive Language-Image Pre-Training) model. Furthermore, LBMKGC adaptively fuses multimodal embeddings with relational guidance by distinguishing between the perceptual and conceptual attributes of triples. Finally, extensive experiments conducted against 21 state-of-the-art baselines demonstrate that LBMKGC achieves superior performance across diverse datasets and scenarios while maintaining efficiency and generalizability. Our code and data are publicly available at: https: //github. com/guoynow/LBMKGC.

NeurIPS Conference 2025 Conference Paper

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

  • Boyu Gou
  • Zanming Huang
  • Yuting Ning
  • Yu Gu
  • Michael Lin
  • Weijian Qi
  • Andrei Kopanev
  • Botao Yu

Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the growing complexity and open-endedness of agentic search have outpaced existing evaluation benchmarks and methodologies, which largely assume short search horizons and static answers. In this paper, we introduce Mind2Web 2, a benchmark of 130 realistic, high-quality, and long-horizon tasks that require real-time web browsing and extensive information synthesis, constructed with over 1000 hours of human labor. To address the challenge of evaluating time-varying and complex answers, we propose a novel Agent-as-a-Judge framework. Our method constructs task-specific judge agents based on a tree-structured rubric design to automatically assess both answer correctness and source attribution. We conduct a comprehensive evaluation of ten frontier agentic search systems and human performance, along with a detailed error analysis to draw insights for future development. The best-performing system, OpenAI Deep Research, can already achieve 50-70% of human performance while spending half the time, highlighting its great potential. Altogether, Mind2Web 2 provides a rigorous foundation for developing and benchmarking the next generation of agentic search systems.

NeurIPS Conference 2025 Conference Paper

SimSort: A Data-Driven Framework for Spike Sorting by Large-Scale Electrophysiology Simulation

  • Yimu Zhang
  • Dongqi Han
  • Yansen Wang
  • Zhenning Lv
  • Yu Gu
  • Dongsheng Li

Spike sorting is an essential process in neural recording, which identifies and separates electrical signals from individual neurons recorded by electrodes in the brain, enabling researchers to study how specific neurons communicate and process information. Although there exist a number of spike sorting methods which have contributed to significant neuroscientific breakthroughs, many are heuristically designed, making it challenging to verify their correctness due to the difficulty of obtaining ground truth labels from real-world neural recordings. In this work, we explore a data-driven, deep learning-based approach. We begin by creating a large-scale dataset through electrophysiology simulations using biologically realistic computational models. We then present SimSort, a pretraining framework for spike sorting. Trained solely on simulated data, SimSort demonstrates zero-shot generalizability to real-world spike sorting tasks, yielding consistent improvements over existing methods across multiple benchmarks. These results highlight the potential of simulation-driven pretraining to enhance the robustness and scalability of spike sorting in experimental neuroscience.

AAAI Conference 2025 Conference Paper

SSL-STMFormer Self-Supervised Learning Spatio-Temporal Entanglement Transformer for Traffic Flow Prediction

  • Zetao Li
  • Zheng Hu
  • Peng Han
  • Yu Gu
  • Shimin Cai

Traffic flow prediction remains a critical issue in intelligent transport systems. Despite significant efforts in traffic flow modeling, existing approaches exhibit several notable limitations: (i) Most models fail to capture traffic flow similarities over long distances and extended periods; (ii) They struggle to account for spatio-temporal heterogeneity induced by varying traffic flow patterns; (iii) Due to their static modeling approach, they struggle to effectively capture the intricate spatio-temporal entanglement. To address these challenges, we propose a traffic flow prediction framework based on self-supervised learning spatio-temporal entanglement transformer(SSL-STMFormer). This framework adopts a self-supervised learning paradigm, leveraging a transformer architecture that captures richer spatio-temporal information to better represent traffic flow patterns. Specifically, a temporal attention module and a spatial attention module are employed to capture the spatio-temporal dependencies of traffic dynamics, respectively, and spatio-temporal entanglement-aware methods are introduced to allow the model to perceive spatio-temporal entanglement and thus better modelling of real traffic environments. Furthermore, to achieve adaptive spatio-temporal self-supervised learning, adaptive data augmentation is applied to the input traffic flow data, and the traffic flow prediction task is enhanced with temporal heterogeneity module and spatial heterogeneity module. Extensive experimental evaluations conducted on six publicly available real-world transportation datasets demonstrate that our method achieves substantial improvements across these datasets.

AAAI Conference 2024 Conference Paper

FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering

  • Zhenyu Li
  • Sunqi Fan
  • Yu Gu
  • Xiuxing Li
  • Zhichao Duan
  • Bowen Dong
  • Ning Liu
  • Jianyong Wang

Knowledge base question answering (KBQA) is a critical yet challenging task due to the vast number of entities within knowledge bases and the diversity of natural language questions posed by users. Unfortunately, the performance of most KBQA models tends to decline significantly in real-world scenarios where high-quality annotated data is insufficient. To mitigate the burden associated with manual annotation, we introduce FlexKBQA by utilizing Large Language Models (LLMs) as program translators for addressing the challenges inherent in the few-shot KBQA task. Specifically, FlexKBQA leverages automated algorithms to sample diverse programs, such as SPARQL queries, from the knowledge base, which are subsequently converted into natural language questions via LLMs. This synthetic dataset facilitates training a specialized lightweight model for the KB. Additionally, to reduce the barriers of distribution shift between synthetic data and real user questions, FlexKBQA introduces an executionguided self-training method to iterative leverage unlabeled user questions. Furthermore, we explore harnessing the inherent reasoning capability of LLMs to enhance the entire framework. Consequently, FlexKBQA delivers substantial flexibility, encompassing data annotation, deployment, and being domain agnostic. Through extensive experiments on GrailQA, WebQSP, and KQA Pro, we observe that under the few-shot even the more challenging zero-shot scenarios, FlexKBQA achieves impressive results with a few annotations, surpassing all previous baselines and even approaching the performance of supervised models, achieving a remarkable 93% performance relative to the fully-supervised models. We posit that FlexKBQA represents a significant advancement towards exploring better integration of large and lightweight models. Code is available at https://github.com/leezythu/FlexKBQA.

NeurIPS Conference 2024 Conference Paper

HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models

  • Bernal J. Gutiérrez
  • Yiheng Shu
  • Yu Gu
  • Michihiro Yasunaga
  • Yu Su

In order to thrive in hostile and ever-changing natural environments, mammalian brains evolved to store large amounts of knowledge about the world and continually integrate new information while avoiding catastrophic forgetting. Despite the impressive accomplishments, large language models (LLMs), even with retrieval-augmented generation (RAG), still struggle to efficiently and effectively integrate a large amount of new experiences after pre-training. In this work, we introduce HippoRAG, a novel retrieval framework inspired by the hippocampal indexing theory of human long-term memory to enable deeper and more efficient knowledge integration over new experiences. HippoRAG synergistically orchestrates LLMs, knowledge graphs, and the Personalized PageRank algorithm to mimic the different roles of neocortex and hippocampus in human memory. We compare HippoRAG with existing RAG methods on multi-hop question answering (QA) and show that our method outperforms the state-of-the-art methods remarkably, by up to 20%. Single-step retrieval with HippoRAG achieves comparable or better performance than iterative retrieval like IRCoT while being 10-20 times cheaper and 6-13 times faster, and integrating HippoRAG into IRCoT brings further substantial gains. Finally, we show that our method can tackle new types of scenarios that are out of reach of existing methods.

AAAI Conference 2024 Conference Paper

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

  • Qiushi Zhu
  • Jie Zhang
  • Yu Gu
  • Yuchen Hu
  • Lirong Dai

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noises. The efficacy of self-supervised learning for far-field multichannel and multi-modal speech processing has not been well explored. Considering that visual information helps to improve speech recognition performance in noisy scenes, in this work we propose the multichannel multi-modal speech self-supervised learning framework AV-wav2vec2, which utilizes video and multichannel audio data as inputs. First, we propose a multi-path structure to process multi-channel audio streams and a visual stream in parallel, with intra-, and inter-channel contrastive as training targets to fully exploit the rich information in multi-channel speech data. Second, based on contrastive learning, we use additional single-channel audio data, which is trained jointly to improve the performance of multichannel multi-modal representation. Finally, we use a Chinese multichannel multi-modal dataset in real scenarios to validate the effectiveness of the proposed method on audio-visual speech recognition (AVSR), automatic speech recognition (ASR), visual speech recognition (VSR) and audio-visual speaker diarization (AVSD) tasks.

IROS Conference 2024 Conference Paper

Transformer-Based Relationship Inference Model for Household Object Organization by Integrating Graph Topology and Ontology

  • Xiaodong Li
  • Guohui Tian
  • Yongcheng Cui
  • Yu Gu

In domestic environments, the conventional organization of objects by service robots often relies on the inherent properties of each object, such as placing fragile bowls in enclosed cupboards. However, this approach tends to overlook the importance of the orderly arrangement of objects, neglecting the specific placement order of bowls within the cabinet. In practice, effective object organization necessitates consideration of both individual properties and the relationships defined by these properties. In this paper, we have constructed a specialized dataset encompassing the ontological properties of household objects along with their relationships. Furthermore, we have introduced a graph-based model to explicitly represent these relationships and proposed a novel feature extraction technique that integrates the Graph Attention Network (GAT) with the BERT model to predict the relationships among objects. Subsequently, we utilized the Transformer framework to train a model, enabling it to infer relationships between objects. Experimental validation demonstrates the effectiveness of our approach in accurately predicting relationships between household objects, thus facilitating their orderly organization. Our approach significantly augments the object organization capabilities for service robots by accurately predicting the relationships among household objects. Our code is available at: https://github.com/Li-XD-Pro/Household-Object-Organization

NeurIPS Conference 2023 Conference Paper

Mind2Web: Towards a Generalist Agent for the Web

  • Xiang Deng
  • Yu Gu
  • Boyuan Zheng
  • Shijie Chen
  • Sam Stevens
  • Boshi Wang
  • Huan Sun
  • Yu Su

We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either use simulated websites or only cover a limited set of websites and tasks, thus not suitable for generalist web agents. With over 2, 000 open-ended tasks collected from 137 websites spanning 31 domains and crowdsourced action sequences for the tasks, Mind2Web provides three necessary ingredients for building generalist web agents: 1) diverse domains, websites, and tasks, 2) use of real-world websites instead of simulated and simplified ones, and 3) a broad spectrum of user interaction patterns. Based on Mind2Web, we conduct an initial exploration of using large language models (LLMs) for building generalist web agents. While the raw HTML of real-world websites are often too large to be fed to LLMs, we show that first filtering it with a small LM significantly improves the effectiveness and efficiency of LLMs. Our solution demonstrates a decent level of performance, even on websites or entire domains the model has never seen before, but there is still a substantial room to improve towards truly generalizable agents. We open-source our dataset, model implementation, and trained models (https: //osu-nlp-group. github. io/Mind2Web) to facilitate further research on building a generalist agent for the web.

IJCAI Conference 2017 Conference Paper

Single-Pass PCA of Large High-Dimensional Data

  • Wenjian Yu
  • Yu Gu
  • Jian Li
  • Shenghua Liu
  • Yaohang Li

Principal component analysis (PCA) is a fundamental dimension reduction tool in statistics and machine learning. For large and high-dimensional data, computing the PCA (i. e. , the top singular vectors of the data matrix) becomes a challenging task. In this work, a single-pass randomized algorithm is proposed to compute PCA with only one pass over the data. It is suitable for processing extremely large and high-dimensional data stored in slow memory (hard disk) or the data generated in a streaming fashion. Experiments with synthetic and real data validate the algorithm's accuracy, which has orders of magnitude smaller error than an existing single-pass algorithm. For a set of high-dimensional data stored as a 150 GB file, the algorithm is able to compute the first 50 principal components in just 24 minutes on a typical 24-core computer, with less than 1 GB memory cost.

v2026.09.13