Arrow Research search

Author name cluster

Jiaqi Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

21 papers
2 author rows

Possible papers

21

EAAI Journal 2026 Journal Article

FishMotionNet: Integrating hydrological factors and fishing motion patterns for enhanced vessel trajectory prediction

  • Guitong Yang
  • Guiyuan Jiang
  • Jiaqi Yang
  • Feng Hong
  • Peilan He
  • Zhongning Zhao

Fisheries resources are vital to global economies, particularly in coastal regions, yet overfishing and increased maritime accidents present significant challenges. Accurate prediction of fishing vessel trajectories is essential for optimizing operations, ensuring sustainable resource use, and enhancing maritime safety. Existing models primarily focus on commercial vessels and are inadequate for the dynamic and irregular patterns of fishing vessels, which frequently switch between navigation and fishing modes with complex maneuvers. This study proposes FishMotionNet, a novel deep learning model integrating Vessel Monitoring System (VMS) data with similar historical trajectories, regional behavioral features, and hydrological factors. Using 90-minute observations, FishMotionNet leverages similar trajectories to capture intricate vessel motion behaviors and predicts 90-minute future trajectories. The model incorporates environmental influences through sea surface height, temperature, salinity, and ocean currents, with an encoder–decoder architecture enhancing complex trajectory pattern learning. Experiments using a comprehensive VMS dataset from the Beidou satellite comprising 1855 trawlers in the East China Sea (September 2016-December 2017) with hydrological data from Copernicus Climate Database showed that over a 90-minute prediction horizon, FishMotionNet achieved an average prediction error of 0. 815 nautical miles and a final displacement error of 1. 403 nautical miles, achieving 14. 7% and 16. 7% improvements over the best baseline model, with consistent superior performance across 30-minute and 60-minute prediction horizons, significantly outperforming baseline models. FishMotionNet effectively addresses the unique challenges of fishing vessel trajectory prediction, offering a valuable tool for fisheries management and maritime safety, and contributing to the sustainable exploitation of marine resources.

JBHI Journal 2026 Journal Article

Text-Driven Weakly Supervised OCT Lesion Segmentation With Structural Guidance

  • Jiaqi Yang
  • Nitish Mehta
  • Xiaoling Hu
  • Chao Chen
  • Chia-Ling Tsai

Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of supervised learning for large datasets. Weakly Supervised Semantic Segmentation (WSSS) offers a promising alternative by using weaker forms of supervision, such as image-level labels, to reduce the annotation burden. Despite its advantages, weak supervision inherently carries limited information. We propose a novel WSSS framework with only image-level labels for OCT lesion segmentation that integrates structural and text-driven guidance to produce high-quality, pixel-level pseudo labels. The framework employs two visual processing modules: one that processes the original OCT images and another that operates on layer segmentations augmented with anomalous signals, enabling the model to associate lesions with their corresponding anatomical layers. Complementing these visual cues, we leverage large-scale pretrained models to provide two forms of textual guidance: label-derived descriptions that encode local semantics, and domain-agnostic synthetic descriptions that, although expressed in natural image terms, capture spatial and relational semantics useful for generating globally consistent representations. By fusing these visual and textual features in a multi-modal framework, our method aligns semantic meaning with structural relevance, thereby improving lesion localization and segmentation performance. Experiments on three OCT datasets demonstrate state-of-the-art results, highlighting its potential to advance diagnostic accuracy and efficiency in medical imaging.

EAAI Journal 2025 Journal Article

A physics informed convolution neural network for spatiotemporal temperature analysis of concrete dams

  • Jiaqi Yang
  • Jinting Wang
  • Feng Jin
  • Jianwen Pan

Structural health monitoring is indispensable throughout the life cycle of dams, and the loading conditions determines the reliability of the assessment. Among them, temperature plays an important role on the behavior of arch dams, which are sparsely monitored in practice. How to use these sparsely measured data to obtain the accurate spatiotemporal temperature field becomes a critical problem. This study proposes a physics informed convolutional neural network for spatiotemporal temperature field of arch dams. A dual thread convolutional neural network considers the effects of spatiotemporal and temporal variables distinctively. The proposed model is validated using measured data from an existing arch dam. Compared with applied convolutional neural network, the proposed model improves the accuracy of temperature field reconstruction by 18 % and reduces reliance on measured data. Benefit of consideration of the continuity and heat transfer, the spatial distribution of the temperature field is more reasonable in continuity, and can retain accuracy even with limited monitoring data. The proposed model can provide the actual spatiotemporal non-uniform temperature field of the arch dam, providing basic data for the analysis and safety evaluation of arch dams throughout their life-cycle.

AAAI Conference 2025 Conference Paper

DFF: Decision-Focused Fine-Tuning for Smarter Predict-Then-Optimize with Limited Data

  • Jiaqi Yang
  • Enming Liang
  • Zicheng Su
  • Zhichao Zou
  • Peng Zhen
  • Jiecheng Guo
  • Wanjing Ma
  • Kun An

Decision-focused learning (DFL) offers an end-to-end approach to the predict-then-optimize (PO) framework by training predictive models directly on decision loss (DL), enhancing decision-making performance within PO contexts. However, the implementation of DFL poses distinct challenges. Primarily, DL can result in deviation from the physical significance of the predictions under limited data. Additionally, some predictive models are non-differentiable or black-box, which cannot be adjusted using gradient-based methods. To tackle the above challenges, we propose a novel framework, Decision-Focused Fine-tuning (DFF), which embeds the DFL module into the PO pipeline via a novel bias correction module. DFF is formulated as a constrained optimization problem that maintains the proximity of the DL-enhanced model to the original predictive model within a defined trust region. We theoretically prove that DFF strictly confines prediction bias within a predetermined upper bound, even with limited datasets, thereby substantially reducing prediction shifts caused by DL under limited data. Furthermore, the bias correction module can be integrated into diverse predictive models, enhancing adaptability to a broad range of PO tasks. Extensive evaluations on synthetic and real-world datasets, including network flow, portfolio optimization, and resource allocation problems with different predictive models, demonstrate that DFF not only improves decision performance but also adheres to fine-tuning constraints, showcasing robust adaptability across various scenarios.

ICRA Conference 2025 Conference Paper

GS-EVT: Cross-Modal Event Camera Tracking Based on Gaussian Splatting

  • Tao Liu
  • Runze Yuan
  • Yi'ang Ju
  • Xun Xu
  • Jiaqi Yang
  • Xiangting Meng
  • Xavier Lagorce
  • Laurent Kneip

Reliable self-localization is a foundational skill for many intelligent mobile platforms. This paper explores the use of event cameras for motion tracking thereby providing a solution with inherent robustness under difficult dynamics and illumination. In order to circumvent the challenge of event camera-based mapping, the solution is framed in a cross-modal way. It tracks a map representation that comes directly from frame-based cameras. Specifically, the proposed method operates on top of gaussian splatting, a state-of-the-art representation that permits highly efficient and realistic novel view synthesis. The key of our approach consists of a novel pose parametrization that uses a reference pose plus first order dynamics for local differential image rendering. The latter is then compared against images of integrated events in a staggered coarse-to-fine optimization scheme. As demonstrated by our results, the realistic view rendering ability of gaussian splatting leads to stable and accurate tracking across a variety of both publicly available and newly recorded data sequences.

ICML Conference 2025 Conference Paper

Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens

  • Ting-Ji Huang
  • Jiaqi Yang
  • Chunxu Shen
  • Kai-Qi Liu
  • De-Chuan Zhan
  • Han-Jia Ye

Characterizing users and items through vector representations is crucial for various tasks in recommender systems. Recent approaches attempt to apply Large Language Models (LLMs) in recommendation through a question&answer format, where real items (eg, Item No. 2024) are represented with compound words formed from in-vocabulary tokens (eg, “item“, “20“, “24“). However, these tokens are not suitable for representing items, as their meanings are shaped by pre-training on natural language tasks, limiting the model’s ability to capture user-item relationships effectively. In this paper, we explore how to effectively characterize users and items in LLM-based recommender systems from the token construction view. We demonstrate the necessity of using out-of-vocabulary (OOV) tokens for the characterization of items and users, and propose a well-constructed way of these OOV tokens. By clustering the learned representations from historical user-item interactions, we make the representations of user/item combinations share the same OOV tokens if they have similar properties. This construction allows us to capture user/item relationships well (memorization) and preserve the diversity of descriptions of users and items (diversity). Furthermore, integrating these OOV tokens into the LLM’s vocabulary allows for better distinction between users and items and enhanced capture of user-item relationships during fine-tuning on downstream tasks. Our proposed framework outperforms existing state-of-the-art methods across various downstream recommendation tasks.

AAAI Conference 2025 Conference Paper

SPU-IMR: Self-supervised Arbitrary-scale Point Cloud Upsampling via Iterative Mask-recovery Network

  • Ziming Nie
  • Qiao Wu
  • Chenlei Lv
  • Siwen Quan
  • Zhaoshuai Qi
  • Muze Wang
  • Jiaqi Yang

Point cloud upsampling aims to generate dense and uniformly distributed point sets from sparse point clouds. Existing point cloud upsampling methods typically approach the task as an interpolation problem. They achieve upsampling by performing local interpolation between point clouds or in the feature space, then regressing the interpolated points to appropriate positions. By contrast, our proposed method treats point cloud upsampling as a global shape completion problem. Specifically, our method first divides the point cloud into multiple patches. Then a masking operation is applied to remove some patches, leaving visible point cloud patches. Finally, our custom-designed neural network iterative completes the missing sections of the point cloud through the visible parts. During testing, by selecting different mask sequences, we can restore various complete patches. A sufficiently dense upsampled point cloud can be obtained by merging all the completed patches. We demonstrate the superior performance of our method through both quantitative and qualitative experiments, showing overall superiority against both existing self-supervised and supervised methods.

IJCAI Conference 2025 Conference Paper

Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology Relationship

  • Pei An
  • Jiaqi Yang
  • Muyao Peng
  • You Yang
  • Qiong Liu
  • Jie Ma
  • Liangliang Nan

Image-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios.

IROS Conference 2024 Conference Paper

MV-ROPE: Multi-view Constraints for Robust Category-level Object Pose and Size Estimation

  • Jiaqi Yang
  • Yucong Chen
  • Xiangting Meng
  • Chenxin Yan
  • Min Li
  • Ran Cheng
  • Lige Liu
  • Tao Sun

Recently there has been a growing interest in category-level object pose and size estimation, and prevailing methods commonly rely on single view RGB-D images. However, one disadvantage of such methods is that they require accurate depth maps which cannot be produced by consumer-grade sensors. Furthermore, many practical real-world situations involve a moving camera that continuously observes its surroundings, and the temporal information of the input video streams is simply overlooked by single-view methods. We propose a novel solution that makes use of RGB video streams. Our framework consists of three modules: a scale-aware monocular dense SLAM solution, a lightweight object pose predictor, and an object-level pose graph optimizer. The SLAM module utilizes a video stream and additional scale-sensitive readings to estimate camera poses and metric depth. The object pose predictor then generates canonical object representations from RGB images. The object pose is estimated through geometric registration of these canonical object representations with estimated object depth points. All per-view estimates finally undergo optimization within a pose graph, culminating in the output of robust and accurate canonical object poses. Our experimental results demonstrate that when utilizing public dataset sequences with high-quality depth information, the proposed method exhibits comparable performance to state-of-the-art RGB-D methods. We also collect and evaluate on new datasets containing depth maps of varying quality to further quantitatively benchmark the proposed method alongside previous RGB-D based methods. We demonstrate a significant advantage in scenarios where depth input is absent or the quality of depth sensing is limited.

UAI Conference 2024 Conference Paper

RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction

  • Songli Wu
  • Liang Du 0004
  • Jiaqi Yang
  • Yuai Wang
  • De-Chuan Zhan
  • Shuang Zhao
  • Zixun Sun

Click-through rate (CTR) prediction is a critical task in recommendation systems, serving as the ultimate filtering step to sort items for a user. Most recent cutting-edge methods primarily focus on investigating complex implicit and explicit feature interactions; however, these methods neglect the spurious correlation issue caused by confounding factors, thereby diminishing the model’s generalization ability. We propose a CTR prediction framework that REmoves Spurious cORrelations in mulTilevel feature interactions, termed RE-SORT, which has two key components. I. A multilevel stacked recurrent (MSR) structure enables the model to efficiently capture diverse nonlinear interactions from feature spaces at different levels. II. A spurious correlation elimination (SCE) module further leverages Laplacian kernel mapping and sample reweighting methods to eliminate the spurious correlations concealed within the multilevel features, allowing the model to focus on the true causal features. Extensive experiments conducted on four challenging CTR datasets, our production dataset, and an online A/B test demonstrate that the proposed method achieves state-of-the-art performance in both accuracy and speed. The utilized codes, models, and dataset will be released at https: //github. com/RE-SORT.

ICLR Conference 2023 Conference Paper

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

  • Chao Yu 0005
  • Jiaxuan Gao
  • Weilin Liu
  • Botian Xu
  • Hao Tang
  • Jiaqi Yang
  • Yu Wang 0002
  • Yi Wu 0013

There is a recent trend of applying multi-agent reinforcement learning (MARL) to train an agent that can cooperate with humans in a zero-shot fashion without using any human data. The typical workflow is to first repeatedly run self-play (SP) to build a policy pool and then train the final adaptive policy against this pool. A crucial limitation of this framework is that every policy in the pool is optimized w.r.t. the environment reward function, which implicitly assumes that the testing partners of the adaptive policy will be precisely optimizing the same reward function as well. However, human objectives are often substantially biased according to their own preferences, which can differ greatly from the environment reward. We propose a more general framework, Hidden-Utility Self-Play (HSP), which explicitly models human biases as hidden reward functions in the self-play objective. By approximating the reward space as linear functions, HSP adopts an effective technique to generate an augmented policy pool with biased policies. We evaluate HSP on the Overcooked benchmark. Empirical results show that our HSP method produces higher rewards than baselines when cooperating with learned human models, manually scripted policies, and real humans. The HSP policy is also rated as the most assistive policy based on human feedback.

IROS Conference 2023 Conference Paper

Revisiting Event-Based Video Frame Interpolation

  • Jiaben Chen
  • Yichen Zhu
  • Dongze Lian
  • Jiaqi Yang
  • Yifu Wang
  • Renrui Zhang
  • Xinhang Liu
  • Shenhan Qian

Dynamic vision sensors or event cameras provide rich complementary information for video frame interpolation. Existing state-of-the-art methods follow the paradigm of combining both synthesis-based and warping networks. However, few of those methods fully respect the intrinsic characteristics of events streams. Given that event cameras only encode intensity changes and polarity rather than color intensities, estimating optical flow from events is arguably more difficult than from RGB information. We therefore propose to incorporate RGB information in an event-guided optical flow refinement strategy. Moreover, in light of the quasi-continuous nature of the time signals provided by event cameras, we propose a divide-and-conquer strategy in which event-based intermediate frame synthesis happens incrementally in multiple simplified stages rather than in a single, long stage. Extensive experiments on both synthetic and real-world datasets show that these modifications lead to more reliable and realistic intermediate frame results than previous video frame interpolation methods. Our findings underline that a careful consideration of event characteristics such as high temporal density and elevated noise benefits interpolation accuracy.

ICRA Conference 2022 Conference Paper

DEVO: Depth-Event Camera Visual Odometry in Challenging Conditions

  • Yi-Fan Zuo
  • Jiaqi Yang
  • Jiaben Chen
  • Xia Wang 0002
  • Yifu Wang
  • Laurent Kneip

We present a novel real-time visual odometry framework for a stereo setup of a depth and high-resolution event camera. Our framework balances accuracy and robustness against computational efficiency towards strong performance in challenging scenarios. We extend conventional edge-based semi-dense visual odometry towards time-surface maps obtained from event streams. Semi-dense depth maps are generated by warping the corresponding depth values of the extrinsically calibrated depth camera. The tracking module updates the camera pose through efficient, geometric semi-dense 3D-2D edge alignment. Our approach is validated on both public and self-collected datasets captured under various conditions. We show that the proposed method performs comparable to state-of-the-art RGB-D camera-based alternatives in regular conditions, and eventually outperforms in challenging conditions such as high dynamics or low illumination.

NeurIPS Conference 2022 Conference Paper

Generalized Delayed Feedback Model with Post-Click Information in Recommender Systems

  • Jiaqi Yang
  • De-Chuan Zhan

Predicting conversion rate (e. g. , the probability that a user will purchase an item) is a fundamental problem in machine learning based recommender systems. However, accurate conversion labels are revealed after a long delay, which harms the timeliness of recommender systems. Previous literature concentrates on utilizing early conversions to mitigate such a delayed feedback problem. In this paper, we show that post-click user behaviors are also informative to conversion rate prediction and can be used to improve timeliness. We propose a generalized delayed feedback model (GDFM) that unifies both post-click behaviors and early conversions as stochastic post-click information, which could be utilized to train GDFM in a streaming manner efficiently. Based on GDFM, we further establish a novel perspective that the performance gap introduced by delayed feedback can be attributed to a temporal gap and a sampling gap. Inspired by our analysis, we propose to measure the quality of post-click information with a combination of temporal distance and sample complexity. The training objective is re-weighted accordingly to highlight informative and timely signals. We validate our analysis on public datasets, and experimental performance confirms the effectiveness of our method.

ICML Conference 2022 Conference Paper

Phasic Self-Imitative Reduction for Sparse-Reward Goal-Conditioned Reinforcement Learning

  • Yunfei Li 0005
  • Tian Gao
  • Jiaqi Yang
  • Huazhe Xu
  • Yi Wu 0013

It has been a recent trend to leverage the power of supervised learning (SL) towards more effective reinforcement learning (RL) methods. We propose a novel phasic solution by alternating online RL and offline SL for tackling sparse-reward goal-conditioned problems. In the online phase, we perform RL training and collect rollout data while in the offline phase, we perform SL on those successful trajectories from the dataset. To further improve sample efficiency, we adopt additional techniques in the online phase including task reduction to generate more feasible trajectories and a value-difference-based intrinsic reward to alleviate the sparse-reward issue. We call this overall framework, PhAsic self-Imitative Reduction (PAIR). PAIR is compatible with various online and offline RL methods and substantially outperforms both non-phasic RL and phasic SL baselines on sparse-reward robotic control problems, including a particularly challenging stacking task. PAIR is the first RL method that learns to stack 6 cubes with only 0/1 success rewards from scratch.

ICML Conference 2022 Conference Paper

Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning

  • Wei Fu
  • Chao Yu 0005
  • Zelai Xu
  • Jiaqi Yang
  • Yi Wu 0013

Many advances in cooperative multi-agent reinforcement learning (MARL) are based on two common design principles: value decomposition and parameter sharing. A typical MARL algorithm of this fashion decomposes a centralized Q-function into local Q-networks with parameters shared across agents. Such an algorithmic paradigm enables centralized training and decentralized execution (CTDE) and leads to efficient learning in practice. Despite all the advantages, we revisit these two principles and show that in certain scenarios, e. g. , environments with a highly multi-modal reward landscape, value decomposition, and parameter sharing can be problematic and lead to undesired outcomes. In contrast, policy gradient (PG) methods with individual policies provably converge to an optimal solution in these cases, which partially supports some recent empirical observations that PG can be effective in many MARL testbeds. Inspired by our theoretical analysis, we present practical suggestions on implementing multi-agent PG algorithms for either high rewards or diverse emergent behaviors and empirically validate our findings on a variety of domains, ranging from the simplified matrix and grid-world games to complex benchmarks such as StarCraft Multi-Agent Challenge and Google Research Football. We hope our insights could benefit the community towards developing more general and more powerful MARL algorithms.

NeurIPS Conference 2021 Conference Paper

Going Beyond Linear RL: Sample Efficient Neural Function Approximation

  • Baihe Huang
  • Kaixuan Huang
  • Sham Kakade
  • Jason D. Lee
  • Qi Lei
  • Runzhe Wang
  • Jiaqi Yang

Deep Reinforcement Learning (RL) powered by neural net approximation of the Q function has had enormous empirical success. While the theory of RL has traditionally focused on linear function approximation (or eluder dimension) approaches, little is known about nonlinear RL with neural net approximations of the Q functions. This is the focus of this work, where we study function approximation with two-layer neural networks (considering both ReLU and polynomial activation functions). Our first result is a computationally and statistically efficient algorithm in the generative model setting under completeness for two-layer neural networks. Our second result considers this setting but under only realizability of the neural net function class. Here, assuming deterministic dynamics, the sample complexity scales linearly in the algebraic dimension. In all cases, our results significantly improve upon what can be attained with linear (or eluder dimension) methods.

NeurIPS Conference 2021 Conference Paper

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

  • Zihan Zhang
  • Jiaqi Yang
  • Xiangyang Ji
  • Simon S. Du

This paper presents new \emph{variance-aware} confidence sets for linear bandits and linear mixture Markov Decision Processes (MDPs). With the new confidence sets, we obtain the follow regret bounds: For linear bandits, we obtain an $\widetilde{O}(\mathrm{poly}(d)\sqrt{1 + \sum_{k=1}^{K}\sigma_k^2})$ data-dependent regret bound, where $d$ is the feature dimension, $K$ is the number of rounds, and $\sigma_k^2$ is the \emph{unknown} variance of the reward at the $k$-th round. This is the first regret bound that only scales with the variance and the dimension but \emph{no explicit polynomial dependency on $K$}. When variances are small, this bound can be significantly smaller than the $\widetilde{\Theta}\left(d\sqrt{K}\right)$ worst-case regret bound. For linear mixture MDPs, we obtain an $\widetilde{O}(\mathrm{poly}(d, \log H)\sqrt{K})$ regret bound, where $d$ is the number of base models, $K$ is the number of episodes, and $H$ is the planning horizon. This is the first regret bound that only scales \emph{logarithmically} with $H$ in the reinforcement learning with linear function approximation setting, thus \emph{exponentially improving} existing results, and resolving an open problem in \citep{zhou2020nearly}. We develop three technical ideas that may be of independent interest: 1) applications of the peeling technique to both the input norm and the variance magnitude, 2) a recursion-based estimator for the variance, and 3) a new convex potential lemma that generalizes the seminal elliptical potential lemma.

NeurIPS Conference 2021 Conference Paper

Optimal Gradient-based Algorithms for Non-concave Bandit Optimization

  • Baihe Huang
  • Kaixuan Huang
  • Sham Kakade
  • Jason D. Lee
  • Qi Lei
  • Runzhe Wang
  • Jiaqi Yang

Bandit problems with linear or concave reward have been extensively studied, but relatively few works have studied bandits with non-concave reward. This work considers a large family of bandit problems where the unknown underlying reward function is non-concave, including the low-rank generalized linear bandit problems and two-layer neural network with polynomial activation bandit problem. For the low-rank generalized linear bandit problem, we provide a minimax-optimal algorithm in the dimension, refuting both conjectures in \cite{lu2021low, jun2019bilinear}. Our algorithms are based on a unified zeroth-order optimization paradigm that applies in great generality and attains optimal rates in several structured polynomial settings (in the dimension). We further demonstrate the applicability of our algorithms in RL in the generative model setting, resulting in improved sample complexity over prior approaches. Finally, we show that the standard optimistic algorithms (e. g. , UCB) are sub-optimal by dimension factors. In the neural net setting (with polynomial activation functions) with noiseless reward, we provide a bandit algorithm with sample complexity equal to the intrinsic algebraic dimension. Again, we show that optimistic approaches have worse sample complexity, polynomial in the extrinsic dimension (which could be exponentially worse in the polynomial degree).

NeurIPS Conference 2021 Conference Paper

Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature

  • Kefan Dong
  • Jiaqi Yang
  • Tengyu Ma

This paper studies model-based bandit and reinforcement learning (RL) with nonlinear function approximations. We propose to study convergence to approximate local maxima because we show that global convergence is statistically intractable even for one-layer neural net bandit with a deterministic reward. For both nonlinear bandit and RL, the paper presents a model-based algorithm, Virtual Ascent with Online Model Learner (ViOlin), which provably converges to a local maximum with sample complexity that only depends on the sequential Rademacher complexity of the model class. Our results imply novel global or local regret bounds on several concrete settings such as linear bandit with finite or sparse model class, and two-layer neural net bandit. A key algorithmic insight is that optimism may lead to over-exploration even for two-layer neural net model class. On the other hand, for convergence to local maxima, it suffices to maximize the virtual return if the model can also reasonably predict the gradient and Hessian of the real return.

YNIMG Journal 2019 Journal Article

Anterior insular cortex is a bottleneck of cognitive control

  • Tingting Wu
  • Xingchao Wang
  • Qiong Wu
  • Alfredo Spagna
  • Jiaqi Yang
  • Changhe Yuan
  • Yanhong Wu
  • Zhixian Gao

Cognitive control, with a limited capacity, is a core process in human cognition for the coordination of thoughts and actions. Although the regions involved in cognitive control have been identified as the cognitive control network (CCN), it is still unclear whether a specific region of the CCN serves as a bottleneck limiting the capacity of cognitive control (CCC). Here, we used a perceptual decision-making task with conditions of high cognitive load to challenge the CCN and to assess the CCC in a functional magnetic resonance imaging study. We found that the activation of the right anterior insular cortex (AIC) of the CCN increased monotonically as a function of cognitive load, reached its plateau early, and showed a significant correlation to the CCC. In a subsequent study of patients with unilateral lesions of the AIC, we found that lesions of the AIC were associated with a significant impairment of the CCC. Simulated lesions of the AIC resulted in a reduction of the global efficiency of the CCN in a network analysis. These findings suggest that the AIC, as a critical hub in the CCN, is a bottleneck of cognitive control.

v2026.09.13