Arrow Research search

Author name cluster

Bing Yan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

ICML Conference 2025 Conference Paper

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

  • Aaron J. Havens
  • Benjamin Kurt Miller
  • Bing Yan
  • Carles Domingo-Enrich
  • Anuroop Sriram
  • Daniel S. Levine 0003
  • Brandon M. Wood
  • Bin Hu

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry. Code & and benchmarks provided at https: //github. com/facebookresearch/adjoint_sampling.

YNIMG Journal 2025 Journal Article

From simulation to clinic: Assessing the required channel count for effective clinical use of OPM-MEG systems

  • Bing Yan
  • Yuming Peng
  • Yixiang Zhang
  • Yun Zhang
  • Haonan Zhang
  • Yifu Cao
  • Chang Sun
  • Ming Ding

The channel count in an Optically Pumped Magnetometer Magnetoencephalography (OPM-MEG) system plays a pivotal role in determining its overall performance. While existing research consistently highlights that a greater number of channels enhances system capabilities, practical constraints such as sensor placement on the head, inter-channel interference, and cost-efficiency impose limitations on channel scalability. Additionally, the optimal channel count required for clinical applications of OPM-MEG remains unclear. In this study, we systematically investigate the impact of channel count on OPM-MEG performance by integrating simulations, phantom experiments, and human MEG experiments. Four configurations with varying channel counts (16, 32, 64, and 128) are evaluated. Specifically, systems with fewer channels (e.g., 16 channels) encounter significant challenges in meeting the demands of clinical MEG applications. In contrast, a 64-channel OPM-MEG system demonstrates performance metrics-such as signal-to-noise ratio (SNR) and localization accuracy-that are comparable to those of a 306-channel Superconducting Quantum Interference Device MEG (SQUID-MEG) system. Notably, a 128-channel OPM-MEG system surpasses the 306-channel SQUID-MEG system, achieving superior results. This work provides a detailed exploration of the relationship between channel count and OPM-MEG system performance, analyzing how many channels of the OPM-MEG system are suitable for clinical applications. By combining simulation-based evaluations with empirical measurements, we found that it is crucial to carefully select the appropriate number of channels based on the specific usage requirements in clinical applications.

IJCAI Conference 2024 Conference Paper

Predictive Accuracy-Based Active Learning for Medical Image Segmentation

  • Jun Shi
  • Shulan Ruan
  • Ziqi Zhu
  • Minfan Zhao
  • Hong An
  • Xudong Xue
  • Bing Yan

Active learning is considered a viable solution to alleviate the contradiction between the high dependency of deep learning-based segmentation methods on annotated data and the expensive pixel-level annotation cost of medical images. However, most existing methods suffer from unreliable uncertainty assessment and the struggle to balance diversity and informativeness, leading to poor performance in segmentation tasks. In response, we propose an efficient Predictive Accuracy-based Active Learning (PAAL) method for medical image segmentation, first introducing predictive accuracy to define uncertainty. Specifically, PAAL mainly consists of an Accuracy Predictor (AP) and a Weighted Polling Strategy (WPS). The former is an attached learnable module that can accurately predict the segmentation accuracy of unlabeled samples relative to the target model with the predicted posterior probability. The latter provides an efficient hybrid querying scheme by combining predicted accuracy and feature representation, aiming to ensure the uncertainty and diversity of the acquired samples. Extensive experiment results on multiple datasets demonstrate the superiority of PAAL. PAAL achieves comparable accuracy to fully annotated data while reducing annotation costs by approximately 50% to 80%, showcasing significant potential in clinical applications. The code is available at https: //github. com/shijun18/PAAL-MedSeg.

ICML Conference 2024 Conference Paper

Structured Chemistry Reasoning with Large Language Models

  • Siru Ouyang
  • Zhuosheng Zhang 0001
  • Bing Yan
  • Xuan Liu 0009
  • Yejin Choi 0001
  • Jiawei Han 0001
  • Lianhui Qin

Large Language Models (LLMs) excel in diverse areas, yet struggle with complex scientific reasoning, especially in the field of chemistry. Different from the simple chemistry tasks (e. g. , molecule classification) addressed in previous studies, complex chemistry problems require not only vast knowledge and precise calculation, but also compositional reasoning about rich dynamic interactions of different concepts (e. g. , temperature changes). Our study shows that even advanced LLMs, like GPT-4, can fail easily in different ways. Interestingly, the errors often stem not from a lack of domain knowledge within the LLMs, but rather from the absence of an effective reasoning structure that guides the LLMs to elicit the right knowledge, incorporate the knowledge in step-by-step reasoning, and iteratively refine results for further improved quality. On this basis, we introduce StructChem, a simple yet effective prompting strategy that offers the desired guidance and substantially boosts the LLMs’ chemical reasoning capability. Testing across four chemistry areas—quantum chemistry, mechanics, physical chemistry, and kinetics—StructChem substantially enhances GPT-4’s performance, with up to 30% peak improvement. Our analysis also underscores the unique difficulties of precise grounded reasoning in science with LLMs, highlighting a need for more research in this area.

IS Journal 2013 Journal Article

Developing Parallel Control and Management for Urban Traffic Systems

  • Qing-Jie Kong
  • Lefei Li
  • Bing Yan
  • Shu Lin
  • Fenghua Zhu
  • Gang Xiong

A streamlined parallel traffic management system (PtMS) is outlined that works alongside a redesigned intelligent transportation system in Qingdao, China. The PtMS's structure provides enhanced control and management support, with increased versatility for use in real-world scenarios.

v2026.09.13