Arrow Research search

Author name cluster

Hanlin Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

NeurIPS Conference 2025 Conference Paper

AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?

  • Ori Press
  • Brandon Amos
  • Haoyu Zhao
  • Yikai Wu
  • Samuel Ainsworth
  • Dominik Krupke
  • Patrick Kidger
  • Touqir Sajed

Despite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming (SWE-Bench) and mathematics (FrontierMath). We therefore propose testing models' ability to design and implement algorithms in an open-ended benchmark: We task LMs with writing code that efficiently solves computationally challenging problems in computer science, physics, and mathematics. Our AlgoTune benchmark consists of 120 tasks collected from domain experts and a framework for validating and timing LM-synthesized solution code, which is compared to reference implementations from popular open-source packages. In addition, we develop a baseline LM agent, AlgoTuner, and evaluate its performance across a suite of frontier models. AlgoTuner achieves an average 1. 58x speedup against reference solvers, including methods from packages such as SciPy, scikit-learn and CVXPY. However, we find that current models fail to discover algorithmic innovations, instead preferring surface-level optimizations. We hope that AlgoTune catalyzes the development of LM agents exhibiting creative problem solving beyond state-of-the-art human performance.

JBHI Journal 2025 Journal Article

Cortico-Ocular Coupling Analysis for Developmental and Behavioral Disorders: A Review

  • Hanlin Zhang
  • Zhiyong Wang
  • Chunchun Hu
  • Peilian Chi
  • Xiu Xu
  • Honghai Liu

Developmental and behavioral disorders (DBD) have a significant impact on children's neurological activity and behavioral performance. Early diagnosis and treatment are known to be beneficial for improving DBD outcomes, yet existing unimodal neurophysiological assessment tools for DBD yield significant heterogeneity in results, highlighting the urgent need for exploring novel assessment tools. Cortico-ocular coupling (COC) refers to the information interaction between the cerebral cortex and eyes, and COC analysis is a technique for quantitatively measuring the correlation of neural oscillations and eye movements as biomarkers for assessment and mechanism disclosure. This review focuses on COC analysis for DBD from four perspectives: neural substrates, research paradigms, analysis methods, and applications. First, this review provides a comprehensive overview of the neural substrates and evocation paradigms related to COC analysis, aiming at helping target brain region selection, experimental result analysis, and paradigm design. The neural substrates and evocation paradigms are categorized according to functional domains, including social functioning, attention, cognition, early visual processing, and motor function. Then, this review summarizes the EEG and eye-tracking features, the analysis methods, and the validation datasets involved in COC analysis, aiming at helping implement COC analysis. Next, this review presents the applications of COC analysis in DBD, proving the validity and advance of COC analysis. In the end, the limitations, challenges, and future directions of COC analysis are discussed.

NeurIPS Conference 2025 Conference Paper

EvoLM: In Search of Lost Training Dynamics for Language Model Reasoning

  • Zhenting Qi
  • Fan Nie
  • Alexandre Alahi
  • James Zou
  • Himabindu Lakkaraju
  • Yilun Du
  • Eric Xing
  • Sham Kakade

Modern language model (LM) training has been divided into multiple stages, making it difficult for downstream developers to evaluate the impact of design choices made at each stage. We present EvoLM, a model suite that enables systematic and transparent analysis of LMs' training dynamics across pre-training, continued pre-training, supervised fine-tuning, and reinforcement learning. By training over 100 LMs with 1B and 4B parameters from scratch, we rigorously evaluate both upstream (language modeling) and downstream (problem-solving) reasoning capabilities, including considerations of both in-domain and out-of-domain generalization. Key insights highlight the diminishing returns from excessive pre-training and post-training, the importance and practices of mitigating forgetting during domain-specific continued pre-training, the crucial role of continued pre-training in bridging pre-training and post-training phases, and various intricate trade-offs when configuring supervised fine-tuning and reinforcement learning. To facilitate open research and reproducibility, we release all pre-trained and post-trained models, training datasets for all stages, and our entire training and evaluation pipeline.

JBHI Journal 2025 Journal Article

Exploring Eye-Tracking Based Biomarkers to Assess Cognitive Abilities in Autistic Children: A Feasibility Study

  • Hanlin Zhang
  • Chunchun Hu
  • Zhiyong Wang
  • Bingrui Zhou
  • Xinming Wang
  • Wei Nie
  • Qinyi Ye
  • Ruihan Lin

Cognitive assessment can reveal a person's cognitive processing and behavioral patterns, making it an indispensable component of autism intervention and prognosis. Existing machine-assisted cognitive assessment methods primarily focus on children's performance outcomes, overlooking distinctive behavioral models, particularly characteristics of eye movement behavior, which have been demonstrated as the most direct indicators of cognitive abilities. In this study, we explore eye-tracking biomarkers for assisting cognitive assessment through a series of meticulously designed multi-level human-computer interaction protocols, encompassing three cognitive abilities: pairing and categorization, emotion recognition, and social interaction. A platform embedded with an eye-tracking module has been developed to reliably collect and analyze eye movement data, even in the presence of unrestricted large head movements in children. Experimental results indicate that there are significant group differences between autism and typically developing children in the eye-tracking features of total fixation duration, response latency, time to first fixation, mean fixation duration, and visit count in the absence of significant intergroup differences in the Wechsler Preschool and Primary Scale of Intelligence (WPPSI) and Wechsler Intelligence Scale for Children (WISC) assessment results. In addition, certain eye-tracking features in each group are correlated with WPPSI/WISC scale scores, enabling clinical cognitive assessments within each group based on these eye movement features. This study suggests that using eye-tracking features as biomarkers to assist detailed cognitive assessments holds significant potential for the intervention and prognosis of autism.

NeurIPS Conference 2024 Conference Paper

CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training

  • David Brandfonbrener
  • Hanlin Zhang
  • Andreas Kirsch
  • Jonathan Richard Schwarz
  • Sham Kakade

Selecting high-quality data for pre-training is crucial in shaping the downstream task performance of language models. A major challenge lies in identifying this optimal subset, a problem generally considered intractable, thus necessitating scalable and effective heuristics. In this work, we propose a data selection method, CoLoR-Filter (Conditional Loss Reduction Filtering), which leverages an empirical Bayes-inspired approach to derive a simple and computationally efficient selection criterion based on the relative loss values of two auxiliary models. In addition to the modeling rationale, we evaluate CoLoR-Filter empirically on two language modeling tasks: (1) selecting data from C4 for domain adaptation to evaluation on Books and (2) selecting data from C4 for a suite of downstream multiple-choice question answering tasks. We demonstrate favorable scaling both as we subselect more aggressively and using small auxiliary models to select data for large target models. As one headline result, CoLoR-Filter data selected using a pair of 150m parameter auxiliary models can train a 1. 2b parameter target model to match a 1. 2b parameter model trained on 25b randomly selected tokens with 25x less data for Books and 11x less data for the downstream tasks. Code: https: //github. com/davidbrandfonbrener/color-filter-olmoFiltered data: https: //huggingface. co/datasets/davidbrandfonbrener/color-filtered-c4

NeurIPS Conference 2024 Conference Paper

DataComp-LM: In search of the next generation of training sets for language models

  • Jeffrey Li
  • Alex Fang
  • Georgios Smyrnis
  • Maor Ivgi
  • Matt Jordan
  • Samir Gadre
  • Hritik Bansal
  • Etash Guha

We introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad suite of 53 downstream evaluations. Participants in the DCLM benchmark can experiment with data curation strategies such as deduplication, filtering, and data mixing atmodel scales ranging from 412M to 7B parameters. As a baseline for DCLM, we conduct extensive experiments and find that model-based filtering is key to assembling a high-quality training set. The resulting dataset, DCLM-Baseline, enables training a 7B parameter language model from scratch to 63% 5-shot accuracy on MMLU with 2T training tokens. Compared to MAP-Neo, the previous state-of-the-art in open-data language models, DCLM-Baseline represents a 6 percentage point improvement on MMLU while being trained with half the compute. Our results highlight the importance of dataset design for training language models and offer a starting point for further research on data curation. We release the \dclm benchmark, framework, models, and datasets at https: //www. datacomp. ai/dclm/

TMLR Journal 2023 Journal Article

Exploring Transformer Backbones for Heterogeneous Treatment Effect Estimation

  • Yifan Zhang
  • Hanlin Zhang
  • Zachary Chase Lipton
  • Li Erran Li
  • Eric Xing

Previous works on Treatment Effect Estimation (TEE) are not in widespread use because they are predominantly theoretical, where strong parametric assumptions are made but untractable for practical application. Recent works use Multilayer Perceptron (MLP) for modeling casual relationships, however, MLPs lag far behind recent advances in ML methodology, which limits their applicability and generalizability. To extend beyond the single domain formulation and towards more realistic learning scenarios, we explore model design spaces beyond MLPs, i.e., transformer backbones, which provide flexibility where attention layers govern interactions among treatments and covariates to exploit structural similarities of potential outcomes for confounding control. Through careful model design, Transformers as Treatment Effect Estimators (TransTEE) is proposed. We show empirically that TransTEE can: (1) serve as a general-purpose treatment effect estimator which significantly outperforms competitive baselines on a variety of challenging TEE problems (e.g., discrete, continuous, structured, or dosage-associated treatments.) and is applicable to both when covariates are tabular and when they consist of structural data (e.g., texts, graphs); (2) yield multiple advantages: compatibility with propensity score modeling, parameter efficiency, robustness to continuous treatment value distribution shifts, explainable in covariate adjustment, and real-world utility in auditing pre-trained language models.

NeurIPS Conference 2020 Conference Paper

Towards Interpretable Natural Language Understanding with Explanations as Latent Variables

  • Wangchunshu Zhou
  • Jinyi Hu
  • Hanlin Zhang
  • Xiaodan Liang
  • Maosong Sun
  • Chenyan Xiong
  • Jian Tang

Recently generating natural language explanations has shown very promising results in not only offering interpretable explanations but also providing additional information and supervision for prediction. However, existing approaches usually require a large set of human annotated explanations for training while collecting a large set of explanations is not only time consuming but also expensive. In this paper, we develop a general framework for interpretable natural language understanding that requires only a small set of human annotated explanations for training. Our framework treats natural language explanations as latent variables that model the underlying reasoning process of a neural model. We develop a variational EM framework for optimization where an explanation generation module and an explanation-augmented prediction module are alternatively optimized and mutually enhance each other. Moreover, we further propose an explanation-based self-training method under this framework for semi-supervised learning. It alternates between assigning pseudo-labels to unlabeled data and generating new explanations to iteratively improve each other. Experiments on two natural language understanding tasks demonstrate that our framework can not only make effective predictions in both supervised and semi-supervised settings, but is also able to generate good natural language explanations.

v2026.09.13