Arrow Research search

Author name cluster

Yukang Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

AIIM Journal 2025 Journal Article

Anatomical prior-based vertebral landmark detection for spinal disorder diagnosis

  • Yukang Yang
  • Yu Wang
  • Tianyu Liu
  • Miao Wang
  • Ming Sun
  • Shiji Song
  • Wenhui Fan
  • Gao Huang

As one of fundamental ways to interpret spine images, detection of vertebral landmarks is an informative prerequisite for further diagnosis and management of spine disorders such as scoliosis and fractures. Most existing machine learning-based methods for automatic vertebral landmark detection suffer from overlapping landmarks or abnormally long distances between nearby landmarks against anatomical priors, and thus lack sufficient reliability and interpretability. To tackle the problem, this paper systematically utilizes anatomical prior knowledge in vertebral landmark detection. We explicitly formulate anatomical priors of the spine, related to distances among vertebrae and spatial order within the spine, and integrate these geometrical constraints within training loss, inference procedure, and evaluation metrics. First, we introduce an anatomy-constraint loss to regularize the training process with the aforementioned contextual priors explicitly. Second, we propose a simple-yet-effective anatomy-aided inference procedure by employing sequential prediction rather than a parallel counterpart. Third, we provide novel anatomy-related metrics to quantitatively evaluate to which extent landmark predictions follow the anatomical priors, as is not reflected within the widely-used landmark localization error metric. We employ the localization framework on 1410 anterior–posterior radiographic images. Compared with competitive baseline models, we achieve superior landmark localization accuracy and comparable Cobb angle estimation for scoliosis assessment. Ablation studies demonstrate the effectiveness of designed components on the decrease of localization error and improvement of anatomical plausibility. Additionally, we exhibit effective generalization performance by transferring our detection method onto sagittal 2-D slices of CT scans and boost the performance of downstream compression fracture classification at vertebra-level.

NeurIPS Conference 2025 Conference Paper

Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers

  • Andrew Nam
  • Henry Conklin
  • Yukang Yang
  • Tom Griffiths
  • Jonathan D Cohen
  • Sarah-Jane Leslie

We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns them a causal taxonomy—facilitating, interfering, or irrelevant—based on their impact on task performance. Unlike prior approaches in mechanistic interpretability, which are hypothesis-driven and require prompt templates or target labels, CHG applies directly to any dataset using standard next-token prediction. We evaluate CHG across multiple large language models (LLMs) in the Llama 3 model family and diverse tasks, including syntax, commonsense, and mathematical reasoning, and show that CHG scores yield causal, not merely correlational, insight validated via ablation and causal mediation analyses. We also introduce contrastive CHG, a variant that isolates sub-circuits for specific task components. Our findings reveal that LLMs contain multiple sparse task-sufficient sub-circuits, that individual head roles depend on interactions with others (low modularity), and that instruction following and in-context learning rely on separable mechanisms.

ICML Conference 2025 Conference Paper

Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models

  • Yukang Yang
  • Declan Campbell
  • Kaixuan Huang
  • Mengdi Wang 0001
  • Jonathan D. Cohen 0003
  • Taylor Whittington Webb

Many recent studies have found evidence for emergent reasoning capabilities in large language models (LLMs), but debate persists concerning the robustness of these capabilities, and the extent to which they depend on structured reasoning mechanisms. To shed light on these issues, we study the internal mechanisms that support abstract reasoning in LLMs. We identify an emergent symbolic architecture that implements abstract reasoning via a series of three computations. In early layers, symbol abstraction heads convert input tokens to abstract variables based on the relations between those tokens. In intermediate layers, symbolic induction heads perform sequence induction over these abstract variables. Finally, in later layers, retrieval heads predict the next token by retrieving the value associated with the predicted abstract variable. These results point toward a resolution of the longstanding debate between symbolic and neural network approaches, suggesting that emergent reasoning in neural networks depends on the emergence of symbolic mechanisms.

NeurIPS Conference 2025 Conference Paper

Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models

  • Yingqing Guo
  • Yukang Yang
  • Hui Yuan
  • Mengdi Wang

Training-free guidance enables controlled generation in diffusion and flow models, but most methods rely on gradients and assume differentiable objectives. This work focuses on training-free guidance addressing challenges from non-differentiable objectives and discrete data distributions. We propose TreeG: Tree Search-Based Path Steering Guidance, applicable to both continuous and discrete settings in diffusion and flow models. TreeG offers a unified framework for training-free guidance by proposing, evaluating, and selecting candidates at each step, enhanced with tree search over active paths and parallel exploration. We comprehensively investigate the design space of TreeG over the candidate proposal module and the evaluation function, instantiating TreeG into three novel algorithms. Our experiments show that TreeG consistently outperforms top guidance baselines in symbolic music generation, small molecule design, and enhancer DNA design with improvements of 29. 01%, 26. 38%, and 18. 43%. Additionally, we identify an inference-time scaling law showing TreeG's scalability in inference-time computation.

NeurIPS Conference 2024 Conference Paper

Gradient Guidance for Diffusion Models: An Optimization Perspective

  • Yingqing Guo
  • Hui Yuan
  • Yukang Yang
  • Minshuo Chen
  • Mengdi Wang

Diffusion models have demonstrated empirical successes in various applications and can be adapted to task-specific needs via guidance. This paper studies a form of gradient guidance for adapting a pre-trained diffusion model towards optimizing user-specified objectives. We establish a mathematical framework for guided diffusion to systematically study its optimization theory and algorithmic design. Our theoretical analysis spots a strong link between guided diffusion models and optimization: gradient-guided diffusion models are essentially sampling solutions to a regularized optimization problem, where the regularization is imposed by the pre-training data. As for guidance design, directly bringing in the gradient of an external objective function as guidance would jeopardize the structure in generated samples. We investigate a modified form of gradient guidance based on a forward prediction loss, which leverages the information in pre-trained score functions and provably preserves the latent structure. We further consider an iteratively fine-tuned version of gradient-guided diffusion where guidance and score network are both updated with newly generated samples. This process mimics a first-order optimization iteration in expectation, for which we proved $\tilde{\mathcal{O}}(1/K)$ convergence rate to the global optimum when the objective function is concave. Our code is released at https: //github. com/yukang123/GGDMOptim. git.

NeurIPS Conference 2023 Conference Paper

GlyphControl: Glyph Conditional Control for Visual Text Generation

  • Yukang Yang
  • Dongnan Gui
  • Yuhui Yuan
  • Weicong Liang
  • Haisong Ding
  • Han Hu
  • Kai Chen

Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method.

NeurIPS Conference 2023 Conference Paper

Rank-DETR for High Quality Object Detection

  • Yifan Pu
  • Weicong Liang
  • Yiduo Hao
  • Yuhui Yuan
  • Yukang Yang
  • Chao Zhang
  • Han Hu
  • Gao Huang

Modern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e. g. , H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https: //github. com/LeapLabTHU/Rank-DETR}.

AIIM Journal 2022 Journal Article

A multi-scale keypoint estimation network with self-supervision for spinal curvature assessment of idiopathic scoliosis from the imperfect dataset

  • Tianyu Liu
  • Yu Wang
  • Yukang Yang
  • Ming Sun
  • Wenhui Fan
  • Cody Bunger
  • Cheng Wu

Idiopathic scoliosis (IS) is a common lifetime disease, which exhibits an obvious deformity of spinal curvature to seriously affect heart and lung function. Accurate radiographic assessment of spinal curvature is vitally important for the clinical diagnosis and treatment planning of idiopathic scoliosis. Deep learning algorithms have been widely adopted to the medical image analysis with the remarkable advancement in computer vision. The automated methods can improve the efficiency of clinical diagnosis to relieve the burden of doctors, which have advantage in dealing with the tedious and repetitive tasks. However, existing methods usually require sufficiently large training datasets with strict annotation, which are costly and laborious especially for medical images. Moreover, the medical images of serious IS always contain the blurry and occlusive parts, which would make the accurate and robust estimation of the spinal curvature more difficult. In this paper, a dot annotation approach is presented to train the spinal curvature assessment model, rather than using strict annotation of IS X-ray images. We develop a multi-scale keypoint estimation network to reduce the requirement for large training datasets, in which the Squeeze-and-Excitation (SE) blocks are incorporated to improve the representational capacity of the model. Then, a self-supervision module is designed to alleviate the blurry and occlusive problem, and we use the two-view radiographic assessments of IS to generate a 3D spinal curvature. Finally, extensive experiments are conducted on a collected clinical dataset, in which we obtain 81. 5 AP and the average E d between the predicted keypoints and the ground truths is 0. 43, making an improvement over the mainstream approaches.

v2026.09.13