Arrow Research search

Author name cluster

Yangming Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
2 author rows

Possible papers

22

ICML Conference 2025 Conference Paper

Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks

  • Yixin Cheng
  • Hongcheng Guo
  • Yangming Li
  • Leonid Sigal

Text watermarking aims to subtly embeds statistical signals into text by controlling the Large Language Model (LLM)’s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The robustness of these watermarking algorithms has become a key factor in evaluating their effectiveness. Current text watermarking algorithms embed watermarks in high-entropy tokens to ensure text quality. In this paper, we reveal that this seemingly benign design can be exploited by attackers, posing a significant risk to the robustness of the watermark. We introduce a generic efficient paraphrasing attack, the Self-Information Rewrite Attack (SIRA), which leverages the vulnerability by calculating the self-information of each token to identify potential pattern tokens and perform targeted attack. Our work exposes a widely prevalent vulnerability in current watermarking algorithms. The experimental results show SIRA achieves nearly 100% attack success rates on seven recent watermarking methods with only $0. 88 per million tokens cost. Our approach does not require any access to the watermark algorithms or the watermarked LLM and can seamlessly transfer to any LLM as the attack model even mobile-level models. Our findings highlight the urgent need for more robust watermarking.

ICLR Conference 2025 Conference Paper

Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples

  • Yangming Li
  • Max Ruiz Luyten
  • Mihaela van der Schaar

Diffusion models are mainly studied on image data. However, non-image data (e.g., tabular data) are also prevalent in real applications and tend to be noisy due to some inevitable factors in the stage of data collection, degrading the generation quality of diffusion models. In this paper, we consider a novel problem setting where every collected sample is paired with a vector indicating the data quality: risk vector. This setting applies to many scenarios involving noisy data and we propose risk-sensitive SDE, a type of stochastic differential equation (SDE) parameterized by the risk vector, to address it. With some proper coefficients, risk-sensitive SDE can minimize the negative effect of noisy samples on the optimization of diffusion models. We conduct systematic studies for both Gaussian and non-Gaussian noise distributions, providing analytical forms of risk-sensitive SDE. To verify the effectiveness of our method, we have conducted extensive experiments on multiple tabular and time-series datasets, showing that risk-sensitive SDE permits a robust optimization of diffusion models with noisy samples and significantly outperforms previous baselines.

ICLR Conference 2024 Conference Paper

On Error Propagation of Diffusion Models

  • Yangming Li
  • Mihaela van der Schaar

Although diffusion models (DMs) have shown promising performances in a number of tasks (e.g., speech synthesis and image generation), they might suffer from error propagation because of their sequential structure. However, this is not certain because some sequential models, such as Conditional Random Field (CRF), are free from this problem. To address this issue, we develop a theoretical framework to mathematically formulate error propagation in the architecture of DMs, The framework contains three elements, including modular error, cumulative error, and propagation equation. The modular and cumulative errors are related by the equation, which interprets that DMs are indeed affected by error propagation. Our theoretical study also suggests that the cumulative error is closely related to the generation quality of DMs. Based on this finding, we apply the cumulative error as a regularization term to reduce error propagation. Because the term is computationally intractable, we derive its upper bound and design a bootstrap algorithm to efficiently estimate the bound for optimization. We have conducted extensive experiments on multiple image datasets, showing that our proposed regularization reduces error propagation, significantly improves vanilla DMs, and outperforms previous baselines.

ICLR Conference 2024 Conference Paper

Soft Mixture Denoising: Beyond the Expressive Bottleneck of Diffusion Models

  • Yangming Li
  • Boris van Breugel
  • Mihaela van der Schaar

Because diffusion models have shown impressive performances in a number of tasks, such as image synthesis, there is a trend in recent works to prove (with certain assumptions) that these models have strong approximation capabilities. In this paper, we show that current diffusion models actually have an expressive bottleneck in backward denoising and some assumption made by existing theoretical guarantees is too strong. Based on this finding, we prove that diffusion models have unbounded errors in both local and global denoising. In light of our theoretical studies, we introduce soft mixture denoising (SMD), an expressive and efficient model for backward denoising. SMD not only permits diffusion models to well approximate any Gaussian mixture distributions in theory, but also is simple and efficient for implementation. Our experiments on multiple image datasets show that SMD significantly improves different types of diffusion models (e.g., DDPM), espeically in the situation of few backward iterations.

ICLR Conference 2021 Conference Paper

Empirical Analysis of Unlabeled Entity Problem in Named Entity Recognition

  • Yangming Li
  • Lemao Liu
  • Shuming Shi 0001

In many scenarios, named entity recognition (NER) models severely suffer from unlabeled entity problem, where the entities of a sentence may not be fully annotated. Through empirical studies performed on synthetic datasets, we find two causes of performance degradation. One is the reduction of annotated entities and the other is treating unlabeled entities as negative instances. The first cause has less impact than the second one and can be mitigated by adopting pretraining language models. The second cause seriously misguides a model in training and greatly affects its performances. Based on the above observations, we propose a general approach, which can almost eliminate the misguidance brought by unlabeled entities. The key idea is to use negative sampling that, to a large extent, avoids training NER models with unlabeled entities. Experiments on synthetic datasets and real-world datasets show that our model is robust to unlabeled entity problem and surpasses prior baselines. On well-annotated datasets, our model is competitive with the state-of-the-art method.

AAAI Conference 2021 Conference Paper

Interpretable NLG for Task-oriented Dialogue Systems with Heterogeneous Rendering Machines

  • Yangming Li
  • Kaisheng Yao

End-to-end neural networks have achieved promising performances in natural language generation (NLG). However, they are treated as black boxes and lack interpretability. To address this problem, we propose a novel framework, heterogeneous rendering machines (HRM), that interprets how neural generators render an input dialogue act (DA) into an utterance. HRM consists of a renderer set and a mode switcher. The renderer set contains multiple decoders that vary in both structure and functionality. For every generation step, the mode switcher selects an appropriate decoder from the renderer set to generate an item (a word or a phrase). To verify the effectiveness of our method, we have conducted extensive experiments on 5 benchmark datasets. In terms of automatic metrics (e. g. , BLEU), our model is competitive with the current state-of-the-art method. The qualitative analysis shows that our model can interpret the rendering process of neural generators well. Human evaluation also confirms the interpretability of our proposed approach.

ICRA Conference 2021 Conference Paper

Learning Surgical Motion Pattern from Small Data in Endoscopic Sinus and Skull Base Surgeries

  • Yangming Li
  • Randall A. Bly
  • Sarah Akkina
  • Fangbo Qin
  • Rajeev C. Saxena
  • Ian Humphreys
  • Mark Whipple
  • Kris S. Moe

Existing studies demonstrated that surgical motion patterns are strongly correlated with surgical outcomes. Real surgeries are complicated and it is expensive to harvest surgical data. Consequently, existing researches on surgical motion patterns focus on specific concise surgical tasks or simple surgical procedures. The paper presents a surgical motion pattern modeling technique that uses small data but can be applied to virtually any Endoscopic Sinus and Skull Base Surgeries (ESSBSs). The proposed method decreases the dimensionalities of the feature space through projecting surgical instrument motions into the endoscope coordinate, based on human expert domain knowledge. Furthermore, the method uses kinematic features and learns the motion pattern with Gaussian Process learning techniques. Comparing with existing surgical motion pattern modeling methods, the proposed method: 1, learns the motion model from small data; 2, can be generally applied to ESSBSs because it neither assumes nor depends on specific surgical tasks; 3, provides informative results in a real-time manner for optimizing surgical motions for improving surgical outcomes. The proposed method was verified by predicting surgical skill levels on cadaver surgeries. The results show the real-time prediction precision is higher than 81% and the offline accumulated precision reach 100%.

AAAI Conference 2020 Conference Paper

DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment Classification

  • Libo Qin
  • Wanxiang Che
  • Yangming Li
  • Mingheng Ni
  • Ting Liu

In dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers’ intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately (Kim and Kim 2018). Most of the existing systems either treat them as separate tasks or just jointly model the two tasks by sharing parameters in an implicit way without explicitly modeling mutual interaction and relation. To address this problem, we propose a Deep Co-Interactive Relation Network (DCR-Net) to explicitly consider the cross-impact and model the interaction between the two tasks by introducing a co-interactive relation layer. In addition, the proposed relation layer can be stacked to gradually capture mutual knowledge with multiple steps of interaction. Especially, we thoroughly study different relation layers and their effects. Experimental results on two public datasets (Mastodon and Dailydialog) show that our model outperforms the state-of-the-art joint model by 4. 3% and 3. 4% in terms of F1 score on dialog act recognition task, 5. 7% and 12. 4% on sentiment classification respectively. Comprehensive analysis empirically verifies the effectiveness of explicitly modeling the relation between the two tasks and the multi-steps interaction mechanism. Finally, we employ the Bidirectional Encoder Representation from Transformer (BERT) in our framework, which can further boost our performance in both tasks.

IROS Conference 2020 Conference Paper

LC-GAN: Image-to-image Translation Based on Generative Adversarial Network for Endoscopic Images

  • Shan Lin
  • Fangbo Qin
  • Yangming Li
  • Randall A. Bly
  • Kris S. Moe
  • Blake Hannaford

Intelligent vision is appealing in computer-assisted and robotic surgeries. Vision-based analysis with deep learning usually requires large labeled datasets, but manual data labeling is expensive and time-consuming in medical problems. We investigate a novel cross-domain strategy to reduce the need for manual data labeling by proposing an image-to-image translation model live-cadaver GAN (LC-GAN) based on generative adversarial networks (GANs). We consider a situation when a labeled cadaveric surgery dataset is available while the task is instrument segmentation on an unlabeled live surgery dataset. We train LC-GAN to learn the mappings between the cadaveric and live images. For live image segmentation, we first translate the live images to fake-cadaveric images with LC-GAN and then perform segmentation on the fake-cadaveric images with models trained on the real cadaveric dataset. The proposed method fully makes use of the labeled cadaveric dataset for live image segmentation without the need to label the live dataset. LC-GAN has two generators with different architectures that leverage the deep feature representation learned from the cadaveric image based segmentation task. Moreover, we propose the structural similarity loss and segmentation consistency loss to improve the semantic consistency during translation. Our model achieves better image-to-image translation and leads to improved segmentation performance in the proposed cross-domain segmentation task.

AAAI Conference 2020 Conference Paper

Span-Based Neural Buffer: Towards Efficient and Effective Utilization of Long-Distance Context for Neural Sequence Models

  • Yangming Li
  • Kaisheng Yao
  • Libo Qin
  • Shuang Peng
  • Yijia Liu
  • Xiaolong Li

Neural sequence model, though widely used for modeling sequential data such as the language model, has sequential recency bias (Kuncoro et al. 2018) to the local context, limiting its full potential to capture long-distance context. To address this problem, this paper proposes augmenting sequence models with a span-based neural buffer that efficiently represents long-distance context, allowing a gate policy network to make interpolated predictions from both the neural buffer and the underlying sequence model. Training this policy network to utilize long-distance context is however challenging due to the simple sentence dominance problem (Marvin and Linzen 2018). To alleviate this problem, we propose a novel training algorithm that combines an annealed maximum likelihood estimation with an intrinsic reward-driven reinforcement learning. Sequence models with the proposed span-based neural buffer significantly improve the state-of-the-art perplexities on the benchmark Penn Treebank and WikiText-2 datasets to 43. 9 and 35. 2 respectively. We conduct extensive analysis and confirm that the proposed architecture and the training algorithm both contribute to the improvements.

ICRA Conference 2019 Conference Paper

Surgical Instrument Segmentation for Endoscopic Vision with Data Fusion of rediction and Kinematic Pose

  • Fangbo Qin
  • Yangming Li
  • Yun-Hsuan Su
  • De Xu
  • Blake Hannaford

The real-time and robust surgical instrument segmentation is an important issue for endoscopic vision. We propose an instrument segmentation method fusing the convolutional neural networks (CNN) prediction and the kinematic pose information. First, the CNN model ToolNet-C is designed, which cascades a convolutional feature extractor trained over numerous unlabeled images and a pixel-wise segmentor trained on few labeled images. Second, the silhouette projection of the instrument body onto the endoscopic image is implemented based on the measured kinematic pose. Third, the particle filter with the shape matching likelihood and the weight suppression is proposed for data fusion, whose estimate refines the kinematic pose. The refined pose determines an accurate silhouette mask, which is the final segmentation output. The experiments are conducted with a surgical navigation system, several animal-tissue backgrounds, and a debrider instrument.

ICRA Conference 2018 Conference Paper

A Novel Recurrent Neural Network for Improving Redundant Manipulator Motion Planning Completeness

  • Yangming Li
  • Shuai Li 0002
  • Blake Hannaford

Recurrent Neural Networks (RNNs) demonstrated advantages on control precision, system robustness and computational efficiency, and have been widely applied to redundant manipulator control optimization. Existing RNN control schemes locally optimize trajectories and are efficient and reliable on obstacle avoidance. However, for motion planning, they suffer from local minimum and do not have planning completeness. This work explained the cause of the planning incompleteness and addressed the problem with a novel RNN control scheme. The paper presented the proposed method in detail and analyzed the global stability and the planning completeness in theory. The proposed method was compared with other three control schemes on the precision, the robustness and the planning completeness in software simulation and the results shows the proposed method has improved precision and robustness, and planning completeness.

IROS Conference 2018 Conference Paper

Soft-obstacle Avoidance for Redundant Manipulators with Recurrent Neural Network

  • Yangming Li
  • Blake Hannaford

Compressing soft-obstacles secondary to a controlled motion task is common for human beings. While these tasks are nearly trivial for teleoperated robots, they remain a challenging problem in robotic autonomy. Addressing the problem is significant. For example, in Minimally Invasive Surgeries (MISs), safely compressing soft tissues ensures the surgical safety and decreases tissue removal, thus dramatically decreases surgical trauma and operating room time, and leads to improved surgical outcomes. In this work, we define the problem of soft-obstacle avoidance and project the safety motion constraints into the task space and the velocity space. We illustrate the significance of addressing this problem in the robotic surgery scenario. We present a Recurrent Neural Networks (RNNs) based solution, which formulates the problem as an inequality constrained optimization problem and solves it in its dual space. The application of the proposed method was demonstrated in the Raven II surgical robot. Experimental results demonstrated that the proposed method is effective in addressing the soft-obstacle avoidance problem.

IROS Conference 2017 Conference Paper

Improving control precision and motion adaptiveness for surgical robot with recurrent neural network

  • Yangming Li
  • Shuai Li 0002
  • David E. Caballero
  • Muneaki Miyasaka
  • Andrew Lewis 0001
  • Blake Hannaford

Surgical robot research is driven by the desire of improving surgical outcomes. This paper proposed a Recurrent Neural Network based controller to address two problems: 1) improving control precision, 2) increasing adaptiveness for robot motion (explained in Section I). RNN was adopted in this work mainly because 1) the problem formulation naturally matches RNN structure, 2) RNN has advantages as an biologically inspired method. The proposed method was explained in detail and analysis shows that the proposed method is able to dynamically regulate outputs to increase the adaptiveness and the control precision. This paper uses Raven II surgical robot as an example to show the application of the proposed method, and the numeral simulation results from the proposed method and three other controllers show that the proposed method has improved precision, improved high robustness against noise and increased movement smoothness, and it keeps the manipulator links as far away as possible from physical boundaries, which potentially increases surgical safety and leads to improved surgical outcomes.

ICRA Conference 2017 Conference Paper

Roboscope: A flexible and bendable surgical robot for single portal Minimally Invasive Surgery

  • Jacob Rosen 0001
  • Laligam N. Sekhar
  • Daniel Glozman
  • Muneaki Miyasaka
  • Jesse Dosher
  • Brian Dellon
  • Kris S. Moe
  • Aylin Kim

Minimally Invasive Surgery (MIS) can reduce iatrogenic injury and decrease the possibility of surgical complications. This paper presents a novel flexible and bendable endoscopic device, “Roboscope”, which delivers two instruments, two miniature scanning fiber endoscopes, and a suction/irrigation port to the operation site through a single portal. Compared with existing bendable and steerable robotic surgical systems, Roboscope provides two bending degrees of freedom for its outer sheath and two insertion degrees of freedom, while simultaneously delivering two instruments and two endoscopes to the surgical site. Each bending axis and insertion freedom of Roboscope is independently controllable via an external actuation pack. Surgical tools can be changed without retracting the robot arm. This paper presents the design of the Roboscope mechanical system, electrical system, and control and software systems, design requirements and prototyping validation as well as analysis of Roboscope workspece.

ICRA Conference 2016 Conference Paper

Dynamic modeling of cable driven elongated surgical instruments for sensorless grip force estimation

  • Yangming Li
  • Muneaki Miyasaka
  • Mohammad Haghighipanah
  • Lei Cheng
  • Blake Hannaford

Haptic feedback plays a key role in surgeries, but it is still a missing component in robotic Minimally Invasive Surgeries. This paper proposes a dynamic model-based sensorless grip force estimation method to address the haptic perception problem for commonly used elongated cable-driven surgical instruments. Cable and cable-pulley properties are studied for dynamic modeling; grip forces, along with driven motor and gripper jaw positions and velocities are jointly estimated with Unscented Kalman Filter and only motor encoder readings and motor output torques are assumed to be known. A bounding filter is used to compensate for model inaccuracy and to improve method robustness. The proposed method was validated on a 10mm gripper which is driven by a Raven-II surgical robot. The gripper was equipped with 1-dimensional force sensors which served as ground truth data. The experimental results showed that the proposed method provides sufficiently good grip force estimation, while only motor encoder and the motor torques are used as observations.

ICRA Conference 2016 Conference Paper

Hysteresis model of longitudinally loaded cable for cable driven robots and identification of the parameters

  • Muneaki Miyasaka
  • Mohammad Haghighipanah
  • Yangming Li
  • Blake Hannaford

In this paper, we propose model of longitudinally loaded cable based on the Bouc-Wen hysteresis model and within the framework of the Duhem operator. By optimizing the 9 hysteresis model parameters with a genetic algorithm, the proposed model is shown to be capable of representing quasi-static response of two different diameter cables, 0. 61 mm (thin) and 1. 19 mm (thick), used for the RAVEN II surgical robotic surgery platform. The construction of the cable is 7 strands with 19 individual wires per strand. Furthermore, it is shown that the dynamic response of the cables are captured by adding a linear damping term. The hysteresis model and linear damper with the optimized parameters accurately models a longitudinal vibration test result in terms of frequency, steady state stretch, and logarithmic decrement. Energy dissipation due solely to the hysteresis term is approximately calculated to be 57 and 71% of the total energy loss for the thin and thick cables respectively. The proposed model may be used for cables with different contraction and diameter and can be applied for control of cable driven robots in which cables are stretched longitudinally without large excitation of other modes.

ICRA Conference 2016 Conference Paper

Unscented Kalman Filter and 3D vision to improve cable driven surgical robot joint angle estimation

  • Mohammad Haghighipanah
  • Muneaki Miyasaka
  • Yangming Li
  • Blake Hannaford

Cable driven manipulators are popular in surgical robots due to compact design, low inertia, and remote actuation. In these manipulators, encoders are usually mounted on the motor, and joint angles are estimated based on transmission kinematics. However, due to non-linear properties of cables such as cable stretch, lower stiffness, and uncertainties in kinematic model parameters, the precision of joint angle estimation is limited with transmission kinematics approach. To improve the positioning of these manipulators, we use a pair of low cost stereo camera as the observation for joint angles and we input these noisy measurements into an Unscented Kalman Filter (UKF) for state estimation. We use the dual UKF to estimate cable parameters and states offline. We evaluated the effectiveness of the proposed method on a Raven-II experimental surgical research platform. Additional encoders at the joint output were employed as a reference system. From the experiments, the UKF improved the accuracy of joint angle estimation by 33– 72%. Also, we tested the reliability of state estimation under camera occlusion. We found that when the system dynamics is tuned with offline UKF parameter estimation, the camera occlusion has no effect on the online state estimation.

IROS Conference 2015 Conference Paper

Improving position precision of a servo-controlled elastic cable driven surgical robot using Unscented Kalman Filter

  • Mohammad Haghighipanah
  • Yangming Li
  • Muneaki Miyasaka
  • Blake Hannaford

Cable driven power transmission is popular in many manipulator applications including medical arms. In spite of advantages obtained by removing motors from the mechanism, cable transmission introduces higher non-linearity and more uncertainties such as cable stretch and cable coupling. In order to improve the control precision and robustness of the Raven-II surgical robot, particularly for automation applications, the Unscented Kalman Filter (UKF) was adopted for state estimation. The UKF estimated state variables of the Raven-II dynamic model from sensor data. The dual UKF was used offline to estimate cable coupling parameters. The experimental results showed that the proposed method improved joint position estimation precision and the estimation consistency, especially on the more elastic links. The improvements for links 2 and 3 of the Raven were 36. 76%, and 62. 99%, respectively. For link 1 the improvement was 1. 43% because the transmission is very stiff.

IROS Conference 2012 Conference Paper

IPJC: The Incremental Posterior Joint Compatibility test for fast feature cloud matching

  • Yangming Li
  • Edwin Olson

One of the fundamental challenges in robotics is data-association: determining which sensor observations correspond to the same physical object. A common approach is to consider groups of observations simultaneously: a constellation of observations can be significantly less ambiguous than the observations considered individually. The Joint Compatibility Branch and Bound (JCBB) test is the gold standard method for these data association problems. But its computational complexity and its sensitivity to non-linearities limit its practical usefulness. We propose the Incremental Posterior Joint Compatibility (IPJC) test. While equivalent to JCBB on linear problems, it is significantly more accurate on non-linear problems. When used for feature-cloud matching (an important special case), IPJC is also dramatically faster than JCBB. We demonstrate the advantages of IPJC over JCBB and other commonly-used methods on both synthetic and real-world datasets.

ICRA Conference 2011 Conference Paper

Structure tensors for general purpose LIDAR feature extraction

  • Yangming Li
  • Edwin Olson

The detection of features from Light Detection and Ranging (LIDAR) data is a fundamental component of feature-based mapping and SLAM systems. Classical approaches are often tied to specific environments, computationally expensive, or do not extract precise features. We describe a general purpose feature detector that is not only efficient, but also applicable to virtually any environment. Our method shares its mathematical foundation with feature detectors from the computer vision community, where structure tensor based methods have been successful. Our resulting method is capable of identifying stable and repeatable features at a variety of spatial scales, and produces uncertainty estimates for use in a state estimation algorithm. We verify the proposed method on standard datasets, including the Victoria Park dataset and the Intel Research Center dataset.

ICRA Conference 2010 Conference Paper

Extracting general-purpose features from LIDAR data

  • Yangming Li
  • Edwin Olson

The detection of features from Light Detection and Ranging (LIDAR) data is a fundamental component of feature-based mapping and SLAM systems. Existing detectors tend to exploit characteristics of specific environments: corners and lines from indoor (rectilinear) environments, and trees from outdoor environments. While these detectors work well in their intended environments, their performance in different environments can be very poor. We describe a general purpose feature detector for LIDAR data that is applicable to virtually any environment. Our methods adapt classic feature detection methods from the image processing literature, specifically the multi-scale Kanade-Tomasi corner detector. Our resulting method is capable of identifying stable features at a variety of spatial scales and produces uncertainty estimates for use in a state estimation algorithm. We present results on standard datasets, including Victoria Park and Intel Research Center (both 2D), and the MIT DARPA Urban Challenge dataset (3D).

v2026.09.13