Arrow Research search

Author name cluster

Zhe Xu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

AAAI Conference 2026 Conference Paper

Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention

  • Zhe Xu
  • Zhicai Wang
  • Junkang Wu
  • Jinda Lu
  • Xiang Wang

Large Vision-Language Models (LVLMs) often suffer from object hallucination, making erroneous judgments about the presence of objects in images. We propose this primarily stems from spurious correlations arising when models strongly associate highly co-occurring objects during training, leading to hallucinated objects influenced by visual context. Current benchmarks mainly focus on hallucination detection but lack a formal characterization and quantitative evaluation of spurious correlations in LVLMs. To address this, we introduce causal analysis into the object recognition scenario of LVLMs, establishing a Structural Causal Model (SCM). Utilizing the language of causality, we formally define spurious correlations arising from co-occurrence bias. To quantify the influence induced by these spurious correlations, we develop Causal-HalBench, a benchmark specifically constructed with counterfactual samples and integrated with comprehensive causal metrics designed to assess model robustness against spurious correlations. Concurrently, we propose an extensible pipeline for the construction of these counterfactual samples, leveraging the capabilities of proprietary LVLMs and Text-to-Image (T2I) models for their generation. Our evaluations on mainstream LVLMs using Causal-HalBench demonstrate these models exhibit susceptibility to spurious correlations, albeit to varying extents.

AAAI Conference 2026 Conference Paper

Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided Search

  • Shuocheng Li
  • Yihao Liu
  • Silin Du
  • Wenxuan Zeng
  • Zhe Xu
  • Mengyu Zhou
  • Yeye He
  • Haoyu Dong

Large language models (LLMs) have shown great promise in automating data science workflows. However, existing models still struggle with multi-step reasoning and tool use, limiting their effectiveness on complex data analysis tasks. To address this limitation, we propose a scalable pipeline that extracts high-quality, tool-based data analysis tasks and their executable multi-step solutions from real-world Jupyter notebooks and associated data files. Using this pipeline, we introduce NbQA, a large-scale dataset of standardized task–solution pairs that reflect authentic tool-use patterns in practical data science scenarios. To further enhance the multi-step reasoning capabilities, we present Jupiter, a framework that formulates data analysis as a search problem and applies Monte Carlo Tree Search (MCTS) to generate diverse solution trajectories for value model learning. During inference, Jupiter combines the value model and node visit counts to efficiently collect executable multi-step plans with minimal search steps. Experimental results show that Qwen2.5-7B and 14B-Instruct models on NbQA solve 77.82% and 86.38% of tasks on InfiAgent-DABench, respectively—matching or surpassing GPT-4o and advanced agent frameworks. Further evaluations demonstrate improved generalization and stronger tool-use reasoning across diverse multi-step reasoning tasks.

IJCAI Conference 2025 Conference Paper

A Logic-Based Approach to Causal Discovery: Signal Temporal Logic Perspective

  • Nasim Baharisangari
  • Yucheng Ruan
  • Chengcheng Zhao
  • Zhe Xu

Causal discovery in time-series datasets is critical for understanding complex systems, especially when the \textit{effectiveness} of causal relationships depends on both the \textit{duration} and \textit{magnitude} of the cause. We introduce a novel framework for causal discovery based on \textbf{Signal Temporal Logic (STL)}, enabling the extraction of interpretable causal diagrams (STL-CD) that explicitly capture these temporal dynamics. Our method first identifies statistically meaningful time intervals, then infers STL formulas that classify system behaviors, and finally employs transfer entropy to determine direct causal relationships among the formulas. This approach not only uncovers causal structure but also identifies the temporal persistence required for causal influence—an insight missed by existing methods. Experimental results on synthetic and real-world datasets demonstrate that our method achieves superior structural accuracy over state-of-the-art baselines, providing more informative and temporally precise causal models.

AAAI Conference 2025 Conference Paper

Compress to One Point: Neural Collapse for Pre-Trained Model-Based Class-Incremental Learning

  • Kun Wei
  • Zhe Xu
  • Cheng Deng

Class-Incremental Learning (CIL) requires an artificial intelligence system to learn different tasks without class overlaps continually. To achieve CIL, some methods introduce the Pre-Trained Model (PTM) and leverage the generalized feature representation of PTM to learn downstream incremental tasks continually. However, the generalized feature representations of PTM are not adaptive and discriminative for these various incremental classes, which may be out of distribution for the pre-trained dataset. In addition, since the incremental classes cannot be learned at once, the class relationship cannot be constructed optimally, leading to undiscriminating feature representation for understream tasks. Thus, we propose a novel Pre-Trained Model-based Class-Incremental Learning (PTM-CIL) method to explore the potential of PTM and obtain optimal class relationships. Inspired by Neural Collapse theory, we introduce the frozen Equiangular Tight Frame classifier to construct optimal classifier structure for all seen classes, guiding the feature representation adaptation for downstream continual tasks. Specifically, Task-Related Adaptation is proposed to modulate the generalized feature representation to bridge the gap between the pre-trained dataset and various downstream datasets. Then, the Feature Compression Module is introduced to compress various features to the specific classifier weights, constructing the feature transfer pattern and satisfying the characteristic of Neural Collapse. Optimal Structural Alignment is designed to supervise the feature compression process, assisting in achieving optimal class relationships across different tasks. Sufficient experiments on seven datasets prove the effectiveness of our method.

NeurIPS Conference 2025 Conference Paper

MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?

  • Zhe Xu
  • Daoyuan Chen
  • Zhenqing Ling
  • Yaliang Li
  • Ying Shen

Large foundation models face challenges in acquiring transferable, structured thinking abilities, especially when supervised with rigid templates or crowd-annotated instruction datasets. Unlike prior approaches, we focus on a thinking-centric data synthesis paradigm that enables models to evolve through self-generated, cognitively guided data. We propose MindGYM, a structured and scalable framework for question synthesis, composed of: (1) Cognitive Thinking Process Injection, which infuses high-level reasoning objectives to shape the model’s synthesis behavior; (2) Seed Single-Hop Question Synthesis, generating atomic questions from diverse semantic types to encourage broader thinking; and (3) Challenging Multi-Hop QA Synthesis, composing more complex multi-hop questions based on QA seeds for deeper reasoning. Detailed analysis shows that synthetic data generated by our method achieves 16. 7% higher average quality and 67. 91% lower quality variance compared to baseline sources, highlighting that both high-quality and self-contained data are essential for effective, thinking-oriented fine-tuning. MindGYM improves performance on six reasoning benchmarks, achieving gains of up to 16% on MathVision using only 400 data samples, and generalizable improvements across different model sizes and architectures. MindGYM underscores the viability of self-challenging mechanisms in refining large model capabilities while minimizing human intervention and resource demands. Code and data are released to promote data-centric research into self-evolving foundation models driven by their internal reasoning capabilities.

JBHI Journal 2025 Journal Article

Seeking Common Ground While Reserving Differences: Multiple Anatomy Collaborative Framework for Undersampled MRI Reconstruction

  • Jiangpeng Yan
  • Chenghui Yu
  • Hanbo Chen
  • Zhe Xu
  • Junzhou Huang
  • Xiu Li
  • Jianhua Yao

Recently, deep neural networks have greatly advanced undersampled Magnetic Resonance Image (MRI) reconstruction, wherein most studies follow the one-anatomy-one-network fashion, i. e. , each expert network is trained and evaluated for a specific anatomy. Apart from inefficiency in training multiple independent models, such convention ignores the shared de-aliasing knowledge across various anatomies which can benefit each other. To explore the shared knowledge, one naive way is to combine all the data from various anatomies to train an all-round network. Unfortunately, despite the existence of the shared de-aliasing knowledge, we reveal that the exclusive knowledge across different anatomies can deteriorate specific reconstruction targets, yielding overall performance degradation. Observing this, in this study, we present a novel deep MRI reconstruction framework with both anatomy-shared and anatomy-specific parameterized learners, aiming to “seek common ground while reserving differences” across different anatomies. Particularly, the primary anatomy-shared learners are exposed to different anatomies to model rich shared de-aliasing knowledge, while the efficient anatomy-specific learners are trained with their target anatomy for exclusive knowledge. Four different implementations of anatomy-specific learners are presented and explored on the top of our framework in two MRI reconstruction networks. Comprehensive experiments on brain, knee and cardiac MRI datasets demonstrate that three of these learners are able to enhance reconstruction performance via multiple anatomy collaborative learning. Extensive studies show that our strategy can also benefit multiple pulse sequence MRI reconstruction by integrating sequence-specific learners.

ICML Conference 2025 Conference Paper

SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking Services

  • Hongcheng Guo
  • Yue Wang
  • Shaosheng Cao
  • Fei Zhao 0012
  • Boyang Wang 0006
  • Lei Li 0039
  • Liang Chen 0024
  • Xinze Lyu

With the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content generation, and communication dynamics. However, recent studies focus on isolated SNS tasks rather than a comprehensive evaluation. In this paper, we introduce SNS-Bench, specially constructed for assessing the abilities of large language models from different Social Networking Services, with a wide range of SNS-related information. SNS-Bench encompasses 8 different tasks such as note classification, query content relevance, and highlight words generation in comments. Finally, 6, 658 questions of social media text, including subjective questions, single-choice, and multiple-choice questions, are concluded in SNS-Bench. Further, we evaluate the performance of over 25+ current diverse LLMs on our SNS-Bench. Models with different sizes exhibit performance variations, yet adhere to the scaling law. Moreover, we hope provide more insights to revolutionize the techniques of social network services with LLMs.

IROS Conference 2025 Conference Paper

TwinTac: A Wide-Range, Highly Sensitive Tactile Sensor with Real-To-Sim Digital Twin Sensor Model

  • Xiyan Huang
  • Zhe Xu
  • Chenxi Xiao

Robot skill acquisition processes driven by reinforcement learning often rely on simulations to efficiently generate large-scale interaction data. However, the absence of simulation models for tactile sensors has hindered the use of tactile sensing in such skill learning processes, limiting the development of effective policies driven by tactile perception. To bridge this gap, we present TwinTac, a system that combines the design of a physical tactile sensor with its digital twin model. Our hardware sensor is designed for high sensitivity and a wide measurement range, enabling high quality sensing data essential for object interaction tasks. Building upon the hardware sensor, we develop the digital twin model using a real-to-sim approach. This involves collecting synchronized cross-domain data, including finite element method results and the physical sensor’s outputs, and then training neural networks to map simulated data to real sensor responses. Through experimental evaluation, we characterized the sensitivity of the physical sensor and demonstrated the consistency of the digital twin in replicating the physical sensor’s output. Furthermore, by conducting an object classification task, we showed that simulation data generated by our digital twin sensor can effectively augment real-world data, leading to improved accuracy. These results highlight TwinTac’s potential to bridge the gap in cross-domain learning tasks.

NeurIPS Conference 2024 Conference Paper

Discrete-state Continuous-time Diffusion for Graph Generation

  • Zhe Xu
  • Ruizhong Qiu
  • Yuzhong Chen
  • Huiyuan Chen
  • Xiran Fan
  • Menghai Pan
  • Zhichen Zeng
  • Mahashweta Das

Graph is a prevalent discrete data structure, whose generation has wide applications such as drug discovery and circuit design. Diffusion generative models, as an emerging research focus, have been applied to graph generation tasks. Overall, according to the space of states and time steps, diffusion generative models can be categorized into discrete-/continuous-state discrete-/continuous-time fashions. In this paper, we formulate the graph diffusion generation in a discrete-state continuous-time setting, which has never been studied in previous graph diffusion models. The rationale of such a formulation is to preserve the discrete nature of graph-structured data and meanwhile provide flexible sampling trade-offs between sample quality and efficiency. Analysis shows that our training objective is closely related to the generation quality and our proposed generation framework enjoys ideal invariant/equivariant properties concerning the permutation of node ordering. Our proposed model shows competitive empirical performance against other state-of-the-art graph generation solutions on various benchmarks while at the same time can flexibly trade off the generation quality and efficiency in the sampling phase.

ICRA Conference 2024 Conference Paper

Optimization Based Dynamic Skateboarding of Quadrupedal Robot

  • Zhe Xu
  • Mohamed Al-Khulaqui
  • Hanxin Ma
  • Jiajun Wang
  • Quanbin Xin
  • Yangwei You
  • Mingliang Zhou 0003
  • Diyun Xiang

Robot skateboarding is a novel and challenging task for legged robots. Accurately modeling the dynamics of dual floating bases and developing effective planning and control methods present significant complexities in accomplishing skateboarding behavior. This paper focuses on enabling the quadrupedal platform CyberDog2 to achieve dynamic balancing and acceleration on a skateboard. An optimization-based control pipeline is developed through careful derivation of the system’s equations of motion, considering both the robot and skateboard dynamics. By accounting for system physical constraints, an advanced offline trajectory optimization method is employed to generate various acceleration trajectories, creating a motion library for the system. An online linear model predictive control with whole body control framework is used to track the generated trajectories and stablize the system in real-time. To validate its effectiveness, we conducted experiments in various scenarios. The quadrupedal robot successfully performed acceleration from a static state to various velocities and demonstrated the ability to balance and steer the skateboard.

AAAI Conference 2023 Conference Paper

Learning Interpretable Temporal Properties from Positive Examples Only

  • Rajarshi Roy
  • Jean-Raphaël Gaglione
  • Nasim Baharisangari
  • Daniel Neider
  • Zhe Xu
  • Ufuk Topcu

We consider the problem of explaining the temporal behavior of black-box systems using human-interpretable models. Following recent research trends, we rely on the fundamental yet interpretable models of deterministic finite automata (DFAs) and linear temporal logic (LTL_f) formulas. In contrast to most existing works for learning DFAs and LTL_f formulas, we consider learning from only positive examples. Our motivation is that negative examples are generally difficult to observe, in particular, from black-box systems. To learn meaningful models from positive examples only, we design algorithms that rely on conciseness and language minimality of models as regularizers. Our learning algorithms are based on two approaches: a symbolic and a counterexample-guided one. The symbolic approach exploits an efficient encoding of language minimality as a constraint satisfaction problem, whereas the counterexample-guided one relies on generating suitable negative examples to guide the learning. Both approaches provide us with effective algorithms with minimality guarantees on the learned models. To assess the effectiveness of our algorithms, we evaluate them on a few practical case studies.

IROS Conference 2023 Conference Paper

Run and Catch: Dynamic Object-Catching of Quadrupedal Robots

  • Yangwei You
  • Tianlin Liu
  • Xiaowei Liang
  • Zhe Xu
  • Mingliang Zhou 0003
  • Zhibin Li 0001
  • Shiwu Zhang

Quadrupedal robots are performing increasingly more real-world capabilities, but are primarily limited to locomotion tasks. To expand their task-level abilities of object acquisition, i. e. , run-to-catch as frisbee catching for dogs, this paper developed a control pipeline using stereo vision for legged robots which allows for dynamic catching balls while the robot is in motion. To achieve high-frame-rate tracking, we designed a ball that can actively emit homogeneous infrared (IR) light and then located the flying ball based on binocular vision positioning using the onboard RealSense D450 camera with an additional IR bandpass filter. The camera was mounted on top of a 2-DoF head to gain a full view of the target ball. A state estimation module was developed to fuse the vision positioning, camera motor readings, localization result of RealSense T265 equipped on the back, and the legged odometry output altogether. With the use of a ballistic model, we achieved a robust estimation of both the ball and robot positions in an inertial coordinate. Additionally, we developed a close-loop catching strategy and employed trajectory prediction so that tracking and run-to-catch were performed simultaneously, which is critical for such drastically dynamic and precise tasks. The proposed approach was validated through both static testing and dynamic catch experiments conducted on the CyberDog robot with a high success rate.

JBHI Journal 2022 Journal Article

All-Around Real Label Supervision: Cyclic Prototype Consistency Learning for Semi-Supervised Medical Image Segmentation

  • Zhe Xu
  • Yixin Wang
  • Donghuan Lu
  • Lequan Yu
  • Jiangpeng Yan
  • Jie Luo
  • Kai Ma
  • Yefeng Zheng

Semi-supervised learning has substantially advanced medical image segmentation since it alleviates the heavy burden of acquiring the costly expert-examined annotations. Especially, the consistency-based approaches have attracted more attention for their superior performance, wherein the real labels are only utilized to supervise their paired images via supervised loss while the unlabeled images are exploited by enforcing the perturbation-based “unsupervised” consistency without explicit guidance from those real labels. However, intuitively, the expert-examined real labels contain more reliable supervision signals. Observing this, we ask an unexplored but interesting question: can we exploit the unlabeled data via explicit real label supervision for semi-supervised training? To this end, we discard the previous perturbation-based consistency but absorb the essence of non-parametric prototype learning. Based on the prototypical networks, we then propose a novel cyclic prototype consistency learning (CPCL) framework, which is constructed by a labeled-to-unlabeled (L2U) prototypical forward process and an unlabeled-to-labeled (U2L) backward process. Such two processes synergistically enhance the segmentation network by encouraging morediscriminative and compact features. In this way, our framework turns previous “unsupervised” consistency into new “supervised” consistency, obtaining the “all-around real label supervision” property of our method. Extensive experiments on brain tumor segmentation from MRI and kidney segmentation from CT images show that our CPCL can effectively exploit the unlabeled data and outperform other state-of-the-art semi-supervised medical image segmentation methods.

AAMAS Conference 2022 Conference Paper

Non-Parametric Neuro-Adaptive Coordination of Multi-Agent Systems

  • Christos K. Verginis
  • Zhe Xu
  • Ufuk Topcu

We develop a learning-based algorithm for the distributed formation control of networked multi-agent systems governed by unknown, nonlinear dynamics. The proposed algorithm integrates neural network-based learning with adaptive control in a two-step procedure. In the first step, each agent learns a controller, represented as a neural network, using training data that correspond to a collection of formation tasks and agent parameters. These parameters and tasks are derived by varying the nominal agent parameters and the formation specifications of the task in hand, respectively. In the second step of the algorithm, each agent incorporates the trained neural network into an online and adaptive control policy in such a way that the behavior of the multi-agent closed-loop system satisfies a user-defined formation task. Both the learning phase and the adaptive control policy are distributed, in the sense that each agent computes its own actions using only local information from its neighboring agents.

AAAI Conference 2021 Conference Paper

Adaptive Teaching of Temporal Logic Formulas to Preference-based Learners

  • Zhe Xu
  • Yuxin Chen
  • Ufuk Topcu

Machine teaching is an algorithmic framework for teaching a target hypothesis via a sequence of examples or demonstrations. We investigate machine teaching for temporal logic formulas—a novel and expressive hypothesis class amenable to time-related task specifications. In the context of teaching temporal logic formulas, an exhaustive search even for a myopic solution takes exponential time (with respect to the time span of the task). We propose an efficient approach for teaching parametric linear temporal logic formulas. Concretely, we derive a necessary condition for the minimal time length of a demonstration to eliminate a set of hypotheses. Utilizing this condition, we propose an efficient myopic teaching algorithm by solving a sequence of integer programming problems. We further show that, under two notions of teaching complexity, the proposed algorithm has near-optimal performance. We evaluate our algorithm extensively under different classes of learners (i. e. , learners with different preferences over hypotheses) and interaction protocols (e. g. , nonadaptive and adaptive). Our results demonstrate the effectiveness of the proposed algorithm in teaching temporal logic formulas; in particular, we show that there are significant gains of teaching efficacy when the teacher adapts to feedback of the learner, or adapts to a (non-myopic) oracle.

AAAI Conference 2021 Conference Paper

Advice-Guided Reinforcement Learning in a non-Markovian Environment

  • Daniel Neider
  • Jean-Raphael Gaglione
  • Ivan Gavran
  • Ufuk Topcu
  • Bo Wu
  • Zhe Xu

We study a class of reinforcement learning tasks in which the agent receives its reward for complex, temporally-extended behaviors sparsely. For such tasks, the problem is how to augment the state-space so as to make the reward function Markovian in an efficient way. While some existing solutions assume that the reward function is explicitly provided to the learning algorithm (e. g. , in the form of a reward machine), the others learn the reward function from the interactions with the environment, assuming no prior knowledge provided by the user. In this paper, we generalize both approaches and enable the user to give advice to the agent, representing the user’s best knowledge about the reward function, potentially fragmented, partial, or even incorrect. We formalize advice as a set of DFAs and present a reinforcement learning algorithm that takes advantage of such advice, with optimal convergence guarantee. The experiments show that using wellchosen advice can reduce the number of training steps needed for convergence to optimal policy, and can decrease the computation time to learn the reward function by up to two orders of magnitude.

AAAI Conference 2021 Conference Paper

Generalized Zero-Shot Learning via Disentangled Representation

  • Xiangyu Li
  • Zhe Xu
  • Kun Wei
  • Cheng Deng

Zero-Shot Learning (ZSL) aims to recognize images belonging to unseen classes that are unavailable in the training process, while Generalized Zero-Shot Learning (GZSL) is a more realistic variant that both seen and unseen classes appear during testing. Most GZSL approaches achieve knowledge transfer based on the features of samples that inevitably contain information irrelevant to recognition, bringing negative influence for the performance. In this work, we propose a novel method, dubbed Disentangled-VAE, which aims to disentangle category-distilling factors and category-dispersing factors from visual as well as semantic features, respectively. In addition, a batch re-combining strategy on latent features is introduced to guide the disentanglement, encouraging the distilling latent features to be more discriminative for recognition. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches on four challenging benchmark datasets.

JBHI Journal 2021 Journal Article

KerNet: A Novel Deep Learning Approach for Keratoconus and Sub-Clinical Keratoconus Detection Based on Raw Data of the Pentacam HR System

  • Ruiwei Feng
  • Zhe Xu
  • Xiangshang Zheng
  • Heping Hu
  • Xiuming Jin
  • Danny Z. Chen
  • Ke Yao
  • Jian Wu

Keratoconus is one of the most severe corneal diseases, which is difficult to detect at the early stage (i. e. , sub-clinical keratoconus) and possibly results in vision loss. In this paper, we propose a novel end-to-end deep learning approach, called KerNet, which processes the raw data of the Pentacam HR system (consisting of five numerical matrices) to detect keratoconus and sub-clinical keratoconus. Specifically, we propose a novel convolutional neural network, called KerNet, containing five branches as the backbone with a multi-level fusion architecture. The five branches receive five matrices separately and capture effectively the features of different matrices by several cascaded residual blocks. The multi-level fusion architecture (i. e. , low-level fusion and high-level fusion) moderately takes into account the correlation among five slices and fuses the extracted features for better prediction. Experimental results show that: (1) our novel approach outperforms state-of-the-art methods on an in-house dataset, by ~1% for keratoconus detection accuracy and ~4 for sub-clinical keratoconus detection accuracy; (2) the attention maps visualized by Grad-CAM show that our KerNet places more attention on the inferior temporal part for sub-clinical keratoconus, which has been proved as the identifying regions for ophthalmologists to detect sub-clinical keratoconus in previous clinical studies. To our best knowledge, we are the first to propose an end-to-end deep learning approach utilizing raw data obtained by the Pentacam HR system for keratoconus and subclinical keratoconus detection. Further, the prediction performance and the clinical significance of our KerNet are well evaluated and proved by two clinical experts. Our code is available at https://github.com/upzheng/Keratoconus.

AAMAS Conference 2021 Conference Paper

Reward Machines for Cooperative Multi-Agent Reinforcement Learning

  • Cyrus Neary
  • Zhe Xu
  • Bo Wu
  • Ufuk Topcu

In cooperative multi-agent reinforcement learning, a collection of agents learns to interact in a shared environment to achieve a common goal. We propose the use of reward machines (RM) — Mealy machines used as structured representations of reward functions — to encode the team’s task. The proposed novel interpretation of RMs in the multi-agent setting explicitly encodes required teammate interdependencies, allowing the team-level task to be decomposed into sub-tasks for individual agents. We define such a notion of RM decomposition and present algorithmically verifiable conditions guaranteeing that distributed completion of the sub-tasks leads to team behavior accomplishing the original task. This framework for task decomposition provides a natural approach to decentralized learning: agents may learn to accomplish their sub-tasks while observing only their local state and abstracted representations of their teammates. We accordingly propose a decentralized q-learning algorithm. Furthermore, in the case of undiscounted rewards, we use local value functions to derive lower and upper bounds for the global value function corresponding to the team task. Experimental results in three discrete settings exemplify the effectiveness of the proposed RM decomposition approach, which converges to a successful team policy an order of magnitude faster than a centralized learner and significantly outperforms hierarchical and independent q-learning approaches.

IJCAI Conference 2019 Conference Paper

Relation Extraction Using Supervision from Topic Knowledge of Relation Labels

  • Haiyun Jiang
  • Li Cui
  • Zhe Xu
  • Deqing Yang
  • Jindong Chen
  • Chenguang Li
  • Jingping Liu
  • Jiaqing Liang

Explicitly exploring the semantics of a relation is significant for high-accuracy relation extraction, which is, however, not fully studied in previous work. In this paper, we mine the topic knowledge of a relation to explicitly represent the semantics of this relation, and model relation extraction as a matching problem. That is, the matching score between a sentence and a candidate relation is predicted for an entity pair. To this end, we propose a deep matching network to precisely model the semantic similarity between a sentence-relation pair. Besides, the topic knowledge also allows us to derive the importance information of samples as well as two knowledge-guided negative sampling strategies in the training process. We conduct extensive experiments to evaluate the proposed framework and observe improvements in AUC of 11. 5% and max F1 of 5. 4% over the baselines with state-of-the-art performance.

IJCAI Conference 2019 Conference Paper

Transfer of Temporal Logic Formulas in Reinforcement Learning

  • Zhe Xu
  • Ufuk Topcu

Transferring high-level knowledge from a source task to a target task is an effective way to expedite reinforcement learning (RL). For example, propositional logic and first-order logic have been used as representations of such knowledge. We study the transfer of knowledge between tasks in which the timing of the events matters. We call such tasks temporal tasks. We concretize similarity between temporal tasks through a notion of logical transferability, and develop a transfer learning approach between different yet similar temporal tasks. We first propose an inference technique to extract metric interval temporal logic (MITL) formulas in sequential disjunctive normal form from labeled trajectories collected in RL of the two tasks. If logical transferability is identified through this inference, we construct a timed automaton for each sequential conjunctive subformula of the inferred MITL formulas from both tasks. We perform RL on the extended state which includes the locations and clock valuations of the timed automata for the source task. We then establish mappings between the corresponding components (clocks, locations, etc. ) of the timed automata from the two tasks, and transfer the extended Q-functions based on the established mappings. Finally, we perform RL on the extended state for the target task, starting with the transferred extended Q-functions. Our implementation results show, depending on how similar the source task and the target task are, that the sampling efficiency for the target task can be improved by up to one order of magnitude by performing RL in the extended state space, and further improved by up to another order of magnitude using the transferred extended Q-functions.

ICRA Conference 2018 Conference Paper

Distributed Multi-Robot Cooperation for Information Gathering Under Communication Constraints

  • Alberto Viseras Ruiz
  • Zhe Xu
  • Luis Merino

Many recent works have proposed algorithms for information gathering that benefit from multi-robot cooperation. However, most algorithms either employ discretization of the state and action spaces, which makes them computationally intractable for robotic systems with complex dynamics; or cannot deal with inter-robot restrictions like e. g. communication constraints. This paper presents an approach for multi-robot information gathering that tackles the two aforementioned issues. To this end we propose an algorithm that combines in an innovative manner Gaussian processes (GPs) to model the physical process of interest, RRT planners to plan paths in a continuous domain, and a distributed decision-making algorithm to achieve multi-robot cooperation. Specifically, we employ the Max-sum algorithm for distributed multi-robot cooperation by defining an information-theoretic utility function together with a path clustering approach. This function maximizes information gathering, subject to inter-robot communication constraints. We validate the proposed approach in simulations, and in a field experiment where three quadcopters explore a simulated wind field. Results demonstrate the effectiveness of the approach.

TCS Journal 2015 Journal Article

Scheduling games on uniform machines with activation cost

  • Fang Xie
  • Zhe Xu
  • Yuzhong Zhang
  • Qingguo Bai

In this paper, we investigate job scheduling games on uniform machines. Activating each machine incurs an activation cost in accordance with its speed. Jobs are self-decision-making players choosing machines with the objective of minimizing their own individual cost. Here, a job's individual cost is defined as the load of its chosen machine plus its activation cost, which is proportionally shared with respect to its size. We first propose an algorithm when there are m kinds of unlimited number of uniform machines, and it is proved to be a Nash Equilibrium (NE) algorithm by using the induction method. Because of equilibria always being far from optimal solutions, the inefficiency of pure NE is measured with the social objective of minimizing the maximum individual cost. Assuming that the longest job is α times as large as the unit activation cost, together with α ≥ a 2, we obtain that both the Price of Anarchy (PoA) and Price of Stability (PoS) are equal to 1. Otherwise, they are bounded by ( α + 1 ) / 2 α. Finally, we provide an example to illustrate the tight bound.

IROS Conference 2014 Conference Paper

Enhanced robotic cleaning with a low-cost tool attachment

  • Zhe Xu
  • Maya Cakmak

Robots that can reliably manipulate human tools can do a diverse range of useful tasks in human environments. However, these tools are often difficult to manipulate, particularly given force requirements for applying the tool. This is often due to the mismatch between the robot's gripper and the tool handle designed for human hands. In this paper, we present the design of a low-cost universal tool attachment that makes the tool gripper-friendly. We demonstrate the performance gain provided by the attachment on 10 different tools in the three stages of tool use: grasping the tool, applying the tool, and placing the tool. Our experiments demonstrate that the attachment performs significantly better in all three stages of tool use.

ICRA Conference 2013 Conference Paper

Decentralised coordination of mobile robots for target tracking with learnt utility models

  • Zhe Xu
  • Robert Fitch
  • Salah Sukkarieh

This paper addresses the coordination of a decentralised robot team for target tracking. In many approaches to coordination, robots jointly plan their actions through negotiation, which incurs communication costs. Previous work examined the use of learning to reduce the need for negotiations in a network of static robots. Robots incrementally learn how each team member impacts the team utility and can thus make coordinated, team-wide decisions. In this paper, we extend the concept of learning utility models to a team of mobile robots. We also propose a mechanism by which robots switch between negotiating and using the learnt model. This mechanism reduces the communications required for coordination whilst maintaining the same level of tracking performance. Hardware experiments demonstrated that our approach resulted in coordinated behaviours while only negotiating intermittently. Simulation results show that our approach reduced the data communicated for negotiations by up to 70%, without making a statistically significant impact on the tracking performance.

ICRA Conference 2012 Conference Paper

Learning utility models for decentralised coordinated target tracking

  • Zhe Xu
  • Robert Fitch
  • Salah Sukkarieh

In decentralised target tracking, a set of sensors observes moving targets. When the sensors are static but steerable, each sensor must dynamically choose which target to observe in a decentralised manner. We show that the information exchanged by the sensors to synchronise their beliefs can be exploited to learn a model of the utility function that drives each others' decisions. Instead of communicating utilities to enable negotiation, each sensor regresses on the learnt model to predict the utilities of other team members. This approach bridges the gap between coordinating implicitly, a locally-greedy solution, and negotiating explicitly. We validated our approach in both hardware and simulations, and found that it out-performed implicit coordination by a statistically significant margin with both ideal and limited communications.

ICRA Conference 2011 Conference Paper

Decentralised control of robot teams with discrete and continuous decision variables

  • Zhe Xu
  • Salah Sukkarieh

The ability to coordinate can offer significant performance advantages. This paper proposes an algorithm for decentralised control that is a function of both discrete and continuous variables. The algorithm is applied to a multi robot, multi-target mapping scenario, where decisions that couple the continuous motion controls and discrete robot-to target assignments are made. The algorithm out-performed benchmarks including implicit coordination, best-response, and the decoupling of continuous and discrete decision variables.

ICRA Conference 2009 Conference Paper

1000 Trials: An empirically validated end effector that robustly grasps objects from the floor

  • Zhe Xu
  • Travis Deyle
  • Charles C. Kemp

Unstructured, human environments present great challenges and opportunities for robotic manipulation and grasping. Robots that reliably grasp household objects with unknown or uncertain properties would be especially useful, since these robots could better generalize their capabilities across the wide variety of objects found within domestic environments. Within this paper, we address the problem of picking up an object sitting on a plane in isolation, as can occur when someone drops an object on the floor - a common problem for motor- impaired individuals. We assume that the robot has the ability to coarsely position itself in front of the object, but otherwise grasps the object with an open-loop strategy that does not vary from object to object. We present a novel end effector that is capable of robustly picking up a diverse array of everyday handheld objects given these conditions. This straight-forward, inexpensive, nonpre- hensile end effector combines a compliant finger with a thin planar component with a leading wedge that slides underneath the object. We empirically validated the efficacy of this design through a set of 1096 trials over which we systematically varied the object location, object type, object configuration, and floor characteristics. Our implementation, which we mounted on a iRobot Create, had a success rate of 94. 71 % on 680 trials, which used 4 floor types with 34 objects of particular relevance to assistive applications in 5 different poses each (4x34x5=680). The robot also had strong performance with objects that would be difficult to grasp using a traditional end effector, such as a dollar bill, a pill, a cloth, a credit card, a coin, keys, and a watch. Prior to this test, we performed 416 trials in order to assess the performance of the end effector with respect to variations in object position.

v2026.09.13