Arrow Research search

Author name cluster

Xiaoping Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

39 papers
2 author rows

Possible papers

39

YNIMG Journal 2025 Journal Article

Sleep indicators and staging: A functional near-infrared spectroscopy study in healthy young adults

  • Yong Cao
  • Xingwei An
  • Wenxiao Zhong
  • Jin Jiang
  • Hongzuo Chu
  • Xuejun Jiao
  • Xiaoping Chen
  • Yufeng Ke

Functional near-infrared spectroscopy(fNIRS)-based sleep staging has attracted considerable interest due to its portability and limited interference with sleep. However, few studies have systematically examined sleep indicators or formulated sleep staging models based on fNIRS features labelled by polysomnography(PSG). This study aimed to address these shortcomings and promote the application of fNIRS in sleep monitoring. 37 volunteers participated in our experiment, with 6-channel prefrontal fNIRS data and standard PSG data collected simultaneously. Sleep indicators were extracted from time-domain, frequency-domain, and entropy perspectives. Sleep staging was developed based on these indicators using human-scored PSG as reference. Our findings indicated deeper sleep was correlated with a decrease in amplitude of time-domain features, while entropy features showed a contrasting trend. The fNIRS-based sleep staging achieved a Cohen's kappa(κ) of 0.76±0.12, 0.72±0.09, 0.71±0.07, with accuracies of 94.2 ± 2.4 %, 87.8 ± 3.2 %, and 82.2 ± 4.1 %, for 2-class(Wake/Sleep), 3-class(Wake/NREM/REM), 4-class (Wake/N1+N2/N3/REM) classifications, respectively. Sleep statistics derived from fNIRS closely aligned with those from PSG, with differences in sleep onset latency, wake after sleep onset, total wake/sleep time within 5 min and sleep efficiency below 3 %. The substantial agreement in both detailed (epoch-by-epoch) and comprehensive (total) sleep statistics with PSG suggests fNIRS is a reliable tool for long-term sleep monitoring in everyday settings.

ICRA Conference 2024 Conference Paper

Kinematic Modeling and Control of a Soft Robotic Arm with Non-constant Curvature Deformation

  • Zhanchi Wang
  • Gaotian Wang
  • Xiaoping Chen
  • Nikolaos M. Freris

The passive compliance of soft robotic arms renders the development of accurate kinematic models and model-based controllers challenging. The most widely used model in soft robotic kinematics assumes Piecewise Constant Curvature (PCC). However, PCC introduces errors when the robot is subject to external forces or even gravity. In this paper, we establish a three-dimensional (3D) kinematic representation of a soft robotic arm with pseudo universal and prismatic joints that are capable of capturing non-constant curvature deformations of the soft segments. We theoretically demonstrate that this constitutes a more general methodology than PCC. Simulations and experiments on the real robot attest to the superior modeling accuracy of our approach in 3D motions with unknown loads. The maximum position/rotation error of the proposed model is verified 6. 7×/4. 6× lower than the PCC model considering gravity and external forces. Furthermore, we devise an inverse kinematic controller that is capable of positioning the tip, tracking trajectories, as well as performing interactive tasks in the 3D space.

ICRA Conference 2023 Conference Paper

Automatic Generation of Robot Facial Expressions with Preferences

  • Bing Tang
  • Rongyun Cao
  • Rongya Chen
  • Xiaoping Chen
  • Bei Hua
  • Feng Wu

The capability of humanoid robots to generate facial expressions is crucial for enhancing interactivity and emotional resonance in human-robot interaction. However, humanoid robots vary in mechanics, manufacturing, and ap-pearance. The lack of consistent processing techniques and the complexity of generating facial expressions pose significant challenges in the field. To acquire solutions with high confidence, it is necessary to enable robots to explore the solution space automatically based on performance feedback. To this end, we designed a physical robot with a human-like appearance and developed a general framework for automatic expression generation using the MAP-Elites algorithm. The main advan-tage of our framework is that it does not only generate facial expressions automatically but can also be customized according to user preferences. The experimental results demonstrate that our framework can efficiently generate realistic facial expressions without hard coding or prior knowledge of the robot kinematics. Moreover, it can guide the solution-generation process in accordance with user preferences, which is desirable in many real-world applications.

IROS Conference 2021 Conference Paper

Crowd-Aware Robot Navigation for Pedestrians with Multiple Collision Avoidance Strategies via Map-based Deep Reinforcement Learning

  • Shunyi Yao
  • Guangda Chen
  • Quecheng Qiu
  • Jun Ma 0034
  • Xiaoping Chen
  • Jianmin Ji

It is challenging for a mobile robot to navigate through human crowds. Existing approaches usually assume that pedestrians follow a predefined collision avoidance strategy, like social force model (SFM) or optimal reciprocal collision avoidance (ORCA). However, their performances commonly need to be further improved for practical applications, where pedestrians follow multiple different collision avoidance strategies. In this paper, we propose a map-based deep reinforcement learning approach for crowd-aware robot navigation with various pedestrians. We use the sensor map to represent the environmental information around the robot, including its shape and observable appearances of obstacles. We also introduce the pedestrian map that specifies the movements of pedestrians around the robot. By applying both maps as inputs of the neural network, we show that a navigation policy can be trained to better interact with pedestrians following different collision avoidance strategies. We evaluate our approach under multiple scenarios both in the simulator and on an actual robot. The results show that our approach allows the robot to successfully interact with various pedestrians and outperforms compared methods in terms of the success rate.

IJCAI Conference 2019 Conference Paper

Capturing Spatial and Temporal Patterns for Facial Landmark Tracking through Adversarial Learning

  • Shi Yin
  • Shangfei Wang
  • Guozhu Peng
  • Xiaoping Chen
  • Bowen Pan

The spatial and temporal patterns inherent in facial feature points are crucial for facial landmark tracking, but have not been thoroughly explored yet. In this paper, we propose a novel deep adversarial framework to explore the shape and temporal dependencies from both appearance level and target label level. The proposed deep adversarial framework consists of a deep landmark tracker and a discriminator. The deep landmark tracker is composed of a stacked Hourglass network as well as a convolutional neural network and a long short-term memory network, and thus implicitly capture spatial and temporal patterns from facial appearance for facial landmark tracking. The discriminator is adopted to distinguish the tracked facial landmarks from ground truth ones. It explicitly models shape and temporal dependencies existing in ground truth facial landmarks through another convolutional neural network and another long short-term memory network. The deep landmark tracker and the discriminator compete with each other. Through adversarial learning, the proposed deep adversarial landmark tracking approach leverages inherent spatial and temporal patterns to facilitate facial landmark tracking from both appearance level and target label level. Experimental results on two benchmark databases demonstrate the superiority of the proposed approach to state-of-the-art work.

AAAI Conference 2019 Conference Paper

Goal-Oriented Dialogue Policy Learning from Failures

  • Keting Lu
  • Shiqi Zhang
  • Xiaoping Chen

Reinforcement learning methods have been used for learning dialogue policies. However, learning an effective dialogue policy frequently requires prohibitively many conversations. This is partly because of the sparse rewards in dialogues, and the very few successful dialogues in early learning phase. Hindsight experience replay (HER) enables learning from failures, but the vanilla HER is inapplicable to dialogue learning due to the implicit goals. In this work, we develop two complex HER methods providing different tradeoffs between complexity and performance, and, for the first time, enabled HER-based dialogue policy learning. Experiments using a realistic user simulator show that our HER methods perform better than existing experience replay methods (as applied to deep Q-networks) in learning rate.

AAAI Conference 2018 Conference Paper

Privacy-Preserving Policy Iteration for Decentralized POMDPs

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

We propose the first privacy-preserving approach to address the privacy issues that arise in multi-agent planning problems modeled as a Dec-POMDP. Our solution is a distributed message-passing algorithm based on trials, where the agents’ policies are optimized using the cross-entropy method. In our algorithm, the agents’ private information is protected using a public-key homomorphic cryptosystem. We prove the correctness of our algorithm and analyze its complexity in terms of message passing and encryption/decryption operations. Furthermore, we analyze several privacy aspects of our algorithm and show that it can preserve the agent privacy of non-neighbors, model privacy, and decision privacy. Our experimental results on several common Dec-POMDP benchmark problems confirm the effectiveness of our approach.

ICRA Conference 2017 Conference Paper

A two-level approach for solving the inverse kinematics of an extensible soft arm considering viscoelastic behavior

  • Hao Jiang 0015
  • Zhanchi Wang
  • Xinghua Liu
  • Xiaotong Chen
  • Yusong Jin
  • Xuanke You
  • Xiaoping Chen

Soft compliant materials and novel actuation mechanisms ensure flexible motions and high adaptability for soft robots, but also increase the difficulty and complexity of constructing control systems. In this work, we provide an efficient control algorithm for a multi-segment extensible soft arm in 2D plane. The algorithm separate the inverse kinematics into two levels. The first level employs gradient descent to select optimized arm's pose (from task space to configuration space) according to designed cost functions. With consideration of viscoelasticity, the second level utilizes neural networks to figure out the pressures from each segment's pose (from configuration space to actuation space). In experiments with a physical prototype, the control accuracy and effectiveness are validated, where the control algorithm is further improved by an optional feedback strategy.

IJCAI Conference 2017 Conference Paper

Integrating Answer Set Programming with Semantic Dictionaries for Robot Task Planning

  • Dongcai Lu
  • Yi Zhou
  • Feng Wu
  • Zhao Zhang
  • Xiaoping Chen

In this paper, we propose a novel integrated task planning system for service robot in domestic domains. Given open-ended high-level user instructions in natural language, robots need to generate a plan, i. e. , a sequence of low-level executable actions, to complete the required tasks. To address this, we exploit the knowledge on semantic roles of common verbs defined in semantic dictionaries such as FrameNet and integrate it with Answer Set Programming --- a task planning framework with both representation language and solvers. In the experiments, we evaluated our approach using common benchmarks on service tasks and showed that it can successfully handle much more tasks than the state-of-the-art solution. Notably, we deployed the proposed planning system on our service robot for the annual RoboCup@Home competitions and achieved very encouraging results.

IROS Conference 2017 Conference Paper

Leveraging commonsense reasoning and multimodal perception for robot spoken dialog systems

  • Dongcai Lu
  • Shiqi Zhang 0001
  • Peter Stone 0001
  • Xiaoping Chen

Probabilistic graphical models, such as partially observable Markov decision processes (POMDPs), have been used in stochastic spoken dialog systems to handle the inherent uncertainty in speech recognition and language understanding. Such dialog systems suffer from the fact that only a relatively small number of domain variables are allowed in the model, so as to ensure the generation of good-quality dialog policies. At the same time, the non-language perception modalities on robots, such as vision-based facial expression recognition and Lidar-based distance detection, can hardly be integrated into this process. In this paper, we use a probabilistic commonsense reasoner to “guide” our POMDP-based dialog manager, and present a principled, multimodal dialog management (MDM) framework that allows the robot's dialog belief state to be seamlessly updated by both observations of human spoken language, and exogenous events such as the change of human facial expressions. The MDM approach has been implemented and evaluated both in simulation and on a real mobile robot using guidance tasks.

IROS Conference 2017 Conference Paper

Model-free control for soft manipulators based on reinforcement learning

  • Xuanke You
  • Yixiao Zhang
  • Xiaotong Chen
  • Xinghua Liu
  • Zhanchi Wang
  • Hao Jiang 0015
  • Xiaoping Chen

Most control methods of soft manipulators are developed based on physical models derived from mathematical analysis or learning methods. However, due to internal nonlinearity and external uncertain disturbances, it is difficult to build an accurate model, further, these methods lack robustness and portability among different prototypes. In this work, we propose a model-free control method based on reinforcement learning and implement it on a multi-segment soft manipulator in 2D plane, which focuses on the learning of control strategy rather than the physical model. The control strategy is validated to be effective and robust in prototype experiments, where we design a simulation method to speed up the training process.

IROS Conference 2017 Conference Paper

Model-less feedback control for soft manipulators

  • Yusong Jin
  • Yufei Wang
  • Xiaotong Chen
  • Zhanchi Wang
  • Xinghua Liu
  • Hao Jiang 0015
  • Xiaoping Chen

Soft manipulators have been a rising focus of soft robotics research. Taking advantage of soft materials and flexible, continuous movements, they have promising applicable prospect. However, their highly internal nonlinearity and unpredictable deformation caused by environmental effects make it difficult to build an exact model for control. In this work, we propose a generalized controller for soft manipulators using an estimated Jacobian-based model derived from structural analysis. The model can be simplified from reasonable assumptions of manipulator structure, and updated to balance conformity to reality and stability. In prototype experiments on an 3D multi-segment soft manipulator, the control method exhibits accuracy as well as adaptability to self gravity and external loads.

IJCAI Conference 2017 Conference Paper

Multi-Agent Planning with Baseline Regret Minimization

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

We propose a novel baseline regret minimization algorithm for multi-agent planning problems modeled as finite-horizon decentralized POMDPs. It guarantees to produce a policy that is provably better than or at least equivalent to the baseline policy. We also propose an iterative belief generation algorithm to effectively and efficiently minimize the baseline regret, which only requires necessary iterations to converge to the policy with minimum baseline regret. Experimental results on common benchmark problems confirm its advantage comparing to the state-of-the-art approaches.

IJCAI Conference 2016 Conference Paper

Coordinating Human-UAV Teams in Disaster Response

  • Feng Wu
  • Sarvapali D. Ramchurn
  • Xiaoping Chen

We consider a disaster response scenario where emergency responders have to complete rescue tasks in dynamic and uncertain environment with the assistance of multiple UAVs to collect information about the disaster space. To capture the uncertainty and partial observability of the domain, we model this problem as a POMDP. However, the resulting model is computationally intractable and cannot be solved by most existing POMDP solvers due to the large state and action spaces. By exploiting the problem structure we propose a novel online planning algorithm to solve this model. Specifically, we generate plans for the responders based on Monte-Carlo simulations and compute actions for the UAVs according to the value of information. Our empirical results confirm that our algorithm significantly outperforms the state-of-the-art both in time and solution quality.

IJCAI Conference 2016 Conference Paper

Planning with Task-Oriented Knowledge Acquisition for a Service Robot

  • Kai Chen
  • Fangkai Yang
  • Xiaoping Chen

We propose a framework for a service robot to behave intelligently in domains that contain incomplete information, underspecified goals and dynamic change. Human robot interaction (HRI), sensing actions and physical actions are uniformly formalized in action language BC. An answer set solver is called to generate plans that guide the robot to acquire task-oriented knowledge and execute actions to achieve its goal, including interacting with human to gather information and sensing the environment to help motion planning. By continuously interpreting and grounding useful sensing information, robot is able to use contingent knowledge to adapt to unexpected changes and faults. We evaluate the approach on service robot KeJia that serves drink to guests, a testing benchmark for general-purpose service robot proposed by RoboCup@Home competition.

IJCAI Conference 2016 Conference Paper

Robot Scavenger Hunt: A Standardized Framework for Evaluating Intelligent Mobile Robots

  • Shiqi Zhang
  • Dongcai Lu
  • Xiaoping Chen
  • Peter Stone

In recent years, many different types of intelligent mobile robots have been developed in research and industrial labs. Although there are significant differences in both hardware and software over these robots, many of them share a common set of AI capabilities, e. g. , planning, learning, vision and natural language processing. At the same time, almost all of them are equipped with traditional robotic capabilities such as mapping, localization, and navigation. However, to date it has been difficult to compare and contrast their capabilities in any controlled way. The main goal of the Robot Scavenger Hunt is to provide a standardized framework that includes a set of standardized tasks for evaluating the AI and robotic capabilities of medium-sized intelligent mobile robots. Compared to existing benchmarks, e. g. , RoboCup@Home, Robot Scavenger Hunt aims at evaluations in larger spaces (multi-floor buildings vs. rooms) over longer periods of time (hours vs. minutes) while interacting with real human residents.

IROS Conference 2015 Conference Paper

Building temporal consistent semantic maps for indoor scenes

  • Zhe Zhao 0004
  • Xiaoping Chen

In this paper, we propose a novel approach to generate a temporal consistent semantic map for 3D indoor scenes. In contrast to previous techniques which generate the semantic map on the whole global scene immediately or smooth the semantic predictions by a graph model, we intend to discover temporal information over RGB-D images and leverage it to enforce the label consistency. Our method contains two key components: a low-level component and a high-level component. On the low level, temporal segmentation method is adopted to find the correspondence of the superpixels incrementally. On the high level, temporal information is treated as the higher order cliques and a higher-order Dense Conditional Random Fields (CRF) is utilized to jointly infer the object category and structural class of the global point cloud. On the experiments, we compare our temporal consistent segmentation algorithm with the state-of-the-art approach and generate the semantic maps from the NYU v2 dataset. Our experiments demonstrate that temporal consistent constraints are significant for the semantic mapping procedure and can improve the precision of the semantic mapping results.

TIST Journal 2015 Journal Article

Online Planning for Large Markov Decision Processes with Hierarchical Decomposition

  • Aijun Bai
  • Feng Wu
  • Xiaoping Chen

Markov decision processes (MDPs) provide a rich framework for planning under uncertainty. However, exactly solving a large MDP is usually intractable due to the “curse of dimensionality”— the state space grows exponentially with the number of state variables. Online algorithms tackle this problem by avoiding computing a policy for the entire state space. On the other hand, since online algorithm has to find a near-optimal action online in almost real time, the computation time is often very limited. In the context of reinforcement learning, MAXQ is a value function decomposition method that exploits the underlying structure of the original MDP and decomposes it into a combination of smaller subproblems arranged over a task hierarchy. In this article, we present MAXQ-OP—a novel online planning algorithm for large MDPs that utilizes MAXQ hierarchical decomposition in online settings. Compared to traditional online planning algorithms, MAXQ-OP is able to reach much more deeper states in the search tree with relatively less computation time by exploiting MAXQ hierarchical decomposition online. We empirically evaluate our algorithm in the standard Taxi domain—a common benchmark for MDPs—to show the effectiveness of our approach. We have also conducted a long-term case study in a highly complex simulated soccer domain and developed a team named WrightEagle that has won five world champions and five runners-up in the recent 10 years of RoboCup Soccer Simulation 2D annual competitions. The results in the RoboCup domain confirm the scalability of MAXQ-OP to very large domains.

IROS Conference 2014 Conference Paper

Semantic mapping for object category and structural class

  • Zhe Zhao 0004
  • Xiaoping Chen

Intelligent robots require a semantic map of the surroundings for applications such as navigation and object localization. With this information, a robot can make task planning, object manipulation and human-robot interaction. However, it still remains an open problem although considerable emphasis has been given. In this paper, we propose a novel approach to generate a dense semantic map for 3D indoor scene. Our approach integrates a robust image labeling algorithm with simultaneous localization and mapping method (SLAM) to generate the semantic map. Scene information, semantic context and geometric context are encoded into a CRF model. Our CRF model computes a simultaneous labeling of image regions into semantic classes (e. g. , bed, table, chair) and structural object classes (Ground, Furniture, Structure, Props). Then semantic labeling results in single images are fused into the 3D map using the estimated camera poses by SLAM. We report our labeling performance on NYU v2 dataset and demonstrate that our algorithm is comparable to and in many cases superior to the previous method. Also we generate our semantic map on the NYU v2 video dataset.

ICAPS Conference 2014 Conference Paper

Thompson Sampling Based Monte-Carlo Planning in POMDPs

  • Aijun Bai
  • Feng Wu 0001
  • Zongzhang Zhang
  • Xiaoping Chen

Monte-Carlo tree search (MCTS) has been drawing great interest in recent years for planning under uncertainty. One of the key challenges is the trade-off between exploration and exploitation. To address this, we introduce a novel online planning algorithm for large POMDPs using Thompson sampling based MCTS that balances between cumulative and simple regrets. The proposed algorithm Dirichlet-Dirichlet-NormalGamma based Partially Observable Monte-Carlo Planning (D2NG-POMCP) treats the accumulated reward of performing an action from a belief state in the MCTS search tree as a random variable following an unknown distribution with hidden parameters. Bayesian method is used to model and infer the posterior distribution of these parameters by choosing the conjugate prior in the form of a combination of two Dirichlet and one NormalGamma distributions. Thompson sampling is exploited to guide the action selection in the search tree. Experimental results confirmed that our algorithm outperforms the state-of-the-art approaches on several common benchmark problems.

NeurIPS Conference 2013 Conference Paper

Bayesian Mixture Modelling and Inference based Thompson Sampling in Monte-Carlo Tree Search

  • Aijun Bai
  • Feng Wu
  • Xiaoping Chen

Monte-Carlo tree search is drawing great interest in the domain of planning under uncertainty, particularly when little or no domain knowledge is available. One of the central problems is the trade-off between exploration and exploitation. In this paper we present a novel Bayesian mixture modelling and inference based Thompson sampling approach to addressing this dilemma. The proposed Dirichlet-NormalGamma MCTS (DNG-MCTS) algorithm represents the uncertainty of the accumulated reward for actions in the MCTS search tree as a mixture of Normal distributions and inferences on it in Bayesian settings by choosing conjugate priors in the form of combinations of Dirichlet and NormalGamma distributions. Thompson sampling is used to select the best action at each decision node. Experimental results show that our proposed algorithm has achieved the state-of-the-art comparing with popular UCT algorithm in the context of online planning for general Markov decision processes.

AAAI Conference 2013 Conference Paper

Goal-Oriented Euclidean Heuristics with Manifold Learning

  • Wenlin Chen
  • Yixin Chen
  • Kilian Weinberger
  • Qiang Lu
  • Xiaoping Chen

Recently, a Euclidean heuristic (EH) has been proposed for A* search. EH exploits manifold learning methods to construct an embedding of the state space graph, and derives an admissible heuristic distance between two states from the Euclidean distance between their respective embedded points. EH has shown good performance and memory efficiency in comparison to other existing heuristics such as differential heuristics. However, its potential has not been fully explored. In this paper, we propose a number of techniques that can significantly improve the quality of EH. We propose a goal-oriented manifold learning scheme that optimizes the Euclidean distance to goals in the embedding while maintaining admissibility and consistency. We also propose a state heuristic enhancement technique to reduce the gap between heuristic and true distances. The enhanced heuristic is admissible but no longer consistent. We then employ a modified search algorithm, known as B0 algorithm, that achieves optimality with inconsistent heuristics using consistency check and propagation. We demonstrate the effectiveness of the above techniques and report un-matched reduction in search costs across several non-trivial benchmark search problems.

IJCAI Conference 2013 Conference Paper

Handling Open Knowledge for Service Robots

  • Xiaoping Chen
  • Jianmin Ji
  • Zhiqiang Sui
  • Jiongkun Xie

Users may ask a service robot to accomplish various tasks so that the designer of the robot cannot program each of the tasks beforehand. As more and more open-source knowledge resources become available, it is worthwhile trying to make use of open-source knowledge resources for service robots. The challenge lies in the autonomous identification, acquisition and utilization of missing knowledge about a user task at hand. In this paper, the core problem is formalized and the complexity results of the main reasoning issues are provided. A mechanism for task planning with open-knowledge rules which are provided by non-experts in semistructured natural language and thus generally underspecified are introduced. Techniques for translating the semi-structured knowledge from a large open-source knowledge base are also presented. Experiments showed a remarkable improvement of the system performance on a test set consisting of hundreds of user desires from the open-source knowledge base.

AAAI Conference 2012 Conference Paper

Covering Number as a Complexity Measure for POMDP Planning and Learning

  • Zongzhang Zhang
  • Michael Littman
  • Xiaoping Chen

Finding a meaningful way of characterizing the difficulty of partially observable Markov decision processes (POMDPs) is a core theoretical problem in POMDP research. State-space size is often used as a proxy for POMDP difficulty, but it is a weak metric at best. Existing work has shown that the covering number for the reachable belief space, which is a set of belief points that are reachable from the initial belief point, has interesting links with the complexity of POMDP planning, theoretically. In this paper, we present empirical evidence that the covering number for the reachable belief space (or just “covering number”, for brevity) is a far better complexity measure than the state-space size for both planning and learning POMDPs on several small-scale benchmark problems. We connect the covering number to the complexity of learning POMDPs by proposing a provably convergent learning algorithm for POMDPs without reset given knowledge of the covering number.

UAI Conference 2012 Conference Paper

FHHOP: A Factored Hybrid Heuristic Online Planning Algorithm for Large POMDPs

  • Zongzhang Zhang
  • Xiaoping Chen

Planning in partially observable Markov decision processes (POMDPs) remains a challenging topic in the artificial intelligence community, in spite of recent impressive progress in approximation techniques. Previous research has indicated that online planning approaches are promising in handling large-scale POMDP domains efficiently as they make decisions “on demand” instead of proactively for the entire state space. We present a Factored Hybrid Heuristic Online Planning (FHHOP) algorithm for large POMDPs. FHHOP gets its power by combining a novel hybrid heuristic search strategy with a recently developed factored state representation. On several benchmark problems, FHHOP substantially outperformed state-of-theart online heuristic search approaches in terms of both scalability and quality.

AAMAS Conference 2012 Conference Paper

Online Planning for Large MDPs with MAXQ Decomposition

  • Aijun Bai
  • Feng Wu
  • Xiaoping Chen

Markov decision processes (MDPs) provide an expressive framework for planning in stochastic domains. However, exactly solving a large MDP is often intractable due to the curse of dimensionality. Online algorithms help overcome the high computational complexity by avoiding computing a policy for each possible state. Hierarchical decomposition is another promising way to help scale MDP algorithms up to large domains by exploiting their underlying structure. In this paper, we present an effort on combining the benefits of a general hierarchical structure based on MAXQ value function decomposition with the power of heuristic and approximate techniques for developing an online planning framework, called MAXQ-OP. The proposed framework provides a principled approach for programming autonomous agents in a large stochastic domain. We have been conducting a long-term case-study with the RoboCup soccer simulation 2D domain, which is extremely larger than domains usually studied in literature, as the major benchmark to this research. The case-study showed that the agents developed with this framework and the related techniques reached outstanding performances, showing its high scalability to very large domains.

ECAI Conference 2012 Conference Paper

Sample-Based Policy Iteration for Constrained DEC-POMDPs

  • Feng Wu 0001
  • Nicholas R. Jennings
  • Xiaoping Chen

We introduce constrained DEC-POMDPs - an extension of the standard DEC-POMDPs that includes constraints on the optimality of the overall team rewards. Constrained DEC-POMDPs present a natural framework for modeling cooperative multi-agent problems with limited resources. To solve such DEC-POMDPs, we propose a novel sample-based policy iteration algorithm. The algorithm builds on multi-agent dynamic programming and benefits from several recent advances in DEC-POMDP algorithms such as MB-DP [12] and TBDP [13]. Specifically, it improves the joint policy by solving a series of standard nonlinear programs (NLPs), thereby building on recent advances in NLP solvers. Our experimental results confirm the algorithm can efficiently solve constrained DEC-POMDPs that cause general DEC-POMDP algorithms to fail.

IJCAI Conference 2011 Conference Paper

Online Planning for Ad Hoc Autonomous Agent Teams

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

We propose a novel online planning algorithm for ad hoc team settings - challenging situations in which an agent must collaborate with unknown teammates without prior coordination. Our approach is based on constructing and solving a series of stage games, and then using biased adaptive play to choose actions. The utility function in each stage game is estimated via Monte-Carlo tree search using the UCT algorithm. We establish analytically the convergence of the algorithm and show that it performs well in a variety of ad hoc team domains.

AIJ Journal 2011 Journal Article

Online planning for multi-agent systems with bounded communication

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

We propose an online algorithm for planning under uncertainty in multi-agent settings modeled as DEC-POMDPs. The algorithm helps overcome the high computational complexity of solving such problems offline. The key challenges in decentralized operation are to maintain coordinated behavior with little or no communication and, when communication is allowed, to optimize value with minimal communication. The algorithm addresses these challenges by generating identical conditional plans based on common knowledge and communicating only when history inconsistency is detected, allowing communication to be postponed when necessary. To be suitable for online operation, the algorithm computes good local policies using a new and fast local search method implemented using linear programming. Moreover, it bounds the amount of memory used at each step and can be applied to problems with arbitrary horizons. The experimental results confirm that the algorithm can solve problems that are too large for the best existing offline planning algorithms and it outperforms the best online method, producing much higher value with much less communication in most cases. The algorithm also proves to be effective when the communication channel is imperfect (periodically unavailable). These results contribute to the scalability of decision-theoretic planning in multi-agent settings.

AAMAS Conference 2011 Conference Paper

Towards Robot Incremental Learning Constraints from Comparative Demonstration

  • Rong Zhang
  • Shangfei Wang
  • Xiaoping Chen
  • Dong Yin
  • Shijia Chen
  • Min Cheng
  • Yanpeng Lv
  • Jianmin Ji

This paper presents an attempt on incremental robot learning from demonstration. Based on previously learnt knowledge about a task in simpler situations, a robot learns to fulfill the same task properly in a more complicated situation through analyzing comparative demonstrations and extracting new knowledge, especially the constraints that the task in the new situation imposes on the robot's behaviors.

AAMAS Conference 2010 Conference Paper

Developing High-level Cognitive Functions for Service Robots

  • Xiaoping Chen
  • Jianmin Ji
  • Jiehui Jiang
  • Guoqiang Jin
  • Feng Wang
  • Jiongkun Xie

The primary target of this work is human-robot collaboration, especially for service robots in complicated applicationscenarios. Three assumptions and four requirements areidentified. State-of-the-art, general-purpose Natural Language Processing (NLP), Commonsense Reasoning (in particular, ASP), and Robotics techniques are integrated in alayered architecture. The architecture and mechanisms havebeen implemented on a service robot, Ke Jia. Instead ofcommand languages, small limited segments of natural languages are employed in spoken dialog between Ke Jia and itsusers. The information in the dialog is extracted, classifiedand transferred into inner representation by Ke Jia's NLPmechanism, and further used autonomously in problem-solvingand planning. A series of case study was conducted onKe Jia with positive results, verifying its ability of acquiringknowledge through spoken dialog with users, autonomoussolving problems by virtue of acquired causal knowledge, and autonomous planning for complex tasks.

AAMAS Conference 2010 Conference Paper

Point-Based Policy Generation for Decentralized POMDPs

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

Memory-bounded techniques have shown great promise in solving complex multi-agent planning problems modeled as DEC-POMDPs. Much of the performance gains can be attributed to pruning techniques that alleviate the complexity of the exhaustive backup step of the original MBDP algorithm. Despite these improvements, state-of-the-art algorithms can still handle a relative small pool of candidate policies, which limits the quality of the solution in some benchmark problems. We present a new algorithm, Point-Based Policy Generation, which avoids altogether searching the entire joint policy space. The key observation is that the best policy for each reachable belief state can be constructed directly, instead of producing first a large set of candidates. We also provide an efficient approximate implementation of this operation. The experimental results show that our solution technique improvesthe performance significantly in terms of both runtime and solution quality.

UAI Conference 2010 Conference Paper

Rollout Sampling Policy Iteration for Decentralized POMDPs

  • Feng Wu 0001
  • Shlomo Zilberstein
  • Xiaoping Chen

We present decentralized rollout sampling policy iteration (DecRSPI) — a new algorithm for multi-agent decision problems formalized as DEC-POMDPs. DecRSPI is designed to improve scalability and tackle problems that lack an explicit model. The algorithm uses Monte- Carlo methods to generate a sample of reachable belief states. Then it computes a joint policy for each belief state based on the rollout estimations. A new policy representation allows us to represent solutions compactly. The key benefits of the algorithm are its linear time complexity over the number of agents, its bounded memory usage and good solution quality. It can solve larger problems that are intractable for existing planning algorithms. Experimental results confirm the effectiveness and scalability of the approach.

AAAI Conference 2010 Conference Paper

Trial-Based Dynamic Programming for Multi-Agent Planning

  • Feng Wu
  • Shlomo Zilberstein
  • Xiaoping Chen

Trial-based approaches offer an efficient way to solve singleagent MDPs and POMDPs. These approaches allow agents to focus their computations on regions of the environment they encounter during the trials, leading to significant computational savings. We present a novel trial-based dynamic programming (TBDP) algorithm for DEC-POMDPs that extends these benefits to multi-agent settings. The algorithm uses trial-based methods for both belief generation and policy evaluation. Policy improvement is implemented efficiently using linear programming and a sub-policy reuse technique that helps bound the amount of memory. The results show that TBDP can produce significant value improvements and is much faster than the best existing planning algorithms.

ICAPS Conference 2009 Conference Paper

Multi-Agent Online Planning with Communication

  • Feng Wu 0001
  • Shlomo Zilberstein
  • Xiaoping Chen

We propose an online algorithm for planning under uncertainty in multi-agent settings modeled as DEC-POMDPs. The algorithm helps overcome the high computational complexity of solving such problems off-line. The key challenge is to produce coordinated behavior using little or no communication. When communication is allowed but constrained, the challenge is to produce high value with minimal communication. The algorithm addresses these challenges by communicating only when history inconsistency is detected, allowing communication to be postponed if necessary. Moreover, it bounds the memory usage at each step and can be applied to problems with arbitrary horizons. The experimental results confirm that the algorithm can solve problems that are too large for the best existing off-line planning algorithms and it outperforms the best online method, producing higher value with much less communication in most cases.

KR Conference 2008 Conference Paper

Computing Loops with at Most One External Support Rule

  • Xiaoping Chen
  • Jianmin Ji
  • Fangzhen Lin

If a loop has no external support rules, then its loop formula is equivalent to a set of unit clauses; and if it has exactly one external support rule, then its loop formula is equivalent to a set of binary clauses. In this paper, we consider how to compute these loops and their loop formulas in a normal logic program, and use them to derive consequences of a logic program. We show that an iterative procedure based on unit propagation, the program completion and the loop formulas of loops with no external support rules can compute the same consequences as the "Expand" operator in smodels, which is known to compute the well-founded model when the given normal logic program has no constraints. We also show that using the loop formulas of loops with at most one external support rule, the same procedure can compute more consequences, and these extra consequences can help ASP solvers such as cmodels to find answer sets of certain logic programs.

KR Conference 2004 Conference Paper

Partial Implication Semantics for Desirable Propositions

  • Xiaoping Chen
  • Yi Zhou

Motivational attitudes play an important role in investigations into intelligent agents. One of the key problems of representing and reasoning about motivational attitudes is which propositions are the desirable ones. The answer based on classical logic is that the propositions that logically imply the goal are desirable and the others are not. We argue that this criterion is inadequate for the incomplete knowledge about environments for an agent. In this paper, we present a simple and intuitive semantics---partial implication---for the characterization of desirable propositions. In this semantics, Proposition P is a desirable one with respect to a given goal Q if and only if P is "useful" and "harmless" to Q in any situation. Partial implication is an extension of classical implication. We investigate some fundamental properties of partial implication and discuss some of the potential applications.

IJCAI Conference 1999 Conference Paper

A Logic of Intention

  • Xiaoping Chen
  • Guiquan Liu

There is a lot of research on formalization of intention. The common idea of these theories is to interprete intention as an unary modal operator in Kripkean semantics. These theories suffer from the side-effect problem seriously. We introduce an alternative approach by establishing a nonclassical logic of intention. This logic is based on a novel non-Kripkean semantics which embodies some cognitive features. We show that this logic does provide a formal specification and a decidable inference mechanism of intention consequences. All and only the instances of sideeffects, except ones in absorbent forms, are forbidden in the logic.

v2026.09.13