Arrow Research search

Author name cluster

Edward Johns

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

ICLR Conference 2025 Conference Paper

Instant Policy: In-Context Imitation Learning via Graph Diffusion

  • Vitalis Vosylius
  • Edward Johns

Following the impressive capabilities of in-context learning with large transformers, In-Context Imitation Learning (ICIL) is a promising opportunity for robotics. We introduce Instant Policy, which learns new tasks instantly from just one or two demonstrations, achieving ICIL through two key components. First, we introduce inductive biases through a graph representation and model ICIL as a graph generation problem using a learned diffusion process, enabling structured reasoning over demonstrations, observations, and actions. Second, we show that such a model can be trained using pseudo-demonstrations – arbitrary trajectories generated in simulation – as a virtually infinite pool of training data. Our experiments, in both simulation and reality, show that Instant Policy enables rapid learning of various everyday robot tasks. We also show how it can serve as a foundation for cross-embodiment and zero-shot transfer to language-defined tasks.

NeurIPS Conference 2025 Conference Paper

Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions

  • Naoki Kiyohara
  • Edward Johns
  • Yingzhen Li

Stochastic differential equations (SDEs) are well suited to modelling noisy and/or irregularly-sampled time series, which are omnipresent in finance, physics, and machine learning applications. Traditional approaches require costly simulation of numerical solvers when sampling between arbitrary time points. We introduce Neural Stochastic Flows (NSFs) and their latent dynamic versions, which learns (latent) SDE transition laws directly using conditional normalising flows, with architectural constraints that preserve properties inherited from stochastic flow. This enables sampling between arbitrary states in a single step, providing up to two orders of magnitude speedup for distant time points. Experiments on synthetic SDE simulations and real-world tracking and video data demonstrate that NSF maintains distributional accuracy comparable to numerical approaches while dramatically reducing computation for arbitrary time-point sampling, enabling applications where numerical solvers remain prohibitively expensive.

ICRA Conference 2025 Conference Paper

One-Shot Dual-Arm Imitation Learning

  • Yilong Wang
  • Edward Johns

We introduce One-Shot Dual-Arm Imitation Learning (ODIL), which enables dual-arm robots to learn precise and coordinated everyday tasks from just a single demonstration of the task. ODIL uses a new three-stage visual servoing (3-VS) method for precise alignment between the end-effector and target object, after which replay of the demonstration trajectory is sufficient to perform the task. This is achieved without requiring prior task or object knowledge, or additional data collection and training following the single demonstration. Furthermore, we propose a new dual-arm coordination paradigm for learning dual-arm tasks from a single demonstration. ODIL was tested on a real-world dual-arm robot, demonstrating state-of-the-art performance across six precise and coordinated tasks in both 4-DoF and 6-DoF settings, and showing robustness in the presence of distractor objects and partial occlusions. Videos are available at: https://www.robot-learning.uk/one-shot-dual-arm.

ICRA Conference 2025 Conference Paper

R+X: Retrieval and Execution from Everyday Human Videos

  • Georgios Papagiannis
  • Norman Di Palo
  • Pietro Vitiello
  • Edward Johns

We present $\mathbf{R}+\mathbf{X}$, a framework which enables robots to learn skills from long, unlabelled, first-person videos of humans performing everyday tasks. Given a language command from a human, $\mathbf{R}+\mathbf{X}$ first retrieves short video clips containing relevant behaviour, and then executes the skill by conditioning an in-context imitation learning method (KAT) on this behaviour. By leveraging a Vision Language Model (VLM) for retrieval, $\mathbf{R}+\mathbf{X}$ does not require any manual annotation of the videos, and by leveraging in-context learning for execution, robots can perform commanded skills immediately, without requiring a period of training on the retrieved videos. Experiments studying a range of everyday household tasks show that $\mathbf{R}+\mathbf{X}$ succeeds at translating unlabelled human videos into robust robot skills, and that $\mathbf{R}+\mathbf{X}$ outperforms several recent alternative methods. Appendix and videos are available at https://www.robot-learning.uk/r-plus-x.

IROS Conference 2024 Conference Paper

Adapting Skills to Novel Grasps: A Self-Supervised Approach

  • Georgios Papagiannis
  • Kamil Dreczkowski
  • Vitalis Vosylius
  • Edward Johns

In this paper, we study the problem of adapting manipulation trajectories involving grasped objects (e. g. tools) defined for a single grasp pose to novel grasp poses. A common approach to address this is to define a new trajectory for each possible grasp explicitly, but this is highly inefficient. Instead, we propose a method to adapt such trajectories directly while only requiring a period of self-supervised data collection, during which a camera observes the robot’s end-effector moving with the object rigidly grasped. Importantly, our method requires no prior knowledge of the grasped object (such as a 3D CAD model), it can work with RGB images, depth images, or both, and it requires no camera calibration. Through a series of real-world experiments involving 1360 evaluations, we find that self-supervised RGB data consistently outperforms alternatives that rely on depth images including several state-of-the-art pose estimation methods. Compared to the best-performing baseline, our method results in an average of 28. 5% higher success rate when adapting manipulation trajectories to novel grasps on several everyday tasks. The appendix accompanying the paper and videos of the experiments are available on our webpage at www. robot-learning.uk/adapting-skills.

ICRA Conference 2024 Conference Paper

DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation Models

  • Norman Di Palo
  • Edward Johns

We propose DINOBot, a novel imitation learning framework for robot manipulation, which leverages the image-level and pixel-level capabilities of features extracted from Vision Transformers trained with DINO. When interacting with a novel object, DINOBot first uses these features to retrieve the most visually similar object experienced during human demonstrations, and then uses this object to align its endeffector with the novel object to enable effective interaction. Through a series of real-world experiments on everyday tasks, we show that exploiting both the image-level and pixel-level properties of vision foundation models enables unprecedented learning efficiency and generalisation. Videos and code are available at https://www.robot-learning.uk/dinobot.

ICRA Conference 2024 Conference Paper

Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models

  • Ivan Kapelyukh
  • Yifei Ren
  • Ignacio Alzugaray
  • Edward Johns

We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the scene, where objects can be rearranged virtually and an image of the resulting arrangement rendered. These renders are evaluated by a VLM, so that the arrangement which best satisfies the user instruction is selected and recreated in the real world with pick-and-place. This enables language-conditioned rearrangement to be performed zero-shot, without needing to collect a training dataset of example arrangements. Results on a series of real-world tasks show that this framework is robust to distractors, controllable by language, capable of understanding complex multi-object relations, and readily applicable to both tabletop and 6-DoF rearrangement tasks. Videos are available on our webpage at: https://www.robot-learning.uk/dream2real.

ICRA Conference 2024 Conference Paper

Open X-Embodiment: Robotic Learning Datasets and RT-X Models: Open X-Embodiment Collaboration

  • Abby O'Neill
  • Abdul Rehman
  • Abhiram Maddukuri
  • Abhishek Gupta 0004
  • Abhishek Padalkar
  • Abraham Lee
  • Acorn Pooley
  • Agrim Gupta

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x. github.io.

TMLR Journal 2024 Journal Article

Prismer: A Vision-Language Model with Multi-Task Experts

  • Shikun Liu
  • Linxi Fan
  • Edward Johns
  • Zhiding Yu
  • Chaowei Xiao
  • Anima Anandkumar

Recent vision-language models have shown impressive multi-modal generation capabilities. However, typically they require training huge models on massive datasets. As a more scalable alternative, we introduce Prismer, a data- and parameter-efficient vision-language model that leverages an ensemble of task-specific experts. Prismer only requires training of a small number of components, with the majority of network weights inherited from multiple readily-available, pre-trained experts, and kept frozen during training. By leveraging experts from a wide range of domains, we show Prismer can efficiently pool this expert knowledge and adapt it to various vision-language reasoning tasks. In our experiments, we show that Prismer achieves fine-tuned and few-shot learning performance which is competitive with current state-of-the-arts, whilst requiring up to two orders of magnitude less training data. Code is available at https://github.com/NVlabs/prismer.

ICRA Conference 2023 Conference Paper

Learning Tethered Perching for Aerial Robots

  • Fabian Hauf
  • Basaran Bahadir Kocer
  • Alan Slatter
  • Hai-Nguyen Nguyen
  • Oscar Pang
  • Ronald Clark
  • Edward Johns
  • Mirko Kovac

Aerial robots have a wide range of applications, such as collecting data in hard-to-reach areas. This requires the longest possible operation time. However, because currently available commercial batteries have limited specific energy of roughly 300 W h kg -1, a drone's flight time is a bottleneck for sustainable long-term data collection. Inspired by birds in nature, a possible approach to tackle this challenge is to perch drones on trees, and environmental or man-made structures, to save energy whilst in operation. In this paper, we propose an algorithm to automatically generate trajectories for a drone to perch on a tree branch, using the proposed tethered perching mechanism with a pendulum-like structure. This enables a drone to perform an energy-optimised, controlled 180° flip to safely disarm upside down. To fine-tune a set of reachable trajectories, a soft actor critic-based reinforcement algorithm is used. Our experimental results show the feasibility of the set of trajectories with successful perching. Our findings demonstrate that the proposed approach enables energy-efficient landing for long-term data collection tasks.

TMLR Journal 2022 Journal Article

Auto-Lambda: Disentangling Dynamic Task Relationships

  • Shikun Liu
  • Stephen James
  • Andrew Davison
  • Edward Johns

Understanding the structure of multiple related tasks allows for multi-task learning to improve the generalisation ability of one or all of them. However, it usually requires training each pairwise combination of tasks together in order to capture task relationships, at an extremely high computational cost. In this work, we learn task relationships via an automated weighting framework, named Auto-Lambda. Unlike previous methods where task relationships are assumed to be fixed, i.e., task should either be trained together or not trained together, Auto-Lambda explores continuous, dynamic task relationships via task-specific weightings, and can optimise any choice of combination of tasks through the formulation of a meta-loss; where the validation loss automatically influences task weightings throughout training. We apply the proposed framework to both multi-task and auxiliary learning problems in computer vision and robotics, and show that AutoLambda achieves state-of-the-art performance, even when compared to optimisation strategies designed specifically for each problem and data domain. Finally, we observe that Auto-Lambda can discover interesting learning behaviors, leading to new insights in multi-task learning. Code is available at https://github.com/lorenmt/auto-lambda.

ICLR Conference 2022 Conference Paper

Bootstrapping Semantic Segmentation with Regional Contrast

  • Shikun Liu
  • Shuaifeng Zhi
  • Edward Johns
  • Andrew J. Davison

We present ReCo, a contrastive learning framework designed at a regional level to assist learning in semantic segmentation. ReCo performs pixel-level contrastive learning on a sparse set of hard negative pixels, with minimal additional memory footprint. ReCo is easy to implement, being built on top of off-the-shelf segmentation networks, and consistently improves performance, achieving more accurate segmentation boundaries and faster convergence. The strongest effect is in semi-supervised learning with very few labels. With ReCo, we achieve high quality semantic segmentation model, requiring only 5 examples of each semantic class.

IROS Conference 2022 Conference Paper

Demonstrate Once, Imitate Immediately (DOME): Learning Visual Servoing for One-Shot Imitation Learning

  • Eugene Valassakis
  • Georgios Papagiannis
  • Norman Di Palo
  • Edward Johns

We present DOME, a novel method for one-shot imitation learning, where a task can be learned from just a single demonstration and then be deployed immediately, without any further data collection or training. DOME does not require prior task or object knowledge, and can perform the task in novel object configurations and with distractors. At its core, DOME uses an image-conditioned object segmentation network followed by a learned visual servoing network, to move the robot's end-effector to the same relative pose to the object as during the demonstration, after which the task can be completed by replaying the demonstration's end-effector velocities. We show that DOME achieves near 100% success rate on 7 real-world everyday tasks, and we perform several studies to thoroughly understand each individual component of DOME. Videos and supplementary material are available at: https://www.robot-learning.uk/dome.

ICRA Conference 2021 Conference Paper

Benchmarking Domain Randomisation for Visual Sim-to-Real Transfer

  • Raghad Alghonaim
  • Edward Johns

Domain randomisation is a very popular method for visual sim-to-real transfer in robotics, due to its simplicity and ability to achieve transfer without any real-world images at all. Nonetheless, a number of design choices must be made to achieve optimal transfer. In this paper, we perform a comprehensive benchmarking study on these different choices, with two key experiments evaluated on a real-world object pose estimation task. First, we study the rendering quality, and nd that a small number of high-quality images is superior to a large number of low-quality images. Second, we study the type of randomisation, and nd that both distractors and textures are important for generalisation to novel environments.

IROS Conference 2021 Conference Paper

Coarse-to-Fine for Sim-to-Real: Sub-Millimetre Precision Across Wide Task Spaces

  • Eugene Valassakis
  • Norman Di Palo
  • Edward Johns

In this paper, we study the problem of zero-shot sim-to-real when the task requires both highly precise control with sub-millimetre error tolerance, and wide task space generalisation. Our framework involves a coarse-to-fine controller, where trajectories begin with classical motion planning using ICP-based pose estimation, and transition to a learned end-to-end controller which maps images to actions and is trained in simulation with domain randomisation. In this way, we achieve precise control whilst also generalising the controller across wide task spaces, and keeping the robustness of vision-based, end-to-end control. Real-world experiments on a range of different tasks show that, by exploiting the best of both worlds, our framework significantly outperforms purely motion planning methods, and purely learning-based methods. Furthermore, we answer a range of questions on best practices for precise sim-to-real transfer, such as how different image sensor modalities and image feature representations perform.

ICRA Conference 2021 Conference Paper

Coarse-to-Fine Imitation Learning: Robot Manipulation from a Single Demonstration

  • Edward Johns

We introduce a simple new method for visual imitation learning, which allows a novel robot manipulation task to be learned from a single human demonstration, without requiring any prior knowledge of the object being interacted with. Our method models imitation learning as a state estimation problem, with the state defined as the end-effector’s pose at the point where object interaction begins, as observed from the demonstration. By modelling a manipulation task as a coarse, approach trajectory followed by a fine, interaction trajectory, this state estimator can be trained in a self-supervised manner, by automatically moving the end-effector’s camera around the object. At test time, the end-effector is moved to the estimated state through a linear path, at which point the demonstration’s end-effector velocities are simply repeated, enabling convenient acquisition of a complex interaction trajectory without actually needing to explicitly learn a policy. Real-world experiments on 8 everyday tasks show that our method can learn a diverse range of skills from just a single human demonstration, whilst also yielding a stable and interpretable controller. Videos at: www. robot-learning.uk/coarse-to-fine-imitation-learning.

IROS Conference 2021 Conference Paper

Hybrid ICP

  • Kamil Dreczkowski
  • Edward Johns

ICP algorithms typically involve a fixed choice of data association method and a fixed choice of error metric. In this paper, we propose Hybrid ICP, a novel and flexible ICP variant which dynamically optimises both the data association method and error metric based on the live image of an object and the current ICP estimate. We show that when used for object pose estimation, Hybrid ICP is more accurate and more robust to noise than other commonly used ICP variants. We also consider the setting where ICP is applied sequentially with a moving camera, and we study the trade-off between the accuracy of each ICP estimate and the number of ICP estimates available within a fixed amount of time.

IROS Conference 2020 Conference Paper

Crossing the Gap: A Deep Dive into Zero-Shot Sim-to-Real Transfer for Dynamics

  • Eugene Valassakis
  • Zihan Ding
  • Edward Johns

Zero-shot sim-to-real transfer of tasks with complex dynamics is a highly challenging and unsolved problem. A number of solutions have been proposed in recent years, but we have found that many works do not present a thorough evaluation in the real world, or underplay the significant engineering effort and task-specific fine tuning that is required to achieve the published results. In this paper, we dive deeper into the sim-to-real transfer challenge, investigate why this is such a difficult problem, and present objective evaluations of a number of transfer methods across a range of real-world tasks. Surprisingly, we found that a method which simply injects random forces into the simulation performs just as well as more complex methods, such as those which randomise the simulator's dynamics parameters, or adapt a policy online using recurrent network architectures.

IROS Conference 2020 Conference Paper

Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning

  • Guillermo Garcia-Hernando
  • Edward Johns
  • Tae-Kyun Kim 0001

Dexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e. g. object weights and surface frictions, the absence of force feedback makes the task challenging, as even slight inaccuracies on finger tips or contact points from HPE may make the interactions fail. Prior arts simply generate contact forces in the direction of the fingers' closures, when finger joints penetrate virtual objects. Although useful for simple grasping scenarios, they cannot be applied to dexterous manipulations such as inhand manipulation. Existing reinforcement learning (RL) and imitation learning (IL) approaches train agents that learn skills by using task-specific rewards, without considering any online user input. In this work, we propose to learn a model that maps noisy input hand poses to target virtual poses, which introduces the needed contacts to accomplish the tasks on a physics simulator. The agent is trained in a residual setting by using a model-free hybrid RL+IL approach. A 3D hand pose estimation reward is introduced leading to an improvement on HPE accuracy when the physics-guided corrected target poses are remapped to the input space. As the model corrects HPE errors by applying minor but crucial joint displacements for contacts, this helps to keep the generated motion visually close to the user input. Since HPE sequences performing successful virtual interactions do not exist, a data generation scheme to train and evaluate the system is proposed. We test our framework in two applications that use hand pose estimates for dexterous manipulations: hand-object interactions in VR and hand-object motion reconstruction in-the-wild. Experiments show that the proposed method outperforms various RL/IL baselines and the simple prior art of enforcing hand closure, both in task success and hand pose accuracy.

ICRA Conference 2020 Conference Paper

Sim-to-Real Transfer for Optical Tactile Sensing

  • Zihan Ding
  • Nathan F. Lepora
  • Edward Johns

Deep learning and reinforcement learning methods have been shown to enable learning of flexible and complex robot controllers. However, the reliance on large amounts of training data often requires data collection to be carried out in simulation, with a number of sim-to-real transfer methods being developed in recent years. In this paper, we study these techniques for tactile sensing using the TacTip optical tactile sensor, which consists of a deformable tip with a camera observing the positions of pins inside this tip. We designed a model for soft body simulation which was implemented using the Unity physics engine, and trained a neural network to predict the locations and angles of edges when in contact with the sensor. Using domain randomisation techniques for sim-to-real transfer, we show how this framework can be used to accurately predict edges with less than 1 mm prediction error in real-world testing, without any real-world data at all.

NeurIPS Conference 2019 Conference Paper

Self-Supervised Generalisation with Meta Auxiliary Learning

  • Shikun Liu
  • Andrew Davison
  • Edward Johns

Learning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be improved without requiring access to any further data. The approach is to train two neural networks: a label-generation network to predict the auxiliary labels, and a multi-task network to train the primary task alongside the auxiliary task. The loss for the label-generation network incorporates the loss of the multi-task network, and so this interaction between the two networks can be seen as a form of meta learning with a double gradient. We show that our proposed method, Meta AuXiliary Learning (MAXL), outperforms single-task learning on 7 image datasets, without requiring any additional data. We also show that MAXL outperforms several other baselines for generating auxiliary labels, and is even competitive when compared with human-defined auxiliary labels. The self-supervised nature of our method leads to a promising new direction towards automated generalisation. Source code can be found at \url{https: //github. com/lorenmt/maxl}.

ICRA Conference 2017 Conference Paper

Application-oriented design space exploration for SLAM algorithms

  • Sajad Saeedi 0001
  • Luigi Nardi
  • Edward Johns
  • Bruno Bodin
  • Paul H. J. Kelly
  • Andrew J. Davison

In visual SLAM, there are many software and hardware parameters, such as algorithmic thresholds and GPU frequency, that need to be tuned; however, this tuning should also take into account the structure and motion of the camera. In this paper, we determine the complexity of the structure and motion with a few parameters calculated using information theory. Depending on this complexity and the desired performance metrics, suitable parameters are explored and determined. Additionally, based on the proposed structure and motion parameters, several applications are presented, including a novel active SLAM approach which guides the camera in such a way that the SLAM algorithm achieves the desired performance metrics. Real-world and simulated experimental results demonstrate the effectiveness of the proposed design space and its applications.

IROS Conference 2016 Conference Paper

Deep learning a grasp function for grasping under gripper pose uncertainty

  • Edward Johns
  • Stefan Leutenegger
  • Andrew J. Davison

This paper presents a new method for parallel-jaw grasping of isolated objects from depth images, under large gripper pose uncertainty. Whilst most approaches aim to predict the single best grasp pose from an image, our method first predicts a score for every possible grasp pose, which we denote the grasp function. With this, it is possible to achieve grasping robust to the gripper's pose uncertainty, by smoothing the grasp function with the pose uncertainty function. Therefore, if the single best pose is adjacent to a region of poor grasp quality, that pose will no longer be chosen, and instead a pose will be chosen which is surrounded by a region of high grasp quality. To learn this function, we train a Convolutional Neural Network which takes as input a single depth image of an object, and outputs a score for each grasp pose across the image. Training data for this is generated by use of physics simulation and depth image simulation with 3D object meshes, to enable acquisition of sufficient data without requiring exhaustive real-world experiments. We evaluate with both synthetic and real experiments, and show that the learned grasp score is more robust to gripper pose uncertainty than when this uncertainty is not accounted for.

ICRA Conference 2013 Conference Paper

Dynamic scene models for incremental, long-term, appearance-based localisation

  • Edward Johns
  • Guang-Zhong Yang

In this paper we present a new appearance-based localisation system that is able to deal with dynamic elements in the scene. By independently modelling the properties of local features observed in a scene over long periods of time, we show that feature appearances and geometric relationships can be learned more accurately than when representing a location by a single image. We also present a new dataset consisting of a 6 km outdoor path traversed once per month for a period of 5 months, which contains several challenges including short-term and long-term dynamic behaviour, lateral deviations in the path, repetitive scene appearances and strong illumination changes. We show superior performance of the dynamic mapping system compared to state-of-the-art techniques on our dataset.

ICRA Conference 2013 Conference Paper

Feature Co-occurrence Maps: Appearance-based localisation throughout the day

  • Edward Johns
  • Guang-Zhong Yang

In this paper we present a new method, Feature Co-occurrence Maps, for appearance-based localisation over the course of a day. We show that by quantising local features in both feature and image space, discriminative statistics can be learned on the co-occurrences of features at different times of the day. This allows for matching at any time, without requiring individual images to be stored representing each time of day, and matching is performed efficiently by simultaneously matching to the entire database. We further show how matching along image sequences can be incorporated into the system and adapt existing methods by allowing for non-zero acceleration. Results on a 20km outdoor dataset show improved performance in precision-recall over state of the art.

IROS Conference 2011 Conference Paper

A scene-associated training method for mobile robot speech recognition in multisource reverberated environments

  • Jindong Liu
  • Edward Johns
  • Guang-Zhong Yang

In this paper, we present a new technique for social mobile robot speech recognition based on scene-associated training models. The key contribution of the paper is a real-time framework that reduces the effect of room reverberation and ambient noise, a challenging problem in speech recognition. In classical approaches, anechoic sound is used to train the model, with the main focus on removing reverberation or noise from the sound. Our technique differs in that we train a number of speech recognizers directly from the reverberated sound, by associating each recognizer with a unique visual scene, to deal with the varying reverberation properties of different rooms. By extracting local features from a captured image and recognizing a scene, the robot can use the appropriate speech recognizer that is trained for the particular structural properties of that scene. We tested our method by using a baseline speech recognition model (HTK) across a variety of rooms and different levels of background noise. The results show that the association between a visual scene and a corresponding speech recognizer greatly improves the robot's speech recognition accuracy, together with increasing the computational speed of recognition, compared to competing techniques.

ICRA Conference 2011 Conference Paper

Global localization in a dense continuous topological map

  • Edward Johns
  • Guang-Zhong Yang

Vision-based topological maps for mobile robot localization traditionally consist of a set of images captured along a path, with a query image then compared to every individual map image. This paper introduces a new approach to topological mapping, whereby the map consists of a set of landmarks that are detected across multiple images, spanning the continuous space between nodal images. Matches are then made to landmarks, rather than to individual images, enabling a topological map of far greater density than traditionally possible, without sacrificing computational speed. Furthermore, by treating each landmark independently, a probabilistic approach to localization can be employed by taking into account the learned discriminative properties of each landmark. An optimization stage is then used to adjust the map according to speed and localization accuracy requirements. Results for global localization show a greater positive location identification rate compared to the traditional topological map, together with enabling a greater localization resolution in the denser topological map, without requiring a decrease in frame rate.

IROS Conference 2010 Conference Paper

Scene association for mobile robot navigation

  • Edward Johns
  • Guang-Zhong Yang

Accurate, efficient and robust location recognition is a fundamental task for any mobile robot. This paper presents a new approach using visual features to efficiently represent a series of locations along a path in an indoor environment. In the training stage, local features which are detected across multiple images from a single tour are combined to represent a real-world landmark, modelled by the expected variance of its descriptor. Those landmarks which represent the scene in the most efficient and discriminative manner are then retained, and this selection is optimized with respect to the scale of the environment. In the recognition stage, features detected in an image are matched to the landmarks in memory, based upon a novel similarity measure drawing from feature co-occurrence statistics.

v2026.09.13