Arrow Research search

Author name cluster

Andrew Calway

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

20 papers
1 author row

Possible papers

20

IROS Conference 2024 Conference Paper

Geolocation on Cartographic Maps with Multi-Modal Fusion

  • Mengjie Zhou
  • Liu Liu 0009
  • Yiran Zhong
  • Andrew Calway

We explore the geolocation problem, aiming to localize ground-view images on cartographic maps, without the need of any GPS priors. This task mimics the human wayfinding ability and offers high scalability and robustness by using the compact and semantic representations of maps. Current methods often rely on 2D maps to encode dense contextual information for ground-to-map matching. In this paper, we lift ground-to-map matching to a 2. 5D space, where heights of structures (e. g. buildings) provide richer geometric information to guide the matching process. We propose a new approach to learning representative embeddings from multi-modal data. Specifically, we establish a projection relationship between 2D and 2. 5D space. The projection is further used to combine multi-modal features from the 2D and 2. 5D maps using an effective pixel-to-point fusion method. By encoding crucial geometric cues, our method learns discriminative location embeddings for matching panoramic images and maps. Additionally, we construct the first large-scale multi-modal geolocation dataset to validate our method and facilitate future research. Both single-image based and route based geolocation experiments are conducted to test our method. Extensive experiments demonstrate that the proposed method achieves significantly higher geolocation accuracy and faster convergence than previous 2D map-based approaches.

IROS Conference 2024 Conference Paper

Object-based SLAM Using Superquadrics

  • Yifan Xing
  • Noe Samano
  • Wen Fan 0001
  • Andrew Calway

Visual SLAM uses visual information, typically point features, to localise a camera and, at the same time, map the environment. In recent years, there has been interest in using scene-understanding capabilities to enhance the mapping process and object-level SLAM systems have appeared in response. However, most of the previous work is limited to prestored object models or pre-trained networks to represent the objects, which limits working scenarios or uses representations with limited scope, such as cubes or quadrics. To address this, we propose to use superquadrics as the object representation and, in this paper, present a proof of principle SLAM system in which object-based mapping is fully integrated with camera tracking via keyframe optimisation. The system was tested on simulated and real datasets, and the results show that the system can achieve lightweight and comparatively good object representation whilst also giving good camera trajectories estimates under certain scenarios.

IROS Conference 2022 Conference Paper

CGiS-Net: Aggregating Colour, Geometry and Implicit Semantic Features for Indoor Place Recognition

  • Yuhang Ming 0001
  • Xingrui Yang 0001
  • Guofeng Zhang 0001
  • Andrew Calway

We describe a novel approach to indoor place recognition from RGB point clouds based on aggregating low-level colour and geometry features with high-level implicit semantic features. It uses a 2-stage deep learning framework, in which the first stage is trained for the auxiliary task of semantic segmentation and the second stage uses features from layers in the first stage to generate discriminate descriptors for place recognition. The auxiliary task encourages the features to be semantically meaningful, hence aggregating the geometry and colour in the RGB point cloud data with implicit semantic information. We use an indoor place recognition dataset derived from the ScanNet dataset for training and evaluation, with a test set comprising 3, 608 point clouds generated from 100 different rooms. Comparison with a traditional feature-based method and four state-of-the-art deep learning methods demonstrate that our approach significantly outperforms all five methods, achieving, for example, a top-3 average recall rate of 75% compared with 41% for the closest rival method. Our code is available at: https://github.com/YuhangMing/Semantic-Indoor-Place-Recognition

ICRA Conference 2022 Conference Paper

FD-SLAM: 3-D Reconstruction Using Features and Dense Matching

  • Xingrui Yang 0001
  • Yuhang Ming 0001
  • Zhaopeng Cui
  • Andrew Calway

It is well known that visual SLAM systems based on dense matching are locally accurate but are also susceptible to long-term drift and map corruption. In contrast, feature matching methods can achieve greater long-term consistency but can suffer from inaccurate local pose estimation when feature information is sparse. Based on these observations, we propose an RGB-D SLAM system that leverages the advantages of both approaches: using dense frame-to-model odometry to build accurate sub-maps and on-the-fly feature-based matching across sub-maps for global map optimisation. In addition, we incorporate a learning-based loop closure component based on 3-D features which further stabilises map building. We have evaluated the approach on indoor sequences from public datasets, and the results show that it performs on par or better than state-of-the-art systems in terms of map reconstruction quality and pose estimation. The approach can also scale to large scenes where other systems often fail.

IROS Conference 2021 Conference Paper

Efficient Localisation Using Images and OpenStreetMaps

  • Mengjie Zhou
  • Xieyuanli Chen
  • Noe Samano
  • Cyrill Stachniss
  • Andrew Calway

The ability to localise is key for robot navigation. We describe an efficient method for vision-based localisation, which combines sequential Monte Carlo tracking with matching ground-level images to 2-D cartographic maps such as OpenStreetMaps. The matching is based on a learned embedded space representation linking images and map tiles, encoding the common semantic information present in both and providing potential for invariance to changing conditions. Moreover, the compactness of 2-D maps supports scalability. This contrasts with the majority of previous approaches based on matching with single-shot geo-referenced images or 3-D reconstructions. We present experiments using the StreetLearn and Oxford RobotCar datasets and demonstrate that the method is highly effective, giving high accuracy and fast convergence.

ICRA Conference 2021 Conference Paper

Global Aerial Localisation Using Image and Map Embeddings

  • Noe Samano
  • Mengjie Zhou
  • Andrew Calway

We present a purely vision based geolocation method for aircraft flying over urban and suburban environments. The method is based on matching aerial images with geolocated map tiles using a shared low dimensional embedded space of descriptors. The Euclidean distance between descriptors is used as a similarity measure between domains. The similarity between the observation and map locations is then integrated with visual odometry to track the aircraft’s position and yaw using a particle filter. Furthermore, we propose an efficient method to generate map descriptors in testing time based on interpolation, allowing compact representation of large areas giving the potential for high levels of scalability. We experimented in different cities with areas above 20 km 2 in size and preliminary results based on a database of aerial imagery demonstrate that the method gives good results.

IROS Conference 2021 Conference Paper

Object-Augmented RGB-D SLAM for Wide-Disparity Relocalisation

  • Yuhang Ming 0001
  • Xingrui Yang 0001
  • Andrew Calway

We propose a novel object-augmented RGB-D SLAM system that is capable of constructing a consistent object map and performing relocalisation based on centroids of objects in the map. The approach aims to overcome the view dependence of appearance-based relocalisation methods using point features or images. During the map construction, we use a pre-trained neural network to detect objects and estimate 6D poses from RGB-D data. An incremental probabilistic model is used to aggregate estimates over time to create the object map. Then in relocalisation, we use the same network to extract objects-of-interest in the ‘lost’ frames. Pairwise geometric matching finds correspondences between map and frame objects, and probabilistic absolute orientation followed by application of iterative closest point to dense depth maps and object centroids gives relocalisation. Results of experiments in desktop environments demonstrate very high success rates even for frames with widely different viewpoints from those used to construct the map, significantly outperforming two appearance- based methods.

ICRA Conference 2019 Conference Paper

Improving drone localisation around wind turbines using monocular model-based tracking

  • Oliver Moolan-Feroze
  • Konstantinos Karachalios
  • Dimitrios N. Nikolaidis
  • Andrew Calway

We present a novel method of integrating image-based measurements into a drone navigation system for the automated inspection of wind turbines. We take a model-based tracking approach, where a 3D skeleton representation of the turbine is matched to the image data. Matching is based on comparing the projection of the representation to that inferred from images using a convolutional neural network. This enables us to find image correspondences using a generic turbine model that can be applied to a wide range of turbine shapes and sizes. To estimate 3D pose of the drone, we fuse the network output with GPS and IMU measurements using a pose graph optimiser. Results illustrate that the use of the image measurements significantly improves the accuracy of the localisation over that obtained using GPS and IMU alone.

IROS Conference 2019 Conference Paper

Simultaneous Drone Localisation and Wind Turbine Model Fitting During Autonomous Surface Inspection

  • Oliver Moolan-Feroze
  • Konstantinos Karachalios
  • Dimitrios N. Nikolaidis
  • Andrew Calway

We present a method for simultaneous localisation and wind turbine model fitting for a drone performing an automated surface inspection. We use a skeletal parameterisation of the turbine that can be easily integrated into a non-linear least squares optimiser, combined with a pose graph representation of the drone’s 3-D trajectory, allowing us to optimise both sets of parameters simultaneously. Given images from an onboard camera, we use a CNN to infer projections of the skeletal model, enabling correspondence constraints to be established through a cost function. This is then coupled with GPS/IMU measurements taken at key frames in the graph to allow successive optimisation as the drone navigates around the turbine. We present two variants of the cost function, one based on traditional 2D point correspondences and the other on direct image interpolation within the inferred projections. Results from experiments on simulated and real-world data show that simultaneous optimisation provides improvements to localisation over only optimising the pose and that combined use of both cost functions proves most effective.

IROS Conference 2018 Conference Paper

Automated Map Reading: Image Based Localisation in 2-D Maps Using Binary Semantic Descriptors

  • Pilailuck Panphattarasap
  • Andrew Calway

We describe a novel approach to image based localisation in urban environments which uses semantic matching between images and a 2-D cartographic map. This contrasts with the majority of existing approaches which use image to image database matching. We use highly compact binary descriptors to represent locations, indicating the presence or not of semantic features, which significantly increases scalability and has the potential for greater invariance to variable imaging conditions. The approach is also more akin to human map reading, making it better suited to human-system interaction. In this initial study we use semantic features relating to buildings and road junctions in discrete viewing directions. CNN classifiers are used to detect the features in images and we match descriptor estimates with location tagged descriptors derived from the 2-D map to give localisation. The descriptors are not sufficiently discriminative on their own, but when concatenated sequentially along a route, their combination becomes highly distinctive and allows localisation even when using non-perfect classifiers. Performance is further improved by taking into account left or right turns over a route. Experimental results obtained using Google StreetView and OpenStreetMap data show that the approach has considerable potential, achieving localisation accuracy of around 85% using routes corresponding to approximately 200 meters.

IROS Conference 2018 Conference Paper

Predicting Out-of-View Feature Points for Model-Based Camera Pose Estimation

  • Oliver Moolan-Feroze
  • Andrew Calway

In this work we present a novel framework that uses deep learning to predict object feature points that are out-of-view in the input image. This system was developed with the application of model-based tracking in mind, particularly in the case of autonomous inspection robots, where only partial views of the object are available. Out-of-view prediction is enabled by applying scaling to the feature point labels during network training. This is combined with a recurrent neural network architecture designed to provide the final prediction layers with rich feature information from across the spatial extent of the input image. To show the versatility of these out-of-view predictions, we describe how to integrate them in both a particle filter tracker and an optimisation based tracker. To evaluate our work we compared our framework with one that predicts only points inside the image. We show that as the amount of the object in view decreases, being able to predict outside the image bounds adds robustness to the final pose estimation.

ICRA Conference 2016 Conference Paper

Absolute pose estimation using multiple forms of correspondences from RGB-D frames

  • Shuda Li
  • Andrew Calway

We describe a new approach to absolute pose estimation from noisy and outlier contaminated matching point sets for RGB-D sensors. We show that by integrating multiple forms of correspondence based on 2-D and 3-D points and surface normals gives more precise, accurate and robust pose estimates. This is because it gives more constraints than using one form alone and increases the available measurements, especially when dealing with sparse matching sets. We demonstrate the approach by incorporating it within a RANSAC algorithm and introduce a novel direct least-square approach to calculate pose estimates. Results from experiments on synthetic and real data demonstrate improved performance over existing methods.

IROS Conference 2015 Conference Paper

Improving MAV control by predicting aerodynamic effects of obstacles

  • John Bartholomew
  • Andrew Calway
  • Walterio W. Mayol-Cuevas

Building on our previous work [1], in this paper we demonstrate how it is possible to improve flight control of a MAV that experiences aerodynamic disturbances caused by objects on its path. Predictions based on low resolution depth images taken at a distance are incorporated into the flight control loop on the throttle channel as this is adjusted to target undisrupted level flight. We demonstrate that a statistically significant improvement (p ≪ 0. 001) is possible for some common obstacles such as boxes and steps, compared to using conventional feedback-only control. Our approach and results are encouraging toward more autonomous MAV exploration strategies.

ICRA Conference 2015 Conference Paper

RGBD relocalisation using pairwise geometry and concise key point sets

  • Shuda Li
  • Andrew Calway

We describe a novel RGBD relocalisation algorithm based on key point matching. It combines two components. First, a graph matching algorithm which takes into account the pairwise 3-D geometry amongst the key points, giving robust relocalisation. Second, a point selection process which provides an even distribution of the ‘most matchable’ points across the scene based on non-maximum suppression within voxels of a volumetric grid. This ensures a bounded set of matchable key points which enables tractable and scalable graph matching at frame rate. We present evaluations using a public dataset and our own more difficult dataset containing large pose changes, fast motion and non-stationary objects. It is shown that the method significantly out performs state-of-the-art methods.

ICRA Conference 2014 Conference Paper

Learning to predict obstacle aerodynamics from depth images for Micro Air Vehicles

  • John Bartholomew
  • Andrew Calway
  • Walterio W. Mayol-Cuevas

Many applications of Micro Air Vehicles (MAVs) require them to operate in cluttered environments, flying in constrained spaces and close to obstacles. Such obstacles affect the airflow around the MAV and can thereby affect its flight characteristics. We describe a system for predicting these effects at a distance, using depth images obtained from an RGB-D sensor. Predictions are based on learning from prior experience gathered during training flights. We show that aerodynamic effects caused by obstacles are consistent, and demonstrate that it is practical to make predictions from experience without running a computationally expensive aerodynamic simulation. Our approach uses a Gaussian process regression, it requires minimal parameter tuning and is able to predict the acceleration that will be expected at a distance in the future. The method produces estimates within 12ms without any code optimisation and the results indicate good prediction ability with mean errors within 4–10cm/s 2 on a database of various obstacles.

IROS Conference 2013 Conference Paper

Enhancing 6D visual relocalisation with depth cameras

  • José Martínez-Carranza
  • Andrew Calway
  • Walterio W. Mayol-Cuevas

Relocalisation in 6D is relevant to a variety of Robotics applications and in particular to agile cameras exploring a 3D environment. While the use of geometry has commonly helped to validate appearance as a back-end process in several relocalisation systems before, we are interested in using 3D information to assist fast pose relocalisation computation as part of a front-end task. Our approach rapidly searches for a reduced number of visual descriptors, previously observed and stored in a database, that can be used to effectively compute the camera pose corresponding to the current view. We guide the search by means of constructing validated candidate sets using a 3D test involving the depth information obtained with an RGB-D camera (e. g. stereo of with structured light). Our experiments demonstrate that this process returns a compact quality set that works better for the pose estimation stage than when using a typical Nearest-Neighbor search over appearance only. The improvements are observed in terms of percentage of relocalised frames and speed, where the latter goes up to two orders of magnitude w. r. t. the conventional search.

ICRA Conference 2013 Conference Paper

Visual mapping using learned structural priors

  • Osian Haines
  • José Martínez-Carranza
  • Andrew Calway

We investigate a new approach to vision based mapping, in which single image structure recognition is used to derive strong priors for initialisation of higher-level primitives in the map. This can reduce state size and speed up the building of more meaningful maps. We focus on plane mapping and use a recognition algorithm to detect and estimate the 3D orientation of planar structures in key frames, which are then used as priors for initialising planes in the map. The recognition algorithm learns the relationship between such structure and appearance from training examples offline. We demonstrate the approach in the context of an EKF based visual odometry system. Preliminary results of experiments in urban environments show that the system is able to build large maps with significant planar structure at average frames rates of around 60 fps whilst maintaining good trajectory estimation. The results suggest that the approach has considerable potential.

ICRA Conference 2012 Conference Paper

Efficient visual odometry using a structure-driven temporal map

  • José Martínez-Carranza
  • Andrew Calway

We describe a method for visual odometry using a single camera based on an EKF framework. Previous work has shown that filtering based approaches can achieve accuracy performance comparable to that of optimisation methods providing that large numbers of features are used. However, computational requirements are signicantly increased and frame rates are low. We address this by employing higher level structure - in the form of planes - to efficiently parameterise features and so reduce the filter state size and computational load. Moreover, we extend a 1-point RANSAC outlier rejection method to the case of features lying on planes. Results of experiments with both simulated and real-world data demonstrate that the method is effective, achieving comparable accuracy whilst running at significantly higher frame rates.

IROS Conference 2012 Conference Paper

Egocentric Real-time Workspace Monitoring using an RGB-D camera

  • Dima Damen
  • Andrew P. Gee
  • Walterio W. Mayol-Cuevas
  • Andrew Calway

We describe an integrated system for personal workspace monitoring based around an RGB-D sensor. The approach is egocentric, facilitating full flexibility, and operates in real-time, providing object detection and recognition, and 3D trajectory estimation whilst the user undertakes tasks in the workspace. A prototype on-body system developed in the context of work-flow analysis for industrial manipulation and assembly tasks is described. The system is evaluated on two tasks with multiple users, and results indicate that the method is effective, giving good accuracy performance.

IROS Conference 2012 Conference Paper

Predicting Micro Air Vehicle landing behaviour from visual texture

  • John Bartholomew
  • Andrew Calway
  • Walterio W. Mayol-Cuevas

We introduce a framework to predict the landing behaviour of a Micro Air Vehicle (MAV) from the appearance of the landing surface. We approach this problem by learning a mapping from visual texture observed from an onboard camera to the landing behaviour on a set of sample materials. In this case we exemplify our framework by predicting the yaw angle of the MAV after landing. Our framework demonstrates the applicability of established texture classification methods usually tested on stationary camera setups for the more challenging case of textures observed from a MAV. Results for supervised training demonstrate good estimation of the landing behaviour and motivate future work to implement autonomous decision making strategies and other behaviour predictions based on imagery.

v2026.09.13