Arrow Research search

Author name cluster

Bertrand Douillard

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
1 author row

Possible papers

15

ICRA Conference 2022 Conference Paper

MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction

  • Balakrishnan Varadarajan
  • Ahmed Hefny
  • Avikalp Srivastava
  • Khaled S. Refaat
  • Nigamaa Nayakanti
  • Andre Cornman
  • Kan Chen
  • Bertrand Douillard

Predicting the future behavior of road users is one of the most challenging and important problems in autonomous driving. Applying deep learning to this problem requires fusing heterogeneous world state in the form of rich perception signals and map information, and inferring highly multi-modal distributions over possible futures. In this paper, we present MultiPath++, a future prediction model that achieves state-of-the-art performance on popular benchmarks. MultiPath++ improves the MultiPath architecture [34] by revisiting many design choices. The first key design difference is a departure from dense image-based encoding of the input world state in favor of a sparse encoding of heterogeneous scene elements: MultiPath++ consumes compact and efficient polylines to describe road features, and raw agent state information directly (e. g. , position, velocity, acceleration). We propose a context-aware fusion of these elements and develop a reusable multi-context gating fusion component. Second, we reconsider the choice of pre-defined static anchors, and develop a way to learn latent anchor embeddings end-to-end in the model. Lastly, we explore ensembling and output aggregation techniques—common in other ML domains—and find effective variants for our probabilistic multimodal output representation. We perform an extensive ablation on these design choices, and show that our proposed model achieves state-of-the-art performance on the Argoverse Motion Forecasting Competition [10] and the Waymo Open Dataset Motion Prediction Challenge [13].

ICRA Conference 2022 Conference Paper

Narrowing the coordinate-frame gap in behavior prediction models: Distillation for efficient and accurate scene-centric motion forecasting

  • DiJia Andy Su
  • Bertrand Douillard
  • Rami Al-Rfou
  • Cheol Park
  • Benjamin Sapp

Behavior prediction models have proliferated in recent years, especially in the popular real-world robotics application of autonomous driving, where representing the distribution over possible futures of moving agents is essential for safe and comfortable motion planning. In these models, the choice of coordinate frames to represent inputs and outputs has crucial trade offs which broadly fall into one of two categories. Agent-centric models transform inputs and perform inference in agent-centric coordinates. These models are intrinsically invari-ant to translation and rotation between scene elements, are best-performing on public leaderboards, but scale quadratically with the number of agents and scene elements. Scene-centric models use a fixed coordinate system to process all agents. This gives them the advantage of sharing representations among all agents, offering efficient amortized inference computation which scales linearly with the number of agents. However, these models have to learn invariance to translation and rotation between scene elements, and typically underperform agent-centric models. In this work, we develop knowledge distillation techniques between probabilistic motion forecasting models, and apply these techniques to close the gap in performance between agent-centric and scene-centric models. This improves scene-centric model performance by 13. 2% on the public Argoverse benchmark, 7. 8% on Waymo Open Dataset and up to 9. 4% on a large In-House dataset. These improved scene-centric models rank highly in public leaderboards and are up to 15 times more efficient than their agent-centric teacher counterparts in busy scenes.

ICRA Conference 2015 Conference Paper

Supervised Remote Robot with Guided Autonomy and Teleoperation (SURROGATE): A framework for whole-body manipulation

  • Paul Hebert
  • Jeremy Ma
  • James Borders
  • Alper Aydemir
  • Max Bajracharya
  • Nicolas Hudson
  • Krishna Shankar
  • Sisir Karumanchi

The use of the cognitive capabilties of humans to help guide the autonomy of robotics platforms in what is typically called “supervised-autonomy” is becoming more commonplace in robotics research. The work discussed in this paper presents an approach to a human-in-the-loop mode of robot operation that integrates high level human cognition and commanding with the intelligence and processing power of autonomous systems. Our framework for a “Supervised Remote Robot with Guided Autonomy and Teleoperation” (SURROGATE) is demonstrated on a robotic platform consisting of a pan-tilt perception head, two 7-DOF arms connected by a single 7-DOF torso, mounted on a tracked-wheel base. We present an architecture that allows high-level supervisory commands and intents to be specified by a user that are then interpreted by the robotic system to perform whole body manipulation tasks autonomously. We use a concept of “behaviors” to chain together sequences of “actions” for the robot to perform which is then executed real time.

ICRA Conference 2014 Conference Paper

Crowdsourced saliency for mining robotically gathered 3D maps using multitouch interaction on smartphones and tablets

  • Matthew Johnson-Roberson
  • Mitch Bryson
  • Bertrand Douillard
  • Oscar Pizarro
  • Stefan B. Williams

This paper presents a system for crowdsourcing saliency interest points for robotically gathered 3D maps rendered on smartphones and tablets. An app was created that is capable of interactively rendering 3D reconstructions gathered with an Autonomous Underwater Vehicle. Through hundreds of thousands of logged user interactions with the models we attempt to data-mine salient interest points. To this end we propose two models for calculating saliency from human interaction with the data. The first uses the view frustum of the camera to track the amount of time points are on screen. The second treats the camera's path as a time series and uses a Hidden Markov model to learn the classification of salient and non-salient points. To provide a comparison to existing techniques, several traditional visual saliency approaches are applied to orthographic views of the models' photo-texturing. The results of all approaches are validated with human attention ground truth gathered using a remote gaze-tracking system that recorded the locations of the person's attention while exploring the models.

ICRA Conference 2014 Conference Paper

Multimodal learning for autonomous underwater vehicles from visual and bathymetric data

  • Dushyant Rao
  • Mark De Deuge
  • Navid Nourani-Vatani
  • Bertrand Douillard
  • Stefan B. Williams
  • Oscar Pizarro

Autonomous Underwater Vehicles (AUVs) gather large volumes of visual imagery, which can help monitor marine ecosystems and plan future surveys. One key task in marine ecology is benthic habitat mapping, the classification of large regions of the ocean floor into broad habitat categories. Since visual data only covers a small fraction of the ocean floor, traditional habitat mapping is performed using shipborne acoustic multi-beam data, with visual data as ground truth. However, given the high resolution and rich textural cues in visual data, an ideal approach should explicitly utilise visual features in the classification process. To this end, we propose a multimodal model which utilises visual data and shipborne multi-beam bathymetry to perform both classification and sampling tasks. Our algorithm learns the relationship between both modalities, but is also effective when visual data is missing. Our results suggest that by performing multimodal learning, classification performance is improved in scenarios where visual data is unavailable, such as the habitat mapping scenario. We also demonstrate empirically that the model is able to perform generative tasks, producing plausible samples from the underlying data-generating distribution.

IROS Conference 2014 Conference Paper

Small body surface mobility with a limbed robot

  • Daniel M. Helmick
  • Bertrand Douillard
  • Max Bajracharya

This paper describes the development of hard- ware, software, and algorithms for a prototype limbed robot capable of surface mobility on small bodies (asteroids and comets). It also describes the development of a laboratory testbed capable of simulating the micro-gravity and terrain of small bodies. A path following algorithm that uses visual odometry, robot body kinematics, and a variety of specialized gaits was used to demonstrate micro-gravity mobility with as few as 12 actuators. A mapping algorithm is also demonstrated that will enable path planning, limb trajectory planning, mobile grasping, and foot placement in future work. The results of this paper demonstrate that robust, stable, and precise small body mobility is feasible with a limbed robot.

ICRA Conference 2013 Conference Paper

Multi-sensor identity tracking with event graphs

  • Peter Morton
  • Bertrand Douillard
  • James Patrick Underwood

The ability to track moving objects is a key part of autonomous robot operation in real-world environments. Whilst for many tasks knowing the positions of objects may be sufficient, tracking the identity of targets may also be desirable. When objects are well separated preserving identities is trivial, however, the identities of objects that pass close to one another may become confused. This paper considers methods to maintain the identities of tracked objects using a combination of LIDAR and video data. When objects are well separated, they are tracked using location information from the LIDAR. When objects move together and their identities cannot be resolved, interactions are recorded and later resolved using appearance models. A vision based approach is adapted for use with LIDAR data and a new method for identity reasoning is proposed. The methods are validated on a dataset comprising a total of 37906 manually labelled point cloud segments.

ICRA Conference 2012 Conference Paper

An occlusion-aware feature for range images

  • Alastair James Quadros
  • James Patrick Underwood
  • Bertrand Douillard

This paper presents a novel local feature for 3D range image data called `the line image'. It is designed to be highly viewpoint invariant by exploiting the range image to efficiently detect 3D occupancy, producing a representation of the surface, occlusions and empty spaces. We also propose a strategy for defining keypoints with stable orientations which define regions of interest in the scan for feature computation. The feature is applied to the task of object classification on sparse urban data taken with a Velodyne laser scanner, producing good results.

ICRA Conference 2012 Conference Paper

Scan segments matching for pairwise 3D alignment

  • Bertrand Douillard
  • Alastair James Quadros
  • Peter Morton
  • James Patrick Underwood
  • Mark De Deuge
  • S. Hugosson
  • M. Hallstrom
  • Tim Bailey

This paper presents a method for pairwise 3D alignment which solves data association by matching scan segments across scans. Generating accurate segment associations allows to run a modified version of the Iterative Closest Point (ICP) algorithm where the search for point-to-point correspondences is constrained to associated segments. The novelty of the proposed approach is in the segment matching process which takes into account the proximity of segments, their shape, and the consistency of their relative locations in each scan. Scan segmentation is here assumed to be given (recent studies provide various alternatives [10], [19]). The method is tested on seven sequences of Velodyne scans acquired in urban environments. Unlike various other standard versions of ICP, which fail to recover correct alignment when the displacement between scans increases, the proposed method is shown to be robust to displacements of several meters. In addition, it is shown to lead to savings in computational times which are potentially critical in real-time applications.

IROS Conference 2011 Conference Paper

Combining radar and vision for self-supervised ground segmentation in outdoor environments

  • Annalisa Milella
  • Giulio Reina
  • James Patrick Underwood
  • Bertrand Douillard

Ground segmentation is critical for a mobile robot to successfully accomplish its tasks in challenging environments. In this paper, we propose a self-supervised radar-vision classification system that allows an autonomous vehicle, operating in natural terrains, to automatically construct online a visual model of the ground and perform accurate ground segmentation. The system features two main phases: the training phase and the classification phase. The training stage relies on radar measurements to drive the selection of ground patches in the camera images, and learn online the visual appearance of the ground. In the classification stage, the visual model of the ground can be used to perform high level tasks such as image segmentation and terrain classification, as well as to solve radar ambiguities. The proposed method leads to the following main advantages: (a) a self-supervised training of the visual classifier, where the radar allows the vehicle to automatically acquire a set of ground samples, eliminating the need for time-consuming manual labeling; (b) the ground model can be continuously updated during the operation of the vehicle, thus making it feasible the use of the system in long range and long duration navigation applications. This paper details the proposed system and presents the results of experimental tests conducted in the field by using an unmanned vehicle.

ICRA Conference 2011 Conference Paper

On the segmentation of 3D LIDAR point clouds

  • Bertrand Douillard
  • James Patrick Underwood
  • Noah Kuntz
  • Vsevolod Vlaskine
  • Alastair James Quadros
  • Peter Morton
  • Alon Frenkel

This paper presents a set of segmentation methods for various types of 3D point clouds. Segmentation of dense 3D data (e. g. Riegl scans) is optimised via a simple yet efficient voxelisation of the space. Prior ground extraction is empirically shown to significantly improve segmentation performance. Segmentation of sparse 3D data (e. g. Velodyne scans) is addressed using ground models of non-constant resolution either providing a continuous probabilistic surface or a terrain mesh built from the structure of a range image, both representations providing close to real-time performance. All the algorithms are tested on several hand labeled data sets using two novel metrics for segmentation evaluation.

IROS Conference 2010 Conference Paper

Hybrid elevation maps: 3D surface models for segmentation

  • Bertrand Douillard
  • James Patrick Underwood
  • Narek Melkumyan
  • Surya P. N. Singh
  • Shrihari Vasudevan
  • Christopher Joseph Brunner
  • Alastair James Quadros

This paper presents an algorithm for segmenting 3D point clouds. It extends terrain elevation models by incorporating two types of representations: (1) ground representations based on averaging the height in the point cloud, (2) object models based on a voxelisation of the point cloud. The approach is deployed on Riegl data (dense 3D laser data) acquired in a campus type of environment and compared against six other terrain models. Amongst elevation models, it is shown to provide the best fit to the data as well as being unique in the sense that it jointly performs ground extraction, overhang representation and 3D segmentation. We experimentally demonstrate that the resulting model is also applicable to path planning.

IROS Conference 2008 Conference Paper

A self-supervised architecture for moving obstacles classification

  • Roman Katz
  • Bertrand Douillard
  • Juan I. Nieto 0001
  • Eduardo M. Nebot

This work introduces a self-supervised, multi-sensor architecture that performs automatic moving obstacles classification. Our approach presents a hierarchical scheme that relies on the ldquostabilityrdquo of a subset of features given by a sensor to perform an initial robust classification based on unsupervised techniques. The obtained results are used as labels to train a set of supervised classifiers, which can be then combined to improve the final classification accuracy. The proposed architecture is general and can be instantiated in a variety of ways, using different sensors and classifiers. The applicability and validity of the proposed architecture is evaluated for a particular realization based on range and visual information that achieves 83% accuracy without using manually labeled data. Experimental results also demonstrate how accuracy can be maintained through self-training capabilities when working conditions change.

IROS Conference 2007 Conference Paper

A spatio-temporal probabilistic model for multi-sensor object recognition

  • Bertrand Douillard
  • Dieter Fox
  • Fabio Ramos 0001

This paper presents a general framework for multi-sensor object recognition through a discriminative probabilistic approach modelling spatial and temporal correlations. The algorithm is developed in the context of Conditional Random Fields (CRFs) trained with virtual evidence boosting. The resulting system is able to integrate arbitrary sensor information and incorporate features extracted from the data. The spatial relationships captured by are further integrated into a smoothing algorithm to improve recognition over time. We demonstrate the benefits of modelling spatial and temporal relationships for the problem of detecting cars using laser and vision data in outdoor environments.

IROS Conference 2006 Conference Paper

Hierarchical Environment Model for Fusing Information from Human Operators and Robots

  • Tobias Kaupp
  • Bertrand Douillard
  • Ben Upcroft
  • Alexei Makarenko

This paper considers the problem of building environment models by fusing information gathered by robotic platforms with human perceptual information. Rich environment models are required in real applications for both autonomous operation of robots and to support human decision making. Hierarchical models are well suited to represent complex environments because they: offer multiple abstractions of the available information to support analysis and decision-making, and permit the incorporation of higher-level human observations. The contributions of this paper are two-fold: (1) development of a probabilistic three-level environment model for distributed information gathering, and (2) experimental demonstration of fully decentralized, cooperative human-robot information gathering using an outdoor sensor network comprised of an unmanned air vehicle, a ground vehicle, and two human operators. Several information exchange patterns are presented which qualitatively demonstrate human-robot information fusion

v2026.09.13