Arrow Research search

Author name cluster

Alexandre M. Bayen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

AAMAS Conference 2026 Conference Paper

Average Unfairness in Routing Games

  • Pan-Yang Su
  • Arwa Alanqary
  • Bryce L Ferguson
  • Manxi Wu
  • Alexandre M. Bayen
  • Shankar Sastry

We propose average unfairness as a new measure of fairness in routing games, defined as the ratio between the average latency and the minimum latency experienced by users. This measure is a natural complement to two existing unfairness notions: loaded unfairness, which compares maximum and minimum latencies of routes with positive flow, and user equilibrium (UE) unfairness, which compares maximum latency with the latency of a Nash equilibrium. We show that the worst-case values of all three unfairness measures coincide and are characterized by a steepness parameter intrinsic to the latency function class. We show that average unfairness is always no greater than loaded unfairness, and the two measures are equal only when the flow is fully fair. Besides that, we offer a complete comparison of the three unfairness measures, which, to the best of our knowledge, is the first theoretical analysis in this direction. Finally, we study the constrained system optimum (CSO) problem, where one seeks to minimize total latency subject to an upper bound on unfairness. We prove that, for the same tolerance level, the optimal flow under an average unfairness constraint achieves lower total latency than any flow satisfying a loaded unfairness constraint. We show that such improvement is always strict in parallel-link networks and establish sufficient conditions for general networks. We further illustrate the latter with numerical examples. Our results provide theoretical guarantees and valuable insights for evaluating fairness-efficiency tradeoffs in network routing.

ICRA Conference 2025 Conference Paper

Decentralized Vehicle Coordination: The Berkeley DeepDrive Drone Dataset and Consensus-Based Models

  • Fangyu Wu 0003
  • Dequan Wang
  • Minjune Hwang
  • Chenhui Hao
  • Jiawei Lu
  • Jiamu Zhang
  • Christopher Chou
  • Trevor Darrell

A significant portion of roads, particularly in densely populated developing countries, lacks explicitly defined right-of-way rules. These understructured roads pose substantial challenges for autonomous vehicle motion planning, where efficient and safe navigation relies on understanding decentralized human coordination for collision avoidance. This coordination, often termed “social driving etiquette, ” remains underexplored due to limited open-source empirical data and suitable modeling frameworks. In this paper, we present a novel dataset and modeling framework designed to study motion planning in these understructured environments. The dataset includes 20 aerial videos of representative scenarios, an image dataset for training vehicle detection models, and a development kit for vehicle trajectory estimation. We demonstrate that a consensus-based modeling approach can effectively explain the emergence of priority orders observed in our dataset, and is therefore a viable framework for decentralized collision avoidance planning.

ICRA Conference 2022 Conference Paper

Deploying Traffic Smoothing Cruise Controllers Learned from Trajectory Data

  • Nathan Lichtlé
  • Eugene Vinitsky
  • Matthew Nice
  • Benjamin Seibold
  • Dan Work
  • Alexandre M. Bayen

Autonomous vehicle-based traffic smoothing con-trollers are often not transferred to real-world use due to challenges in calibrating many-agent traffic simulators. We show a pipeline to sidestep such calibration issues by collecting trajectory data and learning controllers directly from trajectory data that are then deployed zero-shot onto the highway. We construct a dataset of 772. 3 kilometers of recorded drives on the I–24. We then construct a simple simulator using the recorded drives as the lead vehicle in front of a simulated platoon consisting of one autonomous vehicle and five human followers. Using policy-gradient methods with an asymmetric critic to learn the controller, we show that we are able to improve average MPG by 11% in simulation on congested trajectories. We deploy this controller to a mixed platoon of 4 autonomous Toyota RAV-4's and 7 human drivers in a validation experiment and demonstrate that the expected time-gap of the controller is maintained in the real world test. Finally, we release the driving dataset [1], the simulator, and the trained controller at https://github.com/nathanlct/trajectory-training-icra.

AAMAS Conference 2022 Conference Paper

Learning Generalizable Multi-Lane Mixed-Autonomy Behaviors in Single Lane Representations of Traffic

  • Abdul Rahman Kreidieh
  • Yibo Zhao
  • Samyak Parajuli
  • Alexandre M. Bayen

This paper tackles the problem of learning generalizable congestionmitigation strategies in simple representations of traffic. In particular, we look to mixed-autonomy ring roads as depictions of instabilities common to many generic settings, and ask the question: What features are needed to ensure that policies here can be adapted to typical multi-lane highways? To answer this, we study the implications of the scale of the source task and the modeling of pseudo-lane change events within it on the transferability of policies learned to complex networks. Our findings suggest that negating the effects of boundary conditions and introducing lane changes that approximately match trends in more complex systems can significantly improve the generalizability of learned behaviors.

AAMAS Conference 2022 Conference Paper

Solving N-Player Dynamic Routing Games with Congestion: A Mean-Field Approach

  • Theophile Cabannes
  • Mathieu Laurière
  • Julien Perolat
  • Raphael Marinier
  • Sertan Girgin
  • Sarah Perrin
  • Olivier Pietquin
  • Alexandre M. Bayen

The recent emergence of navigational tools has changed traffic patterns and has now enabled new types of congestion-aware routing control like dynamic road pricing. Using the fundamental diagram of traffic flows – applied in macroscopic and mesoscopic traffic modeling – the article introduces a new 𝑁-player dynamic routing game with explicit congestion dynamics. The model is well-posed and can reproduce heterogeneous departure times and congestion spill back phenomena. However, as Nash equilibrium computations are PPAD-complete, solving the game becomes intractable for large but realistic numbers of vehicles 𝑁. Therefore, the corresponding mean field game is also introduced. Experiments were performed on several classical benchmark networks of the traffic community: the Pigou, Braess, and Sioux Falls networks with heterogeneous origin, destination and departure time tuples. The Pigou and the Braess examples reveal that the mean field approximation is generally very accurate and computationally efficient as soon as the number of vehicles exceeds a few dozen. On the Sioux Falls network (76 links, 100 time steps), this approach enables learning traffic dynamics with more than 14, 000 vehicles.

ICRA Conference 2021 Conference Paper

Reachability Analysis for FollowerStopper: Safety Analysis and Experimental Results

  • Fang-Chieh Chou
  • Marsalis T. Gibson
  • Rahul Bhadani
  • Alexandre M. Bayen
  • Jonathan Sprinkle

Motivated by earlier work and the developer of a new algorithm, the FollowerStopper, this article uses reachability analysis to verify the safety of the FollowerStopper algorithm, which is a controller designed for dampening stop-and-go traffic waves. With more than 1100 miles of driving data collected by our physical platform, we validate our analysis results by comparing it to human driving behaviors. The FollowerStopper controller has been demonstrated to dampen stop-and-go traffic waves at low speed, but previous analysis on its relative safety has been limited to upper and lower bounds of acceleration. To expand upon previous analysis, reachability analysis is used to investigate the safety at the speeds it was originally tested and also at higher speeds. Two formulations of safety analysis with different criteria are shown: distance-based and time headway-based. The FollowerStopper is considered safe with distance-based criterion. However, simulation results demonstrate that the FollowerStopper is not representative of human drivers - it follows too closely behind vehicles, specifically at a distance human would deem as unsafe. On the other hand, under the time headway-based safety analysis, the FollowerStopper is not considered safe anymore. A modified FollowerStopper is proposed to satisfy time-based safety criterion. Simulation results of the proposed FollowerStopper shows that its response represents human driver behavior better.

TIST Journal 2020 Journal Article

BISTRO

  • Sidney A. Feygin
  • Jessica R. Lazarus
  • Edward H. Forscher
  • Valentine Golfier-Vetterli
  • Jonathan W. Lee
  • Abhishek Gupta
  • Rashid A. Waraich
  • Colin J. R. Sheppard

The current trend toward urbanization and adoption of flexible and innovative mobility technologies will have complex and difficult-to-predict effects on urban transportation systems. Comprehensive methodological frameworks that account for the increasingly uncertain future state of the urban mobility landscape do not yet exist. Furthermore, few approaches have enabled the massive ingestion of urban data in planning tools capable of offering the flexibility of scenario-based design. This article introduces Berkeley Integrated System for Transportation Optimization (BISTRO), a new open source transportation planning decision support system that uses an agent-based simulation and optimization approach to anticipate and develop adaptive plans for possible technological disruptions and growth scenarios. The new framework was evaluated in the context of a machine learning competition hosted within Uber Technologies, Inc., in which over 400 engineers and data scientists participated. For the purposes of this competition, a benchmark model, based on the city of Sioux Falls, South Dakota, was adapted to the BISTRO framework. An important finding of this study was that in spite of rigorous analysis and testing done prior to the competition, the two top-scoring teams discovered an unbounded region of the search space, rendering the solutions largely uninterpretable for the purposes of decision-support. On the other hand, a follow-on study aimed to fix the objective function. It served to demonstrate BISTRO’s utility as a human-in-the-loop cyberphysical system: one that uses scenario-based optimization algorithms as a feedback mechanism to assist urban planners with iteratively refining objective function and constraints specification on intervention strategies. The portfolio of transportation intervention strategy alternatives eventually chosen achieves high-level regional planning goals developed through participatory stakeholder engagement practices.

ICRA Conference 2018 Conference Paper

Stabilizing Traffic with Autonomous Vehicles

  • Cathy Wu 0002
  • Alexandre M. Bayen
  • Ankur Mehta

Autonomous vehicles promise safer roads, energy savings, and more efficient use of existing infrastructure, among many other benefits. Although the effect of autonomous vehicles has been studied in the limits (near-zero or full penetration), the transition range requires new formulations, mathematical modeling, and control analysis. In this article, we study the ability of small numbers of autonomous vehicles to stabilize a single-lane system of human-driven vehicles. We formalize the problem in terms of linear string stability, derive optimality conditions from frequency-domain analysis, and pose the resulting nonlinear optimization problem. In particular, we introduce two conditions which simultaneously stabilize traffic while imposing a safety constraint on the autonomous vehicle and limiting degradation of performance. With this optimal linear controller in a system with typical human driver behavior, we can numerically determine that only a 6% uniform penetration of autonomously controlled vehicles (i. e. one per string of up to 16 human-driven vehicles) is necessary to stabilize traffic across all traffic conditions.

ICLR Conference 2018 Conference Paper

Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

  • Cathy Wu 0002
  • Aravind Rajeswaran
  • Yan Duan
  • Vikash Kumar
  • Alexandre M. Bayen
  • Sham M. Kakade
  • Igor Mordatch
  • Pieter Abbeel

Policy gradient methods have enjoyed great success in deep reinforcement learning but suffer from high variance of gradient estimates. The high variance problem is particularly exasperated in problems with long horizons or high-dimensional action spaces. To mitigate this issue, we derive a bias-free action-dependent baseline for variance reduction which fully exploits the structural form of the stochastic policy itself and does not make any additional assumptions about the MDP. We demonstrate and quantify the benefit of the action-dependent baseline through both theoretical analysis as well as numerical results, including an analysis of the suboptimality of the optimal state-dependent baseline. The result is a computationally efficient policy gradient algorithm, which scales to high-dimensional control problems, as demonstrated by a synthetic 2000-dimensional target matching task. Our experimental results indicate that action-dependent baselines allow for faster learning on standard reinforcement learning benchmarks and high-dimensional hand manipulation and synthetic tasks. Finally, we show that the general idea of including additional information in baselines for improved variance reduction can be extended to partially observed and multi-agent tasks.

ICML Conference 2015 Conference Paper

The Hedge Algorithm on a Continuum

  • Walid Krichene
  • Maximilian Balandat
  • Claire J. Tomlin
  • Alexandre M. Bayen

We consider an online optimization problem on a subset S of R^n (not necessarily convex), in which a decision maker chooses, at each iteration t, a probability distribution x^(t) over S, and seeks to minimize a cumulative expected loss, where each loss is a Lipschitz function revealed at the end of iteration t. Building on previous work, we propose a generalized Hedge algorithm and show a O(\sqrtt \log t) bound on the regret when the losses are uniformly Lipschitz and S is uniformly fat (a weaker condition than convexity). Finally, we propose a generalization to the dual averaging method on the set of Lebesgue-continuous distributions over S.

ICML Conference 2014 Conference Paper

On the convergence of no-regret learning in selfish routing

  • Walid Krichene
  • Benjamin Drighès
  • Alexandre M. Bayen

We study the repeated, non-atomic routing game, in which selfish players make a sequence of routing decisions. We consider a model in which players use regret-minimizing algorithms as the learning mechanism, and study the resulting dynamics. We are concerned in particular with the convergence to the set of Nash equilibria of the routing game. No-regret learning algorithms are known to guarantee convergence of a subsequence of population strategies. We are concerned with convergence of the actual sequence. We show that convergence holds for a large class of online learning algorithms, inspired from the continuous-time replicator dynamics. In particular, the discounted Hedge algorithm is proved to belong to this class, which guarantees its convergence.

ICRA Conference 2011 Conference Paper

Autonomous river navigation using the Hamilton-Jacobi framework for underactuated vehicles

  • Kevin Weekly
  • Leah Anderson
  • Andrew Tinka
  • Alexandre M. Bayen

Motorized floating sensors have distinct advantages over their non-actuated counterparts. A motorized unit can prevent the sensor from washing ashore or heading into dangerous areas, expanding the mission regions in which they can be feasibly operated. In this article, we present a control frame work and describe the physically realized system used to prove its effectiveness. The controller uses two minimum-time-to-reach (MTTR) functions-one giving the time to reach the center of the river and one giving the time to reach the shoreline. The MTTR functions are constructed from solutions to Hamilton Jacobi-Bellman-Isaacs (HJBI) Equations. Contours along these functions are used to define the state transition thresholds for an on-off controller. The first MTTR function is also used to construct the optimal bearing to travel back to the center of the river. We investigate the effectiveness of the controller using a software-in-the-loop (SIL) simulator. Using prototypes built at UC Berkeley, results from a field operational test in the Sacramento-San Joaquin River Delta are then presented to validate the simulation results.

v2026.09.13