Arrow Research search

Author name cluster

Pete Warden

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

NeurIPS Conference 2021 Conference Paper

MLPerf Tiny Benchmark

  • Colby Banbury
  • Vijay Janapa Reddi
  • Peter Torelli
  • Nat Jeffries
  • Csaba Kiraly
  • Jeremy Holleman
  • Pietro Montino
  • David Kanter

Advancements in ultra-low-power tiny machine learning (TinyML) systems promise to unlock an entirely new class of smart applications. However, continued progress is limited by the lack of a widely accepted and easily reproducible benchmark for these systems. To meet this need, we present MLPerf Tiny, the first industry-standard benchmark suite for ultra-low-power tiny machine learning systems. The benchmark suite is the collaborative effort of more than 50 organizations from industry and academia and reflects the needs of the community. MLPerf Tiny measures the accuracy, latency, and energy of machine learning inference to properly evaluate the tradeoffs between systems. Additionally, MLPerf Tiny implements a modular design that enables benchmark submitters to show the benefits of their product, regardless of where it falls on the ML deployment stack, in a fair and reproducible manner. The suite features four benchmarks: keyword spotting, visual wake words, image classification, and anomaly detection.

NeurIPS Conference 2021 Conference Paper

Multilingual Spoken Words Corpus

  • Mark Mazumder
  • Sharad Chitlangia
  • Colby Banbury
  • Yiping Kang
  • Juan Ciro
  • Keith Achorn
  • Daniel Galvez
  • Mark Sabini

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4. 0. The dataset contains more than 340, 000 keywords, totaling 23. 4 million 1-second spoken examples (over 6, 000 hours). The dataset has many use cases, ranging from voice-enabled consumer devices to call center automation. We generate this dataset by applying forced alignment on crowd-sourced sentence-level audio to produce per-word timing estimates for extraction. All alignments are included in the dataset. We provide a detailed analysis of the contents of the data and contribute methods for detecting potential outliers. We report baseline accuracy metrics on keyword spotting models trained from our dataset compared to models trained on a manually-recorded keyword dataset. We conclude with our plans for dataset maintenance, updates, and open-sourced code.

v2026.09.13