Arrow Research search

Author name cluster

Renato De Mori

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

JMLR Journal 2024 Journal Article

Open-Source Conversational AI with SpeechBrain 1.0

  • Mirco Ravanelli
  • Titouan Parcollet
  • Adel Moumen
  • Sylvain de Langen
  • Cem Subakan
  • Peter Plantinga
  • Yingzhi Wang
  • Pooneh Mousavi

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete recipes of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2024. ( edit, beta )

NeurIPS Conference 1991 Conference Paper

Neural Network - Gaussian Mixture Hybrid for Speech Recognition or Density Estimation

  • Yoshua Bengio
  • Renato De Mori
  • Giovanni Flammia
  • Ralf Kompe

The subject of this paper is the integration of multi-layered Artificial Neu(cid: 173) ral Networks (ANN) with probability density functions such as Gaussian mixtures found in continuous density Hidden Markov Models (HMM). In the first part of this paper we present an ANN/HMM hybrid in which all the parameters of the the system are simultaneously optimized with respect to a single criterion. In the second part of this paper, we study the relationship between the density of the inputs of the network and the density of the outputs of the networks. A few experiments are presented to explore how to perform density estimation with ANNs.

NeurIPS Conference 1989 Conference Paper

Speaker Independent Speech Recognition with Neural Networks and Speech Knowledge

  • Yoshua Bengio
  • Renato De Mori
  • Régis Cardin

We attempt to combine neural networks with knowledge from speech science to build a speaker independent speech recogni(cid: 173) tion system. This knowledge is utilized in designing the preprocessing, input coding, output coding, output supervision and architectural constraints. To handle the temporal aspect of speech we combine delays, copies of activations of hidden and output units at the input level, and Back-Propagation for Sequences (BPS), a learning algorithm for networks with local self-loops. This strategy is demonstrated in several experi(cid: 173) ments, in particular a nasal discrimination task for which the application of a speech theory hypothesis dramatically im(cid: 173) proved generalization.

NeurIPS Conference 1988 Conference Paper

Use of Multi-Layered Networks for Coding Speech with Phonetic Features

  • Yoshua Bengio
  • Régis Cardin
  • Renato De Mori
  • Piero Cosi

Preliminary results on speaker-independant speech recognition are reported. A method that combines expertise on neural networks with expertise on speech recognition is used to build the recognition systems. For transient sounds, event(cid: 173) driven property extractors with variable resolution in the time and frequency domains are used. For sonorant speech, a model of the human auditory system is preferred to FFT as a front-end module.

AAAI Conference 1982 Conference Paper

An Expert System for Interpreting Speech Patterns

  • Renato De Mori
  • Lorenza Saitta

Efficient syllabic hypothesization in continuous speech has been so far an unsolved problem. A novel solution based on the extraction of acoustic cues is proposed in this paper. This extraction is performed by parallel processes implementing an expert system represented by a grammar of frames.

v2026.09.13