LLM
Practical Guides
Multimodalilty
Text-Audio
- Moshi: a speech-text foundation model for real-time dialogue
- Spirit LM: Interleaved Spoken and Written Language Model
- Hertz Dev
Diffusion
Signal Processing
- Discrete-Time Signal Processing - Oppenheim & Schafer
- Introduction to Digital Audio Coding and Standards
- Spectral Audio Processing
Speech Processing
Books
- Speech and Language Processing
- Speech Processing Book - Aalto. University
- Edinborough’s Speech Zone Courses
Speech Recognition
Courses/Talks
- CMU Speech Recognition - Shinji’s course
- Recent Advances in End-to-End Automatic Speech Recognition - Jinyu Li’s Talk
Recent Papers
- Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
- Sequence Transduction with Recurrent Neural Networks - RNNT
- Listen, Attend and Spell
- Conformer: Convolution-augmented Transformer for Speech Recognition
Text-to-speech (TTS)
MFCC Extraction - Practical Cryptography
Recent Papers
Neural Audio Codecs
- Mimi Codec from Moshi
- Descript Audio Codec: High-Fidelity Audio Compression with Improved RVQGAN
- SoundStream: An End-to-End Neural Audio Codec
- Encodec: High Fidelity Neural Audio Compression
Acoustic Models
LM Based
- Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
- AudioLM: a Language Modeling Approach to Audio Generation
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Diffusion Based
- Audiobox: Unified Audio Generation with Natural Language Prompts
- Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
- E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Regression Based
- FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
- Tacotron 2: Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
Vocoders
Non-Autoregressive GAN Based
- BigVGAN: A Universal Neural Vocoder with Large-Scale Training
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
Non-Autoregressive Diffusion Based
- WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
- WaveGrad: Estimating Gradients for Waveform Generation
Non-Autoregressive Flow Based
- WaveGlow: A Flow-based Generative Network for Speech Synthesis
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis