Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

DiffWave Diffusion Vocoder

DiffWave is a diffusion-based vocoder that synthesizes audio by iteratively denoising random noise into a waveform, conditioned on a mel-spectrogram.

2 min readRead
Audio AI

Bark Generative Audio Model

Bark is an open-source text-to-audio model from Suno that generates not just speech but laughter, sighs, music, and sound effects directly from text prompts.

2 min readRead
Audio AI

Tortoise TTS Autoregressive Synthesis

Tortoise TTS is an open-source text-to-speech system prized for unusually natural, emotionally rich voices and strong voice cloning from just a few short…

2 min readRead
Audio AI

XTTS Cross-Lingual Voice Cloning

XTTS is Coqui's multilingual text-to-speech model that can clone a voice from a short clip and then speak in many different languages while preserving…

2 min readRead
Audio AI

SoundStorm Parallel Audio Generation

SoundStorm is a Google audio generation model that produces speech and sound in parallel rather than one token at a time, making high-quality audio synthesis…

2 min readRead
Audio AI

AudioGen Text-to-Audio Synthesis

AudioGen is a Meta model that turns text descriptions into realistic environmental sounds and sound effects, like 'dog barking while birds chirp.

2 min readRead
Audio AI

Stable Audio Latent Diffusion

Stable Audio is Stability AI's text-to-audio system that uses latent diffusion to generate music and sound effects, with explicit control over clip length.

2 min readRead
Audio AI

Kaldi Speech Recognition Toolkit

Kaldi is a free, open-source toolkit that became the dominant research platform for building speech recognition systems.

2 min readRead
Audio AI

Wav2Letter Convolutional ASR

Wav2Letter is an end-to-end speech recognition system from Facebook AI that used only convolutional neural networks, no recurrence.

2 min readRead
Audio AI

Jasper and QuartzNet ASR

Jasper and QuartzNet are NVIDIA's end-to-end convolutional speech recognition models, with QuartzNet being a dramatically smaller, efficient redesign…

2 min readRead
Audio AI

Diffusion Models for Audio

Diffusion models generate audio by learning to reverse a step-by-step noising process, turning random noise into coherent speech, music, or sound effects.

2 min readRead
Audio AI

Sound Event Detection

Sound event detection (SED) identifies what sounds occur in an audio stream and exactly when they start and stop.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.