Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

EnCodec Audio Compression

EnCodec is Meta's high-fidelity neural audio codec that compresses speech and music at very low bitrates with quality rivaling far heavier formats.

2 min readRead
Audio AI

Residual Vector Quantization

Residual vector quantization (RVQ) is the technique that turns continuous audio embeddings into a compact stack of discrete codes by repeatedly quantizing…

2 min readRead
Audio AI

ECAPA-TDNN Speaker Recognition

ECAPA-TDNN is a neural network architecture that turns any speech clip into a compact 'voiceprint' embedding, enabling machines to tell who is speaking.

2 min readRead
Audio AI

Speaker Anti-Spoofing and ASVspoof

Anti-spoofing is the defensive layer that detects fake or replayed voices trying to fool voice-authentication systems.

2 min readRead
Audio AI

Speech Denoising with RNNoise

RNNoise is a tiny, fast neural network that strips background noise from speech in real time.

2 min readRead
Audio AI

Audio Embeddings and Representation Learning

Audio embeddings turn sound into compact numerical vectors that capture meaning, so machines can compare, search, and classify audio the way humans recognize…

2 min readRead
Audio AI

Acoustic Echo Cancellation

Acoustic echo cancellation (AEC) is the technology that stops you from hearing your own voice bounce back during a call.

2 min readRead
Audio AI

Speech Separation and the Cocktail Party Problem

Speech separation is the task of pulling individual voices apart from a recording where several people talk at once.

2 min readRead
Audio AI

Permutation Invariant Training

Permutation invariant training (PIT) is a clever training trick that lets a model separate multiple voices without caring which output slot each voice lands…

2 min readRead
Audio AI

Music Information Retrieval

Music Information Retrieval (MIR) is the field that teaches computers to analyze, understand, and search music from audio signals and scores.

2 min readRead
Audio AI

Demucs and HT-Demucs Music Source Separation

Learn how Demucs and HT-Demucs separate music into vocals and instrument stems using waveform, spectrogram, and transformer modeling.

2 min readRead
Audio AI

Audio Chord Recognition

Audio chord recognition is the task of automatically labeling the chords played throughout a song directly from its audio.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.