Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Cover Song Identification

Cover song identification detects when two very different-sounding recordings are actually the same underlying song — a live acoustic version, a remix…

2 min readRead
Audio AI

DDSP Differentiable Audio Synthesis

DDSP (Differentiable Digital Signal Processing) fuses classic synthesizer building blocks with neural networks, so deep learning can control oscillators…

2 min readRead
Audio AI

VITS End-to-End Speech Synthesis

VITS is a text-to-speech model that turns text directly into raw audio waveforms in a single trained system, skipping the usual two-stage pipeline.

2 min readRead
Audio AI

FastSpeech and Non-Autoregressive TTS

FastSpeech generates an entire speech spectrogram in parallel rather than one frame at a time, making synthesis dramatically faster and more stable.

2 min readRead
Audio AI

NaturalSpeech and Latent Diffusion TTS

NaturalSpeech is a line of Microsoft TTS research aiming for human-level speech quality, with later versions using latent diffusion to generate rich, natural…

2 min readRead
Audio AI

Speech Emotion Recognition

Speech Emotion Recognition (SER) is AI that detects a speaker's emotional state — anger, joy, sadness, frustration — from the sound of their voice, not just…

2 min readRead
Audio AI

Moshi Full-Duplex Speech

Moshi is an open-source, real-time voice AI from Kyutai that talks and listens at the same time — full-duplex — instead of taking strict turns.

2 min readRead
Audio AI

Speech-to-Speech Translation

Speech-to-Speech Translation (S2ST) takes spoken words in one language and produces spoken words in another — ideally preserving the speaker's voice, tone…

2 min readRead
Audio AI

HiFi-GAN and GAN Vocoders

HiFi-GAN is a generative-adversarial vocoder that turns a mel-spectrogram into a raw audio waveform almost instantly, producing studio-quality speech far…

2 min readRead
Audio AI

Grapheme-to-Phoneme Conversion

Grapheme-to-phoneme (G2P) conversion translates written letters into the sounds a speech system should actually pronounce.

2 min readRead
Audio AI

Text Normalization for Speech

Text normalization is the front-end step that rewrites raw written text into fully spoken-out words before a speech system says it.

2 min readRead
Audio AI

SoundStream Neural Codec

SoundStream is Google's end-to-end neural audio codec that compresses speech and music to extremely low bitrates while preserving quality.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.