Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Mimi Streaming Audio Codec

Mimi is a neural audio codec that compresses speech into a tiny stream of discrete tokens in real time, so AI models can listen and speak with very low…

2 min readRead
Audio AI

Listen Attend and Spell

Listen, Attend and Spell (LAS) is a landmark 2015 neural network that transcribes speech directly into characters, with no hand-built pronunciation…

2 min readRead
Audio AI

SpecAugment for Speech Recognition

SpecAugment is a simple but powerful data augmentation method that masks and warps the spectrogram of speech to make recognition models more robust.

2 min readRead
Audio AI

Music Auto-Tagging

Music auto-tagging uses machine learning to listen to a song and automatically attach descriptive labels like genre, mood, instruments, and tempo.

2 min readRead
Audio AI

Musical Timbre Transfer

Timbre transfer reshapes the 'tone color' of audio so one instrument sounds like another, turning a hummed melody into a violin or a trumpet line…

2 min readRead
Audio AI

Acoustic Scene Classification

Acoustic scene classification (ASC) trains machines to recognize the environment a recording was made in, a busy street, a quiet park, a train, a cafe…

2 min readRead
Audio AI

NVIDIA Riva and NeMo Speech

NVIDIA Riva is a GPU-accelerated SDK for production speech AI (ASR, TTS, and translation), while NeMo is the open-source toolkit for training and fine-tuning…

2 min readRead
Audio AI

DeepSpeech Architecture

DeepSpeech is an end-to-end speech recognition model introduced by Baidu in 2014 that maps raw audio features directly to text using a recurrent neural…

2 min readRead
Audio AI

X-Vector Speaker Embeddings

X-vectors are fixed-length numerical fingerprints of a speaker's voice produced by a neural network, used to tell who is speaking regardless of what they say.

2 min readRead
Audio AI

Glow-TTS Monotonic Alignment

Glow-TTS is a text-to-speech model that learns to align text to speech on its own using a clever search trick, removing the need for a separate aligner.

2 min readRead
Audio AI

Parallel WaveGAN Vocoder

Parallel WaveGAN is a fast neural vocoder that turns a mel-spectrogram into a raw audio waveform using a small GAN, generating all samples at once.

2 min readRead
Audio AI

WaveGlow Flow-Based Vocoder

WaveGlow is a flow-based neural vocoder from NVIDIA that synthesizes speech waveforms from mel-spectrograms in a single pass without autoregression.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.