Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Spleeter Stem Separation

Spleeter is an open-source tool from Deezer that splits a finished song into separate tracks (vocals, drums, bass, and more) using deep learning.

2 min readRead
Audio AI

CREPE Pitch Estimation

CREPE is a deep-learning model that estimates the fundamental frequency (pitch) of a monophonic audio signal directly from its raw waveform.

2 min readRead
Audio AI

PESQ and STOI Speech Quality Metrics

PESQ and STOI are standard objective metrics that score how good processed speech sounds and how understandable it is, without needing human listeners.

2 min readRead
Audio AI

Source-Filter Vocoding and WORLD

A vocoder is a tool that takes speech apart into its building blocks and rebuilds it.

2 min readRead
Audio AI

Mean Opinion Score Evaluation

Mean Opinion Score (MOS) is a 1-to-5 average rating from human listeners that measures how good synthesized or transmitted audio sounds.

2 min readRead
Audio AI

Constant-Q Transform for Audio

The Constant-Q Transform (CQT) is a frequency analysis that uses logarithmically spaced bins matched to musical pitch, instead of the evenly spaced bins…

2 min readRead
Audio AI

Filterbank and PLP Features

Filterbank and Perceptual Linear Prediction (PLP) features are ways of summarizing a speech signal into compact, perceptually meaningful numbers that machine…

2 min readRead
Audio AI

StyleTTS 2 Style Diffusion

StyleTTS 2 is a text-to-speech model that treats voice 'style' — prosody, emotion, and speaker timbre — as a random variable sampled with a diffusion model…

2 min readRead
Audio AI

FastPitch Pitch-Controllable TTS

FastPitch is a fast, non-autoregressive text-to-speech model that explicitly predicts the pitch (fundamental frequency) of every input token, letting you…

2 min readRead
Audio AI

Voicebox Flow-Matching Speech Generation

Voicebox is Meta's text-guided speech generation model trained with a flow-matching objective to 'fill in' masked audio, letting one model do zero-shot voice…

2 min readRead
Audio AI

Deep Noise Suppression Challenge

The Deep Noise Suppression (DNS) Challenge is a Microsoft-run competition that pushes researchers to build neural networks that strip background noise…

2 min readRead
Audio AI

Spectral Subtraction and Wiener Filtering

Spectral subtraction and Wiener filtering are the classic, pre-deep-learning workhorses of noise reduction.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.