Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Audio Captioning

Audio captioning generates a natural-language sentence describing the content of an audio clip, such as 'a train horn blares as it passes a level crossing.

2 min readRead
Audio AI

Whisper Speech Recognition

Whisper is OpenAI's open-source automatic speech recognition system that turns audio into text across 90+ languages.

2 min readRead
Audio AI

Wav2Vec 2.0

Wav2Vec 2.0 is Meta AI's self-supervised speech model that learns powerful audio representations from raw, unlabeled recordings.

2 min readRead
Audio AI

HuBERT Self-Supervised Speech

HuBERT (Hidden-Unit BERT) is Meta AI's self-supervised speech model that learns by predicting clustered audio units for masked segments, BERT-style.

2 min readRead
Audio AI

Neural Vocoders

A neural vocoder is a model that turns a compact acoustic representation, usually a mel-spectrogram, into an actual audible waveform.

2 min readRead
Audio AI

Neural Audio Codecs

Neural audio codecs use deep learning to compress sound into tiny streams of discrete tokens and reconstruct it with high fidelity.

2 min readRead
Audio AI

VALL-E and Codec Language Models

VALL-E reframed text-to-speech as a language modeling problem over audio codec tokens, enabling voice cloning from just three seconds of a sample.

2 min readRead
Audio AI

OpenAI Whisper

Whisper is OpenAI's open-source automatic speech recognition system that transcribes and translates spoken audio across dozens of languages.

2 min readRead
Audio AI

Automatic Music Transcription

Automatic Music Transcription (AMT) converts a raw audio recording of music into a symbolic notation like sheet music, MIDI, or a piano roll.

2 min readRead
Audio AI

Beat and Tempo Tracking

Beat and tempo tracking is the task of finding the steady pulse in music: where each beat falls and how fast the song moves in beats per minute (BPM).

2 min readRead
Audio AI

Audio Fingerprinting

Audio fingerprinting creates a compact, noise-resistant digital signature of a sound so it can be recognized later, even through background noise…

2 min readRead
Audio AI

Jukebox

Jukebox is OpenAI's 2020 neural network that generates raw music audio — complete with singing voices, instruments, and even lyrics in the style of specific…

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.