Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Beamforming and Microphone Arrays

Beamforming uses multiple microphones to listen in a chosen direction, amplifying sound from a target while suppressing everything else.

2 min readRead
Audio AI

Riffusion Spectrogram Diffusion

Riffusion is a clever hack that generates music by treating sound as a picture: it fine-tunes the Stable Diffusion image model to paint spectrograms, then…

2 min readRead
Audio AI

MusicLM: Hierarchical Semantic and Acoustic Tokens

Learn how Google MusicLM uses hierarchical semantic and acoustic tokens, SoundStream, and MuLan to turn text prompts into coherent music.

2 min readRead
Audio AI

Noise2Noise Speech Enhancement

Noise2Noise is a training trick that lets a model learn to remove noise without ever seeing a clean reference, by learning from pairs of differently-noisy…

2 min readRead
Audio AI

Conv-TasNet Time-Domain Separation

Conv-TasNet is a neural network that separates mixed audio (like two people talking at once) by working directly on the raw sound waveform instead…

2 min readRead
Audio AI

Dual-Path RNN Separation

Dual-Path RNN (DPRNN) is an audio separation architecture that splits a very long sequence of audio features into short overlapping chunks and processes them…

2 min readRead
Audio AI

Open-Unmix Music Separation

Open-Unmix (UMX) is an open-source deep learning system that splits a song into its parts: vocals, drums, bass, and other instruments.

2 min readRead
Audio AI

Whisper Timestamped Word Alignment

Whisper word alignment pins each transcribed word to an exact start and end time in the audio.

2 min readRead
Audio AI

Music Tagging with Transformers

Music tagging uses transformer models to listen to a song and predict descriptive labels like genre, mood, instruments, and tempo.

2 min readRead
Audio AI

Onset Detection in Audio

Onset detection finds the precise moments when notes, beats, or sounds begin in an audio signal.

2 min readRead
Audio AI

MelGAN Generative Vocoder

MelGAN is a fully convolutional GAN-based vocoder that turns mel-spectrograms into raw audio waveforms in a single fast forward pass.

2 min readRead
Audio AI

UnivNet Multi-Resolution Vocoder

UnivNet is a GAN vocoder that judges generated audio using multiple spectrograms computed at different STFT resolutions, sharpening high-frequency detail.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.