Free AI library

Audio AI guidesFree forever.

117 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

117Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

117 of 1019 guides shown. Filter by track or search above.

Audio AI

Suno and Udio

Suno and Udio are the two leading consumer AI music generators that turn a short text prompt into a full, near-studio-quality song — complete with vocals…

2 min readRead
Audio AI

Symbolic Music Generation

Symbolic music generation creates music as structured notation — notes, pitches, durations, and timing (often as MIDI) — rather than as raw audio.

2 min readRead
Audio AI

Mel-Frequency Cepstral Coefficients

Mel-Frequency Cepstral Coefficients (MFCCs) are a compact set of numbers that summarize the shape of a sound's frequency spectrum the way human ears perceive…

2 min readRead
Audio AI

WaveNet

WaveNet, introduced by DeepMind in 2016, was a breakthrough neural network that generates raw audio one sample at a time, producing strikingly natural speech…

2 min readRead
Audio AI

Tacotron 2

Tacotron 2 is an end-to-end text-to-speech system from Google (2017) that turns written text directly into a mel-spectrogram, which a neural vocoder converts…

2 min readRead
Audio AI

Audio Deepfake Detection

Audio deepfake detection is the set of techniques used to tell whether a voice recording was spoken by a real human or synthesized/cloned by AI.

2 min readRead
Audio AI

Prosody Modeling

Prosody modeling teaches machines the melody of speech, the rhythm, pitch, stress, and pacing that ride on top of the words.

2 min readRead
Audio AI

Emotional Speech Synthesis

Emotional speech synthesis generates voices that sound happy, sad, angry, or calm, not just intelligible but believably felt.

2 min readRead
Audio AI

Voice Conversion

Voice conversion transforms one person's recorded speech so it sounds like it was spoken by someone else, while keeping the original words and timing.

2 min readRead
Audio AI

Speaker Diarization

Speaker diarization answers the question "who spoke when?" by splitting an audio recording into segments labeled by speaker identity.

2 min readRead
Audio AI

Speaker Verification

Speaker verification confirms whether a voice matches a specific claimed identity, acting as a voice-based password.

2 min readRead
Audio AI

Voice Activity Detection

Voice Activity Detection (VAD) decides, moment by moment, whether an audio signal contains human speech or just silence and noise.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.