Free AI library

Visual AI guidesFree forever.

133 plain-English guides, structured learning paths, and an open library — built by an independent 501(c)(3) nonprofit so anyone can understand modern AI.

133Free guides
1Topic tracks
~2 minPer guide
~4hReading time

Start here

Five outcome-based courses

Each course includes explicit outcomes, mapped competencies, practice activities, and an applied capstone.

Topic tracks

Browse by track

Jump into the area you care about. Every track has multiple plain-English guides.

Full library

All guides

133 of 1019 guides shown. Filter by track or search above.

Visual AI

Image Harmonization and Compositing

Image harmonization automatically adjusts a pasted foreground object so its color, lighting, and tone match the new background, making composites look real.

2 min readRead
Visual AI

SDXL and Cascaded Diffusion

SDXL is Stability AI's high-resolution text-to-image model that pairs a powerful base generator with a refiner, while cascaded diffusion chains multiple…

2 min readRead
Visual AI

GLIGEN Grounded Generation

GLIGEN (Grounded-Language-to-Image Generation) lets you control exactly where objects appear in a generated image by feeding the model bounding boxes…

2 min readRead
Visual AI

T2I-Adapter for Multi-Conditional Diffusion Control

Learn T2I-Adapter for multi-conditional diffusion control: how it steers Stable Diffusion with edges, depth, pose, and other conditions using a lightweight…

2 min readRead
Visual AI

Null-Text Inversion

Null-text inversion is a technique that lets you edit a real photo with a text-driven diffusion model like Stable Diffusion while keeping everything you…

2 min readRead
Visual AI

Custom Diffusion Multi-Concept Tuning

Custom Diffusion is a lightweight fine-tuning method that teaches a text-to-image model new personal concepts, like your dog or a specific chair, from just…

2 min readRead
Visual AI

Latent Blending and Image Interpolation

Latent blending mixes images by combining their compressed representations inside a model's latent space rather than averaging raw pixels.

2 min readRead
Visual AI

Vision-Language-Action Models for Robotics

Vision-Language-Action (VLA) models are large neural networks that take in camera images plus a written instruction and directly output robot motor commands.

2 min readRead
Visual AI

Diffusion Policy for Robot Control

Diffusion Policy applies the same denoising idea behind image generators like Stable Diffusion to robot control: instead of predicting a single next action…

2 min readRead
Visual AI

Make-A-Video Text-to-Video

Make-A-Video is Meta's 2022 system that turns a text prompt into a short video clip without ever training on labeled text-video pairs.

2 min readRead
Visual AI

Imagen Video Cascades

Imagen Video is Google's 2022 text-to-video system that builds a clip through a cascade of seven diffusion models, each adding more frames or more resolution.

2 min readRead
Visual AI

CogVideo and CogVideoX

CogVideo (2022) was the first large-scale open text-to-video model, and CogVideoX (2024) is its far more capable open-source successor from Tsinghua/Zhipu AI.

2 min readRead

Finished reading? Prove it.

Check what you learned with topic quizzes, then explore our structured courses and current certification requirements. Core guides remain free to read.