技术指南

Adversarial Diffusion Distillation and Turbo Models

Adversarial Diffusion Distillation (ADD) is a training method that turns a slow, many-step diffusion model into a student that generates images in one to four steps.

  • 4 分钟阅读
  • 最后更新
在本页4 分钟阅读
  1. 概述
  2. 深入探讨
  3. 战略影响
  4. The Future of Adversarial Diffusion Distillation and Turbo Models
  5. 现实世界的实施
  6. 风险与防护栏
  7. 实施路线图
  8. 不断探索
  9. 常见问题

概述

It combines an adversarial loss from a discriminator with guidance from the original teacher model. ADD produced SDXL Turbo, and a latent variant was used for fast models such as SD3 Turbo and, according to Black Forest Labs, Flux.1 schnell. The result is near-real-time image generation on ordinary hardware, at some cost in diversity and fine control.

深入探讨

A standard diffusion model creates an image by removing noise over many steps, commonly 20 to 50, and each step is a full pass through a large network. Distillation trains a student network to reach a similar result in far fewer steps. Stability AI introduced Adversarial Diffusion Distillation in a November 2023 paper by Axel Sauer and colleagues, released alongside SDXL Turbo. The student starts from the pretrained SDXL weights and is trained with two losses. The adversarial loss comes from a discriminator that tries to tell the student's outputs from real images. In ADD, this discriminator is built on a frozen pretrained vision backbone (DINOv2) with small trainable heads. The distillation loss uses the original SDXL as a teacher. The student's output is noised again, the teacher denoises it, and the student is pushed toward the teacher's prediction, a form of score distillation. The adversarial term keeps single-step images sharp, and the teacher term keeps them faithful to what the large model knows. A later variant, Latent Adversarial Diffusion Distillation (LADD), runs the discriminator in latent space using the teacher's own features, which avoids expensive decoding to pixels. It was used for SD3 Turbo, and Black Forest Labs has said Flux.1 schnell was trained this way. The trade-offs are real. Distilled models usually lose some sample diversity, so different seeds for the same prompt can look alike. Classifier-free guidance is typically built in, so the guidance scale and negative prompts behave differently or barely work. Fine detail, text rendering and unusual compositions may lag behind the teacher. A common misconception is that a turbo model is just the full model run with fewer steps. A normal model run at one step gives a blurry mess, and the speed comes from retraining. Other routes to fast sampling include consistency models, LCM, progressive distillation and SDXL Lightning.

战略影响

成本与预算

多年来,架构决策决定着性能和运营成本。

更清晰的判决

技术教育帮助团队选择正确的堆栈,而不仅仅是最新的堆栈。

质量控制

更好的工程选择可以减少生产中的可靠性事故。

The Future of Adversarial Diffusion Distillation and Turbo Models

Few-step generation has become an expected option for new image models, and similar distillation ideas are being applied to video, where each saved step matters even more. Research continues on closing the gaps in diversity and prompt control, for example by combining adversarial losses with distribution-matching objectives. Whether one-step models can fully match their teachers is still an open question, and the answer depends on how quality is measured. For users, the practical result is more choice: fast distilled models for drafts, previews and interactive tools, and slower full models when detail and control matter most.

现实世界的实施

A live drawing app regenerates the image every time the user types or edits a sketch, using SDXL Turbo at one step so results appear almost instantly.

A game studio prototyping concept art generates hundreds of quick variations with Flux.1 schnell at four steps, then refines the chosen ones with a slower full model.

A developer finds that negative prompts and high guidance values do little on a Turbo model, because guidance was built in during distillation. They adjust the prompt wording instead.

A small nonprofit runs a local image tool on a mid-range GPU and picks a distilled model because four-step generation keeps waits short enough for live workshops.

风险与防护栏

  • 优化一项基准测试可以隐藏更广泛的系统弱点。

  • 基础设施和维护成本常常被低估。

  • 随着系统变得更加复杂,安全性和可观察性差距可能会扩大。

实施路线图

  1. 在实施之前定义延迟、质量和成本目标。

  2. 在实际负载和数据条件下进行基准测试。

  3. 仪器监控错误、漂移和用户影响。

  4. 在扩展之前准备回滚和事件响应路径。

不断探索

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Adversarial Diffusion Distillation and Turbo Models quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

开始测验

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

常见问题

What is Adversarial Diffusion Distillation and Turbo Models?

Adversarial Diffusion Distillation (ADD) is a training method that turns a slow, many-step diffusion model into a student that generates images in one to four steps. It combines an adversarial loss from a discriminator with guidance from the original teacher model. ADD produced SDXL Turbo, and a latent variant was used for fast models such as SD3 Turbo and, according to Black Forest Labs, Flux.1 schnell. The result is near-real-time image generation on ordinary hardware, at some cost in diversity and fine control.

Which two losses does ADD combine?

ADD pairs a discriminator's adversarial loss, which keeps images sharp, with a score-distillation loss from the teacher, which keeps them faithful to the original model.

Which model was released alongside the ADD paper?

Stability AI released SDXL Turbo with the November 2023 ADD paper by Sauer and colleagues.

What backbone does ADD's discriminator use?

The discriminator uses frozen pretrained DINOv2 features and trains only small heads on top of them.

How does LADD differ from ADD?

LADD moves discrimination into latent space using the teacher's own features, which avoids expensive decoding to pixels.

Why do negative prompts often have little effect on turbo models?

Negative prompts act through classifier-free guidance. Distilled models usually absorb guidance during training and run without it, so negative prompts lose their effect.