Apple researchers propose STARFlow2 for unified text-and-image generation
Apple researchers describe STARFlow2, an architecture that combines a pretrained vision-language model with autoregressive normalizing flows to generate interleaved text and images through one causal mechanism.