Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

AI Summary
Google DeepMind has developed DiffusionGemma, which retrofits the existing Gemma 4 model into a diffusion model, utilizing less than 10% of the original training budget. While DiffusionGemma can generate 256 tokens simultaneously at a speed of about 1,500 tokens per second, its quality still falls short compared to the original autoregressive model, particularly in reasoning tasks.
From the source
Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks. The article Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model appeared
The full text couldn't be loaded here (the source may require a subscription).
View original at The Decoder