12 Sources
[1]
Google's latest DiffusionGemma open AI model comes with a 4x speed boost
Another day, another AI model from Google. This time, Google DeepMind has released a new member of the Gemma 4 open model family, but it's fundamentally different from the rest of the lineup. DiffusionGemma doesn't generate outputs linearly like most AI models. Instead, it can produce an entire
[2]
Google's new open-weights AI model uses image-generation methods to output text faster
The boffins on Google's DeepMind team unveiled an experimental new language model this week that uses techniques originally developed for AI image generators to boost text output performance by as much as 4x when running on resource-constrained consumer hardware. It's free to download and you can
[3]
Google unveils DiffusionGemma, an AI model that breaks free of left-to-right processing
Rather than generating text word by word, Google's experimental open-source model drafts entire passages simultaneously using diffusion, resulting in up to 4x faster inference. Extremely powerful large language models (LLMs) still operate as though they're typing on a keyboard, processing
[4]
DiffusionGemma is Google's fastest AI yet, but it comes with a big trade-off
Output quality is still inferior to Gemma 4, so it's more of an experimental tool than a finished product. Google has released DiffusionGemma, an experimental AI model that takes a very different approach to how most chatbots generate text today. Instead of writing one word after another in a
[5]
NVIDIA Accelerates Google DeepMind's DiffusionGemma for Local AI
The new Diffusion Gemma open model generates text in parallel -- not one token at a time -- and is optimized to run on the NVIDIA RTX PRO platform, NVIDIA DGX Spark systems and GeForce RTX GPUs. Today, Google DeepMind released DiffusionGemma -- an experimental open model built for exceptionally
[6]
Google's new DiffusionGemma model speeds up text generation by 4x
Google has unveiled DiffusionGemma, a new experimental AI model that generates text using diffusion rather than the autoregressive approach used by most large language models today. The company says the model can deliver up to four times faster text generation on dedicated GPUs while running on
[7]
Google's DiffusionGemma AI Hits 1,000 Tokens Per Second -- And It's Free
On NVIDIA NIM, the model arrived preconfigured at 8,192 tokens of context -- below the 64,000-token floor that agentic frameworks like Hermes Agent require -- meaning autonomous workflows won't run without manual reconfiguration. Google dropped DiffusionGemma today, an open model AI that generates
[8]
DiffusionGemma: 4x faster text generation
You can improve DiffusionGemma's performance on specific tasks through fine-tuning. In the example below, Unsloth fine-tuned DiffusionGemma to play Sudoku -- a task autoregressive models struggle with because each token depends on future tokens. DiffusionGemma's bi-directional attention makes this
[9]
Google open-sources speedy DiffusionGemma text diffusion model
Google open-sources speedy DiffusionGemma text diffusion model Google LLC today released DiffusionGemma, a large language model based on an emerging machine learning approach known as text diffusion. The company says that the algorithm can generate text four times faster than traditional LLMs.
[10]
NVIDIA Delivers Day-1 Support For DeepMind's DiffusionGemma Open Model Across RTX & DGX Platforms, 150 Tokens/s With DGX Spark
NVIDIA's entire RTX/DGX lineups are getting full support for Google DeepMind's DiffusionGemma Open AI model. Google Intros Its Newest Open AI Model: DiffusionGemma - NVIDIA Offers Full Support Across Its DGX & RTX Families The DiffusionGemma model is an open model designed to offer speedy text
[11]
Google unveils DiffusionGemma open AI model with up to 4x faster text generation
Google has introduced DiffusionGemma, an experimental open-weight AI model that explores diffusion-based text generation. Released under the Apache 2.0 license, the 26-billion-parameter Mixture-of-Experts (MoE) model moves beyond the sequential token-by-token generation used by traditional
[12]
Google launches DiffusionGemma with 4x faster text generation By Investing.com
Investing.com - Google (NASDAQ:GOOGL) released DiffusionGemma today, an experimental open model that generates text up to four times faster than traditional language models on dedicated GPUs. The 26 billion parameter Mixture of Experts model uses text diffusion to generate entire blocks of text
Share
Copy Link
Google DeepMind unveiled DiffusionGemma, an experimental open-source AI model that generates text in parallel rather than sequentially. The model achieves over 1,000 tokens per second on NVIDIA H100 GPUs and runs on consumer hardware with just 18GB of memory. However, the speed gains come with a notable trade-off in output quality compared to traditional models.
Google DeepMind has released DiffusionGemma, an experimental open-source AI model that fundamentally reimagines how language models generate text. Unlike conventional autoregressive models that produce text one token at a time from left to right, DiffusionGemma employs image generation techniques borrowed from systems like Stable Diffusion to create entire blocks of text simultaneously
2
. This latest addition to the Gemma 4 family marks a significant departure from traditional language model design, prioritizing speed and efficiency for local hardware deployment over the sequential processing that has dominated the field3
.
Source: Ars Technica
The model operates through a denoising process that starts with a canvas of random placeholder tokens and progressively refines them across multiple passes until coherent text emerges. DiffusionGemma can denoise up to 256 tokens per step instead of predicting one at a time, enabling parallel text generation that shifts the computational bottleneck from memory bandwidth to compute
5
. This architectural choice makes the model particularly well-suited for single-user scenarios where traditional models often leave GPUs underutilized.Built as a Mixture-of-Experts model with 26 billion total parameters, DiffusionGemma activates only 3.8 billion parameters during inference, allowing it to fit within the 18GB memory footprint of high-end consumer GPUs. In testing on an RTX 5090, the model delivers approximately 700 tokens per second, while a single NVIDIA H100 accelerator achieves over 1,000 tokens per second. These figures represent roughly four times the inference speed of similarly sized autoregressive models running in the same single-user regime
5
.
Source: NVIDIA
NVIDIA has optimized DiffusionGemma to run across its full hardware lineup, from GeForce RTX GPUs to the RTX PRO platform and DGX Spark systems
5
. The model's parallel generation approach transforms text generation from a memory-bound problem into a compute-bound workload, playing directly to the strengths of NVIDIA GPUs and their Tensor Cores5
. This makes DiffusionGemma particularly effective for developers and researchers running latency-sensitive, single-user workloads where interactive responsiveness matters.While DiffusionGemma achieves impressive speed gains, Google acknowledges the model does not match the output quality of standard Gemma 4 models
4
. The writing can be less stable and less refined, with higher error rates than traditional approaches. In the GPQA-Diamond benchmark, the 26 billion-parameter model falls just behind Gemma 4 12B, with its primary advantage being output speed rather than accuracy2
.This trade-off stems from fundamental differences between language and images. While a single mispredicted pixel in image diffusion models doesn't render the output useless, language is discrete—an equivalent error in text can make an entire block of tokens meaningless and force regeneration. Additionally, diffusion models waste resources when the desired output is only a few tokens long, requiring extensive parallel work that autoregressive models complete more efficiently in just a few steps.
Related Stories
Despite quality limitations, DiffusionGemma excels at non-linear tasks where its ability to see and refine entire blocks simultaneously provides advantages. The model performs well on in-line editing, molecular sequencing, mathematical graphing, and structured formats like JSON. Google demonstrated how DiffusionGemma was tuned to solve Sudoku puzzles, a notoriously challenging task for standard autoregressive AI models because each token depends on future tokens. The model's capacity to continuously self-correct large sets of tokens makes such logic-heavy problems more tractable.

Source: SiliconANGLE
Google has released DiffusionGemma under the permissive Apache 2.0 license, with model weights available for download on Hugging Face. Support has already been merged into popular inference engines including vLLM, MLX, and Hugging Face Transformers, with llama.cpp support coming soon
2
. The model can be tested for free using NVIDIA-hosted APIs at build.nvidia.com5
. While positioned as an experimental tool rather than a production-ready solution, DiffusionGemma signals a potential direction for AI text generation where models draft and refine entire passages rather than typing them out sequentially.Summarized by
Navi
[2]
[3]
[4]