Google's DiffusionGemma: A Radically Different Approach to Text Generation

Most local LLMs follow a predictable pattern. Download a model, point your application at it, ask a question, and watch text stream across your screen one token at a time. Some models perform better than others, but the core experience stays fundamentally the same. DiffusionGemma breaks that mold—especially when you enable visual mode. Google's experimental Gemma variant doesn't type answers left to right like traditional models. Instead, it processes entire blocks of text at once, progressively replacing and refining tokens until a complete response emerges. It feels like watching an image denoising tool clean up a photograph, which is essentially what diffusion is. The experience couldn't be more different from the token-by-token generation you're accustomed to.
One developer tested it on an M4 Pro MacBook using 4-bit GGUF quantization through a custom Unsloth fork of llama.cpp. Performance didn't beat Google's standard Gemma 4 26B-A4B on the same hardware, and the Mac did get noticeably bogged down compared to running conventional LLMs. But that's beside the point. What matters is the genuinely strange—and genuinely interesting—experience of watching something so visually different from traditional autoregressive language models.
What Exactly Is DiffusionGemma?
DiffusionGemma is Google's open-weights experimental model for text generation, built on a fundamentally different idea. Instead of composing text word-by-word like virtually every language model you've used, it drafts and refines entire text blocks in parallel. Google claims this approach can accelerate text generation up to 4x faster on GPUs. The model is an Apache 2.0 licensed variant based on the Gemma 4 family—a 26-billion parameter Mixture-of-Experts architecture with roughly 4 billion parameters active during inference. It accepts text, images, and video as input while producing text output.
How DiffusionGemma Reshapes Text Generation

DiffusionGemma feels strange because the output doesn't resemble normal text generation. With visual mode enabled, you watch a 256-token canvas continuously rewritten as the model works. Text appears almost as placeholder content before shifting and morphing, gradually becoming more coherent. It's nothing like the typical word-after-word progression you see everywhere else. That alone makes it feel like an entirely different category of local model.
You don't need to watch the generation process for the model to be useful—most local LLM interfaces are actually better at hiding these details. But here, the visual feedback perfectly illustrates what makes DiffusionGemma different. You can read all the technical papers you want about text diffusion, but watching text constantly reshape itself clarifies the concept far better than words alone.
A typical autoregressive model must commit to its next token, then the one after that, then the one after that. It can plan loosely—good models definitely do this—but tokens generated now can't directly depend on tokens it will generate 50 steps later because those tokens don't exist yet. DiffusionGemma flips the script. It works across a block with bidirectional attention inside that canvas. It uses later portions of the block to refine earlier portions, which is why the output appears refined rather than typed.
That's the conceptual advantage of diffusion-based language models, speed considerations aside. A 256-token canvas gives the model a temporary drafting space where the beginning and end of a text block can influence each other before that block gets finalized. This is why diffusion architectures become genuinely interesting for tasks like direct editing, code completion, structured text processing, and cases where left-to-right sequential generation isn't always optimal.
That's also why DiffusionGemma feels so distinctly different from the local models people typically use. We're accustomed to seeing Qwen, Gemma, Llama, and others generate text in a way that feels like they're actually writing. DiffusionGemma in visual mode creates the impression that it's editing a draft right in front of you—just with the added oddity of seeing every strange intermediate state along the way.
Google's Speed Claims Deserve Context

DiffusionGemma's main selling point is speed. In the launch announcement, Google stated the model can generate text up to 4x faster on dedicated GPUs—over 1,000 tokens per second on an Nvidia H100 and over 700 tokens per second on an RTX 5090. They also noted that quantized versions fit within 18GB VRAM on high-end consumer GPUs.
M4 Pro results told a different story. While exact token-per-second figures weren't captured, a footer screenshot showed 137.9 seconds total, 123 denoising steps, and 9 blocks—working out to roughly 1.121 seconds per step. Since each block is a 256-token canvas, that's about 2,304 canvas positions across 123 steps, or roughly 18.7 token positions per denoising step.
Hardware matters enormously here. The Mac experienced system-wide slowdown during execution, and the actual experience wasn't noticeably faster than running Google's standard Gemma 4 26B-A4B locally. Google explicitly cautioned that Apple Silicon Macs might not achieve similar speedups because unified memory systems typically hit memory bandwidth limits during inference, while DiffusionGemma's gains depend on offloading heavier computational workloads to dedicated accelerators.
That doesn't invalidate Google's speed claims—it just means the real value isn't raw throughput. The actual value lies in seeing a model employ a distinctly different generation process and observing how that changes the experience of interacting with a linear local model.
Running It Locally Is Still Early-Stage and Somewhat Cumbersome
The setup method is Unsloth's GGUF build, which depends on a DiffusionGemma branch from an open llama.cpp pull request. Unsloth's documentation requires building a dedicated llama-diffusion-cli runner because standard llama-cli and llama-server paths can't yet generate from this model.
That distinction matters if you're used to dropping models into Ollama or standard llama.cpp and treating them like any other GGUF. This isn't that kind of model. It needs the right branch, the right runner, and the --diffusion-visual flag if you want the visual component. The command to run it with visual output, after compilation, is:
./llama-diffusion-cli -m ./diffusiongemma-26B-A4B-it-Q4_K_M.gguf -ngl 99 -cnv -n 4096 --diffusion-visualQuantized files are at least practical for consumer hardware. Unsloth lists a 16GB Q4KM variant as the smallest option, with larger versions at 18GB, 21GB, 25GB, and 47GB. That puts it in the same ballpark as other large local models you can run on consumer GPUs with reasonable VRAM.
This remains experimental infrastructure, though. The real questions now are around user support, operational stability, and model quality—not minor rough edges around an otherwise conventional boring model. If you've read about diffusion-based models and want hands-on experience, this is your opportunity.
DiffusionGemma Isn't a Direct Gemma 4 Upgrade
The name suggests another Gemma family member, and it is—but with very different goals. Google describes it as an experimental open-weights model based on the Gemma 4 26B A4B Mixture of Experts architecture, totaling roughly 26 billion parameters with about 4 billion active during inference. The key difference is the diffusion-based, block-wise generation approach rather than the fundamental MoE architecture itself.
Google is explicit: standard autoregressive Gemma 4 models remain their recommendation for maximum output quality. DiffusionGemma prioritizes speed and parallel block generation. Published benchmarks typically show it trailing the standard Gemma 4 26B A4B across reasoning, programming, vision, and long-context tests.
At least one practical test worked fine. A user asked it to build a Flappy Bird-style game in Python that runs in a browser via Flask, and the generated project actually worked. The gravity felt overpowered—gameplay wasn't comfortable—but it produced the Flask application, HTML, CSS, and JavaScript needed for a functional in-browser game. You can see the full output in a public Gist.
DiffusionGemma is still experimental, still early-stage, and not comparable to a standard local LLM. Watching the denoising process unfold is genuinely odd, slightly distracting, but genuinely useful for understanding what Google's attempting—and it makes diffusion-based models far easier to grasp than any written explanation could manage.
Description: Explore Google's experimental DiffusionGemma model that generates text through diffusion instead of token-by-token prediction.
Related Articles
- Creating AI-Powered Flashcards with NoteGPT: A Complete Guide
- The Best AI Tools for Teachers in 2026: A Practical Guide
- 8 Essential Safety Tips for Using ChatGPT, Gemini, and Other AI Tools Securely
- 5 Safety Guardrails Built Into Claude Code to Stop Costly Terminal Mistakes
- Best AI Tools for Microsoft Excel in 2026: Data Analysis, Formulas & Charts
No Comment to " Google's DiffusionGemma: A Radically Different Approach to Text Generation "