Skip to content

DiffusionGemma 26B A4BDisabled since Aug 19, 2026

Chat completions return empty content for this discrete-diffusion model under normal chat token budgets. It is unlisted rather than advertised as a live chat model.

DiffusionGemma 26B A4B IT is an open-weights multimodal model from Google DeepMind that generates text via discrete diffusion. Built on the Gemma 4 26B A4B MoE architecture (25.2B total / 3.8B active), it emits tokens in parallel 256-token blocks for high-throughput generation, with a 256K context window, configurable thinking mode, native function calling, and 35+ language support.

Providers
Capabilities
NVIDIA
nvidia
$0
$0
NVIDIA
nvidia-byok
$0
$0
Unavailable
Usage analytics

Loading usage…

API & code
Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

Share cards
DiffusionGemma 26B A4B share card
DiffusionGemma 26B A4B
NVIDIA upstream share card
NVIDIA upstream
NVIDIA upstream share card
NVIDIA upstream
Credits
Use your own key

Run DiffusionGemma 26B A4B on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Text generation
Context length262,144 tokens
Max output209,715 tokens
ArchitectureTransformer
Categorytext
ReleasedJun 10, 2026
Modalities
→
Capabilities