Models published by Google, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.
Google's lightest Gemini 2.5 Flash variant, designed for high-throughput, cost-efficient inference with a 1M-token context window and fast response times.
Google's fast and cost-efficient Gemini model with strong reasoning capabilities, designed for high-throughput tasks.
Google's advanced Gemini 2.5 Pro model offering deep reasoning, long context, and multimodal capabilities for complex research, coding, and analytical workflows.
Google's Gemini 3 Flash model combining fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.
Google's lightest and most cost-efficient Gemini model for high-throughput tasks.
Google's most intelligent Gemini model with improved reasoning, a medium thinking level, and a 1M token context window.
Google's Gemini 3.5 Flash Lite is a cost-efficient, high-throughput multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Optimized for agentic workflows, complex coding, and long-horizon tasks across a 1M-token context window.
Google's high-performance multimodal AI model designed for agentic workflows, complex coding, and long-horizon tasks.
Google's Gemini 3.6 Flash is a frontier-level multimodal model available as a partner-hosted SKU on Cloudflare Workers AI. Combines fast inference with strong reasoning, coding, and multimodal understanding across a 1M-token context window.
Google's Gemini 3.7 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.
Google's Gemini 3.8 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.
Gemini Embedding 2 is Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into one 3,072-dimension vector space for cross-modal semantic search, document retrieval, and recommendations over 100+ languages. Upstream accepts 8,192 input tokens and MRL-truncates to 128–3,072 dimensions. AnyRouter exposes the text path only — the OpenAI-compatible /embeddings body is a text `input`, so the extra upstream modalities are deliberately not declared here.
Gemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.
Gemma 4 31B is Google's dense frontier-level open reasoning model, optimized for complex reasoning, agentic workflows, coding, and multimodal understanding.
Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.