Skip to content

venice

6 modelsModel creator

Models published by venice, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.

256K

Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.

131K

Hermes 3 Llama 3.1 405B is Nous Research's flagship 405B parameter model designed for advanced reasoning, coding, and instruction following. Built on Llama 3.1, it excels at complex problem-solving, multi-step reasoning, and agentic workflows with support for function calling and structured output.

198K

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with advanced agentic capabilities through multi-agent collaboration. Features strong reasoning, function calling, and structured output capabilities with 198K context window.

1M

NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context. Designed for advanced function calling, structured output, and complex reasoning tasks.

400K

GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.

128K

Optimized for creative roleplay scenarios with maximum freedom. Designed for immersive storytelling, character interactions, and open-ended creative writing. Features vision capabilities, function calling, and structured output with 128K context window.