venice
Models published by venice, available through the AnyRouter API. Each can route across multiple upstream providers for availability and price.
Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.
Hermes 3 Llama 3.1 405B is Nous Research's flagship 405B parameter model designed for advanced reasoning, coding, and instruction following. Built on Llama 3.1, it excels at complex problem-solving, multi-step reasoning, and agentic workflows with support for function calling and structured output.
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with advanced agentic capabilities through multi-agent collaboration. Features strong reasoning, function calling, and structured output capabilities with 198K context window.
NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context. Designed for advanced function calling, structured output, and complex reasoning tasks.
GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.
Optimized for creative roleplay scenarios with maximum freedom. Designed for immersive storytelling, character interactions, and open-ended creative writing. Features vision capabilities, function calling, and structured output with 128K context window.