Skip to content
All use cases
Evals

Run one eval across many models

Loop the same prompts over a list of catalog ids with one key and one client. Compare answers, latency, and cost side by side.

01

The problem

What this is for

Evaluating models from different vendors means one SDK and one key per provider, and cost lands on separate invoices you have to add up.

02

What to do

Three steps, then copy

  1. 1.Run anyr models (or open /models) and pick the ids you want to compare.
  2. 2.Send the same prompt to each id through one AI SDK provider.
  3. 3.Compare outputs and token usage, then check cost per request in Dashboard → Logs.
Eval loop
ts
import { createAnyRouter } from "@anyr/ai-sdk-provider"
import { generateText } from "ai"

const anyrouter = createAnyRouter()
const models = ["openai/gpt-5.4-mini", "anthropic/claude-sonnet-4.6", "z-ai/glm-4.7-flash"]

for (const id of models) {
  const started = Date.now()
  const { text, usage } = await generateText({
    model: anyrouter(id),
    prompt: "Summarize the CAP theorem in two sentences.",
  })
  console.log(id, Date.now() - started, "ms", usage.totalTokens, "tokens\n", text)
}
anyr CLI
bash
anyr models

Ready to try it

The guide has the full walkthrough. Sign up if you still need a key.