# Moderations API

> POST /api/v1/moderations — classify text for harmful content. OpenAI-compatible, served by guard models.


# Moderations API

Check whether text is harmful before you send it to a model or show it to a user. The request and response follow OpenAI's `POST /v1/moderations`, so the OpenAI SDKs work by changing the base URL and key. Guard models such as NVIDIA Nemotron Content Safety and Llama Guard do the classification.

::endpoint{method="POST" path="/api/v1/moderations"}

:::note
Authenticate with an **LLM API key** (`sk-ar-v1-…`) from [/dashboard/keys](https://anyrouter.dev/dashboard/keys), sent as `Authorization: Bearer …`.
:::

## Request

| Field | Type | Required | Description |
|---|---|---|---|
| `input` | string \| string[] \| `{type:"text", text}`[] | yes | Up to 32 texts, each up to 32,000 characters. Images are not supported. |
| `model` | string | no | A moderation model id. Defaults to `nvidia/nemotron-3.5-content-safety`. OpenAI ids such as `omni-moderation-latest` use the default too. |
| `session_id` | string | no | Groups related requests in [Request Logs](/features/request-logs#sessions). |

```bash
curl https://anyrouter.dev/api/v1/moderations \
  -H "Authorization: Bearer $ANYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input": ["What is a good pasta recipe?", "How do I buy a gun with no background check?"]}'
```

## Response

```json
{
  "id": "modr-req_01HXYZ",
  "model": "nvidia/nemotron-3.5-content-safety",
  "results": [
    {
      "flagged": false,
      "categories": { "harassment": false, "illicit/violent": false, "violence": false },
      "category_scores": { "harassment": 0, "illicit/violent": 0, "violence": 0 },
      "guard_categories": []
    },
    {
      "flagged": true,
      "categories": { "harassment": false, "illicit/violent": true, "violence": false },
      "category_scores": { "harassment": 0, "illicit/violent": 1, "violence": 0 },
      "guard_categories": ["guns_and_illegal_weapons"]
    }
  ],
  "usage": { "prompt_tokens": 955, "completion_tokens": 24, "total_tokens": 979, "cost": 0 }
}
```

The example is shortened. Every result lists all 13 OpenAI categories: `harassment`, `harassment/threatening`, `hate`, `hate/threatening`, `illicit`, `illicit/violent`, `self-harm`, `self-harm/intent`, `self-harm/instructions`, `sexual`, `sexual/minors`, `violence`, `violence/graphic`.

| Field | Description |
|---|---|
| `results[]` | One result per input, in input order. |
| `flagged` | `true` when the guard model judged the text unsafe. |
| `categories` | OpenAI categories the guard's verdict maps to. |
| `category_scores` | `1` for a flagged category, else `0`. Guard models return a verdict, not a probability. |
| `guard_categories` | The guard model's own labels. Some (privacy, defamation, elections, …) have no OpenAI category: the text is still `flagged`, but no OpenAI category is set. |
| `usage` | Tokens summed over every input. Priced at the model's per-1M token rates. |

## Models

Moderation models appear under the **Moderation** tab at [anyrouter.dev/models](https://anyrouter.dev/models):

| Model | Notes |
|---|---|
| `nvidia/nemotron-3.5-content-safety` | Default. Returns categories. |
| `meta/llama-guard-4-12b` | MLCommons hazard codes S1–S14. Requires your own DeepInfra or OpenRouter key ([BYOK](/features/byok)). |
| `openai/gpt-oss-safeguard-20b` | Policy-following reasoning classifier. Requires your own key (Groq, OpenRouter, Vercel, Hugging Face, or Nous). |

## Errors

| Status | Code | When |
|---|---|---|
| 400 | `invalid_request` | `input` is missing, empty, contains an image, or exceeds the limits. |
| 400 | `model_not_moderation_compatible` | The model is not a moderation model. |
| 502 | `upstream_invalid_response` | The guard model replied without a verdict. |
| 502 | `upstream_error` | Every provider for the model failed. |
