Llama Guard 4 is a 12B natively multimodal safety classifier for content safety classification on LLM inputs and responses. It acts as an LLM itself: it generates text indicating whether a prompt or response is safe or unsafe and, if unsafe, lists the violated categories (MLCommons hazards taxonomy S1–S14). Use it as a guardrail/judge hop in front of or behind chat models.
Share cards3 images


