GovernanceCore
Risk & assurance

Red teaming

Working definition

Reviewed 30 July 2026

AI red teaming is a structured, adversarial evaluation in which testers probe a model or system to expose undesirable behavior, misuse pathways, security weaknesses, unsafe interactions, or failures that ordinary testing may miss.

Context

Why it matters

Generative and general-purpose systems can fail across open-ended contexts that fixed benchmark suites do not cover. Red teaming broadens discovery but cannot prove that a system is safe.

Operating note

What this looks like in practice

  1. 01Set threat models, target harms, rules of engagement, and safe handling before testing.
  2. 02Use diverse domain experts and affected perspectives alongside automated attacks.
  3. 03Track findings through remediation, retest, risk acceptance, and monitoring.

Sources & further research

Primary authority anchors the definition. Research links add conceptual or operational depth. External sources may update independently; always verify legal duties against the current official text.

  1. Official source

    Artificial Intelligence Risk Management Framework: Generative AI Profile

    Includes adversarial testing and red teaming among suggested actions for generative-AI risk management.

    NIST
    2024
    Open source ↗
  2. Research paper

    Red Teaming Language Models with Language Models

    Demonstrates automated adversarial test generation for discovering harmful language-model behavior.

    Perez et al.
    2022
    Open source ↗