Jailbreaking

https://taxonomy.eticas.ai/risk/jailbreaking

Maturity: established

Techniques used to bypass safety controls, content filters, or usage restrictions to trigger prohibited behaviours or outputs.

Also known as: Jailbreakability

System type: Large language models (LLM)
Lifecycle stages: Post Processing

Mappings to external frameworks

Standards & frameworks

Framework Reference
AIUC-1 — AI Underwriting Company Standard Prevent jailbreaks (safety controls)

Taxonomies & vocabularies

Framework Reference
MIT AI Risk Repository AI system security vulnerabilities and attacks

Source: HRA project taxonomy