Harmful content and toxicity

https://taxonomy.eticas.ai/risk/harmful-content-toxicity

Maturity: emerging

Offensive, hateful, or otherwise harmful outputs that damage user well-being and erode trust.

This subcategory is emerging. It has not yet been validated through established assessment methods.

Also known as: Toxicity · Harmful content prevention

System type: Large language models (LLM)
Lifecycle stages: Post Processing

Mappings to external frameworks

Standards & frameworks

Framework Reference
AIUC-1 — AI Underwriting Company Standard Prevent harmful outputs
NIST AI 600-1 — Generative AI Risk Profile Dangerous, Violent, or Hateful Content

Taxonomies & vocabularies

Framework Reference
MIT AI Risk Repository Exposure to toxic content
AIR 2024 Content Safety Risks
IBM AI Risk Atlas Output → Toxicity / Harmful output

Source: HRA project taxonomy