https://taxonomy.eticas.ai/risk/evasion-attacks
Maturity: established
Inputs crafted so that a deployed system acts on a request its safeguards should have blocked. For LLM deployments, the robustness of the guardrail or filter layer to input transformations that keep the request readable to the model while defeating the guardrail; for classifier-based systems, adversarial perturbations that force misclassification at inference time. Poisoning of training or fine-tuning data is data-poisoning.
Also known as: Evasion attacks · Guardrail robustness · Adversarial evasion
System type: ADM and LLM systems
Lifecycle stages: In Processing, Post Processing
| Framework | Reference |
|---|---|
| AIUC-1 — AI Underwriting Company Standard | Detect adversarial input |
| Framework | Reference |
|---|---|
| W3C Data Privacy Vocabulary — AI Extension | Adversarial Attack |
| MIT AI Risk Repository | AI system security vulnerabilities and attacks |
| IBM AI Risk Atlas | Inference → Robustness → Adversarial robustness |