Skip to main content

Explore the Complete Risk Framework

View Detailed AI Risk Categories →

Understanding the Landscape of AI Risks

Enkrypt AI organises AI risk into 8 top-level categories containing 35 sub-categories. This taxonomy is the shared vocabulary across the platform: it is what you scope a Red Team run with, and it is how results are reported back to you.
These IDs are the ones you use in the risk_categories field of a Red Team payload. For configuration detail see the Risk Category Catalog.

The eight categories

Content that materially harms people or enables physical harm, irrespective of legality.
Disclosure, inference, or aggregation of personal data — memorisation, cross-session leakage, redaction bypass, re-identification, and unsafe data egress.
Material uplift to financial fraud, identity theft, credential theft, social engineering, malware or exploit development, or attacks against third-party systems.
Dignitary, allocative, or representational harm — stereotyping, identity-based attacks, and biased decisions across protected attributes.
Damage to the deploying brand’s standing — off-voice output, fabricated brand claims, competitor disparagement, and content the brand would never endorse.
Users pushing your system outside its intended operating scope — out-of-domain hijacking, role override, and resource abuse.
Quality failures under non-malicious load — incorrect, inconsistent, or ill-formed output.

Compliance frameworks

The same taxonomy underpins compliance reporting. Each framework’s controls map onto the categories and sub-categories above, so a single run can be re-expressed against any of them.

Risk Assessment Framework

Detection Capabilities

  • Real-time Analysis — instant detection across risk categories
  • Context Understanding — nuanced interpretation grounded in your system description
  • Multi-language Support — detection and adversarial testing across languages
  • Multi-modal Coverage — text, image and audio inputs

Mitigation Strategies

  • Content Filtering — automatic blocking of harmful content via Guardrails
  • Risk Scoring — attack success rate per category, sub-category and attack method
  • Generated Remediation — hardened system prompts and Guardrails policies derived from run results
  • Compliance Reporting — results mapped onto the frameworks above for audit

Risk Category Catalog

Full detail on what each category and sub-category tests for.

Red Team Quickstart

Scope a run against these categories and read the report.

Attack Methods

The 21 techniques used to probe each category.

Guardrails

Detect and block these risks at runtime.