Files
Anthropic-Cybersecurity-Skills/skills/red-teaming-llms-with-garak/references/standards.md
T
mukul975 8cae0648ec Add 55 new skills across 3 new domains + 6 undercovered areas (762 -> 817)
Demand-driven expansion targeting the fastest-growing 2025-2026 threat and
skills categories (ISC2/WEF/CrowdStrike/Mandiant signals):

- AI Security (NEW domain, 12 skills): LLM red-teaming with garak/PyRIT,
  prompt injection (direct/indirect/RAG), MCP tool-poisoning, agentic tool
  invocation, guardrails, model/data poisoning, system-prompt leakage,
  embedding/vector weaknesses, model extraction, continuous red-teaming
- Supply Chain Security (NEW domain, 5 skills): SBOMs, dependency confusion,
  malicious-npm triage, typosquatting, SLSA/Sigstore provenance
- Hardware & Firmware Security (NEW domain, 4 skills): CHIPSEC/UEFI audit,
  Secure Boot bypass, TPM measured-boot attestation, ESP bootkit hunting
- Identity (10): Entra ID/ROADtools, GraphRunner, AADInternals, ADCS/Certipy,
  shadow credentials, coercion, BloodHound CE, device-code phishing, SSO abuse
- Cloud-native (8): Stratus, Pacu, CloudFox, container escape, K8s RBAC,
  Falco, Trivy, kube-bench
- Offensive C2 (6): Sliver, Havoc, NetExec, DPAPI, NTLM relay ESC8, redirectors
- DFIR (6): Hayabusa, Chainsaw, KAPE, Velociraptor, EZ Tools, Plaso
- Backfill (4): OpenCTI, MISP, honeytokens, post-quantum crypto migration

Each skill follows the repo taxonomy (SKILL.md + references/{standards,api-reference}.md
+ scripts/agent.py + LICENSE), with researched real tool commands (no placeholders),
complete frontmatter, and ATT&CK/ATLAS + NIST CSF mappings. Updates README domain
table, skill count, and index.json.
2026-06-22 19:08:16 +02:00

1.4 KiB

Standards and Framework Mapping — Red-Teaming LLMs with garak

MITRE ATLAS (Adversarial Threat Landscape for AI Systems)

ID Name Rationale
AML.T0051 LLM Prompt Injection garak's promptinject and latentinjection probes craft inputs that override intended model instructions; the scanner measures how often the target obeys the injected directive.
AML.T0054 LLM Jailbreak garak's dan and related probes attempt to push the model past its safety guardrails so it produces restricted output; the detector verdict measures jailbreak success.

Reference: https://atlas.mitre.org/

NIST AI Risk Management Framework (AI RMF 1.0)

ID Function/Subcategory Rationale
MEASURE-2.7 AI system security and resilience are evaluated and documented garak produces repeatable, quantitative measurements (per-probe hit rates) of an LLM's resistance to injection, jailbreak, and leakage, directly evidencing this subcategory.

Reference: https://www.nist.gov/itl/ai-risk-management-framework

OWASP Top 10 for LLM Applications (cross-reference)

OWASP ID Risk garak probe family
LLM01:2025 Prompt Injection promptinject, latentinjection, encoding, dan
LLM02:2025 Sensitive Information Disclosure leakreplay, xss
LLM07:2025 System Prompt Leakage leakreplay

Reference: https://genai.owasp.org/