Files
Anthropic-Cybersecurity-Skills/skills/defending-llms-with-guardrails/references/standards.md
T
mukul975 8cae0648ec Add 55 new skills across 3 new domains + 6 undercovered areas (762 -> 817)
Demand-driven expansion targeting the fastest-growing 2025-2026 threat and
skills categories (ISC2/WEF/CrowdStrike/Mandiant signals):

- AI Security (NEW domain, 12 skills): LLM red-teaming with garak/PyRIT,
  prompt injection (direct/indirect/RAG), MCP tool-poisoning, agentic tool
  invocation, guardrails, model/data poisoning, system-prompt leakage,
  embedding/vector weaknesses, model extraction, continuous red-teaming
- Supply Chain Security (NEW domain, 5 skills): SBOMs, dependency confusion,
  malicious-npm triage, typosquatting, SLSA/Sigstore provenance
- Hardware & Firmware Security (NEW domain, 4 skills): CHIPSEC/UEFI audit,
  Secure Boot bypass, TPM measured-boot attestation, ESP bootkit hunting
- Identity (10): Entra ID/ROADtools, GraphRunner, AADInternals, ADCS/Certipy,
  shadow credentials, coercion, BloodHound CE, device-code phishing, SSO abuse
- Cloud-native (8): Stratus, Pacu, CloudFox, container escape, K8s RBAC,
  Falco, Trivy, kube-bench
- Offensive C2 (6): Sliver, Havoc, NetExec, DPAPI, NTLM relay ESC8, redirectors
- DFIR (6): Hayabusa, Chainsaw, KAPE, Velociraptor, EZ Tools, Plaso
- Backfill (4): OpenCTI, MISP, honeytokens, post-quantum crypto migration

Each skill follows the repo taxonomy (SKILL.md + references/{standards,api-reference}.md
+ scripts/agent.py + LICENSE), with researched real tool commands (no placeholders),
complete frontmatter, and ATT&CK/ATLAS + NIST CSF mappings. Updates README domain
table, skill count, and index.json.
2026-06-22 19:08:16 +02:00

1.8 KiB

Standards and Framework Mapping

NIST AI Risk Management Framework (AI RMF 1.0 / GenAI Profile NIST AI 600-1)

ID Name Rationale
MANAGE-2.1 Resources required to manage AI risks are documented and put into action Deploying Llama Guard / NeMo / LLM Guard is the operational control that manages identified LLM safety risks at runtime.

MITRE ATLAS

ID Name Rationale
AML.T0054 LLM Jailbreak The guardrail layer is the primary mitigation that detects and blocks jailbreak attempts before/after model inference.
AML.T0051 LLM Prompt Injection Input rails and the PromptInjection scanner block direct injection attempts.
AML.T0051.001 LLM Prompt Injection: Indirect Retrieval/input scanning blocks injection embedded in retrieved or tool-returned content.
AML.T0057 LLM Data Leakage Output scanners (Sensitive, Secrets, Deanonymize) prevent leakage of PII, secrets, and instructions.

OWASP Top 10 for LLM Applications (2025)

ID Name Rationale
LLM01 Prompt Injection Guardrails are the recommended runtime mitigation for direct and indirect injection.
LLM02 Sensitive Information Disclosure Output PII/secrets scanners prevent disclosure.
LLM07 System Prompt Leakage Input/output rails detect attempts to extract and leak the system prompt.

MLCommons Hazard Taxonomy (Llama Guard 3 categories)

S1 Violent Crimes · S2 Non-Violent Crimes · S3 Sex-Related Crimes · S4 Child Sexual Exploitation · S5 Defamation · S6 Specialized Advice · S7 Privacy · S8 Intellectual Property · S9 Indiscriminate Weapons · S10 Hate · S11 Suicide & Self-Harm · S12 Sexual Content · S13 Elections · S14 Code Interpreter Abuse.