14 procédures ai security issues de Anthropic-Cybersecurity-Skills, chargées par les agents CyberForge. Chaque skill est un fichier lisible, versionné, retirable.
| Skill | Ce qu'il fait |
|---|---|
| assessing vector and embedding weaknesses | Test vector stores for embedding inversion, cross-tenant leakage, and poisoning. |
| auditing mcp servers for tool poisoning | Scan Model Context Protocol servers and tool metadata for poisoning, SSRF, and unauthenticated exposure. |
| continuous llm red teaming with promptfoo | Wire Promptfoo and DeepTeam into CI/CD for automated regression red-teaming of LLM apps against OWASP LLM Top 10 and OWASP Agentic presets, failing the build when jailbreak or injection vulnerabilities regress. |
| defending llms with guardrails | Deploy Llama Guard, NeMo Guardrails, and LLM Guard input/output scanners as runtime defenses. |
| detecting ai model prompt injection attacks | Detects prompt injection attacks targeting LLM-based applications using a multi-layered defense combining regex pattern matching for known attack signatures, heuristic scoring for structural anomalies, and transformer-based classification with DeBERTa models. |
| detecting data and model poisoning | Identify poisoned training data and backdoored models across the ML pipeline. |
| detecting indirect prompt injection | Detect and defend against prompt injection hidden in documents, web pages, and images consumed by an agent. |
| detecting model extraction attacks | Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability. |
| implementing llm guardrails for security | Implements input and output validation guardrails for LLM-powered applications to prevent prompt injection, data leakage, toxic content generation, and hallucinated outputs. Builds a security validation pipeline using NVIDIA NeMo Guardrails Colang definitions, |
| orchestrating llm attacks with pyrit | Build multi-turn, Crescendo, and Tree-of-Attacks-with-Pruning (TAP) automated attack chains against conversational LLM agents using Microsoft PyRIT, with adversarial chat and scorer feedback loops. |
| red teaming llms with garak | Run NVIDIA garak probe suites against an LLM endpoint to test for jailbreaks, prompt injection, data leakage, and toxic generation, then interpret the hit-rate report for triage and reporting. |
| securing agentic ai tool invocation | Apply least-privilege tool allowlisting, identity binding, and human-in-the-loop controls for agent tool calls. |
| testing for system prompt leakage | Extract and defend system prompts plus embedded secrets and routing logic. |
| testing prompt injection in rag pipelines | Probe RAG applications for prompt injection via poisoned retrieved context and embedding manipulation. |
Source : Anthropic-Cybersecurity-Skills (Apache-2.0). Hub Forge AI indexe et exécute ; il n'est pas l'auteur de ces skills.
Ces skills, exécutés par vos agents
CyberForge déployé sur votre infrastructure
Cette bibliothèque est ce que les agents CyberForge chargent quand ils travaillent chez vous : à jour chaque semaine, en lecture seule par défaut, avec validation humaine sur toute action sensible.