prompt-guard (AI-Research-SKILLs)
Safety skill for deploying Prompt Guard to detect prompt-injection and jailbreak inputs.
🛡️ AgentReady threat assessment
MAESTRO 7-layer threat model + OWASP AIVSS risk score for prompt-guard (AI-Research-SKILLs), derived from its capabilities.
These scores are auto-generated from public information (the agent's own listing, docs, and repository) using the canonical OWASP AIVSS formula and the MAESTRO framework — an estimate for guidance, not a penetration test, audit, or certification. See the scoring methodology — every score is re-derived by the same automated method as an agent's public evidence changes.
Overview
A safety-alignment skill covering Meta's Prompt Guard classifier to detect prompt-injection and jailbreak attempts on LLM inputs. Surface: injects deployment guidance and writes/runs classifier code — directly relevant to agent security defenses.
Key features and capabilities
- Prompt-injection/jailbreak detection
- Prompt Guard classifier integration
- Part of the safety-alignment skill set
Use cases
- Add injection detection to an LLM app
- Filter malicious agent inputs