code-safety-monitor
DSPy-powered AI safety monitor that detects backdoors and malicious behavior in code with ~90% detection rate.
🛡️ AgentReady threat assessment
MAESTRO 7-layer threat model + OWASP AIVSS risk score for code-safety-monitor, derived from its capabilities.
These scores are auto-generated from public information (the agent's own listing, docs, and repository) using the canonical OWASP AIVSS formula and the MAESTRO framework — an estimate for guidance, not a penetration test, audit, or certification. See the scoring methodology — every score is re-derived by the same automated method as an agent's public evidence changes.
Overview
A Claude Code plugin that hooks into the agent workflow to scan generated or reviewed code for backdoors and malicious patterns using a DSPy classifier. It adds detection commands and audit checkpoints, flagging suspicious behavior before code is committed. Runs as part of an agentic pipeline, giving it real inspection surface over what the agent writes.
Key features and capabilities
- DSPy-based backdoor/malware classifier
- ~90% detection rate on injected malicious code
- Audit checkpoints in the dev loop
Use cases
- Screen AI-generated code for backdoors
- Guard agentic pipelines against malicious injection