constitutional-ai (AI-Research-SKILLs)
Safety-alignment skill for applying Constitutional AI methods in LLM training.
🛡️ AgentReady threat assessment
MAESTRO 7-layer threat model + OWASP AIVSS risk score for constitutional-ai (AI-Research-SKILLs), derived from its capabilities.
These scores are auto-generated from public information (the agent's own listing, docs, and repository) using the canonical OWASP AIVSS formula and the MAESTRO framework — an estimate for guidance, not a penetration test, audit, or certification. See the scoring methodology — every score is re-derived by the same automated method as an agent's public evidence changes.
Overview
A safety/alignment skill from the AI-Research-SKILLs library covering Constitutional AI — self-critique and revision against a constitution during training/fine-tuning. Surface: injects methodology and writes/runs training code.
Key features and capabilities
- Constitutional AI self-critique loop
- Alignment-focused training guidance
- Sibling to LlamaGuard/NeMo-Guardrails skills
Use cases
- Apply Constitutional AI to a model
- Design an RLAIF-style alignment pipeline