promptfoo-evaluation
Configure and run LLM evaluations with Promptfoo — configs, Python assertions, llm-rubric judging.
🛡️ AgentReady threat assessment
MAESTRO 7-layer threat model + OWASP AIVSS risk score for promptfoo-evaluation, derived from its capabilities.
These scores are auto-generated from public information (the agent's own listing, docs, and repository) using the canonical OWASP AIVSS formula and the MAESTRO framework — an estimate for guidance, not a penetration test, audit, or certification. See the scoring methodology — every score is re-derived by the same automated method as an agent's public evidence changes.
Overview
Community Agent Skill for prompt testing with the Promptfoo framework. Generates promptfooconfig.yaml, writes Python custom assertions, implements llm-rubric LLM-as-judge, and manages few-shot examples. Runs Promptfoo and Python evaluation scripts on the host.
Key features and capabilities
- promptfooconfig.yaml generation
- Python custom assertions
- llm-rubric LLM-as-judge
Use cases
- Setting up a Promptfoo eval suite
- Adding LLM-as-judge assertions to prompts