transformer-lens (AI-Research-SKILLs)
Mechanistic-interpretability skill for using TransformerLens to probe LLM internals.
🛡️ AgentReady threat assessment
MAESTRO 7-layer threat model + OWASP AIVSS risk score for transformer-lens (AI-Research-SKILLs), derived from its capabilities.
These scores are auto-generated from public information (the agent's own listing, docs, and repository) using the canonical OWASP AIVSS formula and the MAESTRO framework — an estimate for guidance, not a penetration test, audit, or certification. See the scoring methodology — every score is re-derived by the same automated method as an agent's public evidence changes.
Overview
Part of Orchestra-Research's AI research skills library — a skill for mechanistic interpretability work with the TransformerLens library (activation caching, hooks, circuit analysis). Surface: injects library-specific guidance and writes/runs Python interpretability code.
Key features and capabilities
- TransformerLens hooks and activation caching
- Circuit-analysis workflows
- One of a large AI-research skill library
Use cases
- Probe an LLM's internal activations
- Run a mechanistic-interpretability experiment