eval-performance
Investigates expensive globs, deep import chains, file I/O in property functions, and repeated project evaluations. Covers evaluation phases, DefaultItemExcludes, and import analysis through preprocessed output.
Your agent can do more. Find the skill that makes it happen.
Find your next superpower If you are an agent, refer to our llms.txt for full access.Investigates expensive globs, deep import chains, file I/O in property functions, and repeated project evaluations. Covers evaluation phases, DefaultItemExcludes, and import analysis through preprocessed output.
Design automated evaluation pipelines for LLM and agent systems — combining deterministic checks, statistical metrics, and LLM-as-judge scoring into repeatable, CI-integrated eval suites. Load when the user asks to set up automated evals, design an eval pipeline, integrate evals into CI/CD, create an eval suite, do eval-driven development, or says "automate my evals", "CI eval integration", "evalu…
Show a team skill's eval history — committed receipts per version and this machine's local runs — as a Markdown board from the terum-skills CLI, without fetching. Use when the user asks how a shared skill scored, which version passed, or what its receipts say.
Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "ev…
Performs heuristic evaluations, cognitive walkthroughs, anti-pattern detection, and task success analysis. Produces quantitative assessments and prioritized findings, then routes issues to Intent skills for resolution.
Helps weigh competing priorities and assess long-term value in complex product decisions. Covers build-versus-buy reasoning, technical debt evaluation, and communicating tradeoffs to senior leadership.
Interpret plugin quality dimensions and diagnose low scores. Supports improving skill triggering and orchestration fitness, calibrating marketplace scoring thresholds, and explaining quality badges to partners.
Extracts people from event pages, filters companies against the user's ideal customer profile, and researches matching speakers. Produces a person-focused HTML report explaining each prospect's relevance, with public links and a direct-message opener.
Build event sourcing infrastructure and implement event persistence patterns. Supports choosing event store technologies and designing their implementation.
Creates, reads, and manages events and reminders with calendar selection, alarms, and recurrence rules. Covers access requests, calendar-change observation, event storage, and system calendar editing interfaces.
Covers hosting, sponsoring, exhibiting, speaking, and attending events. Supports webinars, conferences, trade shows, dinners, workshops, and virtual summits, including follow-up and ROI planning.
Provides an agent harness performance system for Claude Code and other AI coding agents. It includes skills, instincts, memory, hooks, commands, and security scanning.
Skills give your agent reusable instructions for a specific job. Pick one, read what it does, and bring it into your workflow.
A name only tells half the story. Search the full description and instructions to find the right fit.
Read the skill, visit its source, and see exactly what you’re adding to your agent.
Copy the install command from a skill page and run it in your project.
npx skillycli add owner/repo --skill name