Anthropic shipped claude plugin eval in Claude Code v2.1.269 on September 11, and it quietly answers a question every team building internal Claude Code plugins has been dodging for months: does this skill actually do anything, or does it just look busy in the transcript?

I’ve built and reviewed enough internal skills by now to know the failure mode by heart. Someone writes a SKILL.md with a nice description, tests it manually three or four times, watches Claude call the right tool, ships it. Three weeks later a teammate reports the skill “doesn’t seem to fire anymore.” Nobody can say for certain whether it ever reliably fired, because there was never a number attached to “reliably” — just a few anecdotal runs and a vague sense that it worked last time someone checked.

The core idea: score against a baseline, not against a vibe

The command that matters is claude plugin eval init, which asks about your plugin, proposes prompts and graders, and writes a suite under evals/ in your plugin directory. Each case is a prompt a real user might type, plus one or more graders that check whether the result was actually correct — a regex over the output, whether a specific tool was called (tool_used), the order tools fired in (tool_order), whether a file got created (file_exists), or a judge-model rubric (llm or baseline) for fuzzier correctness checks.

The part that actually changes how you think about plugin quality is the no-plugin baseline. By default, every case runs three times with your plugin loaded and three times without it. You get two scores — WITH and W/OUT — and the delta between them, Δ, is what your plugin actually contributed:

CASE        WITH  W/OUT Δ      RUNS COST    NOTES
first-case  1.00  0.33  +0.67  6    $0.41

This is the detail worth sitting with. A plugin can score 1.0 on every case and still be worthless, if Claude would have produced the same correct answer without it. I’ve seen this exact thing happen with a “changelog from diff” skill I reviewed for a client: the skill scored perfectly, but so did the no-plugin baseline, because Claude was already competent enough at that task on its own. The skill wasn’t wrong. It was redundant. Without the baseline, that distinction is invisible — you just see a green checkmark and move on.

Wiring it into CI

The part I actually care about as a lead is the exit-code contract, because that’s what turns this from a nice local dev tool into an actual release gate:

claude plugin eval . \
  --trust-plugin \
  --json results.json \
  --threshold 0.8 \
  --model claude-sonnet-5 \
  --judge-model claude-haiku-4-5 \
  --no-publish \
  --max-cost-usd 20
  • --threshold 0.8 fails the build (exit code 1) if any case’s with-plugin score drops below 0.8. Default is 1.0 — perfect score required — which is unrealistic for most real plugins and worth overriding deliberately rather than leaving at the default and being surprised.
  • --max-cost-usd 20 puts a hard ceiling on the run’s list-price cost estimate. Once it’s spent, no new runs start, in-flight ones finish, and the command exits 2 with partial: true in the results — a different exit code from a real failure, which matters if your CI script branches on it.
  • --trust-plugin skips the interactive trust prompt, which otherwise refuses to run at all in a non-interactive CI shell.
  • Pinning both --model and --judge-model matters more than it looks. Without pinning, a silent model upgrade on Anthropic’s side can shift your scores and get misread as a plugin regression — exactly the kind of false alarm that erodes trust in a CI gate within a month of turning it on.

One practical wrinkle: tool_used: Skill graders are automatically excluded from the baseline comparison’s score, because “the skill was invoked” can never be true in the no-plugin arm — counting it would artificially inflate Δ. They still show up as pass/fail indicators, which is the right call, but it means you can’t naively average every grader’s pass rate into your gate threshold without understanding which graders are being scored in which arm. Read aggregate-result.json, not just the summary table, before you wire an automated gate to it.

Where this actually pays off

The concrete finding I’d expect most teams to hit in week one: a case with Δ near zero and a failing tool_used: Skill grader. That combination means Claude isn’t reliably choosing your skill on natural phrasing — a description problem, not a logic problem. That’s a fixable, specific, testable claim, which is a different category of bug report than “the skill feels flaky.”

The honest caveat: this doesn’t replace integration testing or catch every regression class. It’s non-deterministic by nature — each case defaults to three runs precisely because one run of an agent tells you close to nothing — and llm graders can disagree with themselves between runs on nuanced rubrics, especially with a small judge model. Anthropic’s own guidance is to grade long, structured output with regex against file contents rather than an llm grader against a wall of text, and to reserve judge-model grading for short, concrete PASS/FAIL rubrics. That’s good advice, and it’s advice you’ll only find useful if you’ve already been burned by a flaky judge grader once.

If you maintain more than two or three internal Claude Code plugins, this is worth adopting this week, not next quarter. The cost is real — every eval run and every judge call counts against your plan or API bill — but a --max-cost-usd ceiling keeps that bounded, and the alternative is what most teams have today: plugins that either work by reputation or get quietly abandoned when someone finally asks for proof.

Export for reading

Comments