Since the system already runs the tests and calculates pass rates for me, is looking at the benchmark.json numbers enough, without needing to look at the actual outputs myself?
No. This is exactly what the system's design specifically warns about — the analyst pass exists precisely because pass-rate numbers alone can be an illusion: if a given assertion passes regardless of whether the Skill was used, that "100% pass rate" looks great but proves nothing about the Skill actually making a difference, since it has zero discriminating power with respect to whether the Skill was present.
The habit actually worth building is checking the analyst pass before looking at benchmark numbers — did it flag any "non-discriminating" assertions or "high-variance" tests? If it did, the pass-rate numbers tied to those test cases should be discounted heavily, or the test case itself should be redesigned entirely. The numbers always need reading alongside that analysis layer — looking at a bare pass-rate percentage on its own is easy to be misled by a poorly designed assertion.
How does the description optimization's 60/40 train/test split differ from simply running it multiple times and averaging?
The difference is in what problem each one guards against. Simply running multiple times and averaging solves "single-run instability"; description optimization faces a different risk — overfitting: if you use the same batch of queries both to improve the description and to score how good that description is, it's easy to end up with a description that happens to pass this particular batch of queries but turns out inaccurate against a fresh batch, because the optimization process will unconsciously tune toward the specific wording of that particular batch.
The 60% train / 40% held-out test split addresses exactly that: the description-improvement process only sees the training set, and which version ultimately gets selected as best_description is based on its score against the "unseen" test set, not the training score. The documentation states this directly: it's done "to avoid overfitting" — if you picked the final version based solely on training score, you could easily end up with a description that merely "memorized this particular batch of questions" and falls apart the moment real users phrase things differently.
The documentation says "simple queries like 'read file X' won't reliably trigger skills even with perfect description matches" — what does that mean?
It means the triggering mechanism itself has a hidden threshold: writing a sufficiently precise description doesn't guarantee a trigger — it also depends on whether the request itself is complex enough to be "worth consulting a Skill for." If a request is simple enough to handle in a couple sentences, Claude's decision mechanism may simply choose not to trigger any Skill at all, even if that Skill's description matches the request's keywords perfectly.
This has a direct impact on designing eval queries: if you test whether a Skill should trigger using an overly simple query like "read file X," the result you get may be unreliable from the start — not because the description is poorly written, but because the query itself doesn't carry enough weight to be "worth triggering a Skill for." The documentation states the causal chain directly: "bad eval queries → bad descriptions" — if the test cases you fed in from the start picked the wrong difficulty level, no number of optimization rounds afterward will produce a genuinely reliable description.
I'm using skill-creator in Claude.ai, and some evaluation features just don't run at all — why?
Because the Claude.ai environment itself lacks capabilities certain features depend on, and the documentation lists the corresponding limitations explicitly: Claude.ai has no subagents, so test cases can only run serially (less rigorous than parallel execution), and there's no way to run a baseline — quantitative benchmarking and blind comparison both only make sense with an independent baseline, so both get skipped entirely; description optimization requires the claude -p command-line interface, which only exists in Claude Code, so Claude.ai skips that too.
In other words, what you get in Claude.ai is effectively a stripped-down version — running test cases and producing qualitative feedback presented directly in the conversation, rather than through a browser-based viewer. If you need full quantitative benchmarking, blind A/B comparison, or automatic description optimization, those features explicitly require the Claude Code environment and subagent capability — it's not a misconfiguration in skill-creator, it's the actual capability boundary of the environment itself.
Before skill-creator added its eval functionality, judging whether a Skill actually made Claude perform better came down to trial and gut feeling. The official skill-creator SKILL.md now bakes in a full evaluation workflow — running test cases, producing quantitative metrics, running blind A/B comparisons, and even auto-optimizing a Skill's description field. This piece pulls that system apart to see what it actually tells you, and where you still need to supply your own judgment.
The core logic isn't complicated: for the same task prompt, two subagents spin up at once — one running with your Skill, one without it entirely (or, if you're improving an existing Skill, running the old version instead) — with each side's output saved to its own separate folder. The official documentation then asks you to draft assertions while the runs are still in progress, defining objectively verifiable checks for each test case up front, rather than figuring out scoring after the fact once results are already in.
There's a line the documentation is explicit about here: assertions must be "objectively verifiable." Subjective output — writing style, design aesthetics — is better evaluated qualitatively; forcing a quantitative assertion onto something fundamentally subjective produces a score that ends up more misleading than useful.
Once tests finish, scripts.aggregate_benchmark compiles results into benchmark.json and benchmark.md, listing pass rate, time, and token usage per configuration with mean, standard deviation, and delta. What's genuinely interesting is the analyst pass that follows — the documentation specifically calls for an analysis to surface assertions that "always pass regardless of whether the Skill was used" (meaning the assertion has no discriminating power — it can't tell you whether the Skill made a difference at all), and evals with "unusually high variance" (meaning the test itself may be inherently flaky and unreliable). This step is arguably the most honest part of the whole system: it proactively flags cases where "this metric isn't actually telling you anything," rather than letting a wall of clean-looking pass rates mislead you.
For a more rigorous comparison of which version is genuinely better, the system offers a blind comparison — an independent agent judges two outputs on quality alone, without being told which came from which version, and then analyzes why the winner won. The documentation's own positioning of this feature is notably conservative: "this is optional, requires subagents, and most users won't need it — the human review loop is usually sufficient." That honest self-assessment is worth noting — the system's own designers concluded that most scenarios don't need this heavy a verification process; simply having a human look at outputs and give feedback is already enough.