Bible Network Crypto DeFi Onchain RWA AI Agent Stablecoin Chain SAFU CryptoTax DeFAI AGI Claude Me Claude Skill Claude Design Claude Cowork
Independent Media
Not affiliated with any project
Learn Claude Skills. Do Everything Better.
claudeskill-me.com
LATEST
Does skill-creator's Eval System Actually Tell You Anything? Testing Its A/B Comparison and Description Optimizer  ·  Your rm -rf Protection May Have Never Actually Worked: Wrapping It in bash -c Bypasses Claude Code's Safety Prompt Entirely  ·  Claude Mods Aren't Just Another Plugin — They Can Draw Panes and Intercept Screen Rendering, Which External Extensions Simply Can't  ·  Claude Suddenly Stopped Using Your Skills After /compact? It's Not Broken — It's Designed Not to Restore Them Automatically  ·  Claude Code Now Supports AGENTS.md — But Half the Articles Online Describe the Old Rules, and CLAUDE.local.md Can Silently Break It  ·  Your Hook Says "Blocking Error" But the File Still Changed? PostToolUse and PreToolUse Don't Actually Block the Same Way
reviews

Does skill-creator's Eval System Actually Tell You Anything? Testing Its A/B Comparison and Description Optimizer

30-Second Version · For the impatient
The system itself admits most people don't need the blind comparison — a human glance at the outputs is usually enough. Honestly telling you which verification steps are actually skippable is worth more than piling on additional metrics.

Full Explanation +
01 · Why did this happen?

Since the system already runs the tests and calculates pass rates for me, is looking at the benchmark.json numbers enough, without needing to look at the actual outputs myself?

No. This is exactly what the system's design specifically warns about — the analyst pass exists precisely because pass-rate numbers alone can be an illusion: if a given assertion passes regardless of whether the Skill was used, that "100% pass rate" looks great but proves nothing about the Skill actually making a difference, since it has zero discriminating power with respect to whether the Skill was present.

The habit actually worth building is checking the analyst pass before looking at benchmark numbers — did it flag any "non-discriminating" assertions or "high-variance" tests? If it did, the pass-rate numbers tied to those test cases should be discounted heavily, or the test case itself should be redesigned entirely. The numbers always need reading alongside that analysis layer — looking at a bare pass-rate percentage on its own is easy to be misled by a poorly designed assertion.

02 · What is the mechanism?

How does the description optimization's 60/40 train/test split differ from simply running it multiple times and averaging?

The difference is in what problem each one guards against. Simply running multiple times and averaging solves "single-run instability"; description optimization faces a different risk — overfitting: if you use the same batch of queries both to improve the description and to score how good that description is, it's easy to end up with a description that happens to pass this particular batch of queries but turns out inaccurate against a fresh batch, because the optimization process will unconsciously tune toward the specific wording of that particular batch.

The 60% train / 40% held-out test split addresses exactly that: the description-improvement process only sees the training set, and which version ultimately gets selected as best_description is based on its score against the "unseen" test set, not the training score. The documentation states this directly: it's done "to avoid overfitting" — if you picked the final version based solely on training score, you could easily end up with a description that merely "memorized this particular batch of questions" and falls apart the moment real users phrase things differently.

03 · How does it affect me?

The documentation says "simple queries like 'read file X' won't reliably trigger skills even with perfect description matches" — what does that mean?

It means the triggering mechanism itself has a hidden threshold: writing a sufficiently precise description doesn't guarantee a trigger — it also depends on whether the request itself is complex enough to be "worth consulting a Skill for." If a request is simple enough to handle in a couple sentences, Claude's decision mechanism may simply choose not to trigger any Skill at all, even if that Skill's description matches the request's keywords perfectly.

This has a direct impact on designing eval queries: if you test whether a Skill should trigger using an overly simple query like "read file X," the result you get may be unreliable from the start — not because the description is poorly written, but because the query itself doesn't carry enough weight to be "worth triggering a Skill for." The documentation states the causal chain directly: "bad eval queries → bad descriptions" — if the test cases you fed in from the start picked the wrong difficulty level, no number of optimization rounds afterward will produce a genuinely reliable description.

04 · What should I do?

I'm using skill-creator in Claude.ai, and some evaluation features just don't run at all — why?

Because the Claude.ai environment itself lacks capabilities certain features depend on, and the documentation lists the corresponding limitations explicitly: Claude.ai has no subagents, so test cases can only run serially (less rigorous than parallel execution), and there's no way to run a baseline — quantitative benchmarking and blind comparison both only make sense with an independent baseline, so both get skipped entirely; description optimization requires the claude -p command-line interface, which only exists in Claude Code, so Claude.ai skips that too.

In other words, what you get in Claude.ai is effectively a stripped-down version — running test cases and producing qualitative feedback presented directly in the conversation, rather than through a browser-based viewer. If you need full quantitative benchmarking, blind A/B comparison, or automatic description optimization, those features explicitly require the Claude Code environment and subagent capability — it's not a misconfiguration in skill-creator, it's the actual capability boundary of the environment itself.

Full Content +

Before skill-creator added its eval functionality, judging whether a Skill actually made Claude perform better came down to trial and gut feeling. The official skill-creator SKILL.md now bakes in a full evaluation workflow — running test cases, producing quantitative metrics, running blind A/B comparisons, and even auto-optimizing a Skill's description field. This piece pulls that system apart to see what it actually tells you, and where you still need to supply your own judgment.

The basic flow: run with-skill and without-skill versions simultaneously, then grade

The core logic isn't complicated: for the same task prompt, two subagents spin up at once — one running with your Skill, one without it entirely (or, if you're improving an existing Skill, running the old version instead) — with each side's output saved to its own separate folder. The official documentation then asks you to draft assertions while the runs are still in progress, defining objectively verifiable checks for each test case up front, rather than figuring out scoring after the fact once results are already in.

There's a line the documentation is explicit about here: assertions must be "objectively verifiable." Subjective output — writing style, design aesthetics — is better evaluated qualitatively; forcing a quantitative assertion onto something fundamentally subjective produces a score that ends up more misleading than useful.

Aggregate and analyst pass: surfacing "non-discriminating" assertions and "possibly flaky" tests

Once tests finish, scripts.aggregate_benchmark compiles results into benchmark.json and benchmark.md, listing pass rate, time, and token usage per configuration with mean, standard deviation, and delta. What's genuinely interesting is the analyst pass that follows — the documentation specifically calls for an analysis to surface assertions that "always pass regardless of whether the Skill was used" (meaning the assertion has no discriminating power — it can't tell you whether the Skill made a difference at all), and evals with "unusually high variance" (meaning the test itself may be inherently flaky and unreliable). This step is arguably the most honest part of the whole system: it proactively flags cases where "this metric isn't actually telling you anything," rather than letting a wall of clean-looking pass rates mislead you.

Blind A/B comparison: deliberately designed as optional, and explicitly flagged as unnecessary for most people

For a more rigorous comparison of which version is genuinely better, the system offers a blind comparison — an independent agent judges two outputs on quality alone, without being told which came from which version, and then analyzes why the winner won. The documentation's own positioning of this feature is notably conservative: "this is optional, requires subagents, and most users won't need it — the human review loop is usually sufficient." That honest self-assessment is worth noting — the system's own designers concluded that most scenarios don't need this heavy a verification process; simply having a human look at outputs and give feedback is already enough.

Sources: skill-creator SKILL.md - anthropics/claude-plugins-official, Claude Agent Skills 2.0: The Beginner's Guide to the Updated Skill-Creator
Diagram
skill-creator 的評測迴圈與 description 優化流程上排是平行測試、評分、彙整、變異分析的評測流程;下排是 60/40 切分的 description 優化,以 held-out 測試分數選出最佳版本skill-creator: Eval Loop and Description OptimizationParallel runswith-skill vs baseline(or old vs new version)Grader subagentgrading.jsontext · passed · evidenceAggregatebenchmark.json / .mdmean ± stddev, deltaAnalyst passflags always-pass checksand high-variance evalsHuman review → improve skill → rerun (optional blind A/B comparison)Description optimization (needs claude -p, Claude Code only)20 trigger queries: 8–10 should-trigger, 8–10 tricky near-misses; each run 3×, up to 5 iterationsTrain 60% — tune the descriptionHeld-out test 40% — picks the winnerbest_description is chosen by TEST score, not train score — this is what prevents overfittingClaude Skill Me · claudeskill-me.com
Feel free to share. Please credit the source.
Ask a Question
Please enter at least 10 characters
Related Articles
Subagents Aren't Smarter Mini-Claudes — They Solve Isolation, Not Capability
advanced · Aug 31
Reviewing Superpowers: A TDD Framework That Literally Deletes Code Written Before Tests Exist
reviews · Aug 15
Claude Mods Aren't Just Another Plugin — They Can Draw Panes and Intercept Screen Rendering, Which External Extensions Simply Can't
skill-library · Oct 05
How Long Should SKILL.md Actually Be? The Logic Behind the Official 500-Line Guideline
skill-library · Aug 31
Related News
More Related Topics