Best for
- An evaluation verdict is a regression, underpowered, or "no credible improvement".
- A skill wins in the isolated arm but not in the plugin arm, or is reported "not activated".
- /evaluate reports "Evaluation ran but produced no results".
dotnet/skills/.agents/skills/improve-skill-quality/SKILL.md
Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate reports no results, or when deciding whether a weak skill should be strengthened or retired. Do not use for scaffolding a brand-new skill (use create-skill) or a brand-new eval (use create-skill-test).
Decision brief
Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/dotnet/skills --skill ".agents/skills/improve-skill-quality"Inspect the Agent Skill "improve-skill-quality" from https://github.com/dotnet/skills/blob/73555e9231867c5978db07191514b7beb22cd253/.agents/skills/improve-skill-quality/SKILL.md at commit 73555e9231867c5978db07191514b7beb22cd253. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Read InvestigatingResults.md for how to download artifacts and read results.json. Extract, per failing stimulus:
Read InvestigatingResults.md for how to download artifacts and read results.json. Extract, per failing stimulus:
Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.
See references/eval-triage.md for the full catalogue. The recurring ones:
Run python eng/eval-quality/checkevalquality.py — it blocks eleven defect classes that can cost a real result here. Then confirm by hand:
Permission review
The documentation asks the agent to run terminal commands or scripts.
python eng/eval-quality/check_eval_quality.pyEvidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 95/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 5,241 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.
/evaluate reports "Evaluation ran but produced no results".create-skill.eval.yaml from scratch — use create-skill-test.eng/skill-validator, eng/vally-adapter, evaluation*.yml).| Input | Required | Description |
|---|---|---|
| Verdict evidence | Yes | The /evaluate PR comment, or results.json from the run artifacts |
| Losing trial transcripts | Yes for content fixes | Baseline vs. skilled output plus the judge's stated reason |
| Stimulus-vote W/T/L and repeated-run W/T/L | Yes | Separates cross-task evidence from reliability |
| Activation status per arm | Yes | Isolated and plugin activation are different failures |
Read InvestigatingResults.md for how to
download artifacts and read results.json. Extract, per failing stimulus:
Do not change skill content until you can quote a losing trial and the judge's reason for it. For the other cause classes the evidence is different: harness failures are diagnosed from the job log and the spec, and power problems from the trial record — neither has a losing trial to quote, and demanding one is what sends people rewriting prose instead.
Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.
| Symptom | Real cause class | Go to |
|---|---|---|
| A fixture does not build, is untracked by git, breaks for the wrong reason, or contradicts itself | Fixture | Step 4 |
No results.json, "produced no results", or the spec never loaded | Harness / spec-load | Step 3 |
| Trials errored, timed out, or returned empty output | Reliability | Step 3 |
| Trajectories unmatched, a trial errored, or the summary disagrees — verdict reported inconclusive | Reliability (not power) | Step 3 |
| Positive record (e.g. 16W/8T/1L), comparison conclusive, verdict still not a pass | Statistical power | Step 5 |
| Skilled arm equals baseline arm by construction | Eval design | Step 6 |
| Activated and lost on quality, judge names a concrete defect | Skill content | Step 7 |
| Activated in isolation, not in plugin | Activation / routing | Step 8 |
| Not activated in either arm | Frontmatter description | Step 8 |
| Wins but costs far more than baseline | Scope and cost | Step 7 |
A verdict is only a measured result when the comparison was conclusive: adapt.mjs requires zero
errored trials, zero unmatched trajectories, and an agreeing summary before it will report a pass or
a regression. Confirm that before reading a record as a power problem.
See references/eval-triage.md for the full catalogue. The recurring ones:
config: and defaults: is rejected by vally, the job still exits 0, and
the PR comment blames "transient infrastructure". Merge them into one defaults: block.session.idle
failures look identical from the verdict and need harness fixes, not SDK pins.expect_tools: [bash] on an advisory question forces a restore or build and turns an answer into
a timeout with no quality gain.Run python eng/eval-quality/check_eval_quality.py — it blocks eleven defect classes that can
cost a real result here. Then confirm by hand:
git ls-files), not merely on disk — .gitignore
has silently swallowed committed coverage fixtures;line-rate, summary totals and <line> elements differ is the canonical case — or the
two arms legitimately read different truths.The gate has two independent bars, and confusing them is the usual misdiagnosis:
underpowered — never a pass, never a regression.| discordant stimulus votes | records that pass | p |
|---|---|---|
| ≤ 4 | none, however good the skill | ≥ 0.0625 |
| 5–7 | zero losses only (5W/0L) | 0.031 |
| 8 | one loss survivable (7W/1L) | 0.035 |
So at exactly 5 stimuli a single tie is fatal — it leaves 4 discordant. At 6 stimuli one tie is survivable (5W/1T/0L); at 7, up to two are (5W/2T/0L). A loss is not.
So a positive record with a failing verdict is a power problem, not a content problem. Fix it by
adding discriminating stimuli. Raising runs measures reliability for the same task and cannot
clear the floor.
An eval that compares the skill against itself measures judge noise:
expect_activation: false) must not also set constraints.reject_skills.
That makes the skilled arm skill-free, i.e. identical to baseline. Across four evals the same
guard scored −0.4, +0.4, +0.4 and 0, twice costing a skill its pass.disable-model-invocation: true cannot self-activate, so an eval graded on
activation compares two identical arms. Cover it through a consumer skill, or grade the answer
content instead, as tests/dotnet-test/filter-syntax/eval.yaml and
tests/dotnet-test/platform-detection/eval.yaml do.config is missing its required key enforces nothing, so the stimulus has one
fewer assertion than it appears to.Only now change the skill. Apply the patterns in references/writing-for-baseline-delta.md; the ones that most often flip a loss:
references/ reads and size any
orchestration to the user's scope.Activation failures are frontmatter and routing failures, not body failures. See references/eval-triage.md. Summary:
| Failure | Fix |
|---|---|
| Not activated in any arm | Put the user's own words in description: symptoms, error codes, artifact names, quoted requests |
| A sibling skill wins the prompt | Claim the exact ambiguous words in description, and add matching exclusions on both siblings |
| Model answers with no skill at all | Raise the stakes in the description, de-crowd the plugin menu, verify with the plugin arm |
| Boundary excludes real scenarios | Re-read every "do not use for" clause against every eval prompt and real workflow phase |
| Description at the 1,024-char ceiling | Cut restated body content, not trigger phrases; check the plugin menu budget too |
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill>
Then request the official run by submitting a PR review containing /evaluate (Files changed →
Review changes), which binds the run to the reviewed commit. Before declaring a regression on the
result, confirm the skill payload actually changed — reruns on byte-identical content have shifted
7W/2T/2L to 4W/5T/2L.
check_eval_quality.py and skill-validator check both pass.| Pitfall | Solution |
|---|---|
| Rewriting skill prose in response to an underpowered verdict | Underpowered means too few distinct stimuli; add discriminating stimuli instead |
Adding defaults: runs: to a spec that already has config: | Merge into a single defaults: block; vally rejects specs with both |
Padding runs to clear the stimulus floor | Repeats measure reliability for one task; add stimuli |
| Treating an errored trial as fixture nondeterminism | Read the stderr first; judge-side auth failures need harness fixes |
| Fixing a "wrong" answer that the fixture actually made wrong | Check fixture self-consistency before blaming the response |
| Strengthening a skill nobody uses and nothing passes | Weak eval signal plus thin telemetry is a valid retirement case |
| Landing a fix without re-running | Verify the invoked payload contains the fix; judge noise is real |
results.json. This is the current guide; the similarly-named eng/skill-validator/src/docs/InvestigatingResults.md documents the retired skill-validator evaluate schema and does not describe today's results.Frequently asked questions
Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.
The source record exposes this install command: npx skills add https://github.com/dotnet/skills --skill ".agents/skills/improve-skill-quality". Inspect the command and pinned source before running it.
Static rules flagged exec-script in the source; the page lists the matching lines and excerpts.
Alternatives
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
OpenClaudia/openclaudia-skills
Create email drip campaigns, nurture sequences, and automated email flows. Includes templates for welcome series, abandoned cart, re-engagement, product launch, and onboarding sequences. Trigger phrases: "email sequence", "drip campaign", "nurture sequence", "email flow", "welcome series", "abandoned cart emails", "onboarding emails", "email automation", "product launch emails", "re-engagement campaign", "send email", "send sequence".
aaron-he-zhu/aaron-marketing-skills
Use when the user asks to "run a deliverability pre-flight before I send", "check my SPF/DKIM/DMARC/BIMI", "why am I landing in spam / promotions", or "score my sender reputation and list hygiene"; runs the ONE-TIME pre-send SEND S1 authentication pre-flight and builds the SEND S (Sender-integrity / Deliverability) evidence read — DNS + DMARC-RUA auth, domain/IP reputation, inbox placement, content/link/render, and point-in-time bounce/complaint hygiene — using Pass/Partial/Fail/Unknown/N/A stat
github/awesome-copilot
Build, scaffold, and deploy Power Automate cloud flows using the FlowStudio MCP server. Your agent constructs flow definitions, wires connections, deploys, and tests — all via MCP without opening the portal. Load this skill when asked to: create a flow, build a new flow, deploy a flow definition, scaffold a Power Automate workflow, construct a flow JSON, update an existing flow's actions, patch a flow definition, add actions to a flow, wire up connections, or generate a workflow definition from