Best for
- Two or more judges have run on the same target and their verdict blocks need
- A mixed review spanning code judges + the artifact/defence judges (e.g. a
- /review-changes step 5 (consolidation) wants the canonical synthesis format
event4u-app/agent-config/src/skills/judge-synthesis/SKILL.md
Use to consolidate multiple already-run judge verdicts into one report — consensus, conflicts, must-fix/should-fix with per-judge provenance. Consume-only, no opaque score, never auto-gates.
Decision brief
You are the synthesizer across an already-run set of judges. You do not run the judges and you do not re-judge — you consume their emitted verdict blocks and produce one structured report: a side-by-side verdict table, the consensus findings (flagged by ≥2 judges = highest confi…
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/event4u-app/agent-config --skill "src/skills/judge-synthesis"Inspect the Agent Skill "judge-synthesis" from https://github.com/event4u-app/agent-config/blob/6a5670b7881a676c0da90d2afb950298087c4ccb/src/skills/judge-synthesis/SKILL.md at commit 6a5670b7881a676c0da90d2afb950298087c4ccb. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
First check that ≥2 judge verdict blocks are present (if only one, stop — nothing to synthesize). Then tabulate one row per judge: judge · target · verdict · finding count. Preserve each judge's own verdict word (do not normalise away breached into reject); add a severity tier (…
Two or more judges have run on the same target and their verdict blocks need to become one decision-ready report. A mixed review spanning code judges + the artifact/defence judges (e.g. a PR that ships a roadmap + code + an injection-defense fixture) — no single existing command…
The verdict blocks the judges already emit. Each carries at minimum a judge name, a verdict, and zero or more findings with a severity. The three verdict vocabularies in the suite map onto one ordered severity axis:
First check that ≥2 judge verdict blocks are present (if only one, stop — nothing to synthesize). Then tabulate one row per judge: judge · target · verdict · finding count. Preserve each judge's own verdict word (do not normalise away breached into reject); add a severity tier (…
A finding flagged by ≥2 judges (same file:line / same dimension / same technique) is a consensus finding — the highest-confidence item. List these first. Consensus is by overlap of the finding, never by counting votes for a score.
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 93/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 9 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
You are the synthesizer across an already-run set of judges. You do not run the judges and you do not re-judge — you consume their emitted verdict blocks and produce one structured report: a side-by-side verdict table, the consensus findings (flagged by ≥2 judges = highest confidence), the conflicts (judges that disagree), and a synthesized must-fix / should-fix / advisory split with per-finding judge provenance. No opaque single score. Never auto-gates — the human decides.
/review-changes step 5 (consolidation) wants the canonical synthesis format
instead of an ad-hoc merge.Do NOT use when:
judge-bug-hunter,
judge-artifact-completeness,
judge-injection-defense, etc.).The verdict blocks the judges already emit. Each carries at minimum a judge name, a verdict, and zero or more findings with a severity. The three verdict vocabularies in the suite map onto one ordered severity axis:
| Judge family | Verdict vocabulary |
|---|---|
code judges (judge-bug-hunter, -code-quality, -security-auditor, -test-coverage, architecture-review-lens) | apply / revise / reject |
judge-artifact-completeness | complete / partial / incomplete |
judge-injection-defense | defended / partial / breached |
judge-spec-compliance | per criterion: SATISFIED / PARTIAL / MISSING — plus a criteria_source state that is not a verdict |
Ordered worst→best: reject/incomplete/breached > revise/partial > apply/complete/defended.
judge-spec-compliance is deliberately absent from that axis. Its verdicts
answer a different question — did the change do what was asked — and mapping
MISSING onto reject would let a craft-clean diff average it away, which is
the miss this judge was added to catch. It gets its own dimension in § 4c.
First check that ≥2 judge verdict blocks are present (if only one, stop — nothing
to synthesize). Then tabulate one row per judge: judge · target · verdict ·
finding count. Preserve each
judge's own verdict word (do not normalise away breached into reject); add a
severity tier (worst/mid/clean) only as a sort key.
A sort key, never a filter. Severity orders the synthesis; it never decides
what enters it. Dropping a judge's low-severity finding during synthesis
reproduces the pre-filter defect one layer up — the finding was found, the
reviewer reported it, and the aggregator withheld it. Filtering is the
consumer's pass, after the ledger is whole. Output shape, the separate
Confidence field, and the preserve-an-unverified-S0 rule are specified once in
adversarial-review-protocol
§ 3.
A finding flagged by ≥2 judges (same file:line / same dimension / same technique) is a consensus finding — the highest-confidence item. List these first. Consensus is by overlap of the finding, never by counting votes for a score.
Two judges reaching opposite verdicts on the same target (one apply, another
reject) is a conflict. Surface both verdicts and the disagreement
explicitly — never silently resolve it by averaging or vote-count. The human
adjudicates. The only deterministic rule: for the must-fix list, the most
severe verdict wins (a single reject puts the target in must-fix even if four
judges said apply) — but the conflict is still shown so the human sees it was
contested.
reject / incomplete
/ breached) + every consensus finding at the highest severity.revise / partial) findings.Each entry carries provenance: which judge(s) raised it. Never merge two judges' findings into one unattributed line.
A panelist assertion carrying neither fresh evidence produced this run (a
file:line, a command's output, a diff hunk) nor a citation is marked
uncited where it appears. It still ships.
FLAG THE UNCITED ASSERTION. NEVER DROP IT.
A SUPPRESSED FINDING IS INDISTINGUISHABLE FROM A FINDING NOBODY MADE.
The distinction is load-bearing and it is the reason this is a marking rule rather than a filter: an unevidenced assertion may still be the most valuable line in the report — a judge noticing something it could not yet prove is exactly the signal a human wants. What the reader needs is to know which category they are reading, not to be protected from one of them.
Drop was considered and rejected against this file. The formulation this was adapted from offers "drop or flag"; the drop half contradicts § 4's Advisory tier three paragraphs up — "Emitted in full, never elided … a finding that reached a judge and not the reader was suppressed by the aggregator". Taking it would have put two rules in one skill in direct conflict, with the newer one silently winning. Marking satisfies the same goal (the reader can tell evidence from assertion) at zero information cost.
Scope: this marks panelist assertions inside a synthesis. It does not reach
the reviewed change, and it is not the code-provenance knowledge-layer
obligation, which governs what a durable artefact asserts rather than what a
transient review does.
A SPEC FINDING NEVER BECOMES A CRAFT FINDING.
REPORT THE SPEC DIMENSION SEPARATELY, WITH ITS `criteria_source` STATE.
A CRAFT-CLEAN DIFF THAT MISSES ITS CRITERION IS NOT A CLEAN REVIEW.
NO CRITERIA SUPPLIED IS AN UNVERIFIED DIMENSION, NEVER A PASS.
judge-spec-compliance findings do not enter must-fix / should-fix /
advisory. They form a separate block carrying, in this order: the
criteria_source state, the per-criterion table, and the count of MISSING
plus PARTIAL.
The separation is the whole point. Folded into the craft tiers a MISSING
criterion competes with five other judges' findings and can be outvoted by
their silence; kept apart it cannot be, because there is nothing to average it
against. Consensus (§ 2) and conflict (§ 3) do not apply either: no other judge
reads the criteria, so a spec finding has no possible second voter and its
absence of corroboration says nothing about it.
Three criteria_source states, and each says something different about what
this review established:
| State | What the synthesis reports |
|---|---|
supplied | the dimension was verified; report the per-criterion verdicts |
not_provided | the dimension was not verified — say so, never report it as clean |
supplied_unparseable | an error, not a no-criteria run: criteria were handed over and could not be read, so the reader must know their review silently skipped a dimension it was asked to check |
One sentence, not a number: block (any worst-tier verdict), revise (any
mid-tier, no worst-tier), or proceed (all clean). This is a recommendation the
human acts on — it does not gate anything.
The sentence names the spec dimension explicitly, in every case. A MISSING
criterion is a block and a PARTIAL one is at least a revise, whatever the
craft judges said. Where no criteria were supplied, the sentence says what was
and was not established — "craft quality verified; requirement compliance NOT
verified (no criteria supplied)" — because a bare proceed over an unverified
dimension reads as a full pass, and that reading is the defect: a reviewer
cannot tell "we checked and it complies" from "nobody checked" unless the
sentence distinguishes them.
Synthesis: <N> judges over <target>
Recommendation: block | revise | proceed
Verdicts:
judge-bug-hunter reject (2 findings)
judge-security-auditor apply (0)
judge-artifact-completeness partial (1)
...
Consensus (≥2 judges — highest confidence):
🔴 path:line — <finding> [judge-bug-hunter, judge-code-quality]
Conflicts (judges disagree — human adjudicates):
<target>: judge-bug-hunter=reject vs judge-security-auditor=apply
Must-fix:
🔴 <finding> [judge]
Should-fix:
🟡 <finding> [judge]
Advisory:
🟢 <finding> [judge]
Required fields (ordered):
breached and reject sort the same
but mean different things; keep each judge's own word in the table.apply; but show the
conflict so the human sees it was contested, never bury it./review-changes); synthesis has nothing to do
on empty input.judge-bug-hunter,
judge-code-quality,
judge-security-auditor,
judge-test-coverage,
architecture-review-lens.judge-artifact-completeness,
judge-injection-defense./review-changes
(5 code judges), subagent-orchestration
(parallel judge fan-out).ai-council
(independent breadth), /team (collaborative repo-access
depth) — this skill consolidates in-session same-weights judges.Frequently asked questions
You are the synthesizer across an already-run set of judges. You do not run the judges and you do not re-judge — you consume their emitted verdict blocks and produce one structured report: a side-by-side verdict table, the consensus findings (flagged by ≥2 judges = highest confi…
The source record exposes this install command: npx skills add https://github.com/event4u-app/agent-config --skill "src/skills/judge-synthesis". Inspect the command and pinned source before running it.
Alternatives
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
coreyhaines31/marketingskills
When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
wanshuiyin/Auto-claude-code-research-in-sleep
Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.