Best for
- Use when trade-offs, uncertainty, meaningful downside, or an important second opinion could change the outcome.
smshahbaj/crucible/skills/crucible/SKILL.md
Pressure-test important decisions, plans, proposals, strategies, architectures, and recommendations before acting. Use when trade-offs, uncertainty, meaningful downside, or an important second opinion could change the outcome. Stay lightweight for routine or trivial work.
Decision brief
Runs identically on Claude Code, OpenCode, and Roo Code — the methodology below is host-independent. Only the delegation mechanism (how this orchestrating pass hands work to an independent specialist lens) differs per host; see references/platforms.md for the exact mapping and t…
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/smshahbaj/crucible --skill "skills/crucible"Inspect the Agent Skill "crucible" from https://github.com/smshahbaj/crucible/blob/635f10de39c6850d69da4d31db23175e9561c750/skills/crucible/SKILL.md at commit 635f10de39c6850d69da4d31db23175e9561c750. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Budget controller — before any additional call, price it against what it can buy: Prioritize decision impact × uncertainty × plausibility of being wrong. Review single-point dependencies first. Track calls spent vs. stakes; a QUICK decision that has consumed a REVIEW-sized budge…
Help Claude make better decisions with the least sufficient compute.
important choices and trade-offs;
simple factual lookups;
Frame → Route → Verify → Review only what matters → Challenge → Quality-check → Stop → Act
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 9 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Runs identically on Claude Code, OpenCode, and Roo Code — the
methodology below is host-independent. Only the delegation mechanism (how
this orchestrating pass hands work to an independent specialist lens)
differs per host; see references/platforms.md for the exact mapping and
the fallback procedure if delegation isn't available in the current
environment.
Help Claude make better decisions with the least sufficient compute.
Think of this as a built-in second brain: it should quietly improve important choices without turning every question into a long debate or exposing the machinery. The user should experience a clearer, more trustworthy answer—not an "agent show."
Default UX: concise, direct, plain language. Go deeper only when the stakes, uncertainty, or user request justify it.
Frame → Route → Verify → Review only what matters → Challenge → Quality-check → Stop → Act
Identify only what matters:
If the literal request appears to optimize the wrong thing, answer the underlying decision instead of blindly optimizing the wording (e.g. "which laptop benchmarks highest" usually means "which laptop is best for my workload"). State the reframing briefly, only when it changes the recommendation.
If one missing fact truly blocks the answer, ask one high-value question. Otherwise state an assumption and continue.
Choose the cheapest sufficient mode:
QUICK — one-pass reasoning. Use when stakes/uncertainty are low.
REVIEW — 2 independent lenses. Use when plausible options or material uncertainty remain.
DEEP — 3–5 heterogeneous lenses plus targeted verification/challenge. Use only for high stakes, asymmetric downside, irreversibility, conflicting evidence, or unstable REVIEW results.
Never escalate because a prompt is long. Escalate only for a named unresolved issue.
For decision-critical claims, distinguish:
Prefer a direct check, calculation, authoritative source, or small experiment over another opinion when it can resolve the uncertainty more cheaply.
Anti-anchoring protocol: independent reviewers must form their views before seeing other reviewers' conclusions or a leading answer. Never show a draft verdict to a "second opinion" pass before it has reasoned independently.
Agreement is not evidence. Never use majority vote as truth. Count shared underlying evidence once. Preserve a strong minority view when it identifies a plausible decision-changing failure.
Evidence gate: a consequential recommendation may not rely on an unverified claim that, if false, would flip the action. Either verify it directly, mark the recommendation's confidence down accordingly, or name it explicitly as the key open risk. Fabricated or assumed-true evidence never passes the gate.
Counterfactual robustness: before finalizing, ask "if the strongest assumption were false, would the verdict change?" If yes and the assumption is unverified, that is a blocking uncertainty, not a footnote.
Budget controller — before any additional call, price it against what it can buy:
Prioritize decision impact × uncertainty × plausibility of being wrong.
Review single-point dependencies first. Track calls spent vs. stakes; a QUICK
decision that has consumed a REVIEW-sized budget should stop, not escalate quietly.
One-flip principle: do another call only when it could plausibly:
Otherwise stop.
Tool-use hierarchy — when a check is needed, prefer in this order:
Token-minimization protocol:
Before a consequential recommendation, check for:
Score the draft against four Decision invariants before releasing it:
If a material flaw remains, downgrade confidence or recommend the cheapest useful verification.
For REVIEW/DEEP, identify the single highest-impact plausible failure mode before finalizing — not a generic risk list. For that failure: name the trigger, the earliest observable signal, a mitigation, and the residual risk after mitigation.
Dependency-aware reasoning: map which claims the verdict actually depends on. A recommendation resting on one unverified fact is fragile even if it is surrounded by ten well-supported but non-load-bearing facts.
Independence is a process property, not a headcount. Three lenses that read the same source and reach the same conclusion are one opinion, not three. Independence requires distinct evidence paths or distinct reasoning frames — simply invoking more agents does not create it.
Evidence freshness and scope: before trusting external evidence, check it is current for a time-sensitive question and that its scope (population, geography, version, market) actually matches the decision at hand. Evidence that is authoritative but scope-mismatched is treated as unverified.
Action robustness: prefer the recommendation that holds up across the plausible range of unknowns over the one that is optimal only under the single most likely scenario.
Missing-information discipline: distinguish "unknown but irrelevant to the action" from "unknown and blocking." Only the latter earns a clarifying question or a verification step.
Research stopping contract: every verification call must name, in advance, the question it is meant to resolve and what changes if the answer comes back differently. A call with no pre-stated stopping condition is not made.
Recommendation stability test: the final verdict must survive substituting the strongest alternative's best-case assumptions. If it doesn't survive, either the confidence is too high or the verdict is wrong.
Compare:
Prefer the option that is best-supported and robust across plausible futures, not the option with the most votes.
Return a decision card:
If the evidence is weak, say so plainly. If the answer is obvious, keep it short. Do not expose hidden chain-of-thought or internal agent transcripts.
Evaluate a recommendation on five dimensions:
A recommendation is not "better" merely because it is more detailed.
Do not manufacture numeric confidence.
Instead, locate uncertainty:
Spend verification budget on material/blocking uncertainty first.
When two choices are close, prefer the path that:
This is a preference, not a universal rule: a reversible option is not automatically better if it is materially worse on the user's core objective.
Before another research/review call ask:
"If this check comes back differently, would I actually choose differently?"
If no, stop. If yes, do the cheapest credible check that can answer it.
For consequential actions, distinguish analysis from execution. Do not silently take an irreversible action merely because the analysis recommends it. Surface uncertainty and seek user confirmation when the action or permission boundary requires it.
Never invent evidence, citations, measurements, tool results, or certainty.
Load only the reference needed for the current stage:
references/routing.mdreferences/evidence.mdreferences/adversarial.mdreferences/orchestration-contract.mdreferences/output-schema.mdreferences/quality-gate.mdreferences/ledger.md (only if persisting the decision is relevant)references/platforms.md (only if delegation to a specialist lens behaves unexpectedly, or the host isn't Claude Code — it has the mapping and the no-delegation fallback)Use specialist agents only when the unresolved problem maps to their unique job.
Frequently asked questions
Runs identically on Claude Code, OpenCode, and Roo Code — the methodology below is host-independent. Only the delegation mechanism (how this orchestrating pass hands work to an independent specialist lens) differs per host; see references/platforms.md for the exact mapping and t…
The source record exposes this install command: npx skills add https://github.com/smshahbaj/crucible --skill "skills/crucible". Inspect the command and pinned source before running it.
Alternatives
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
garrytan/gbrain
End-to-end discipline for turning any large data source (audio libraries, email takeouts, document corpora, chat exports, API dumps) into brain pages at scale. The lifecycle spine: SCHEMA → ACCESS → TRIAL → EVALUATE → IMPROVE → CODIFY → TEST → SKILLIFY → BULK → MONITOR. State is tracked in a durable JSON manifest (see MANIFEST-PATTERN.md) so any crash, session boundary, or subagent fan-out resumes from ground truth instead of memory.
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
dotnet/skills
Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP). Use when user asks to "migrate to MTP", "switch from VSTest", "enable Microsoft.Testing.Platform", "use MTP runner", set OutputType=Exe only for test projects in Directory.Build.props, or mentions EnableMSTestRunner, EnableNUnitRunner, or UseMicrosoftTestingPlatformRunner. USE FOR: MTP behavioral differences vs VSTest (exit code 8, zero tests discovered, --ignore-exit-code, TESTINGPLATFORM_EXITCODE_IGNORE); centralizing