Source profileQuality 95/100

NVIDIA-NeMo/nemo-platform/plugins/nemo-eval-author/skills/eval-author-audit/SKILL.md

eval-author-audit

Generate, validate, measure, and report on an audit-spec coverage denominator for Eval Author. Use when the user wants a hand-editable audit.md file derived from Ethos, needs schema enforcement for declared tools, capabilities, failure cases, evidence, and references, wants to measure which audit items one ATIF trace covers, or wants to aggregate coverage across measured traces. Changes none of the user's source, and saves audit artifacts under `.eval-author/`.

Source repository stars
72
Declared platforms
0
Static risk flags
0
Last source update
2026-08-28
Source checked
2026-08-28

Decision brief

What it does: where it fits

Read eval-author for the shared standard, vocabulary, and boundaries. This sub-flow generates and validates a finite coverage denominator from ETHOS.md and reviewed audit items, then can measure one ATIF trace and aggregate coverage reports against it. It does not generate tasks…

Best for

  • Use when the user wants a hand-editable audit.

Not for

  • Tasks that require unconfirmed production actions or broad system permissions.
  • Environments where the pinned source and install steps cannot be inspected.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/NVIDIA-NeMo/nemo-platform --skill "plugins/nemo-eval-author/skills/eval-author-audit"
Safe inspection promptEditorial

Inspect the Agent Skill "eval-author-audit" from https://github.com/NVIDIA-NeMo/nemo-platform/blob/640f78c55bdf902b49fa80868731423352108eef/plugins/nemo-eval-author/skills/eval-author-audit/SKILL.md at commit 640f78c55bdf902b49fa80868731423352108eef. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Step 1: Draft Or Update Audit Items

    Before drafting or updating .eval-author/audit-items.yaml, read templates/audit.md and schemas/audit.schema.json. Use the template as the worked example and the JSON Schema descriptions as the field definitions. Do not use validation as the primary way to discover the format; va…

    Before drafting or updating .eval-author/audit-items.yaml, read templates/audit.md and schemas/audit.schema.json. Use the template as the worked example and the JSON Schema descriptions as the field definitions. Do not…Read ETHOS.md and draft audit items at the level between Ethos and runnable tasks: canonical tools, high-level capabilities, and material failure cases. Keep the list finite. Do not create separate items for prompt para…For tool items, use the names that appear in the actual runtime traces or tool registry, including eval-specific tools that may be more precise than product tools named in Ethos prose. If Ethos describes a generic tool…
  2. 02

    Step 2: Generate Or Reconcile Audit.md

    Create or update .eval-author/audit.md from ETHOS.md and the reviewed item proposals:

    Create or update .eval-author/audit.md from ETHOS.md and the reviewed item proposals:The default mode is reconcile. If .eval-author/audit.md does not exist, it creates the file. If it already exists, the generator parses the existing marked block, updates source metadata such as the Ethos digest, preser…By default, the generator treats .eval-author/audit-items.yaml as a partial update proposal. Missing existing items are not stale in that mode, because the items file may contain only additions or edits. Use --items-mod…
  3. 03

    Step 3: Validate

    Run validation after every generated or hand-edited audit file:

    Run validation after every generated or hand-edited audit file:schemas/audit.schema.json is the canonical structural schema. The Python validator applies that schema first, then checks any source digests that are provided and cross-item references such as requiredtools and appliest…Validation proves only structure and references, not that the denominator is complete or correct.
  4. 04

    Step 4: Measure One ATIF Trace

    After validation, measure one completed trial directory or one ATIF trajectory file. Use --measure to select one or more comma-separated measurement methods; the default is toolcalls.

    After validation, measure one completed trial directory or one ATIF trajectory file. Use --measure to select one or more comma-separated measurement methods; the default is toolcalls.The current --trial-dir reader supports Harbor-style trial directories that normally contain agent/trajectory.json; agents that do not emit ATIF may not have that file. When the trace file is already known, pass it dire…--measure may be passed more than once or as CSV, for example --measure toolcalls,boundary. The script loads the trajectory once, then runs each selected method against the same parsed Harbor trajectory model. Unknown m…
  5. 05

    Step 5: Aggregate Coverage Reports

    After measuring one or more traces, aggregate the per-trace coverage.json files into a coverage report:

    After measuring one or more traces, aggregate the per-trace coverage.json files into a coverage report:Use --coverage for explicit files, or repeat --coverage-dir and --coverage when the inputs are split across directories. The script scans coverage directories recursively for files named coverage.json, validates every i…The aggregate report uses schemas/auditcoveragereport.schema.json. It contains overall and per-kind count summaries, the union of covered item names, warnings, the measured item kinds, and uncovereditems. Each uncovered…

Permission review

Static risk signals and limitations

No configured static risk pattern was detected

This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score95/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars72SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
NVIDIA-NeMo/nemo-platform
Skill path
plugins/nemo-eval-author/skills/eval-author-audit/SKILL.md
Commit
640f78c55bdf902b49fa80868731423352108eef
License
Apache-2.0
Collected
2026-08-28
Default branch
main
View the original SKILL.md

Eval Author: audit

Read eval-author for the shared standard, vocabulary, and boundaries. This sub-flow generates and validates a finite coverage denominator from ETHOS.md and reviewed audit items, then can measure one ATIF trace and aggregate coverage reports against it. It does not generate tasks yet.

The audit-spec approach has three item kinds in v1:

KindMeaning
toolA canonical tool name the agent may call
capabilityA high-level behavior the agent should exercise
failure_caseExpected safe behavior when a capability cannot proceed normally

Every item uses name as its stable coverage key. Names must be unique across the whole file; do not add sequential numeric IDs. Tool references in required_tools, expected_tools, and evidence_required[].tool must match the name of a declared tool item. prohibited_tools may name any syntactically valid tool name, including tools the agent must never call and therefore should not declare as allowed tools.

Write audit artifacts under .eval-author/. Do not edit the customer's source, existing evals, source-of-truth documents, or ETHOS.md.

Scripts

Audit-spec mechanics live under scripts/audit_spec/:

Read scripts/audit_spec/README.md for the current measurement assumptions: ATIF input, Harbor trajectory parsing, v1 tool_calls coverage only, and coverage aggregation from coverage.json files.

ScriptUse it to
scripts/audit_spec/generate.pyCreate, reconcile, replace, or preview .eval-author/audit.md from ETHOS.md and reviewed item proposals
scripts/audit_spec/measure.pyMeasure one ATIF trace or Harbor trial directory against audit.md and write coverage/details files for each selected method
scripts/audit_spec/report.pyAggregate per-trace coverage.json files into one coverage report with uncovered audit items
scripts/audit_spec/validate.pyValidate the marked audit-spec block in audit.md

Shared helpers are private modules in the same tree: scripts/audit_spec/_schema.py and scripts/audit_spec/_markdown.py. Measurement uses Harbor's harbor.models.trajectories.Trajectory to read ATIF files. Measurement methods live under scripts/audit_spec/measurements/; v1 ships scripts/audit_spec/measurements/tool_calls.py. Shared coverage output is defined by schemas/audit_coverage.schema.json. Aggregate coverage output is defined by schemas/audit_coverage_report.schema.json. Tool-call debug output is defined by schemas/audit_tool_calls_details.schema.json. Concrete instances live under examples/schemas/tool_calls.coverage.json, examples/schemas/tool_calls.details.json, and examples/schemas/coverage_report.json. Runtime dependencies are listed in requirements.txt.

Step 1: Draft Or Update Audit Items

Before drafting or updating .eval-author/audit-items.yaml, read templates/audit.md and schemas/audit.schema.json. Use the template as the worked example and the JSON Schema descriptions as the field definitions. Do not use validation as the primary way to discover the format; validation is the enforcement and repair step after drafting.

Read ETHOS.md and draft audit items at the level between Ethos and runnable tasks: canonical tools, high-level capabilities, and material failure cases. Keep the list finite. Do not create separate items for prompt paraphrases, fixture variants, or ordinary happy-path permutations.

For tool items, use the names that appear in the actual runtime traces or tool registry, including eval-specific tools that may be more precise than product tools named in Ethos prose. If Ethos describes a generic tool such as sqlite but measurement traces expose execute_sql and submit_sql, declare the runtime tool names and connect capabilities or failure cases to those names. Do not invent tool names that will not appear in the measurement surface.

Save the reviewed item proposals as .eval-author/audit-items.yaml. The items file may be either a mapping with an items key or the item list itself. It should use the same item shape shown in templates/audit.md and enforced by schemas/audit.schema.json.

For an initial audit, this file should contain the full proposed denominator. For an update, it may contain only the proposed additions or edits. Existing reviewed audit.md items remain the source of truth in reconcile mode.

Capabilities that do not need tools, such as policy refusals or out-of-scope handling, should use required_tools: []. Failure cases attach to capability names through applies_to; tool-level failure expectations stay on the tool item as expected_failure_behavior.

Step 2: Generate Or Reconcile Audit.md

Create or update .eval-author/audit.md from ETHOS.md and the reviewed item proposals:

uv run --with pyyaml --with jsonschema \
  <skill_dir>/scripts/audit_spec/generate.py \
  --ethos ETHOS.md \
  --items .eval-author/audit-items.yaml \
  --out .eval-author/audit.md

The default mode is reconcile. If .eval-author/audit.md does not exist, it creates the file. If it already exists, the generator parses the existing marked block, updates source metadata such as the Ethos digest, preserves existing item bodies by stable name, appends new proposed items, and reports proposed edits without silently rewriting them. Hand-authored prose outside the marked block is preserved.

By default, the generator treats .eval-author/audit-items.yaml as a partial update proposal. Missing existing items are not stale in that mode, because the items file may contain only additions or edits. Use --items-mode full only when the items file is intended to be the complete denominator; then existing items omitted from the proposal are reported as possibly_stale_items.

Use the explicit modes when the default is not what the user wants:

uv run --with pyyaml --with jsonschema \
  <skill_dir>/scripts/audit_spec/generate.py \
  --ethos ETHOS.md \
  --items .eval-author/audit-items.yaml \
  --out .eval-author/audit.md \
  --mode suggest

uv run --with pyyaml --with jsonschema \
  <skill_dir>/scripts/audit_spec/generate.py \
  --ethos ETHOS.md \
  --items .eval-author/audit-items.yaml \
  --out .eval-author/audit.md \
  --mode reconcile \
  --items-mode full

uv run --with pyyaml --with jsonschema \
  <skill_dir>/scripts/audit_spec/generate.py \
  --ethos ETHOS.md \
  --items .eval-author/audit-items.yaml \
  --out .eval-author/audit.md \
  --mode replace

suggest performs the same comparison as reconcile but writes nothing. replace rewrites the whole file from the item proposal file, including prose outside the marked block, and should be used only when the user wants to discard the existing generated audit file.

The generator prints a JSON summary containing written, added_items, conflicting_items, conflicting_items_applied, and possibly_stale_items. Treat conflicting_items as items where the proposal differs from the reviewed audit item; reconcile mode preserves the reviewed item and reports conflicting_items_applied: false so the user can accept the change manually or through a future editor. If reconcile adds items, finds conflicts, reports stale items in full mode, or detects an agent-name change, an approved audit is demoted to status: draft unless the user passes --status approved.

The generator adds an optional sources entry for Ethos with name: ethos, a path relative to audit.md, and a real sha256 digest. It uses the frontmatter name from ETHOS.md when --agent is omitted. If that inferred name differs from the existing audit's agent, reconcile preserves the reviewed audit name and reports agent_change unless the user passes --agent explicitly. source_refs are advisory provenance notes in v1; the validator preserves them but does not resolve them against sources until a future generator or grammar defines that reference format.

Step 3: Validate

Run validation after every generated or hand-edited audit file:

uv run --with pyyaml --with jsonschema \
  <skill_dir>/scripts/audit_spec/validate.py --audit .eval-author/audit.md

schemas/audit.schema.json is the canonical structural schema. The Python validator applies that schema first, then checks any source digests that are provided and cross-item references such as required_tools and applies_to.

Validation proves only structure and references, not that the denominator is complete or correct.

Step 4: Measure One ATIF Trace

After validation, measure one completed trial directory or one ATIF trajectory file. Use --measure to select one or more comma-separated measurement methods; the default is tool_calls.

uv run --with-requirements <skill_dir>/requirements.txt \
  <skill_dir>/scripts/audit_spec/measure.py \
  --audit .eval-author/audit.md \
  --trial-dir <harbor-job-dir>/<trial-dir> \
  --measure tool_calls \
  --out-dir .eval-author/audit-measurements

The current --trial-dir reader supports Harbor-style trial directories that normally contain agent/trajectory.json; agents that do not emit ATIF may not have that file. When the trace file is already known, pass it directly and stamp the task explicitly:

uv run --with-requirements <skill_dir>/requirements.txt \
  <skill_dir>/scripts/audit_spec/measure.py \
  --audit .eval-author/audit.md \
  --trace <path-to>/trajectory.json \
  --task-id <task-id> \
  --run-id <run-id> \
  --measure tool_calls \
  --out-dir .eval-author/audit-measurements

--measure may be passed more than once or as CSV, for example --measure tool_calls,boundary. The script loads the trajectory once, then runs each selected method against the same parsed Harbor trajectory model. Unknown method names fail before the trace is loaded.

The script writes one folder per task, run, and method. Task and run ids are encoded as single path components so ids containing / cannot create nested or escaping paths:

.eval-author/audit-measurements/task=<encoded-task-id>/run=<encoded-run-id>/tool_calls/coverage.json
.eval-author/audit-measurements/task=<encoded-task-id>/run=<encoded-run-id>/tool_calls/details.json

coverage.json uses the shared coverage schema and contains only the stable audit item names this trace covered plus provider-neutral subject identity (trace, trace_format, task_id, and run_id). It also records item_kind_count, the denominator for the measured item kind. Coverage aggregation should consume this file and ignore method-specific debug details. details.json is specific to the selected method and carries traceability data for humans.

The only v1 measurement method is tool_calls. It covers only audit items whose kind is tool by matching each tool item's name against ATIF steps[].tool_calls[].function_name, including embedded subagent trajectories. It does not judge capabilities, failure cases, user intent, output quality, or other evidence kinds.

The script validates coverage.json against schemas/audit_coverage.schema.json and validates details.json against the selected method's details schema before writing. Use the next step to union coverage across tasks and runs.

Step 5: Aggregate Coverage Reports

After measuring one or more traces, aggregate the per-trace coverage.json files into a coverage report:

uv run --with-requirements <skill_dir>/requirements.txt \
  <skill_dir>/scripts/audit_spec/report.py \
  --audit .eval-author/audit.md \
  --coverage-dir .eval-author/audit-measurements \
  --out .eval-author/audit-coverage-report.json

Use --coverage <path-to-coverage.json> for explicit files, or repeat --coverage-dir and --coverage when the inputs are split across directories. The script scans coverage directories recursively for files named coverage.json, validates every input against schemas/audit_coverage.schema.json, and rejects inputs whose audit metadata no longer matches the current audit.md. A status-only mismatch, such as measurements produced while the audit was draft and aggregated after it became approved, is reported as a warning instead of forcing a re-measure.

The aggregate report uses schemas/audit_coverage_report.schema.json. It contains overall and per-kind count summaries, the union of covered item names, warnings, the measured item kinds, and uncovered_items. Each uncovered item includes the original audit item plus generation-oriented context: a stable reason, a one-sentence focus, likely needed_tools, and the item's evidence_required. Use reason: not_measured_by_any_method to distinguish gaps that no included measurement method could close from reason: not_covered_by_any_input_report, which means the item kind was measured but no input report covered that item. Treat that list as the input for a later task-generation step.

Next Steps

  • For audit-generation inputs and reconciliation modes, return to Step 2: Generate Or Reconcile Audit.md.
  • When aggregate uncovered_items includes actionable tools (reason: not_covered_by_any_input_report), hand off to eval-author-task-create to scaffold and prove one gap at a time. Capability or failure-case items with reason: not_measured_by_any_method stay audit findings only in v1.

Frequently asked questions

What to verify before installation and use

What does the eval-author-audit source document cover?

Read eval-author for the shared standard, vocabulary, and boundaries. This sub-flow generates and validates a finite coverage denominator from ETHOS.md and reviewed audit items, then can measure one ATIF trace and aggregate coverage reports against it. It does not generate tasks…

How do I install eval-author-audit?

The source record exposes this install command: npx skills add https://github.com/NVIDIA-NeMo/nemo-platform --skill "plugins/nemo-eval-author/skills/eval-author-audit". Inspect the command and pinned source before running it.

Alternatives

Compare before choosing

Computed 10045,960

coreyhaines31/marketingskills

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

Computed 10045,960

coreyhaines31/marketingskills

churn-prevention

When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o

Computed 10025,136

alirezarezvani/claude-skills

app-store-optimization

App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist

Computed 10015,385

wanshuiyin/Auto-claude-code-research-in-sleep

citation-audit

Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.