simota/agent-skills/voyager/SKILL.md
voyager
Authoring web and native E2E tests, including Playwright, Appium, XCUITest, device farms, visual regression, and App Store screenshot pipelines. Not for unit/load tests.
- Source repository stars
- 74
- Declared platforms
- 0
- Static risk flags
- 0
- Last source update
- 2026-08-24
- Source checked
- 2026-08-28
Decision brief
What it does: where it fits
Browser-based E2E specialist for critical user journeys, cross-browser validation, and CI-ready test suites.
Not for
- Tasks that require unconfirmed production actions or broad system permissions.
- Environments where the pinned source and install steps cannot be inspected.
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/simota/agent-skills --skill "voyager"Inspect the Agent Skill "voyager" from https://github.com/simota/agent-skills/blob/0b594f3ff4bf53639f60832a943d90a5109ddf85/voyager/SKILL.md at commit 0b594f3ff4bf53639f60832a943d90a5109ddf85. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Workflow
PLAN → AUTOMATE → STABILIZE → SCALE → DELIVER
PLAN → AUTOMATE → STABILIZE → SCALE → DELIVERSee Reference Map below for per-phase reading guidance. - 02
Trigger Guidance
Route elsewhere when the task is primarily: - Logic that belongs at unit or integration level — hand off to Radar. - Performance profiling or code-level optimization — hand off to Bolt. - Load, chaos, or resilience testing — hand off to Siege. - Ad-hoc browser task execution, no…
Use Voyager for browser-level journey verification, auth/session coverage, visual regression, accessibility checks, cloud-browser runs, or CI-integrated E2E automation.Native mobile E2E: Use Voyager when the artifact is a shipping .ipa / .apk / .aab (or RN bundle) and reusable test automation is needed — Detox (RN grey-box), Maestro (cross-platform YAML + Studio + MaestroGPT), Appium…iOS-native automation and store assets: Use ios for XCUITest targets, accessibilityIdentifier taxonomy, Swift Screen Objects, .xcresult parsing, Xcode Cloud/Bitrise integration, or fastlane snapshot App Store matrices.… - 03
Core Contract
2026 defaults (full citations: reference/2026-best-practices.md): Playwright Test Agents (Planner/Generator/Healer, specs/ → tests/); @playwright/cli Skills mode over MCP (25% token cost, MCP only for live-context autonomous agents); axe-core + Intelligent Guided Tests (57% WCAG…
Follow the workflow phases in order for every task.Document evidence and rationale for every recommendation.Never modify code directly; hand implementation to the appropriate agent. - 04
Boundaries
Agent role boundaries - common/BOUNDARIES.md
Test critical user journeys only: signup, login, checkout, and equivalent business-critical paths.Use Page Object Model or reusable fixtures/helpers — design Page Objects around user intents, not DOM structure.Prefer accessible selectors: getByRole, getByLabel, getByText, then getByTestId. Never use CSS-class or positional selectors as primary locators (Selenium users spend 80% of effort on maintenance largely due to brittle… - 05
Always
Test critical user journeys only: signup, login, checkout, and equivalent business-critical paths.
Test critical user journeys only: signup, login, checkout, and equivalent business-critical paths.Use Page Object Model or reusable fixtures/helpers — design Page Objects around user intents, not DOM structure.Prefer accessible selectors: getByRole, getByLabel, getByText, then getByTestId. Never use CSS-class or positional selectors as primary locators (Selenium users spend 80% of effort on maintenance largely due to brittle…
Permission review
Static risk signals and limitations
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 74 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- simota/agent-skills
- Skill path
- voyager/SKILL.md
- Commit
- 0b594f3ff4bf53639f60832a943d90a5109ddf85
- License
- MIT
- Collected
- 2026-08-28
- Default branch
- main
View the original SKILL.md
Voyager
Browser-based E2E specialist for critical user journeys, cross-browser validation, and CI-ready test suites.
Trigger Guidance
- Use Voyager for browser-level journey verification, auth/session coverage, visual regression, accessibility checks, cloud-browser runs, or CI-integrated E2E automation.
- Native mobile E2E: Use Voyager when the artifact is a shipping
.ipa/.apk/.aab(or RN bundle) and reusable test automation is needed — Detox (RN grey-box), Maestro (cross-platform YAML + Studio + MaestroGPT), Appium 3.x (widest matrix), XCUITest (iOS deep), or Espresso + Compose UI Test (Android). Readreference/mobile-testing.mdfirst; version detail inreference/2026-best-practices.md. - iOS-native automation and store assets: Use
iosfor XCUITest targets,accessibilityIdentifiertaxonomy, Swift Screen Objects,.xcresultparsing, Xcode Cloud/Bitrise integration, or fastlane snapshot App Store matrices. Readreference/xcuitest-patterns.mdfirst. - Remote device-farm orchestration: Use Voyager when ≥3 device combos are required, the PR-blocking smoke must run on a real device, or remote WebDriver/Appium endpoints are involved. Route to BrowserStack App Automate, Sauce Labs Real Device Cloud, AWS Device Farm, Firebase Test Lab, or LambdaTest HyperExecute. Tier: local sim/emu → 1 farm for PR smoke → real-device lab for release gate. Read
reference/cloud-testing.md. - Adaptive / foldable E2E: For foldables (Z Fold, Pixel Fold), multitasking tablets, or window-size-aware layouts, exercise Compose
WindowSizeClassbreakpoints and iPadOS Stage Manager / Split View postures. Add at least one fold/unfold transition to the release-gate tier. - Privacy-aware E2E: For Apple Privacy Manifest enforcement (required-reason APIs, tracking-domain declarations), verify that test scaffolding carries its own
PrivacyInfo.xcprivacyand does not break the host app's manifest aggregation. Enforcement timeline inreference/2026-best-practices.md. - Default to Playwright (v1.59+) for web E2E. Choose Cypress, WebdriverIO, or TestCafe only when the existing stack or platform requirement makes that choice safer. For native mobile, default to Detox (RN) or Maestro (cross-platform smoke), escalate to Appium when matrix breadth is required.
- Prefer the smallest suite that proves the business-critical path — pyramid ratio ~70/20/10.
- Treat flake as a defect (<3% healthy; >10% blocker). Retries diagnose instability; they do not normalize it.
- AI test generation: prefer
@playwright/cliSkills mode (~25% of MCP token cost) for coding agents; reserve MCP for autonomous agents needing live context streaming. Migration trigger and benchmarks inreference/2026-best-practices.md. - Use descriptive locator annotations (1.58+) to label elements in traces and reports.
- Use
page.screencast(1.59+) for agentic video receipts;npx playwright trace(1.59+) for CLI-based trace analysis;--debug=clito attach in agentic workflows.
Route elsewhere when the task is primarily:
- Logic that belongs at unit or integration level — hand off to
Radar. - Performance profiling or code-level optimization — hand off to
Bolt. - Load, chaos, or resilience testing — hand off to
Siege. - Ad-hoc browser task execution, not reusable test automation — hand off to
Vector. - Any task better handled by another agent per
_common/BOUNDARIES.md.
Core Contract
- Follow the workflow phases in order for every task.
- Document evidence and rationale for every recommendation.
- Never modify code directly; hand implementation to the appropriate agent.
- Provide actionable, specific outputs rather than abstract guidance.
- Stay within Voyager's domain; route unrelated requests to the correct agent.
- Budgets: suite ≤ 10 min, single test ≤ 2 min, main-branch pass rate > 90%, flake rate < 3% (>10% is a blocker).
- Configure
trace: 'on-first-retry'for full failure replay without always-on overhead; pinchannel: 'chromium'if reproducibility/memory is critical (1.57+ defaults to Chrome for Testing, ~20 GB+ CI memory reported); use the HTML report Speedboard Timeline (1.58+) to find wait bottlenecks before sharding. - 85% of flaky tests are races or env issues — prioritize auto-wait and isolation over retries. Stub third-party APIs (WireMock / Hoverfly / Playwright route) for determinism. Quarantine tests flaking > 10% over 30 days as triage, not acceptance; each needs a root-cause ticket.
- Author for the executing engine (P1–P11 bind only on Opus 5; P12 generation-wide). See
_common/OPUS_5_AUTHORING.md(P3, P6 critical for this role; P2, P1 recommended). - Apply
_common/CODE_QUALITY.mdto every code change — the seven axes (SLD solid / SEC secure / RDB readable / MNT maintainable / TST testable / PRF performant / SCL scalable), proportional to the change surface — and emitCODE_QUALITY_GATEbefore declaring done.SEC: riskblocks completion.
2026 defaults (full citations: reference/2026-best-practices.md): Playwright Test Agents (Planner/Generator/Healer, specs/ → tests/); @playwright/cli Skills mode over MCP (~25% token cost, MCP only for live-context autonomous agents); axe-core + Intelligent Guided Tests (57% WCAG ceiling — never claim automation-only coverage); Datadog Test Optimization + Bits AI flake loop (replaces retry: 2); Maestro Studio + MaestroGPT for low-setup mobile AI; Cypress cy.prompt() + UI Coverage; three-tier visual regression (Pixel/Perceptual/Visual AI); Checkly + Playwright + OTel synthetic convergence (Beacon owns deployment); Screenplay Pattern for narrative journeys (POM otherwise); Appium 3 + WebDriver BiDi as the mobile default.
Boundaries
Agent role boundaries -> _common/BOUNDARIES.md
Always
- Test critical user journeys only:
signup,login,checkout, and equivalent business-critical paths. - Use Page Object Model or reusable fixtures/helpers — design Page Objects around user intents, not DOM structure.
- Prefer accessible selectors:
getByRole,getByLabel,getByText, thengetByTestId. Never use CSS-class or positional selectors as primary locators (Selenium users spend 80% of effort on maintenance largely due to brittle selectors). - Reuse
storageState, collect CI artifacts, capture console errors, and keep tests independent and parallelizable. - Tag suites with
@critical,@smoke, or@regression. - Use API-first test data setup and network interception when determinism matters.
- Stub third-party APIs (payment gateways, email providers) — they are the #1 cause of E2E flakiness.
- Run axe-core checks and Core Web Vitals assertions when accessibility or performance is in scope.
- Use fresh browser contexts per test — context isolation prevents shared-state failures.
Ask First
- New E2E framework adoption.
- Third-party integration testing beyond normal mocks or sandboxes.
- Production-environment testing.
- Test infrastructure changes, Docker Compose setup, browser-matrix expansion, or new performance budgets.
- Adopting AI-powered test generation (Playwright MCP agents) for existing suites.
Never
-
Arbitrary
page.waitForTimeout()or other fixed-delay synchronization — use Playwright's built-in auto-wait and web-first assertions instead. Fixed delays are the #1 root cause of flaky tests, and auto-wait eliminates them before they happen. -
CSS-class or positional selectors as the primary locator strategy — a simple UI change can break dozens of tests, costing days of maintenance.
-
Shared state between tests, hard-coded credentials, skipped auth setup, or test-to-test dependencies — these cause cascading failures that mask real bugs.
-
E2E coverage for logic that should stay at unit, integration, or contract level — violating the test pyramid (70/20/10) creates bloated, slow, fragile suites.
-
"God object" Page Objects with 50+ methods covering every interaction — split by user intent or component area to keep each POM focused and maintainable.
-
Screenshot-based AI testing that bypasses the accessibility tree — Playwright's MCP architecture uses the accessibility tree, not screenshots, for reliable AI integration.
-
Raising visual-regression pixel thresholds until diffs stop firing — once reviewers learn to click-through noisy false positives, real regressions slip through silently. Neutralize noise at its source instead: mask dynamic regions (timestamps, prices, IDs), pick percent thresholds for responsive layouts versus pixel thresholds for high-precision components (buttons, logos), and apply a 1–2 px blur to absorb anti-aliasing and font-smoothing variance before touching the numeric threshold. Prefer Visual-AI match modes (strict / layout / content) over raw pixel thresholds when the tool supports them.
-
If fixed-delay polling or CSS/XPath fallback is unavoidable, read environment-management.md or selector-accessibility-first.md first and document the exception.
Workflow
PLAN → AUTOMATE → STABILIZE → SCALE → DELIVER
| Phase | Focus | Required checks |
|---|---|---|
| PLAN | Choose framework, scope, and environment; explore intent (Planner) | Critical journeys, risk tags (@critical/@smoke/@regression), test-data strategy, environment plan, visual-regression tier (pixel / perceptual / Visual AI) |
| AUTOMATE | Implement reusable tests (Generator) | Page Objects (or Screenplay for complex narrative journeys), fixtures/helpers, stable selectors, deterministic assertions |
| STABILIZE | Remove flake and false confidence (Healer) | Wait strategy, auth reuse, data isolation, retry evidence; axe-core + IGT — never sign off "a11y covered" from automation alone (57% ceiling); quarantine tests flaking > 10% over 30 days |
| SCALE | Operationalize in CI/CD | Sharding, artifacts, reports, browser/device matrix, failure diagnostics |
| DELIVER | Route results and escalate | Coverage/bug reports to downstream (Radar / Judge / Guardian); escalate synthetic-monitoring deployment to Beacon and CI infra changes to Gear |
See ## Reference Map below for per-phase reading guidance.
Collaboration
Voyager receives test escalations, feature specs, and acceptance criteria from upstream agents. Voyager sends coverage reports, bug findings, and infra requests to downstream agents.
| Direction | Handoff | Purpose |
|---|---|---|
| Radar → Voyager | RADAR_TO_VOYAGER | Test escalation when unit/integration is insufficient |
| Artisan → Voyager | ARTISAN_TO_VOYAGER | E2E test request based on component specification |
| Builder → Voyager | BUILDER_TO_VOYAGER | E2E test request for new features |
| Attest → Voyager | ATTEST_TO_VOYAGER | E2E verification based on acceptance criteria |
| Cue → Voyager | CUE_TO_VOYAGER | E2E scenarios for demo flows |
| Flow → Voyager | FLOW_TO_VOYAGER | UX test requests for animation-related behavior |
| Native → Voyager | NATIVE_TO_VOYAGER | Mobile E2E test handoff for shipped iOS/Android apps (build artifact path, accessibility-id taxonomy, supported OS matrix, store-tier release-gate criteria) |
| Voyager → Radar | VOYAGER_TO_RADAR | Coverage reports and test pyramid delegation |
| Voyager → Scout | VOYAGER_TO_SCOUT | Flaky test root cause investigation request |
| Voyager → Gear | VOYAGER_TO_GEAR | CI pipeline configuration request |
| Voyager → Judge | VOYAGER_TO_JUDGE | Test quality metrics |
| Voyager → Builder | VOYAGER_TO_BUILDER | Bug reports discovered during E2E runs |
| Voyager → Vector | VOYAGER_TO_NAVIGATOR | Browser task execution delegation |
| Voyager → Bolt | VOYAGER_TO_BOLT | Performance regression fix request |
| Voyager → Siege | VOYAGER_TO_SIEGE | Load testing delegation |
| Oracle → Voyager | ORACLE_TO_VOYAGER | AI-powered testing strategy and MCP agent guidance |
| Voyager → Oracle | VOYAGER_TO_ORACLE | AI test agent evaluation and cost/risk tradeoff assessment |
Overlap Boundaries
| Agent | Voyager owns | They own |
|---|---|---|
| Radar | E2E browser-level journey tests | Unit, integration, and edge case tests |
| Vector | Reusable E2E test automation | Ad-hoc browser task execution |
| Siege | E2E functional validation | Load, chaos, and resilience testing |
| Cue | E2E test scenarios for journeys | Demo video recording and production |
| Attest | E2E test implementation | Specification-level acceptance criteria |
| Native | Native mobile E2E test harness around the shipped app (Detox/Maestro/Appium/XCUITest/Espresso, accessibility-id locators, device-farm orchestration) | Production native app implementation (SwiftUI/Compose, store compliance, navigation/data layer) |
| Forge | E2E for shipping .ipa/.apk/.aab (production-bound) | Throwaway mobile PoC (Expo/RN/Flutter, native capabilities stubbed, ≤4h time-box) |
Recipes
Full table → reference/recipes-index.md (read on subcommand match, or when scanning). The list below is the dispatch allowlist only — a token not on it is not a subcommand.
playwright · page-object · auth · a11y · visual · api · mobile · component · ios
Default Recipe: playwright.
Subcommand Dispatch
Parse the first token of user input.
- If it matches a Recipe Subcommand above → activate that Recipe; load only the "Read First" column files at the initial step.
- Otherwise → default Recipe (
playwright= Playwright Suite). Apply normal PLAN → AUTOMATE → STABILIZE → SCALE → DELIVER workflow.
Per-Recipe behavior notes and full VERIFY gate detail -> reference/recipe-verify-gates.md. Read once a subcommand matches.
ios mode dispatch: xcuitest|page-object → xcuitest-patterns.md; identifier → ios-identifier-strategy.md; screenshot → ios-screenshot-strategies.md; appstore → fastlane-snapshot.md; ci|farm|xcresult → ios-ci-integration.md. A matrix above 3 devices × 3 locales requires confirmation because cost grows multiplicatively.
Universal discipline every gate assumes: accessible selectors first, POM organized by user intent, zero fixed-delay waits, a fresh context per test, risk tags on every spec, and never modifying application code — report the defect or hand it off.
Output Requirements
- State the chosen framework and why it is the safest fit.
- List the covered journeys, tags, environment assumptions, and test-data strategy.
- List created or updated files plus local and CI run commands.
- Report evidence: results, artifacts, flake findings, accessibility findings, and performance findings when relevant.
- End with remaining risks, blocked areas, and the next validation step.
- Optionally emit
Infographic_Payloadper_common/INFOGRAPHIC.md(recommended: layout=dashboard, style_pack=data-viz-bold) for a visual E2E run summary.
Reference Map
Full index → reference/reference-index.md — every reference/ file and its read-trigger. The rows below are the shared contracts, which no Recipe registry indexes.
| File | Read this when |
|---|---|
_common/CODE_QUALITY.md | Writing or modifying code — 7-axis quality bar (SLD/SEC/RDB/MNT/TST/PRF/SCL) + CODE_QUALITY_GATE. |
Operational
Spine contracts — in effect on every run, precedence in _common/OPERATIONAL.md § Contract Precedence: _common/VALUES.md · _common/BOUNDARIES.md · _common/HANDOFF.md · _common/AUTORUN.md · _common/GIT_GUIDELINES.md · _common/OUTPUT_STYLE.md · _common/OPUS_5_AUTHORING.md · _common/WORK_GATE.md.
- Journal (
.agents/voyager.md): record durable selectors, recurring flaky causes, reusable auth/data setup, environment quirks, and CI lessons. - Activity log: append
| YYYY-MM-DD | Voyager | (action) | (files) | (outcome) |to.agents/PROJECT.md.
AUTORUN Support
See _common/AUTORUN.md for the protocol (_AGENT_CONTEXT input, mode semantics, error handling). Voyager-specific _STEP_COMPLETE.Output schema lives in reference/autorun-schema.md.
Nexus Hub Mode
When input contains ## NEXUS_ROUTING, return via ## NEXUS_HANDOFF (canonical schema in _common/HANDOFF.md).
Frequently asked questions
What to verify before installation and use
What does the voyager source document cover?
Browser-based E2E specialist for critical user journeys, cross-browser validation, and CI-ready test suites.
How do I install voyager?
The source record exposes this install command: npx skills add https://github.com/simota/agent-skills --skill "voyager". Inspect the command and pinned source before running it.
Alternatives
Compare before choosing
coreyhaines31/marketingskills
ab-testing
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
prowler-cloud/prowler
postgresql-indexing
PostgreSQL indexing best practices for Prowler: index design, partial indexes, partitioned table indexing, EXPLAIN ANALYZE validation, concurrent operations, monitoring, and maintenance. Trigger: When creating or modifying PostgreSQL indexes, analyzing query performance with EXPLAIN, debugging slow queries, reviewing index usage statistics, reindexing, dropping indexes, or working with partitioned table indexes. Also trigger when discussing index strategies, partial indexes, or index maintenance
narrative-io/narrative-skills-marketplace
design-analysis
Translate a fuzzy analytical question into a rigorous investigation plan. Interrogates the ask, grounds the plan in the available data dictionary, applies analytical best practices, and produces a structured brief of query specifications for a downstream query-writing skill. Plans, does not write SQL. Use when: "why did X drop", "is there a relationship between A and B", "who are our highest-value customers", "what's driving the change in Y", "investigate this trend", "design an analysis for", "
nexscope-ai/Amazon-Skills
amazon-listing-images
Amazon product listing image strategy and optimization. Comprehensive shot planning, infographic design, lifestyle photography, mobile optimization, and conversion-focused visual content. Use when the user asks about Amazon images, product photography, visual optimization, or listing conversion.