Source profileQuality 92/100Review permissions

terrylica/cc-skills/plugins/quality-tools/skills/multi-agent-e2e-validation/SKILL.md

multi-agent-e2e-validation

Multi-agent parallel E2E validation for database refactors. TRIGGERS - E2E validation, schema migration testing, database refactor validation.

Source repository stars
61
Declared platforms
0
Static risk flags
1
Last source update
2026-08-26
Source checked
2026-08-28

Decision brief

What it does: where it fits

Self-Evolving Skill: This skill improves through use. If instructions are wrong, parameters drifted, or a workaround was needed — fix this file immediately, don't defer. Only update for real, reproducible issues.

Best for

  • Database refactors (e.g., v3.x file-based → v4.x QuestDB)
  • Schema migrations requiring validation
  • Bulk data ingestion pipeline testing

Not for

  • 1. Skipping Environment Validation
  • 2. Serial Agent Execution

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/terrylica/cc-skills --skill "plugins/quality-tools/skills/multi-agent-e2e-validation"
Safe inspection promptEditorial

Inspect the Agent Skill "multi-agent-e2e-validation" from https://github.com/terrylica/cc-skills/blob/05f53c5b24a445c1895e9b0590212e66cd70f39e/plugins/quality-tools/skills/multi-agent-e2e-validation/SKILL.md at commit 05f53c5b24a445c1895e9b0590212e66cd70f39e. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Workflow: Step-by-Step

    Input: ADR document (e.g., ADR-0002 QuestDB Refactor) Output: Validation plan with 3-7 agents

    Input: ADR document (e.g., ADR-0002 QuestDB Refactor) Output: Validation plan with 3-7 agents
  2. 02

    Step 1: Create Validation Plan (ADR-Driven)

    Input: ADR document (e.g., ADR-0002 QuestDB Refactor) Output: Validation plan with 3-7 agents

    Input: ADR document (e.g., ADR-0002 QuestDB Refactor) Output: Validation plan with 3-7 agents
  3. 03

    Agent 1: Environment Setup

    Deploy QuestDB via Docker

    Deploy QuestDB via DockerApply schema.sqlValidate connectivity (ILP, PG, HTTP)
  4. 04

    Step 2: Execute Agent 1 (Environment)

    ✅ Container running

    ✅ Container running✅ Ports accessible (9009 ILP, 8812 PG, 9000 HTTP)✅ Schema applied without errors
  5. 05

    Step 3: Execute Agents 2-3 in Parallel

    Agent 3: Query Interface

    Agent 3: Query Interface

Permission review

Static risk signals and limitations

Runs scripts

medium · line 287

The documentation asks the agent to run terminal commands or scripts.

git add src/gapless_crypto_clickhouse/collectors/questdb_bulk_loader.py

Runs scripts

medium · line 288

The documentation asks the agent to run terminal commands or scripts.

git commit -m "fix: prevent pandas from treating first CSV column as index

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score92/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars61SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
terrylica/cc-skills
Skill path
plugins/quality-tools/skills/multi-agent-e2e-validation/SKILL.md
Commit
05f53c5b24a445c1895e9b0590212e66cd70f39e
License
MIT
Collected
2026-08-28
Default branch
main
View the original SKILL.md

Multi-Agent E2E Validation

Self-Evolving Skill: This skill improves through use. If instructions are wrong, parameters drifted, or a workaround was needed — fix this file immediately, don't defer. Only update for real, reproducible issues.

Overview

Prescriptive workflow for spawning parallel validation agents to comprehensively test database refactors. Successfully identified 5 critical bugs (100% system failure rate) in QuestDB migration that would have shipped in production.

When to Use This Skill

Use this skill when:

  • Database refactors (e.g., v3.x file-based → v4.x QuestDB)
  • Schema migrations requiring validation
  • Bulk data ingestion pipeline testing
  • System migrations with multiple validation layers
  • Pre-release validation for database-centric systems

Key outcomes:

  • Parallel agent execution for comprehensive coverage
  • Structured validation reporting (VALIDATION_FINDINGS.md)
  • Bug discovery with severity classification (Critical/Medium/Low)
  • Release readiness assessment

Core Methodology

1. Validation Architecture (3-Layer Model)

Layer 1: Environment Setup

  • Container orchestration (Colima/Docker)
  • Database deployment and schema application
  • Connectivity validation (ILP, PostgreSQL, HTTP ports)
  • Configuration file creation and validation

Layer 2: Data Flow Validation

  • Bulk ingestion testing (CloudFront → QuestDB)
  • Performance benchmarking against SLOs
  • Multi-month data ingestion
  • Deduplication testing (re-ingestion scenarios)
  • Type conversion validation (FLOAT→LONG casts)

Layer 3: Query Interface Validation

  • High-level query methods (get_latest, get_range, execute_sql)
  • Edge cases (limit=1, cross-month boundaries)
  • Error handling (invalid symbols, dates, parameters)
  • Gap detection SQL compatibility

2. Agent Orchestration Pattern

Sequential vs Parallel Execution:

Agent 1 (Environment) → [SEQUENTIAL - prerequisite]
  ↓
Agent 2 (Bulk Loader) → [PARALLEL with Agent 3]
Agent 3 (Query Interface) → [PARALLEL with Agent 2]

Dependency Rule: Environment validation must pass before data flow/query validation

Dynamic Todo Management:

  • Start with high-level plan (ADR-defined phases)
  • Prune completed agents from todo list
  • Grow todo list when bugs discovered (e.g., Bug #5 found by Agent 3)
  • Update VALIDATION_FINDINGS.md incrementally

3. Validation Script Structure

Each agent produces:

  1. Test Script (e.g., test_bulk_loader.py)
    • 5+ test functions with clear pass/fail criteria
    • Structured output (test name, result, details)
    • Summary report at end
  2. Artifacts (logs, config files, evidence)
  3. Findings Report (bugs, severity, fix proposals)

Example Test Structure:

def test_feature(conn):
    """Test 1: Feature description"""
    print("=" * 80)
    print("TEST 1: Feature description")
    print("=" * 80)

    results = {}

    # Test 1a: Subtest name
    print("\n1a. Testing subtest:")
    result_1a = perform_test()
    print(f"   Result: {result_1a}")
    results["subtest_1a"] = result_1a == expected_1a

    # Summary
    print("\n" + "-" * 80)
    all_passed = all(results.values())
    print(f"Test 1 Results: {'✓ PASS' if all_passed else '✗ FAIL'}")
    for test_name, passed in results.items():
        print(f"  - {test_name}: {'✓' if passed else '✗'}")

    return {"success": all_passed, "details": results}

4. Bug Classification and Tracking

Severity Levels:

  • 🔴 Critical: 100% system failure (e.g., API mismatch, timestamp corruption)
  • 🟡 Medium: Degraded functionality (e.g., below SLO performance)
  • 🟢 Low: Minor issues, edge cases

Bug Report Format:

#### Bug N: Descriptive Name (**SEVERITY** - Status)

**Location**: `file/path.py:line`

**Issue**: One-sentence description

**Impact**: Quantified impact (e.g., "100% ingestion failure")

**Root Cause**: Technical explanation

**Fix Applied**: Code changes with before/after

**Verification**: Test results proving fix

**Status**: ✅ FIXED / ⚠️ PARTIAL / ❌ OPEN

5. Release Readiness Decision Framework

Go/No-Go Criteria:

BLOCKER = Any Critical bug unfixed
SHIP = All Critical bugs fixed + (Medium bugs acceptable OR fixed)
DEFER = >3 Medium bugs unfixed OR any High-severity bug

Example Decision:

  • 5 Critical bugs found → all fixed ✅
  • 1 Medium bug (performance 55% below SLO) → acceptable ✅
  • Verdict: RELEASE READY

Workflow: Step-by-Step

Step 1: Create Validation Plan (ADR-Driven)

Input: ADR document (e.g., ADR-0002 QuestDB Refactor) Output: Validation plan with 3-7 agents

Plan Structure:

## Validation Agents

### Agent 1: Environment Setup

- Deploy QuestDB via Docker
- Apply schema.sql
- Validate connectivity (ILP, PG, HTTP)
- Create .env configuration

### Agent 2: Bulk Loader Validation

- Test CloudFront → QuestDB ingestion
- Benchmark performance (target: >100K rows/sec)
- Validate deduplication (re-ingestion test)
- Multi-month ingestion test

### Agent 3: Query Interface Validation

- Test get_latest() with various limits
- Test get_range() with date boundaries
- Test execute_sql() with parameterized queries
- Test detect_gaps() SQL compatibility
- Test error handling (invalid inputs)

Step 2: Execute Agent 1 (Environment)

Directory Structure:

tmp/e2e-validation/
  agent-1-env/
    test_environment_setup.py
    questdb.log
    config.env
    schema-check.txt

Validation Checklist:

  • ✅ Container running
  • ✅ Ports accessible (9009 ILP, 8812 PG, 9000 HTTP)
  • ✅ Schema applied without errors
  • ✅ .env file created

Step 3: Execute Agents 2-3 in Parallel

Agent 2: Bulk Loader

tmp/e2e-validation/
  agent-2-bulk/
    test_bulk_loader.py
    ingestion_benchmark.txt
    deduplication_test.txt

Agent 3: Query Interface

tmp/e2e-validation/
  agent-3-query/
    test_query_interface.py
    gap_detection_test.txt

Execution:

# Terminal 1
cd tmp/e2e-validation/agent-2-bulk
uv run python test_bulk_loader.py

# Terminal 2
cd tmp/e2e-validation/agent-3-query
uv run python test_query_interface.py

Step 4: Document Findings in VALIDATION_FINDINGS.md

Template:

# E2E Validation Findings Report

**Validation ID**: ADR-XXXX
**Branch**: feat/database-refactor
**Date**: YYYY-MM-DD
**Target Release**: vX.Y.Z
**Status**: [BLOCKED / READY / IN_PROGRESS]

## Executive Summary

E2E validation discovered **N critical bugs** that would have caused [impact]:

| Finding | Severity | Status | Impact       | Agent   |
| ------- | -------- | ------ | ------------ | ------- |
| Bug 1   | Critical | Fixed  | 100% failure | Agent 2 |

**Recommendation**: [RELEASE READY / BLOCKED / DEFER]

## Agent 1: Environment Setup - [STATUS]

...

## Agent 2: [Name] - [STATUS]

...

Step 5: Iterate on Fixes

For each bug:

  1. Document in VALIDATION_FINDINGS.md with 🔴/🟡/🟢 severity
  2. Apply fix to source code
  3. Re-run failing test
  4. Update bug status to ✅ FIXED
  5. Commit with semantic message (e.g., fix: correct timestamp parsing in CSV ingestion)

Example Fix Commit:

git add src/gapless_crypto_clickhouse/collectors/questdb_bulk_loader.py
git commit -m "fix: prevent pandas from treating first CSV column as index

BREAKING CHANGE: All timestamps were defaulting to epoch 0 (1970-01)
due to pandas read_csv() auto-indexing. Added index_col=False to
preserve first column as data.

Fixes #ABC-123"

Step 6: Final Validation and Release Decision

Run all tests:

/usr/bin/env bash << 'SKILL_SCRIPT_EOF'
cd tmp/e2e-validation
for agent in agent-*; do
    echo "=== Running $agent ==="
    cd $agent
    uv run python test_*.py
    cd ..
done
SKILL_SCRIPT_EOF

Update VALIDATION_FINDINGS.md status:

  • Count Critical bugs: X fixed, Y open
  • Count Medium bugs: X fixed, Y open
  • Apply decision framework
  • Update Status field to ✅ RELEASE READY or ❌ BLOCKED

Real-World Example: QuestDB Refactor Validation

Context: Migrating from file-based storage (v3.x) to QuestDB (v4.0.0)

Bugs Found:

  1. 🔴 Sender API mismatch - Used non-existent Sender.from_uri() instead of Sender.from_conf()
  2. 🔴 Type conversion - number_of_trades sent as FLOAT, schema expects LONG
  3. 🔴 Timestamp parsing - pandas treating first column as index → epoch 0 timestamps
  4. 🔴 Deduplication - WAL mode doesn't provide UPSERT semantics (needed DEDUP ENABLE UPSERT KEYS)
  5. 🔴 SQL incompatibility - detect_gaps() used nested window functions (QuestDB unsupported)

Impact: Without this validation, v4.0.0 would ship with 100% data corruption and 100% ingestion failure

Outcome: All 5 bugs fixed, system validated, v4.0.0 released successfully

Common Pitfalls

1. Skipping Environment Validation

Bad: Assume Docker/database is working, jump to data ingestion tests ✅ Good: Agent 1 validates environment first, catches port conflicts, schema errors early

2. Serial Agent Execution

Bad: Run Agent 2, wait for completion, then run Agent 3 ✅ Good: Run Agent 2 & 3 in parallel (no dependency between them)

3. Manual Test Reporting

Bad: Copy/paste test output into Slack/email ✅ Good: Structured VALIDATION_FINDINGS.md with severity, status, fix tracking

4. Ignoring Medium Bugs

Bad: "Performance is 55% below SLO, but we'll fix it later" ✅ Good: Document in VALIDATION_FINDINGS.md, make explicit go/no-go decision

5. No Re-validation After Fixes

Bad: Apply fix, assume it works, move on ✅ Good: Re-run failing test, update status in VALIDATION_FINDINGS.md

Resources

scripts/

Not applicable - validation scripts are project-specific (stored in tmp/e2e-validation/)

references/

  • example_validation_findings.md - Complete VALIDATION_FINDINGS.md template
  • agent_test_template.py - Template for creating validation test scripts
  • bug_severity_classification.md - Detailed severity criteria and examples

assets/

Not applicable - validation artifacts are project-specific


Troubleshooting

IssueCauseSolution
Container not startingColima/Docker not runningRun colima start before Agent 1
Port conflictsPorts already in useStop conflicting containers or use different ports
Schema application failsInvalid SQL syntaxCheck schema.sql for database-specific compatibility
Agent 2/3 fail without Agent 1Environment not validatedEnsure Agent 1 completes before starting Agent 2/3
Test script import errorsMissing dependenciesRun uv pip install in agent directory
Bug status not updatingVALIDATION_FINDINGS.md staleManually refresh status after each fix
Parallel agents interferenceShared resources conflictEnsure agents use isolated directories
Decision unclearSeverity mixed Critical/MediumApply Go/No-Go criteria strictly per documentation

Post-Execution Reflection

After this skill completes, reflect before closing the task:

  1. Locate yourself. — Find this SKILL.md's canonical path before editing.
  2. What failed? — Fix the instruction that caused it.
  3. What worked better than expected? — Promote to recommended practice.
  4. What drifted? — Fix any script, reference, or dependency that no longer matches reality.
  5. Log it. — Evolution-log entry with trigger, fix, and evidence.

Do NOT defer. The next invocation inherits whatever you leave behind.

Frequently asked questions

What to verify before installation and use

What does the multi-agent-e2e-validation source document cover?

Self-Evolving Skill: This skill improves through use. If instructions are wrong, parameters drifted, or a workaround was needed — fix this file immediately, don't defer. Only update for real, reproducible issues.

How do I install multi-agent-e2e-validation?

The source record exposes this install command: npx skills add https://github.com/terrylica/cc-skills --skill "plugins/quality-tools/skills/multi-agent-e2e-validation". Inspect the command and pinned source before running it.

Which permission-related actions were detected?

Static rules flagged exec-script in the source; the page lists the matching lines and excerpts.

Alternatives

Compare before choosing

Computed 97198

microsoft/Sico

android-tester

Execute Android UI workflows on a sandbox device, review results, and produce a structured execution report.

Computed 9621

upex-galaxy/agentic-qa-boilerplate

regression-testing

Execute regression test suites via CI/CD, analyze results, classify failures, and produce GO/NO-GO release decisions. Use when running regression, smoke, or sanity suites through GitHub Actions, monitoring workflow runs, downloading Allure or Playwright artifacts, classifying failures (REGRESSION vs FLAKY vs KNOWN vs ENVIRONMENT vs NEW TEST), computing pass-rate and trend metrics, deciding release readiness, generating executive quality reports, or creating regression issues. Triggers on: run re

Computed 9536,049

K-Dense-AI/scientific-agent-skills

simpy

Build, inspect, test, and analyze bounded process-based discrete-event simulations with SimPy, including events, resources, interrupts, monitoring, replications, warm-up, and reproducible output analysis.

Computed 9489

aAAaqwq/AGI-Super-Team

trade-prediction-markets

Build and test Polymarket prediction market trading strategies for YES/NO token trading. Provides 6 tools: get_all_prediction_events (browse markets, $0.001), get_prediction_market_data (analyze price history, $0.001), create_prediction_market_strategy (generate code, $1-$4.50), run_prediction_market_backtest (test performance, $0.001). Trade on real-world events (politics, economics, sports, crypto). Currently simulation only (live deployment coming soon).