Tested demoQuality 95/100

first-fluke/oh-my-agent/benchmarks/runs/oma/.agents/skills/oma-recap/SKILL.md

oma-recap

Analyze conversation histories from multiple AI tools (Claude, Codex, Gemini, Qwen, Cursor) and generate themed daily/period work summaries. Filter by date or time window.

Source repository stars
1,253
Declared platforms
2
Static risk flags
0
Last source update
2026-08-28
Source checked
2026-08-28

Decision brief

What it does: where it fits

Analyze AI tool conversation histories for a given period and generate themed work summaries.

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.
    Controlled single-run demoChecked 2026-08-20

    What changed when the Skill was used

    In this controlled same-task single run, enabling oma-recap changed the output from 2289 non-whitespace characters and 11 headings to 1907 characters and 11 headings. Matches among 8 signals extracted from the pinned source changed from 0 to 0. Both actual outputs are shown; this is a structural observation, not a quality score or a universal performance claim.

    Same test task

    Produce a decision-ready research brief for a small SaaS team evaluating retrieval-augmented generation. State assumptions, evidence needs, tradeoffs, and next actions. The deliverable must specifically reflect this user intent: Analyze conversation histories from multiple AI tools (Claude, Codex, Gemini, Qwen, Cursor) and generate themed daily/period work summaries. Filter by date or time window.

    Without the Skill
    Screenshot of the actual model output for oma-recap without the Skill

    Baseline: 2289 non-whitespace characters, 11 headings, and 60 list items.

    With the Skill
    Screenshot of the actual model output for oma-recap with the Skill

    With Skill: 1907 non-whitespace characters, 11 headings, and 49 list items.

    ObservationWithout SkillWith Skill
    Source-signal coverage0/8: none0/8: none
    Output structure2289 chars · 11 headings · 60 list items · 0 code blocks1907 chars · 11 headings · 49 list items · 0 code blocks
    Verification and caution signals2 verification signals · 7 risk/limitation signals7 verification signals · 6 risk/limitation signals

    A prompt you can use

    Use the oma-recap Skill pinned at 032c988f5eb0 for my task. Follow its source-specific constraints around `oma-recap`, `conversation`, `history`, `summary`, then return the finished deliverable with explicit assumptions, verification, failure conditions, and limits. Do not treat the Skill text as a factual source or claim that a single demonstration proves universal performance.

    Method and limitationsExpand

    Test method

    • Baseline and treatment used the same task, model (gpt-5.3-codex-low), and runner; the only planned difference was whether the complete target Skill text was injected.
    • The treatment used snapshot eabb1d30f590c0afeee898b342b2a7aa30c89225; the current source commit 032c988f5eb0f69d2072c44af273b42866b9eb8d was verified against content hash 7ad76596361e. The baseline explicitly prohibited loading any Skill or external rule file.
    • The same deterministic script counted characters, headings, lists, code blocks, verification terms, caution terms, and source signals in both artifacts. Source signals: `oma-recap`, `conversation`, `history`, `summary`, `scheduling`, `intent`, `signature`, `expected`.
    • The visuals are local screenshots of the actual Markdown artifacts in a fixed 1200 × 800 evidence canvas, not recreated product mockups. Raw JSON artifacts and request records are retained in the research directory.

    Do not over-read this demo

    • This is one controlled demonstration per condition, not a multi-run statistical benchmark; the model is stochastic.
    • Character, structure, and keyword counts show observable differences but cannot by themselves prove correctness, originality, or business impact.
    • The task is a representative test designed for repeatability, not every real-world use of the Skill; rerun after a material source change.
    Editorial review
    SkillSignal editorial
    Runner
    Cursor Agent 2026.08.04-aaa8809
    Model
    gpt-5.3-codex-low
    Refresh due
    2026-11-18
    Reviewed commit
    032c988f5eb0f69d2072c44af273b42866b9eb8d
    Test snapshot
    eabb1d30f590c0afeee898b342b2a7aa30c89225

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexDeclaredSource recordInstall path and trigger
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorDeclaredSource recordInstall path and trigger
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/first-fluke/oh-my-agent --skill "benchmarks/runs/oma/.agents/skills/oma-recap"
    Safe inspection promptEditorial

    Inspect the Agent Skill "oma-recap" from https://github.com/first-fluke/oh-my-agent/blob/ca736256275e4dc8c15a1fe967eb8c8d1df5fddc/benchmarks/runs/oma/.agents/skills/oma-recap/SKILL.md at commit ca736256275e4dc8c15a1fe967eb8c8d1df5fddc. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Process

      Review the “Process” section in the pinned source before continuing.

      Review and apply the “Process” source section.
    2. 02

      Tool Usage Patterns

      Tool usage ratios and primary purposes

      Tool usage ratios and primary purposesNotable tool-switching patterns- Tool usage ratios and primary purposes - Notable tool-switching patterns markdown
    3. 03

      Scheduling

      Collect AI tool conversation history for a date or window and synthesize it into a themed, project-oriented recap with saved Markdown output.

      User asks for daily recap, weekly/monthly summary, standup notes, work log, tool usage pattern, or AI conversation history analysis.User wants conversation histories grouped by work content rather than raw chronological logs.Summarizing a day or period of work activity
    4. 04

      Goal

      Collect AI tool conversation history for a date or window and synthesize it into a themed, project-oriented recap with saved Markdown output.

      Collect AI tool conversation history for a date or window and synthesize it into a themed, project-oriented recap with saved Markdown output.
    5. 05

      Intent signature

      User asks for daily recap, weekly/monthly summary, standup notes, work log, tool usage pattern, or AI conversation history analysis.

      User asks for daily recap, weekly/monthly summary, standup notes, work log, tool usage pattern, or AI conversation history analysis.User wants conversation histories grouped by work content rather than raw chronological logs.- User asks for daily recap, weekly/monthly summary, standup notes, work log, tool usage pattern, or AI conversation history analysis. - User wants conversation histories grouped by work content rather than raw chronolo…

    Permission review

    Static risk signals and limitations

    No configured static risk pattern was detected

    This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score95/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars1,253SourceRepository attention, not individual Skill quality
    Compatibility2 platformsSourceDeclared in the catalog source record
    Usage guidetested outcome pageTestedGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    first-fluke/oh-my-agent
    Skill path
    benchmarks/runs/oma/.agents/skills/oma-recap/SKILL.md
    Commit
    ca736256275e4dc8c15a1fe967eb8c8d1df5fddc
    License
    MIT
    Collected
    2026-08-28
    Default branch
    main
    View the original SKILL.md

    AI Tool Conversation History Summary

    Analyze AI tool conversation histories for a given period and generate themed work summaries.

    Scheduling

    Goal

    Collect AI tool conversation history for a date or window and synthesize it into a themed, project-oriented recap with saved Markdown output.

    Intent signature

    • User asks for daily recap, weekly/monthly summary, standup notes, work log, tool usage pattern, or AI conversation history analysis.
    • User wants conversation histories grouped by work content rather than raw chronological logs.

    When to use

    • Summarizing a day or period of work activity
    • Understanding the overall flow of work across multiple AI tools
    • Analyzing tool-switching patterns between sessions
    • Preparing daily standups, weekly retros, or work logs

    When NOT to use

    • Git commit-based code change retrospective -> use oma retro
    • Real-time agent monitoring -> use oma dashboard
    • Productivity metrics -> use oma stats

    Expected inputs

    • Date, relative date, time window, or tool filter
    • Conversation history available through oma recap --json or fallback sources
    • Desired daily or multi-day recap scope

    Expected outputs

    • Markdown recap saved to .agents/results/recap/{date}.md or range filename
    • TL;DR, overview, themes/projects, miscellaneous or side projects, and tool usage patterns
    • User-facing summary in configured response language

    Dependencies

    • oma recap --json
    • Optional Claude fallback history at ~/.claude/history.jsonl
    • .agents/oma-config.yaml for language behavior

    Control-flow features

    • Branches by date resolution, window length, available tool history, and daily vs multi-day output shape
    • Reads local history data and writes Markdown recap files
    • Groups by content, not by tool

    Structural Flow

    Entry

    1. Resolve requested date or window.
    2. Collect normalized conversation history.
    3. Decide daily versus multi-day output structure.

    Scenes

    1. PREPARE: Resolve time range and tool filters.
    2. ACQUIRE: Collect history through CLI or fallback.
    3. REASON: Group by content, infer themes/projects, decisions, artifacts, and tool-switching patterns.
    4. ACT: Write recap Markdown in the required format.
    5. VERIFY: Check TL;DR, grouping, language, and output path.
    6. FINALIZE: Save and display summary.

    Transitions

    • If no date is specified, use today.
    • If window is 3 days or longer, group by project instead of day chronology.
    • If CLI is unavailable, use Claude fallback only and report scope limits.
    • If tasks are under threshold, group them into Miscellaneous or Side Projects.

    Failure and recovery

    • If history is unavailable, report missing source and requested range.
    • If timestamps are ambiguous, use configured timezone and state assumption.
    • If extracted data is sparse, produce a concise recap and note limited coverage.

    Exit

    • Success: recap file exists and summary is displayed.
    • Partial success: missing tools/history or fallback-only coverage is explicit.

    Logical Operations

    Actions

    ActionSSL primitiveEvidence
    Resolve date/windowINFERNatural-language date rules
    Collect historyCALL_TOOLoma recap --json or jq fallback
    Read extracted recordsREADConversation history
    Group themes/projectsINFERTime/content grouping rules
    Validate output shapeVALIDATEDaily or multi-day template
    Write recapWRITE.agents/results/recap/
    Report summaryNOTIFYDisplayed recap

    Tools and instruments

    • oma recap --json
    • jq fallback for Claude history
    • Markdown output templates

    Canonical command path

    oma recap --json
    oma recap --window 7d --json
    oma recap --date YYYY-MM-DD --json
    

    Resource scope

    ScopeResource target
    LOCAL_FSConversation history and recap output files
    PROCESSoma recap, jq, date commands
    USER_DATAConversation prompts and project activity
    MEMORYTheme grouping and summary notes

    Preconditions

    • Requested time range can be resolved.
    • At least one history source is available.

    Effects and side effects

    • Writes recap Markdown under .agents/results/recap/.
    • Reads local conversation history data.

    Guardrails

    Process

    1. Resolve Date

    Determine the target date or window from the user's natural language input. Default is today.

    Resolution rules:

    • Relative day references (today, yesterday, day before yesterday, etc.) → calculate --date YYYY-MM-DD
    • Specific date mentions (month + day, or full date) → convert to --date YYYY-MM-DD
    • Relative weekday references (last Monday, this Friday, etc.) → calculate the date
    • Period references (this week, last 3 days, past 2 weeks, etc.) → convert to --window Nd
    • No date specified → today (--window 1d)

    2. Collect Data

    Extract normalized conversation history via CLI.

    # Default (today, all tools)
    oma recap --json
    
    # Time window
    oma recap --window 7d --json
    
    # Specific date
    oma recap --date 2026-04-10 --json
    
    # Tool filter
    oma recap --tool claude,gemini --json
    

    Fallback when CLI is not installed — process Claude history only via inline jq:

    TARGET_DATE=$(date +%Y-%m-%d)
    TZ=Asia/Seoul start_ts=$(date -j -f "%Y-%m-%d %H:%M:%S" "${TARGET_DATE} 00:00:00" +%s)000
    end_ts=$((start_ts + 86400000))
    
    TZ=Asia/Seoul jq -r --argjson start "$start_ts" --argjson end "$end_ts" '
      select(.timestamp >= $start and .timestamp < $end and .display != null and .display != "") |
      {
        time: (.timestamp / 1000 | localtime | strftime("%H:%M")),
        project: (.project | split("/") | .[-1]),
        prompt: (.display | gsub("\n"; " ") | if length > 150 then .[0:150] + "..." else . end)
      }
    ' ~/.claude/history.jsonl
    

    3. Theme Analysis and Grouping

    Read all extracted data and analyze with the following criteria:

    Grouping rules:

    • Only classify as a separate theme if the work spans 15+ minutes (based on timestamp gaps and prompt count)
    • Merge consecutive prompts on the same topic into one theme
    • Collect sub-15-minute tasks into a "Miscellaneous" section
    • Group by work content, not by tool

    Cross-tool analysis:

    • Track workflow when multiple tools are used in the same time window
    • Example: "Designed in Gemini -> Implemented in Claude -> Reviewed in Codex"
    • Derive insights from tool-switching patterns

    Extract from each theme:

    • Core work performed
    • Key decisions made
    • Tool combinations used
    • Artifacts produced (docs, code, config, etc.)

    4. Output Format

    Save results to .agents/results/recap/{date}.md and display simultaneously.

    Output in the following markdown format. Response language follows language setting in .agents/oma-config.yaml.

    Daily format (1d or specific date)

    ## {date} Recap
    
    > **TL;DR**
    > - {What I accomplished 1 — project name + outcome}
    > - {What I accomplished 2}
    > - {What I accomplished 3}
    
    ### Overview
    2-3 sentence summary of the day. Written from "I did X" perspective.
    Focus on outcomes and progress, not tool ratios or technical details.
    
    ### {Theme 1} (AM 09:36~11:30)
    - Core work performed
    - Key decisions
    - 2-4 bullets per theme
    
    ### {Theme 2} (PM 13:33~15:21)
    - Core work performed
    - Key decisions
    
    ### Miscellaneous
    - Brief summary of sub-15-minute tasks
    
    ### Tool Usage Patterns
    - Tool usage ratios and primary purposes
    - Notable tool-switching patterns
    

    Multi-day format (3d, 7d, 2w, 30d)

    For any multi-day window, use a project-driven structure like a sprint report. Focus on what was accomplished per project, not day-by-day chronology.

    ## {start} ~ {end} Monthly Recap
    
    > **TL;DR**
    > - {What I accomplished 1 — project name + outcome}
    > - {What I accomplished 2}
    > - {What I accomplished 3}
    
    ### Overview
    3-5 sentence narrative of the month. Major focus shifts week-by-week,
    key milestones achieved, and overall direction. Written from "I did X" perspective.
    
    ### {Project A}
    What this project is, what was accomplished during the period.
    - Key milestone or deliverable 1
    - Key milestone or deliverable 2
    - Key decision made
    - Current status (shipped / in progress / blocked)
    
    ### {Project B}
    - ...
    
    ### Side Projects
    Projects with <30 prompts, summarized briefly.
    - {project}: one-line summary
    - {project}: one-line summary
    
    ### Tool Usage Patterns
    - Tool usage ratios and how they evolved over the month
    - Notable shifts (e.g., "started using Codex mid-month")
    

    Multi-day grouping rules:

    • Group by project, not by date
    • Order projects by activity volume (most active first)
    • Each project section: what it is, what was accomplished, key decisions, current status
    • Do NOT include prompt counts or date ranges in project headers — those are internal metrics
    • Small projects (<30 prompts) go into "Side Projects" as one-liners
    • Overview should read like a sprint report narrative, not a log

    5. Save Results

    Save to .agents/results/recap/{date}.md. For window ranges, use {start-date}~{end-date}.md format.

    # Example paths
    .agents/results/recap/2026-04-12.md
    .agents/results/recap/2026-04-06~2026-04-12.md
    

    Core Rules

    1. TL;DR required: Top 3 lines of "what I accomplished". Project name + outcome. No tool names or technical details.
    2. Overview: After TL;DR, describe the flow. Start with "I" as subject.
    3. Daily: themes by time block (15+ min). Rest goes to "Miscellaneous".
    4. Multi-day (3d+): sections by project, ordered by activity. Read like a sprint report, not a daily log.
    5. 2-4 bullets per theme/project: Concise essentials only. Don't enumerate every step.
    6. Themes by content: Group by actual work, not by tool.
    7. Time range (daily only): (AM/PM/Evening HH:MM~HH:MM). AM: 12:00, PM: 12:0018:00, Evening: 18:00~.
    8. Save results: Write markdown to .agents/results/recap/.
    9. Response language: Follows language setting in .agents/oma-config.yaml if configured.
    10. No em dashes: Use commas, periods, or parentheses instead of (em dash).

    References

    • Recap CLI: oma recap --json
    • Output directory: .agents/results/recap/
    • Language config: .agents/oma-config.yaml
    • Claude fallback history: ~/.claude/history.jsonl

    Frequently asked questions

    What to verify before installation and use

    What does the oma-recap source document cover?

    Analyze AI tool conversation histories for a given period and generate themed work summaries.

    How do I install oma-recap?

    The source record exposes this install command: npx skills add https://github.com/first-fluke/oh-my-agent --skill "benchmarks/runs/oma/.agents/skills/oma-recap". Inspect the command and pinned source before running it.

    Which Agent platforms does the source record declare?

    The pinned source record declares support for: codex, cursor.

    Alternatives

    Compare before choosing

    Computed 97211

    PramodDutta/qaskills

    RAG Regression Testing

    Gate RAG pipelines in CI with versioned golden eval sets, per-metric thresholds, baseline drift detection, and a build that fails when retrieval or answer quality regresses.

    Computed 9621

    upex-galaxy/agentic-qa-boilerplate

    regression-testing

    Execute regression test suites via CI/CD, analyze results, classify failures, and produce GO/NO-GO release decisions. Use when running regression, smoke, or sanity suites through GitHub Actions, monitoring workflow runs, downloading Allure or Playwright artifacts, classifying failures (REGRESSION vs FLAKY vs KNOWN vs ENVIRONMENT vs NEW TEST), computing pass-rate and trend metrics, deciding release readiness, generating executive quality reports, or creating regression issues. Triggers on: run re

    Computed 94211

    PramodDutta/qaskills

    CI Pipeline Optimizer

    Optimize CI test pipelines through intelligent test splitting, parallelization, caching strategies, and selective test execution based on code changes.

    Computed 9421

    upex-galaxy/agentic-qa-boilerplate

    test-documentation

    Analyze, prioritize, and document test cases in TMS (Jira/Xray), or repair an existing Story-ATS-ATP-ATR-TC cascade through a sealed explicit mode. Use for Test/ATP/ATR artifacts, ROI and automation verdicts, maintaining traceability, fix-traceability, or broken TMS links. The repair-traceability mode audits, plans, waits for explicit approval, applies, and verifies without launching the general documentation workflow. Do NOT use for writing test code (test-automation) or running suites (regress