Source profileQuality 93/100Review permissions

eugenelim/agent-ready-repo/packs/atlassian/.apm/skills/confluence-crawler/SKILL.md

confluence-crawler

Crawl an authenticated Confluence space (Atlassian Cloud or on-prem Server/Data Center) by page hierarchy and convert each page to clean Markdown with frontmatter. Handles macros, attachments, internal link rewriting, depth limits, and idempotent re-crawling. Use when the user wants to mirror, export, or ingest Confluence content.

Source repository stars
17
Declared platforms
0
Static risk flags
2
Last source update
2026-08-28
Source checked
2026-08-28

Decision brief

What it does: where it fits

Crawl a Confluence space (Cloud or Server/Data Center) and write each page as Markdown with YAML frontmatter.

Best for

  • Use when the user wants to mirror, export, or ingest Confluence content.

Not for

  • Tasks that require unconfirmed production actions or broad system permissions.
  • Environments where the pinned source and install steps cannot be inspected.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/eugenelim/agent-ready-repo --skill "packs/atlassian/.apm/skills/confluence-crawler"
Safe inspection promptEditorial

Inspect the Agent Skill "confluence-crawler" from https://github.com/eugenelim/agent-ready-repo/blob/12b2c9f36c800761156a1daa149949af4d84986f/packs/atlassian/.apm/skills/confluence-crawler/SKILL.md at commit 12b2c9f36c800761156a1daa149949af4d84986f. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Instructions

    You are a Confluence export agent. The heavy lifting — authentication, REST pagination, macro conversion, link rewriting, idempotency — lives in scripts/. Do not re-implement any of that logic; just invoke the scripts with the right arguments and report the result.

    Atlassian Cloud (.atlassian.net) — Basic auth with email + API token from id.atlassian.com. Base URL must include /wiki (setup adds it automatically).Confluence Server / Data Center — Bearer auth with a Personal Access Token from the user's Confluence profile.Secrets live only in /.agentbundle/credentials.env
  2. 02

    Step 1: Verify the environment

    Check Python dependencies are installed. If not, install them:

    Exit code 0 → authenticated, proceed.Exit code 2 → read the bounded error. On the token path, missing or invalidAny other non-zero → see When a request fails.
  3. 03

    Step 2: Crawl the space

    Invoke the crawler with the user's arguments. Only these flags are supported:

    Invoke the crawler with the user's arguments. Only these flags are supported:
  4. 04

    Step 3: Interpret the output

    The final log line reports wrote N pages (failed: X, skipped: Y). Relay this to the user. If any pages failed, check the log for which IDs — usually permission issues on specific pages.

    /.md per page, flat layout. Each file starts with YAML frontmatter carrying confluenceid, version, spacekey, updated, author, parentid, labels, url, slug./attachments// for downloaded attachments.- /.md per page, flat layout. Each file starts with YAML frontmatter carrying confluenceid, version, spacekey, updated, author, parentid, labels, url, slug. - /attachments// for downloaded attachments.
  5. 05

    Step 4: Re-crawling

    The script is idempotent. On re-run:

    It compares each page's current version.number against the version field in the existing .md frontmatter.Unchanged pages are skipped.Changed pages are re-fetched and overwritten.

Permission review

Static risk signals and limitations

Reads files

low · line 73

The documentation asks the agent to read local files, directories, or repositories.

*Never** read that file, print it, or echo the token.

Reads files

low · line 100

The documentation asks the agent to read local files, directories, or repositories.

the bytes. **Never** read the jar file directly, print its contents, or echo

Runs scripts

medium · line 117

The documentation asks the agent to run terminal commands or scripts.

python -m pip install -r requirements.txt

Runs scripts

medium · line 123

The documentation asks the agent to run terminal commands or scripts.

python '<skill-dir>/scripts/crawl_space.py' --check

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score93/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars17SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
eugenelim/agent-ready-repo
Skill path
packs/atlassian/.apm/skills/confluence-crawler/SKILL.md
Commit
12b2c9f36c800761156a1daa149949af4d84986f
License
Apache-2.0
Collected
2026-08-28
Default branch
main
View the original SKILL.md

Confluence Crawler

Crawl a Confluence space (Cloud or Server/Data Center) and write each page as Markdown with YAML frontmatter.

Output rendering

Key–value / one record — For a single record's fields, use an aligned key: value list, not a two-row table.

Installed entry-point contract

Treat <skill-dir> as the installer-supplied directory containing this active SKILL.md; never infer it from the current working directory, user input, an environment variable, or a profile path. Replace <skill-dir> with that actual validated directory before executing or relaying any command; never send the placeholder to a runtime or user. Before every invocation of crawl_space.py or setup_sso.py:

  1. Canonicalize <skill-dir>, its scripts/ child, and the expected entry point, resolving symlinks. Require the entry point to be a regular file and its resolved path to remain beneath the canonical scripts/ directory.
  2. If the entry is missing, is not a regular file, encounters a symlink loop or resolution error, or escapes that directory, stop before launching Python. Report only error: installed skill entry point is unavailable: <entry>, substituting the basename. Do not expose an absolute, home, profile, environment, or protected path; do not relay raw runtime stderr; and do not offer credential, SSO-capture, token, scope, or dependency remediation.
  3. Invoke with a discrete argument vector, for example ["<python>", "<skill-dir>/scripts/crawl_space.py", "..."], so spaces, both quote characters, $(), backticks, and variable-shaped text cannot be expanded by a shell. Keep the project root as the working directory so user content paths retain their documented meaning.
  4. If only a shell string is available, use a single-quoted literal path on POSIX or PowerShell and refuse paths containing a single quote. On cmd.exe, use a double-quoted path and refuse paths containing ", %, or !. If the adapter cannot represent the path safely, refuse instead of invoking.

Interpret exit codes only after this preflight succeeds and the entry point actually runs.

Instructions

You are a Confluence export agent. The heavy lifting — authentication, REST pagination, macro conversion, link rewriting, idempotency — lives in scripts/. Do not re-implement any of that logic; just invoke the scripts with the right arguments and report the result.

Flavor support

The skill works against both:

  • Atlassian Cloud (*.atlassian.net) — Basic auth with email + API token from id.atlassian.com. Base URL must include /wiki (setup adds it automatically).
  • Confluence Server / Data Center — Bearer auth with a Personal Access Token from the user's Confluence profile.

Flavor is auto-detected from the base URL. Override via CONFLUENCE_FLAVOR=cloud|server if needed.

Configuration location

Credentials are resolved by the build-projected credentials_shim.load_credentials through Tier 1 (env) → Tier 2 (OS keyring) → Tier 3 dotfile. The dotfile lives at ~/.agentbundle/credentials.env. The declared schema is in references/creds-schema.toml:

KeyRequiredNotes
CONFLUENCE_BASE_URLyesCloud: https://<site>.atlassian.net/wiki. Server: https://confluence.corp.example.com.
CONFLUENCE_API_TOKENyesCloud API token or Server PAT.
CONFLUENCE_EMAILCloud onlyAtlassian account email.
CONFLUENCE_FLAVORnocloud or server. Auto-detected from URL host when unset.

Populate any tier by running credential-setup skill.

Security rules (non-negotiable)

  • Secrets live only in ~/.agentbundle/credentials.env (mode 0600 on POSIX; DACL-restricted on Windows), the OS keyring, or process environment variables. Never read that file, print it, or echo the token.
  • Never put the token on the command line. The primitive refuses flags like --token / --api-token / --bearer / --pat / --password and exits — do not work around it.
  • On the token path, if --check reports missing or invalid credentials, tell the user to run credential-setup themselves. It is interactive — do not run it for them. A 403 is a permission failure, not a setup trigger; surface it without starting credential setup.
  • CONFLUENCE_BASE_URL is user-configured. Before invoking the crawler, verify the configured URL resolves to a known Confluence host (e.g. *.atlassian.net for Cloud, the organisation's known on-premises host for Server) — not to a private IP range (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8) or a cloud-metadata endpoint (169.254.0.0/16). If the user supplies an unexpected host, stop and ask them to confirm before running. This is an agent pre-flight check: the scripts validate only the URL scheme (http:// or https://), not the resolved host or IP range. On the token path follow_redirects=True is active, so verify the initial host before invoking.

This skill is dual-auth (auth: sso-cookie with a creds fallback): on a Data Center instance behind corporate SSO it authenticates by a captured web session (cookie jar) resolved through the sso-broker; everywhere else it uses the token (creds) path above. On the SSO-cookie path:

  • The session cookie jar lives only under the broker's 0600 store; the skill reads it in-process via the credbroker resolver, which returns a path, not the bytes. Never read the jar file directly, print its contents, or echo cookie values.
  • Never put a session cookie on the command line. The skill attaches cookies to its HTTP client internally and sends no Authorization header on this path.
  • Run --check first and allow its single headless recovery attempt. The automatic attempt shows no browser window and obtains its sign-in destination only from CredBroker's registered profile. It never uses login_url from sso-config.toml as an automatic destination.
  • Request manual setup with python '<skill-dir>/scripts/setup_sso.py' only when --check says automatic recovery refused or failed. That helper opens a browser for interactive sign-in, so do not run any setup helper for them.

Step 1: Verify the environment

Check Python dependencies are installed. If not, install them:

python -m pip install -r requirements.txt

Then verify connectivity:

python '<skill-dir>/scripts/crawl_space.py' --check
  • Exit code 0 → authenticated, proceed.
  • Exit code 2 → read the bounded error. On the token path, missing or invalid credentials require user-run credential-setup. On the SSO path, request user-run python '<skill-dir>/scripts/setup_sso.py' only when the message says the single headless recovery refused or failed. A 403, malformed configuration, confinement failure, or dependency problem is terminal for this attempt and must be surfaced as written; do not start setup for it. Stop here.
  • Any other non-zero → see When a request fails.

When a request fails

The CLI uses a banded exit-code contract; read the stderr message for the specific cause, then act on the band:

ExitBandWhat to do
0successproceed
1functional error — bad/missing args, server 5xx, transport, a partial crawl (some pages failed), keychain hard-fail, unexpectedsurface the message; for a partial crawl the per-page failures are in the log — report them, don't loop
2user must act — token credentials, SSO recovery refusal/failure, permission/configuration, or dependency problemfollow the bounded message: request the matching manual setup only for missing token credentials or an explicit SSO recovery refusal/failure; surface 403/configuration/confinement/dependency errors without setup, then re-run --check only after the user resolves the named cause
130interrupted (Ctrl-C)the run was cancelled; nothing to fix

Tier2HardFailError (OS keyring unavailable) or an unprojected shim surface as exit 1 with a message naming the cause.

Step 2: Crawl the space

Invoke the crawler with the user's arguments. Only these flags are supported:

FlagMeaning
--space KEYSpace key, e.g. ENG. Required.
--root PAGE_IDStart from a specific page (default: space homepage).
--depth NMax hierarchy depth from root (default: unlimited).
--output DIROutput directory (default: ./confluence-out).
--forceRe-fetch and overwrite all pages, ignoring frontmatter version.
--no-attachmentsSkip attachment downloads.
--concurrency NParallel requests (default: 4).
--min-delay-ms NMinimum ms between requests (default: 100).
--insecureDisable TLS verification. Only if the user explicitly asks.
--verboseDebug logging.

Example:

python '<skill-dir>/scripts/crawl_space.py' --space ENG --depth 3 --output ./out

Step 3: Interpret the output

The script writes:

  • <output>/<slug>.md per page, flat layout. Each file starts with YAML frontmatter carrying confluence_id, version, space_key, updated, author, parent_id, labels, url, slug.
  • <output>/attachments/<page_id>/<filename> for downloaded attachments.

The final log line reports wrote N pages (failed: X, skipped: Y). Relay this to the user. If any pages failed, check the log for which IDs — usually permission issues on specific pages.

Step 4: Re-crawling

The script is idempotent. On re-run:

  • It compares each page's current version.number against the version field in the existing .md frontmatter.
  • Unchanged pages are skipped.
  • Changed pages are re-fetched and overwritten.
  • Pass --force to bypass the version check and re-fetch everything.

Behavior notes

  • Depth is measured in page hierarchy (parent → child), not link hops.
  • Macros in an allowlist (code, info, warning, note, tip, panel, expand, status) are converted to Markdown equivalents. Others are replaced with a visible *[confluence macro not rendered: NAME]* italic marker so reviewers can spot gaps.
  • Internal links to pages that were also crawled become relative .md paths. Links to pages outside the crawl set remain absolute Confluence URLs.
  • Attachments are downloaded alongside the referencing page and linked via relative paths.

Don't

  • Don't read ~/.agentbundle/credentials.env from skill body.
  • Don't print or log the PAT.
  • Don't run credential-setup skill non-interactively or pipe the PAT into it.
  • Don't write your own REST calls to Confluence — extend the scripts instead, and surface the gap to the user if a flag is missing.
  • Don't assume --insecure is safe to add by default. Only when the user explicitly says they accept it.

Edge cases

  • Cloud base URL without /wiki: if the user's config somehow has https://foo.atlassian.net without /wiki, API calls will 404. The setup script appends it automatically; if the user hand-edited the config, have them re-run setup.
  • Space has no homepage: the script exits 2 and asks for --root PAGE_ID. Relay this to the user.
  • Orphaned pages not in the hierarchy: not crawled by design. If the user wants them, they need to pass --root for each, or request a future "full-space" mode.
  • Very large spaces: discovery does a full hierarchy walk first (one listing call per page). Expect a minute or two for thousands of pages. Fetch and convert then runs with bounded concurrency.
  • Title changes between runs: the old <old-slug>.md file remains on disk — the new run writes <new-slug>.md because slugs derive from the current title. Warn the user that old files may linger and let them clean up.
  • Network failures mid-crawl: the .part tempfile pattern prevents half-written .md files. Re-running resumes cleanly.

Frequently asked questions

What to verify before installation and use

What does the confluence-crawler source document cover?

Crawl a Confluence space (Cloud or Server/Data Center) and write each page as Markdown with YAML frontmatter.

How do I install confluence-crawler?

The source record exposes this install command: npx skills add https://github.com/eugenelim/agent-ready-repo --skill "packs/atlassian/.apm/skills/confluence-crawler". Inspect the command and pinned source before running it.

Which permission-related actions were detected?

Static rules flagged read-files, exec-script in the source; the page lists the matching lines and excerpts.