Source profileQuality 94/100Review permissions

K-Dense-AI/scientific-agent-skills/skills/citation-management/SKILL.md

citation-management

Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.

Source repository stars
34,478
Declared platforms
0
Static risk flags
2
Last source update
2026-08-24
Source checked
2026-08-26

Decision brief

What it does: where it fits

Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries.

Best for

  • Searching for specific papers on Google Scholar or PubMed
  • Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
  • Extracting complete metadata for citations (authors, title, journal, year, etc.)

Not for

  • Single source bias: Only using one database
  • Solution: Search at least OpenAlex and PubMed, then merge with

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/citation-management"
Safe inspection promptEditorial

Inspect the Agent Skill "citation-management" from https://github.com/K-Dense-AI/scientific-agent-skills/blob/36d8f13a1e754618794bf42f417884940077b4ae/skills/citation-management/SKILL.md at commit 36d8f13a1e754618794bf42f417884940077b4ae. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Core Workflow

    Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/coreworkflow.md.

    Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/coreworkflow.md.Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.
  2. 02

    Phase 1: Paper Discovery and Search

    Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.

    Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.
  3. 03

    Phase 2: Metadata Extraction

    Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.

    Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.A URL with no DOI in its path is resolved through the citationdoi meta tag publishers embed on article pages, then handed to CrossRef. Every producer in this skill emits the same citation key for the same paper, so entr…
  4. 04

    Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)

    APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then log what was found and where. If a fi…

    APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when…Check the cheap sources first — an OpenAlex or CrossRef record often carries the field that PubMed omitted:Treat extracted metadata as untrusted. Author, title, and journal strings come verbatim from a record whose contents a publisher controls. A title containing $(...), a backtick, or a quote becomes shell syntax the momen…
  5. 05

    Phase 3: BibTeX Formatting

    Produce clean, consistent entries. Entry types and required fields are in references/bibtexformatting.md.

    Produce clean, consistent entries. Entry types and required fields are in references/bibtexformatting.md.Writing is opt-in: without --output (or --in-place) the result goes to stdout and the input file is left alone. Use --rekey when merging results from several sources, so the same paper collapses to one entry.

Permission review

Static risk signals and limitations

Runs scripts

medium · line 41

The documentation asks the agent to run terminal commands or scripts.

python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json

Runs scripts

medium · line 44

The documentation asks the agent to run terminal commands or scripts.

python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json

Network access

medium · line 224

The documentation includes network, browsing, or remote request actions.

`search_openalex.py`: OpenAlex search client (no API key)

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score94/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars34,478SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
K-Dense-AI/scientific-agent-skills
Skill path
skills/citation-management/SKILL.md
Commit
36d8f13a1e754618794bf42f417884940077b4ae
License
MIT
Collected
2026-08-26
Default branch
main
View the original SKILL.md

Citation Management

Overview

Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.

Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.

When to Use This Skill

Use this skill when:

  • Searching for specific papers on Google Scholar or PubMed
  • Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
  • Extracting complete metadata for citations (authors, title, journal, year, etc.)
  • Validating existing citations for accuracy
  • Cleaning and formatting BibTeX files
  • Finding highly cited papers in a specific field
  • Verifying that citation information matches the actual publication
  • Building a bibliography for a manuscript or thesis
  • Checking for duplicate citations
  • Ensuring consistent citation formatting

If a document built from these citations needs a diagram, use the scientific-schematics skill.


Core Workflow

Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/core_workflow.md.

Phase 1: Paper Discovery and Search

Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.

# OpenAlex: ~250M works, every discipline, no API key, documented REST API
python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json

# PubMed: the authority for biomedical and life sciences (35M+ citations)
python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json

# Google Scholar: broadest reach, but scraped -- rate-limited and prone to blocking
python scripts/search_google_scholar.py "CRISPR gene editing" --limit 50 --output scholar.json

Prefer OpenAlex or PubMed as the primary source. Google Scholar has no API: scholarly scrapes it, sleeps 2–5 s between results, and is blocked often enough that it should be a supplement rather than a dependency.

Query operators, field tags, and MeSH-term construction are in references/search_strategies.md.

Phase 2: Metadata Extraction

Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.

python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2         # quick, single DOI
python scripts/extract_metadata.py --pmid 34265844                  # DOI/PMID/PMCID/arXiv/URL
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib

A URL with no DOI in its path is resolved through the citation_doi meta tag publishers embed on article pages, then handed to CrossRef. Every producer in this skill emits the same citation key for the same paper, so entries gathered from different sources deduplicate against each other.

Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)

APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then log what was found and where. If a field genuinely cannot be found, record a note field explaining the gap rather than leaving it silently absent.

Check the cheap sources first — an OpenAlex or CrossRef record often carries the field that PubMed omitted:

python scripts/search_openalex.py "<exact title>" --limit 1

Treat extracted metadata as untrusted. Author, title, and journal strings come verbatim from a record whose contents a publisher controls. A title containing $(...), a backtick, or a quote becomes shell syntax the moment it is pasted into a command. Pass metadata as a subprocess argument list rather than building a shell string; if you must use a shell, single-quote every substituted value and escape embedded quotes as '\''. Validate any citation key against ^[A-Za-z0-9]+$ before it reaches a path.

Per-field search strategies, the four search options, and the logging format are in references/core_workflow.md.

Phase 3: BibTeX Formatting

Produce clean, consistent entries. Entry types and required fields are in references/bibtex_formatting.md.

python scripts/format_bibtex.py references.bib --output clean.bib --deduplicate
python scripts/format_bibtex.py references.bib --output clean.bib --rekey --deduplicate

Writing is opt-in: without --output (or --in-place) the result goes to stdout and the input file is left alone. Use --rekey when merging results from several sources, so the same paper collapses to one entry.

Phase 4: Citation Validation

Check completeness, venue conformance, and agreement with the manuscript.

python scripts/validate_citations.py references.bib --report report.json
python scripts/validate_citations.py references.bib --venue nature
python scripts/validate_citations.py references.bib --manuscript paper.tex
python scripts/validate_citations.py references.bib --check-dois     # slow; hits CrossRef

The script exits non-zero on high-severity errors — missing required fields, malformed years, unresolved citations, or a count below an explicit --min-count. Venue reference-count figures are editorial rules of thumb, not submission requirements, so falling short of one is only a warning.

Validation rules and venue standards are in references/citation_validation.md.

Phase 5: Integration with Writing Workflow

Search, extract, format, validate, then cite. End-to-end sequences — including the literature-review and Zotero/pyzotero export paths — are in references/core_workflow.md and references/example_workflows.md.

Reference Files

Common Pitfalls to Avoid

  1. Single source bias: Only using one database

    • Solution: Search at least OpenAlex and PubMed, then merge with format_bibtex.py --rekey --deduplicate
  2. Accepting metadata blindly: Not verifying extracted information

    • Solution: Spot-check extracted metadata against original sources
  3. Ignoring DOI errors: Broken or incorrect DOIs in bibliography

    • Solution: Run validation before final submission
  4. Inconsistent formatting: Mixed citation key styles, formatting

    • Solution: Use format_bibtex.py to standardize
  5. Duplicate entries: Same paper cited multiple times with different keys

    • Solution: Use duplicate detection in validation
  6. Missing required fields: Incomplete BibTeX entries (volume, pages, DOI missing)

    • Solution: Run Phase 2.5 metadata enrichment — web search for every missing field before proceeding. NEVER leave an @article entry without volume, pages, and DOI.
  7. Outdated preprints: Citing preprint when published version exists

    • Solution: Check if preprints have been published, update to journal version
  8. Special character issues: Broken LaTeX compilation due to characters

    • Solution: Use proper escaping or Unicode in BibTeX
  9. No validation before submission: Submitting with citation errors

    • Solution: Always run validation as final check
  10. Manual BibTeX entry: Typing entries by hand

    • Solution: Always extract from metadata sources using scripts

Integration with Other Skills

Literature Review Skill

Citation Management provides the technical infrastructure for Literature Review:

  • Literature Review: Multi-database systematic search and synthesis
  • Citation Management: Metadata extraction and validation

Combined workflow:

  1. Use literature-review for systematic search methodology
  2. Use citation-management to extract and validate citations
  3. Use literature-review to synthesize findings
  4. Use citation-management to ensure bibliography accuracy

Scientific Writing Skill

Citation Management ensures accurate references for Scientific Writing:

  • Export validated BibTeX for use in LaTeX manuscripts
  • Verify citations match publication standards
  • Format references according to journal requirements

Venue Templates Skill

Citation Management works with Venue Templates for submission-ready manuscripts:

  • Different venues require different citation styles
  • Generate properly formatted references
  • Validate citations meet venue requirements

Resources

Bundled Resources

References (in references/):

  • google_scholar_search.md: Complete Google Scholar search guide
  • pubmed_search.md: PubMed and E-utilities API documentation
  • metadata_extraction.md: Metadata sources and field requirements
  • citation_validation.md: Validation criteria and quality checks
  • bibtex_formatting.md: BibTeX entry types and formatting rules

Scripts (in scripts/):

  • search_openalex.py: OpenAlex search client (no API key)
  • search_pubmed.py: PubMed E-utilities API client
  • search_google_scholar.py: Google Scholar search automation
  • extract_metadata.py: Universal metadata extractor
  • validate_citations.py: Citation validation and verification
  • format_bibtex.py: BibTeX formatter and cleaner
  • doi_to_bibtex.py: Quick DOI to BibTeX converter
  • _common.py: shared BibTeX parser, renderer, and citation-key scheme

Assets (in assets/):

  • bibtex_template.bib: Example BibTeX entries for all types
  • citation_checklist.md: Quality assurance checklist

External Resources

Search Engines:

Metadata APIs:

Tools and Validators:

Citation Styles:

Dependencies

Required Python Packages

uv pip install requests  # HTTP access to CrossRef, PubMed, OpenAlex, arXiv

BibTeX parsing, rendering, deduplication, and validation are standard library (scripts/_common.py), so format_bibtex.py and validate_citations.py run with no third-party packages at all.

Optional

uv pip install scholarly  # only for search_google_scholar.py

Where credentials are sent

This skill needs no API key. The two environment variables it reads are optional identifiers, each sent to the one service it belongs to and nowhere else; no script bundles environment variables together.

VariableSent only toPurpose
NCBI_API_KEYeutils.ncbi.nlm.nih.govRaises Entrez rate limits
NCBI_EMAILeutils.ncbi.nlm.nih.govEntrez caller identification (requested by NCBI)
OPENALEX_EMAILapi.openalex.orgJoins the faster OpenAlex polite pool

api.openalex.org, api.crossref.org, api.datacite.org, export.arxiv.org, and eutils.ncbi.nlm.nih.gov are all queried without credentials when these are unset.

Summary

The citation-management skill provides:

  1. Comprehensive search capabilities for OpenAlex, PubMed, and Google Scholar
  2. Automated metadata extraction from DOI, PMID, PMCID, arXiv ID, URLs
  3. Citation validation with DOI verification and completeness checking
  4. BibTeX formatting with standardization and cleaning tools
  5. Quality assurance through validation and reporting
  6. Integration with scientific writing workflow
  7. Reproducibility through documented search and extraction methods

Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.

Frequently asked questions

What to verify before installation and use

What does the citation-management source document cover?

Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries.

How do I install citation-management?

The source record exposes this install command: npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/citation-management". Inspect the command and pinned source before running it.

Which permission-related actions were detected?

Static rules flagged exec-script, network in the source; the page lists the matching lines and excerpts.

Alternatives

Compare before choosing

Computed 923,338

synthetic-sciences/openscience

citation-management

Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.

Computed 9834,478

K-Dense-AI/scientific-agent-skills

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

Computed 9834,478

K-Dense-AI/scientific-agent-skills

neurokit2

Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.

Computed 975,248

dotnet/skills

test-tagging

Analyzes test suites in any language and tags each test with standardized traits (positive, negative, critical-path, boundary, smoke, regression, integration, performance, security). Use when the user wants to categorize, audit, or label tests with traits. Works across .NET (MSTest/xUnit/NUnit/TUnit), Python (pytest), TS/JS (Jest/Vitest), Java, Go, Ruby, Rust, Swift, Kotlin, PowerShell, and C++ — auto-editing when the framework has canonical tag syntax, otherwise report-only. Do not use for writ