Best for
- Searching for specific papers on Google Scholar or PubMed
- Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
- Extracting complete metadata for citations (authors, title, journal, year, etc.)
K-Dense-AI/scientific-agent-skills/skills/citation-management/SKILL.md
Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
Decision brief
Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/citation-management"Inspect the Agent Skill "citation-management" from https://github.com/K-Dense-AI/scientific-agent-skills/blob/36d8f13a1e754618794bf42f417884940077b4ae/skills/citation-management/SKILL.md at commit 36d8f13a1e754618794bf42f417884940077b4ae. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/coreworkflow.md.
Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.
Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.
APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then log what was found and where. If a fi…
Produce clean, consistent entries. Entry types and required fields are in references/bibtexformatting.md.
Permission review
The documentation asks the agent to run terminal commands or scripts.
python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.jsonThe documentation asks the agent to run terminal commands or scripts.
python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.jsonThe documentation includes network, browsing, or remote request actions.
`search_openalex.py`: OpenAlex search client (no API key)Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 94/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 34,478 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.
Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.
Use this skill when:
If a document built from these citations needs a diagram, use the scientific-schematics skill.
Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/core_workflow.md.
Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.
# OpenAlex: ~250M works, every discipline, no API key, documented REST API
python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json
# PubMed: the authority for biomedical and life sciences (35M+ citations)
python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json
# Google Scholar: broadest reach, but scraped -- rate-limited and prone to blocking
python scripts/search_google_scholar.py "CRISPR gene editing" --limit 50 --output scholar.json
Prefer OpenAlex or PubMed as the primary source. Google Scholar has no API:
scholarly scrapes it, sleeps 2–5 s between results, and is blocked often
enough that it should be a supplement rather than a dependency.
Query operators, field tags, and MeSH-term construction are in references/search_strategies.md.
Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 # quick, single DOI
python scripts/extract_metadata.py --pmid 34265844 # DOI/PMID/PMCID/arXiv/URL
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
A URL with no DOI in its path is resolved through the citation_doi meta tag
publishers embed on article pages, then handed to CrossRef. Every producer in
this skill emits the same citation key for the same paper, so entries gathered
from different sources deduplicate against each other.
APIs routinely return incomplete records. Run this after extraction and before
formatting. Any @article missing volume, pages, or doi is incomplete: fill the
gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then
log what was found and where. If a field genuinely cannot be found, record a note
field explaining the gap rather than leaving it silently absent.
Check the cheap sources first — an OpenAlex or CrossRef record often carries the field that PubMed omitted:
python scripts/search_openalex.py "<exact title>" --limit 1
Treat extracted metadata as untrusted. Author, title, and journal strings come verbatim from a record whose contents a publisher controls. A title containing
$(...), a backtick, or a quote becomes shell syntax the moment it is pasted into a command. Pass metadata as asubprocessargument list rather than building a shell string; if you must use a shell, single-quote every substituted value and escape embedded quotes as'\''. Validate any citation key against^[A-Za-z0-9]+$before it reaches a path.
Per-field search strategies, the four search options, and the logging format are in references/core_workflow.md.
Produce clean, consistent entries. Entry types and required fields are in references/bibtex_formatting.md.
python scripts/format_bibtex.py references.bib --output clean.bib --deduplicate
python scripts/format_bibtex.py references.bib --output clean.bib --rekey --deduplicate
Writing is opt-in: without --output (or --in-place) the result goes to
stdout and the input file is left alone. Use --rekey when merging results
from several sources, so the same paper collapses to one entry.
Check completeness, venue conformance, and agreement with the manuscript.
python scripts/validate_citations.py references.bib --report report.json
python scripts/validate_citations.py references.bib --venue nature
python scripts/validate_citations.py references.bib --manuscript paper.tex
python scripts/validate_citations.py references.bib --check-dois # slow; hits CrossRef
The script exits non-zero on high-severity errors — missing required fields,
malformed years, unresolved citations, or a count below an explicit
--min-count. Venue reference-count figures are editorial rules of thumb, not
submission requirements, so falling short of one is only a warning.
Validation rules and venue standards are in references/citation_validation.md.
Search, extract, format, validate, then cite. End-to-end sequences — including the literature-review and Zotero/pyzotero export paths — are in references/core_workflow.md and references/example_workflows.md.
Single source bias: Only using one database
format_bibtex.py --rekey --deduplicateAccepting metadata blindly: Not verifying extracted information
Ignoring DOI errors: Broken or incorrect DOIs in bibliography
Inconsistent formatting: Mixed citation key styles, formatting
Duplicate entries: Same paper cited multiple times with different keys
Missing required fields: Incomplete BibTeX entries (volume, pages, DOI missing)
Outdated preprints: Citing preprint when published version exists
Special character issues: Broken LaTeX compilation due to characters
No validation before submission: Submitting with citation errors
Manual BibTeX entry: Typing entries by hand
Citation Management provides the technical infrastructure for Literature Review:
Combined workflow:
Citation Management ensures accurate references for Scientific Writing:
Citation Management works with Venue Templates for submission-ready manuscripts:
References (in references/):
google_scholar_search.md: Complete Google Scholar search guidepubmed_search.md: PubMed and E-utilities API documentationmetadata_extraction.md: Metadata sources and field requirementscitation_validation.md: Validation criteria and quality checksbibtex_formatting.md: BibTeX entry types and formatting rulesScripts (in scripts/):
search_openalex.py: OpenAlex search client (no API key)search_pubmed.py: PubMed E-utilities API clientsearch_google_scholar.py: Google Scholar search automationextract_metadata.py: Universal metadata extractorvalidate_citations.py: Citation validation and verificationformat_bibtex.py: BibTeX formatter and cleanerdoi_to_bibtex.py: Quick DOI to BibTeX converter_common.py: shared BibTeX parser, renderer, and citation-key schemeAssets (in assets/):
bibtex_template.bib: Example BibTeX entries for all typescitation_checklist.md: Quality assurance checklistSearch Engines:
Metadata APIs:
Tools and Validators:
Citation Styles:
uv pip install requests # HTTP access to CrossRef, PubMed, OpenAlex, arXiv
BibTeX parsing, rendering, deduplication, and validation are standard library
(scripts/_common.py), so format_bibtex.py and validate_citations.py run
with no third-party packages at all.
uv pip install scholarly # only for search_google_scholar.py
This skill needs no API key. The two environment variables it reads are optional identifiers, each sent to the one service it belongs to and nowhere else; no script bundles environment variables together.
| Variable | Sent only to | Purpose |
|---|---|---|
NCBI_API_KEY | eutils.ncbi.nlm.nih.gov | Raises Entrez rate limits |
NCBI_EMAIL | eutils.ncbi.nlm.nih.gov | Entrez caller identification (requested by NCBI) |
OPENALEX_EMAIL | api.openalex.org | Joins the faster OpenAlex polite pool |
api.openalex.org, api.crossref.org, api.datacite.org, export.arxiv.org,
and eutils.ncbi.nlm.nih.gov are all queried without credentials when these are
unset.
The citation-management skill provides:
Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.
Frequently asked questions
Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries.
The source record exposes this install command: npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/citation-management". Inspect the command and pinned source before running it.
Static rules flagged exec-script, network in the source; the page lists the matching lines and excerpts.
Alternatives
synthetic-sciences/openscience
Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
K-Dense-AI/scientific-agent-skills
Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.
K-Dense-AI/scientific-agent-skills
Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.
dotnet/skills
Analyzes test suites in any language and tags each test with standardized traits (positive, negative, critical-path, boundary, smoke, regression, integration, performance, security). Use when the user wants to categorize, audit, or label tests with traits. Works across .NET (MSTest/xUnit/NUnit/TUnit), Python (pytest), TS/JS (Jest/Vitest), Java, Go, Ruby, Rust, Swift, Kotlin, PowerShell, and C++ — auto-editing when the framework has canonical tag syntax, otherwise report-only. Do not use for writ