Best for
- Searching for Protein Function: Querying functional annotations, GO
- Searching for Protein Sequence: Searching for protein sequences by their
- Understanding Protein/Organism Relationships: Leveraging the Taxonomy
google-deepmind/science-skills/skills/uniprot_database/SKILL.md
Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or retrieving functional annotations and publications. Don't use for sequence alignment, protein folding, or sequence similarity search (use specialized skills for those tasks).
Decision brief
Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/google-deepmind/science-skills --skill "skills/uniprot_database"Inspect the Agent Skill "uniprot-database" from https://github.com/google-deepmind/science-skills/blob/0b42509800f49e6eb7809505d96e20a890ef99bd/skills/uniprot_database/SKILL.md at commit 0b42509800f49e6eb7809505d96e20a890ef99bd. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Copy this checklist and track progress:
1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH. 2. User Notification: If .licenses/uniprotdatabaseLICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https…
Use the Wrapper: Always use the provided Python scripts (e.g.,
Searching for Protein Function: Querying functional annotations, GO
Choose the right tool based on the task type and data volume:
Permission review
The documentation asks the agent to create, modify, or delete local files.
https://www.uniprot.org/help/api_queries, then (2) create the file recordingThe documentation includes network, browsing, or remote request actions.
acid sequence. The REST API `/search` endpoint does not support directThe documentation includes network, browsing, or remote request actions.
PREFIX up: <http://purl.uniprot.org/core/>Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 90/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 2,765 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
uv: Read the uv skill and follow its Setup instructions to ensure
uv is installed and on PATH.Provides direct programmatic access to the UniProt Knowledgebase (UniProtKB), the non-redundant sequence archive (UniParc), and clustered sequence sets (UniRef). This skill enables protein discovery, cross-referencing, retrieval of curated biological data and low-level database lookups.
scripts/uniprot_tools.py) rather than constructing custom curl requests.Choose the right tool based on the task type and data volume:
get: Retrieves metadata and sequence for a specific entry. Best for a
single, known accession.
--dataset unisave), which
is essential for reconciling data from older releases or identifying why
a formerly valid accession no longer appears in search results.search: Searches for entries matching a query. Best for exploration
and discovery.
--limit 5 to verify if a query returns the expected proteins
before committing to a larger download.--limit as it applies to lines, not entries.stream: Streams all matching entries. Best for bulk retrieval of
large datasets (up to 10,000,000 entries).
--limit; always returns the full result set.search with --limit if you need a subset.count: Counts entries matching a query. Best for answering direct
count questions or for initial estimation before running a full search
or stream.sparql: Executes graph queries for complex discovery. Best for
counting, exact sequence matches, and multi-database queries.
map: Converts IDs between UniProt and 100+ databases. Best for ID
mapping tasks.
search vs. map: Try search first before resorting to map if
not explicitly requested by the user. E.g., an external ID might be
searchable in UniParc but fail to map to UniProtKB.Copy this checklist and track progress:
reviewed:true).If a direct query (e.g., gene:SYMBOL) fails:
protein_name:Alpha-crystallin A).[!IMPORTANT] Always prefer
streamorsparqlfor bulk data.searchis suitable for exploration; if results exceed 500 entries, it automatically paginates to provide a stable download.
count: ALWAYS check the result count before running a
search or stream.stream: The primary method for bulk data retrieval (up to
10M entries). Does NOT support --limit; always returns all results.sparql: Best for complex filtering and exact matching
during retrieval.[!IMPORTANT] Use SPARQL when searching for a protein by its full amino acid sequence. The REST API
/searchendpoint does not support direct sequence-string lookups. For any non-exact match use specialized sequence similarity search skills. Use UniParc if you cannot find query in UniProt.
SPARQL Query Pattern (UniProt):
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
SELECT ?protein ?name WHERE {
?protein a up:Protein ;
up:sequence/rdf:value "SEQUENCE_HERE" .
OPTIONAL {
?protein up:recommendedName/up:fullName ?name .
}
}
SPARQL Query Pattern (UniParc):
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
SELECT ?uniparc ?val WHERE {
GRAPH <http://sparql.uniprot.org/uniparc> {
?uniparc a up:Sequence ;
rdf:value ?val .
FILTER (?val = "SEQUENCE_HERE")
}
}
[!IMPORTANT] Use
countorSPARQLfor counting entries (e.g., "How many proteins in Human?").
Counting Pattern (Proteins per Organism):
PREFIX up: <http://purl.uniprot.org/core/>
PREFIX taxon: <http://purl.uniprot.org/taxonomy/>
SELECT (COUNT(?protein) AS ?count) WHERE {
?protein a up:Protein ;
up:reviewed true ;
up:organism taxon:9606 .
}
OR
to separate items.
accession:(P12345 OR P67890)accession:P12345 OR accession:P67890gene:p53 human searches for both.Below are example commands for each mode of uniprot_tools.py.
Count total number of entries for a given query.
uv run scripts/uniprot_tools.py count "taxonomy_id:9606"
Search for entries.
uv run scripts/uniprot_tools.py search "gene:p53 AND reviewed:true" --limit 5
Retrieve a single entry by accession.
uv run scripts/uniprot_tools.py get P04637
Retrieve Historical/Deleted Entry (UniSave).
uv run scripts/uniprot_tools.py get P04637 --dataset unisave
Stream large result sets for bulk retrieval (returns ALL matched entries, no
--limit support).
uv run scripts/uniprot_tools.py stream "taxonomy_id:9606 AND reviewed:true" --format tsv --fields accession,gene_names > human_reviewed.tsv
Map IDs from one database to another.
uv run scripts/uniprot_tools.py map "P04637" --from_db UniProtKB_AC-ID --to_db Gene_Name
Execute graph queries with SPARQL.
uv run scripts/uniprot_tools.py sparql 'PREFIX up: <http://purl.uniprot.org/core/> SELECT ?protein WHERE { ?protein a up:Protein ; up:reviewed true . } LIMIT 5'
name: instead of protein_name:: name: is not a supported
query term, use protein_name: instead.P04637) are
linked to functional metadata; UniParc IDs (UPI...) are for sequences
only. You can find cross-references from UniParc IDs to UniProtKB Accessions
using the ID Mapping tool.UniProtKB instead.search "term") frequently return false positives (e.g., common maintenance
proteins) because UniProt searches full metadata, including publication
titles. ALWAYS prefer field-specific filters like cc_function: or
protein_name: for functional discovery.lanM) can match substrings in organism names (e.g., Lancefieldella) or
other fields. Use quotes and field prefixes (e.g., gene:lanM) to isolate
true hits.search for retrieving
millions of entries if stream or sparql can do the job. Streaming is
more efficient for very large datasets. Note that stream has a hard limit
of 10,000,000 outputs and does NOT support --limit.count before running
a search without --limit or before using stream. Unlimited queries can
take a long time and consume significant resources if millions of entries
are returned.--limit with stream: The stream command does NOT support
--limit. If you need a limited number of results, use search with
--limit instead.scripts/uniprot_tools.py):
get, search, stream, count -> rest.uniprot.org/{dataset}/map -> rest.uniprot.org/idmapping/sparql -> sparql.uniprot.org/sparqlget --dataset unisave -> rest.uniprot.org/unisave/Frequently asked questions
Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.
The source record exposes this install command: npx skills add https://github.com/google-deepmind/science-skills --skill "skills/uniprot_database". Inspect the command and pinned source before running it.
Static rules flagged write-files, network in the source; the page lists the matching lines and excerpts.
Alternatives
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
wanshuiyin/Auto-claude-code-research-in-sleep
Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.
prowler-cloud/prowler
PostgreSQL indexing best practices for Prowler: index design, partial indexes, partitioned table indexing, EXPLAIN ANALYZE validation, concurrent operations, monitoring, and maintenance. Trigger: When creating or modifying PostgreSQL indexes, analyzing query performance with EXPLAIN, debugging slow queries, reviewing index usage statistics, reindexing, dropping indexes, or working with partitioned table indexes. Also trigger when discussing index strategies, partial indexes, or index maintenance
brucesongs/kali-claw
Insecure Design (OWASP A06:2025) focuses on security flaws in system architecture and design phases, rather than code implementation-level bugs.