uspto-database
Access USPTO patent data via PatentsView REST API and Google Patents Public Data (BigQuery). Search by inventor, assignee, CPC, or keywords; download metadata and claims; analyze portfolios; track tech trends.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill uspto-databaseInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of uspto-database
uspto-database scores 91/100 on our quality scale, 273rd of 730 Operations skills we index (top 38%).
Its SKILL.md is 18 KB long, well organised into 30 sections with 13 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so uspto-database is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
uspto-database compared with similar skills
All 4 of these similar skills score higher than uspto-database; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| uspto-database (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.8k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.7k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install uspto-database?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill uspto-database. The install tabs above show the steps for each supported agent. - Which AI agents does uspto-database work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is uspto-database safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is uspto-database still maintained?
- The repository was last updated 37 days ago, so uspto-database is actively maintained.
Skill content
View source on GitHubname: "uspto-database" description: "Access USPTO patent data via PatentsView REST API and Google Patents Public Data (BigQuery). Search by inventor, assignee, CPC, or keywords; download metadata and claims; analyze portfolios; track tech trends. For IP landscape analysis, competitor monitoring, prior art search, and tech forecasting in life sciences and biotech." license: "CC0-1.0"
uspto-database
Overview
The USPTO provides two primary programmatic access points for patent data: the PatentsView API (REST, free, no key required for basic use) for structured queries by inventor, assignee, CPC classification, and keywords; and Google Patents Public Data (BigQuery public dataset) for large-scale analytics across the full patent corpus. Both expose data under the CC0 Public Domain Dedication. This skill covers Python-based access patterns for both, plus basic patent portfolio analytics.
When to Use
- Prior art search: Finding existing patents relevant to a technology before filing or to assess freedom-to-operate.
- Competitor IP landscape analysis: Querying all patents from a specific assignee (company or institution) to map their technology portfolio.
- CPC classification search: Finding patents in a specific technology area using Cooperative Patent Classification codes (e.g., C12N for nucleotides/genetic engineering).
- Inventor network analysis: Identifying prolific inventors in a field and their institutional affiliations.
- Technology trend tracking: Counting patent filings by year and technology category to identify emerging areas.
- Life sciences IP analysis: Searching biotech-specific classifications (A61K for pharmaceuticals, C12N for genetics, G16B for bioinformatics).
- For full-text patent PDF downloads, use the USPTO Bulk Data Storage System (BDSS) or Google Patents direct links.
- Rate limits: PatentsView API allows 45 requests/minute without an API key; request a free key for 45 req/min with higher daily limits.
Prerequisites
- Python packages:
requests,pandas,matplotlib - Optional:
google-cloud-bigqueryfor Google Patents Public Data queries - Data requirements: No account needed for PatentsView basic queries; Google Cloud account required for BigQuery
- Rate limits: PatentsView — 45 requests/minute (unauthenticated), higher with free API key
pip install requests pandas matplotlib
pip install google-cloud-bigquery # optional: for BigQuery access
Quick Start
import requests
import pandas as pd
# Search PatentsView API: patents assigned to "Genentech" in CPC class C12N
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"_and": [
{"_contains": {"assignee_organization": "Genentech"}},
{"_contains": {"cpc_subgroup_id": "C12N"}},
]},
"f": ["patent_number", "patent_title", "patent_date", "assignee_organization"],
"o": {"per_page": 25},
}
resp = requests.post(url, json=payload)
data = resp.json()
df = pd.DataFrame(data["patents"])
print(f"Found: {data['total_patent_count']} patents")
print(df[["patent_number", "patent_title", "patent_date"]].head())
Core API
Query Type 1: Search by Assignee (Company / Institution)
Find all patents granted to a specific organization.
import requests
import pandas as pd
def search_by_assignee(assignee_name: str, per_page: int = 100) -> pd.DataFrame:
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"_contains": {"assignee_organization": assignee_name}},
"f": [
"patent_number", "patent_title", "patent_date",
"patent_abstract", "assignee_organization", "assignee_country",
],
"o": {"per_page": per_page, "sort": [{"patent_date": "desc"}]},
}
resp = requests.post(url, json=payload)
resp.raise_for_status()
data = resp.json()
df = pd.DataFrame(data.get("patents", []))
print(f"Assignee '{assignee_name}': {data.get('total_patent_count', 0)} total patents")
return df
# Example: patents from Broad Institute
df_broad = search_by_assignee("Broad Institute")
print(df_broad[["patent_number", "patent_title", "patent_date"]].head(10))
# Paginate through all results for large portfolios
def search_assignee_all_pages(assignee_name: str, page_size: int = 100) -> pd.DataFrame:
url = "https://api.patentsview.org/patents/query"
all_patents = []
page = 1
while True:
payload = {
"q": {"_contains": {"assignee_organization": assignee_name}},
"f": ["patent_number", "patent_title", "patent_date", "cpc_subgroup_id"],
"o": {"per_page": page_size, "page": page},
}
resp = requests.post(url, json=payload)
data = resp.json()
patents = data.get("patents", [])
if not patents:
break
all_patents.extend(patents)
total = data.get("total_patent_count", 0)
if len(all_patents) >= total:
break
page += 1
df = pd.DataFrame(all_patents)
print(f"Retrieved {len(df)} patents for '{assignee_name}'")
return df
Query Type 2: Search by CPC Classification
CPC (Cooperative Patent Classification) codes organize patents by technology. Life sciences codes include C12N (nucleotides/genetics), A61K (pharmaceuticals), and G16B (bioinformatics).
import requests
import pandas as pd
# Search by CPC subgroup: C12N15 (mutation/genetic engineering)
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"_begins": {"cpc_subgroup_id": "C12N15"}},
"f": [
"patent_number", "patent_title", "patent_date",
"assignee_organization", "cpc_subgroup_id",
],
"o": {"per_page": 50, "sort": [{"patent_date": "desc"}]},
}
resp = requests.post(url, json=payload)
data = resp.json()
df = pd.DataFrame(data["patents"])
print(f"C12N15 patents: {data['total_patent_count']}")
print(df[["patent_number", "patent_title", "assignee_organization"]].head(10))
# Common life sciences CPC codes
CPC_LIFE_SCIENCES = {
"C12N": "Microorganisms / enzymes / compositions",
"C12N15": "Mutation / genetic engineering",
"C12Q": "Measuring / testing involving enzymes or microorganisms",
"A61K": "Preparations for medical use",
"A61P": "Therapeutic activity of chemical compounds",
"G16B": "Bioinformatics",
"G16H": "Healthcare informatics",
"C07K": "Peptides / proteins",
}
for code, desc in CPC_LIFE_SCIENCES.items():
print(f" {code:10s}: {desc}")
Query Type 3: Full-Text Keyword Search
Search patent titles and abstracts for specific terms.
import requests
import pandas as pd
def keyword_search(keyword: str, per_page: int = 50) -> pd.DataFrame:
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"_or": [
{"_text_any": {"patent_title": keyword}},
{"_text_any": {"patent_abstract": keyword}},
]},
"f": [
"patent_number", "patent_title", "patent_date",
"patent_abstract", "assignee_organization",
],
"o": {"per_page": per_page, "sort": [{"patent_date": "desc"}]},
}
resp = requests.post(url, json=payload)
resp.raise_for_status()
data = resp.json()
df = pd.DataFrame(data.get("patents", []))
print(f"Keyword '{keyword}': {data.get('total_patent_count', 0)} patents found")
return df
# Search for CRISPR-related patents
df_crispr = keyword_search("CRISPR")
print(df_crispr[["patent_number", "patent_title", "patent_date"]].head(10))
Query Type 4: Inventor Search
Find patents by inventor name or retrieve an inventor's full publication history.
import requests
import pandas as pd
# Search by inventor name
url = "https://api.patentsview.org/inventors/query"
payload = {
"q": {"_and": [
{"inventor_last_name": "Doudna"},
{"inventor_first_name": "Jennifer"},
]},
"f": ["inventor_id", "inventor_first_name", "inventor_last_name",
"inventor_city", "inventor_state", "inventor_country"],
"o": {"per_page": 10},
}
resp = requests.post(url, json=payload)
data = resp.json()
print(f"Found {data.get('total_inventor_count', 0)} inventors matching 'Jennifer Doudna'")
for inv in data.get("inventors", []):
print(f" ID: {inv['inventor_id']}, Location: {inv.get('inventor_city')}, {inv.get('inventor_country')}")
# Get all patents for a specific inventor by inventor_id
inventor_id = "fl:j_ln:doudna-1" # PatentsView inventor ID format
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"inventor_id": inventor_id},
"f": ["patent_number", "patent_title", "patent_date", "assignee_organization"],
"o": {"per_page": 100, "sort": [{"patent_date": "desc"}]},
}
resp = requests.post(url, json=payload)
data = resp.json()
df = pd.DataFrame(data.get("patents", []))
print(f"Patents for inventor {inventor_id}: {data.get('total_patent_count', 0)}")
print(df.head(5))
Query Type 5: Date Range and Combined Filters
Combine multiple filters for targeted searches.
import requests
import pandas as pd
# Patents in gene therapy (CPC A61K48) filed 2020-2024 by a US assignee
url = "https://api.patentsview.org/patents/query"
payload = {
"q": {"_and": [
{"_begins": {"cpc_subgroup_id": "A61K48"}},
{"_gte": {"patent_date": "2020-01-01"}},
{"_lte": {"patent_date": "2024-12-31"}},
{"_eq": {"assignee_country": "US"}},
]},
"f": [
"patent_number", "patent_title", "patent_date",
"assignee_organization", "patent_num_claims",
],
"o": {"per_page": 100, "sort": [{"patent_date": "desc"}]},
}
resp = requests.post(url, json=payload)
data = resp.json()
df = pd.DataFrame(data.get("patents", []))
print(f"Gene therapy patents 2020-2024 (US assignee): {data.get('total_patent_count', 0)}")
print(df[["patent_number", "patent_title", "patent_date", "assignee_organization"]].head(10))
Query Type 6: Google Patents BigQuery
For large-scale corpus analytics, use the public Google Patents dataset in BigQuery.
from google.cloud import bigquery
client = bigquery.Client(project="YOUR_GCP_PROJECT")
# Count CRISPR patents by year (Google Patents public data)
query = """
SELECT
EXTRACT(YEAR FROM filing_date) AS filing_year,
COUNT(*) AS patent_count,
COUNT(DISTINCT assignee) AS unique_assignees
FROM `patents-public-data.patents.publications`
WHERE
(LOWER(title_localized[SAFE_OFFSET(0)].text) LIKE '%crispr%'
OR LOWER(abstract_localized[SAFE_OFFSET(0)].text) LIKE '%crispr%')
AND filing_date >= '2010-01-01'
AND country_code = 'US'
GROUP BY filing_year
ORDER BY filing_year
"""
df_bq = client.query(query).to_dataframe()
print(df_bq)
print(f"Peak year: {df_bq.loc[df_bq.patent_count.idxmax(), 'filing_year']} "
f"({df_bq.patent_count.max()} patents)")
Key Parameters
| Parameter | Module | Default | Range / Options | Effect |
|-----------|--------|---------|-----------------|--------|
| per_page | PatentsView "o" | 25 | 1–10000 | Results per API call |
| page | PatentsView "o" | 1 | 1–max pages | Page number for pagination |
| sort | PatentsView "o" | API default | any field + "asc"/"desc" | Sort order of results |
| "f" fields | PatentsView | minimal | any valid field list | Fields returned in response (controls payload size) |
| "_begins" | query operator | — | field + prefix string | Prefix match (e.g., CPC code prefix) |
| "_contains" | query operator | — | field + substring | Substring search (case-insensitive) |
| "_text_any" | query operator | — | field + keywords | Full-text search on title/abstract fields |
Best Practices
- Request only the fields you need: The
"f"(fields) parameter controls what
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
