extract-webpage-data
Extract structured data from web pages using AI
Install / Use
npx skills add gooseworks-ai/goose-skills --skill extract-webpage-dataInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of extract-webpage-data
extract-webpage-data scores 85/100 on our quality scale, 1833rd of 4,355 Development & Engineering skills we index (top 43%).
Its SKILL.md is 7.3 KB long, well organised into 22 sections with 10 code examples: a thorough specification that gives an agent plenty to work with.
With 1,222 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 8 days ago, so extract-webpage-data is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-01. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
extract-webpage-data compared with similar skills
All 4 of these similar skills score higher than extract-webpage-data; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| extract-webpage-data (this skill)by gooseworks-ai | 85 | 1.2k | 8d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 87.2k | 15d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.2k | today | CLAUDE.md |
| ai-job-searchby MadsLorentzen | 100 | 44.7k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 1d ago | CLAUDE.md |
Frequently asked questions
- How do I install extract-webpage-data?
- Run
npx skills add gooseworks-ai/goose-skills --skill extract-webpage-data. The install tabs above show the steps for each supported agent. - Which AI agents does extract-webpage-data work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is extract-webpage-data safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is extract-webpage-data still maintained?
- The repository was last updated 8 days ago, so extract-webpage-data is actively maintained.
Skill content
View source on GitHubname: extract-webpage-data description: Extract structured data from web pages using AI source: orthogonal
Extract Webpage Data
Setup
Read your credentials from ~/.gooseworks/credentials.json:
export GOOSEWORKS_API_KEY=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json'))['api_key'])")
export GOOSEWORKS_API_BASE=$(python3 -c "import json;print(json.load(open('$HOME/.gooseworks/credentials.json')).get('api_base','https://api.gooseworks.ai'))")
If ~/.gooseworks/credentials.json does not exist, tell the user to run: npx gooseworks login
All endpoints use Bearer auth: -H "Authorization: Bearer $GOOSEWORKS_API_KEY"
Extract structured data from any web page using AI. Turn messy HTML into clean, organized data.
When to Use
- User wants to extract specific data from a website
- User asks to scrape information from a page
- User needs structured data from unstructured content
- User wants to pull product info, contact details, etc.
- Converting web content to usable data
How It Works
Uses Olostep, Scrapegraph, or Riveter APIs for AI-powered data extraction.
Usage
Simple Scrape with Olostep
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"olostep","path":"/v1/scrapes","body":{"url_to_scrape":"https://example.com/products"}}'
AI-Powered Extraction with Scrapegraph
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"scrapegraph","path":"/v1/smartscraper","body":{"website_url":"https://example.com/team","user_prompt":"Extract all team members with their names, titles, and LinkedIn URLs"}}'
Schema-Based Extraction with Riveter
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://example.com","schema":{"name":"string","price":"number","description":"string"}}}'
Get AI Answer from Web
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"olostep","path":"/v1/answers","body":{"task":"Find the pricing for Notion Teams plan from their website"}}'
Crawl Multiple Pages
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"olostep","path":"/v1/crawls","body":{"start_url":"https://example.com","max_pages":10}}'
Parameters
Olostep Scrape
- url_to_scrape (required) - URL to scrape
- formats - Output formats (markdown, html, text)
Scrapegraph
- website_url (required) - URL to scrape
- user_prompt (required) - Natural language description of what to extract
Riveter
- url (required) - URL to scrape
- schema - JSON schema defining the data structure to extract
Olostep Answer
- task (required) - Natural language task/question
Response
Olostep Response
Returns a scrape object:
- id (string) - Scrape ID (e.g.,
scrape_z926lxxon3) - result.markdown_content (string|null) - Page content as markdown
- result.html_content (string|null) - Raw HTML (if requested via
formats) - result.text_content (string|null) - Plain text (if requested)
- result.markdown_hosted_url (string|null) - S3 URL for large content
- result.links_on_page (array) - Links found on the page
- result.screenshot_hosted_url (string|null) - Screenshot URL (if requested)
- result.page_metadata (object) -
status_codeof the page - credits_consumed (integer) - Credits used for this scrape
Async crawls: POST /v1/crawls returns an id. Poll with GET /v1/crawls/{id} until complete.
Scrapegraph Response
Returns structured extraction result:
- request_id (string) - Unique request identifier
- status (string) -
completedorpending - result (object) - AI-extracted data matching your prompt (dynamic keys)
- error (string) - Empty on success, error message on failure
Note: For large pages, the POST may return status: "pending". Poll with GET /v1/smartscraper/{request_id} until status is completed.
Riveter Response
Returns scrape result:
- request_status (string) -
successorerror - message (string) - Human-readable status
- text (string) - Extracted page text content
- url (string) - URL that was scraped
- status_code (integer) - HTTP status of the page
- run_key (string) - Unique run identifier
- base_url_for_links (string) - Base URL for resolving relative links
- riveter_app_link (string) - Link to view run in Riveter dashboard
- credit_used (integer) - Credits consumed
Examples
User: "Get all the product names and prices from this page"
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"scrapegraph","path":"/v1/smartscraper","body":{"website_url":"https://example.com/products","user_prompt":"Extract all products with name, price, and description"}}'
User: "Scrape the team page and get everyone's info"
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"scrapegraph","path":"/v1/smartscraper","body":{"website_url":"https://example.com/about/team","user_prompt":"Extract team members: name, role, bio, photo URL, LinkedIn"}}'
User: "What are Stripe's API pricing details?"
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"olostep","path":"/v1/answers","body":{"task":"Find Stripe API pricing breakdown from stripe.com/pricing"}}'
User: "Get all blog post titles and dates from this blog"
curl -s -X POST $GOOSEWORKS_API_BASE/v1/proxy/orthogonal/run \
-H "Authorization: Bearer $GOOSEWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"api":"riveter","path":"/v1/scrape","body":{"url":"https://blog.example.com","schema":{"posts":[{"title":"string","date":"string","url":"string"}]}}}'
Error Handling
- 504 - Olostep timeout on slow pages — retry or try a simpler URL
- 400 - Missing required parameters (
url_to_scrapefor Olostep,website_url+user_promptfor Scrapegraph,urlfor Riveter) - Scrapegraph returns
errorfield in response body — check it even on 200 status - Riveter returns
request_status: "error"with details inmessage - Some sites block automated scraping — try a different API if one fails
Tips
- Scrapegraph is best for natural language extraction
- Riveter is best when you know the exact schema you want
- Olostep is great for general scraping and AI answers
- For dynamic sites (JavaScript-heavy), these tools handle rendering
- Be specific in your prompts for better extraction results
- Some sites may block automated access
Related Skills
Agent-Reach
87.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.2kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ai-job-search
44.7kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
