apify-core-workflow-a
'Build a complete web scraping Actor with Crawlee and deploy to Apify.
Install / Use
npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-aInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of apify-core-workflow-a
apify-core-workflow-a scores 90/100 on our quality scale, 962nd of 2,607 Automation skills we index (top 37%).
Its SKILL.md is 6.8 KB long, well organised into 21 sections with 5 code examples: a thorough specification that gives an agent plenty to work with.
With 2,785 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 6 days ago, so apify-core-workflow-a is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
apify-core-workflow-a compared with similar skills
All 4 of these similar skills score higher than apify-core-workflow-a; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| apify-core-workflow-a (this skill)by jeremylongshore | 90 | 2.8k | 6d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.3k | 14d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.1k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.6k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.6k | today | MCP Server |
Frequently asked questions
- How do I install apify-core-workflow-a?
- Run
npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-a. The install tabs above show the steps for each supported agent. - Which AI agents does apify-core-workflow-a work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is apify-core-workflow-a safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is apify-core-workflow-a still maintained?
- The repository was last updated 6 days ago, so apify-core-workflow-a is actively maintained.
Skill content
View source on GitHubname: apify-core-workflow-a description: 'Build a complete web scraping Actor with Crawlee and deploy to Apify.
Use when you need end-to-end web scraping on Apify: defining an input schema, building a router-based Crawlee crawler, extracting structured data, storing results in a dataset, testing locally, and deploying the Actor to the platform.
Trigger with "apify scrape website", "build apify actor", "crawlee scraper", "apify main workflow".
' allowed-tools: Read, Write, Edit, Bash(npm:), Bash(npx:), Bash(apify:*), Grep version: 1.5.0 license: MIT author: Jeremy Longshore jeremy@intentsolutions.io tags:
- saas
- scraping
- automation
- apify compatibility: Designed for Claude Code
Apify Core Workflow A — Build & Deploy a Scraper
Overview
End-to-end workflow: define input schema, build a Crawlee-based Actor, extract structured data, store results in datasets, test locally, and deploy to Apify platform. This is the primary money-path workflow for Apify.
Prerequisites
npm install apify crawleein your projectnpm install -g apify-cliandapify logincompleted- For programmatic retrieval (Step 6), an API token in
APIFY_TOKEN— read it from the environment (process.env.APIFY_TOKEN), never hard-code it - Familiarity with
apify-sdk-patterns
Instructions
Step 1: Define Input Schema
Create .actor/INPUT_SCHEMA.json:
{
"title": "E-Commerce Scraper",
"type": "object",
"schemaVersion": 1,
"properties": {
"startUrls": {
"title": "Start URLs",
"type": "array",
"description": "Product listing page URLs to scrape",
"editor": "requestListSources",
"prefill": [{ "url": "https://example-store.com/products" }]
},
"maxItems": {
"title": "Max items",
"type": "integer",
"description": "Maximum number of products to scrape",
"default": 100,
"minimum": 1,
"maximum": 10000
},
"proxyConfig": {
"title": "Proxy configuration",
"type": "object",
"description": "Select proxy to use",
"editor": "proxy",
"default": { "useApifyProxy": true }
}
},
"required": ["startUrls"]
}
Step 2: Build the Actor with Router Pattern
Use a Crawlee router that splits handling by page type: the default handler
enqueues product links + pagination from listing pages, and a PRODUCT-labeled
handler extracts structured fields from detail pages. The entry point wires proxy
config, concurrency, a failed-request handler, and a run summary into the key-value
store. Skeleton:
// src/main.ts
import { Actor } from 'apify';
import { CheerioCrawler, createCheerioRouter, Dataset, log } from 'crawlee';
const router = createCheerioRouter();
router.addDefaultHandler(async ({ enqueueLinks }) => {
await enqueueLinks({ selector: 'a.product-card', label: 'PRODUCT' });
await enqueueLinks({ selector: 'a.next-page', label: 'LISTING' });
});
router.addHandler('PRODUCT', async ({ request, $ }) => {
await Actor.pushData({ url: request.url, name: $('h1.product-title').text().trim() });
});
await Actor.main(async () => {
const input = await Actor.getInput();
const crawler = new CheerioCrawler({ requestHandler: router, maxRequestsPerCrawl: input?.maxItems ?? 100 });
await crawler.run(input.startUrls.map(s => s.url));
});
The full typed Actor — Product/ProductInput interfaces, proxy configuration,
failedRequestHandler, and the SUMMARY key-value write — is in
implementation.md, Step 2.
Step 3: Configure Dockerfile
Use the apify/actor-node:20 base with a two-stage build (compile TypeScript in a
builder stage, ship only dist/ + production deps). Full Dockerfile:
implementation.md, Step 3.
Step 4: Test Locally
# Create test input
mkdir -p storage/key_value_stores/default
echo '{"startUrls":[{"url":"https://example.com"}],"maxItems":5}' \
> storage/key_value_stores/default/INPUT.json
# Run locally
apify run
# Check results
ls storage/datasets/default/
cat storage/key_value_stores/default/SUMMARY.json
Step 5: Deploy to Apify Platform
# Push to Apify (creates Actor if it doesn't exist)
apify push
# Or push to a specific Actor
apify push username/my-actor
# Run on platform
apify actors call username/my-actor
Step 6: Retrieve Results Programmatically
From any client, use the apify-client SDK to call the deployed Actor, list its
dataset items, and download results (JSON/CSV). The token comes from
process.env.APIFY_TOKEN — never hard-code it. Full retrieval code:
implementation.md, Step 6.
Output
- Deployable Actor with typed input schema
- Router-based crawler handling listing + detail pages
- Structured product data in default dataset
- Run summary in default key-value store
- Failed requests tracked with error messages
Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Actor build failed | Dockerfile/deps issue | Check build logs on platform |
| Selector returns empty | Page structure changed | Update CSS selectors |
| maxRequestsPerCrawl hit | Too many pages enqueued | Increase limit or filter URLs |
| Proxy errors | Anti-bot blocking | Switch to residential proxy |
| TIMED-OUT status | Actor exceeded timeout | Increase timeout or reduce scope |
Examples
A quick example — seed a local input, run the Actor, and check results:
mkdir -p storage/key_value_stores/default
echo '{"startUrls":[{"url":"https://example-store.com/products"}],"maxItems":5}' \
> storage/key_value_stores/default/INPUT.json
apify run
cat storage/key_value_stores/default/SUMMARY.json
Three fuller worked scenarios live in examples.md:
- Scrape a catalog locally, then deploy — the full seed →
apify run→ inspect →apify pushloop, with the expectedSUMMARY.jsonoutput. - Run the deployed Actor and export CSV — call the Actor via
apify-clientand download the dataset as CSV. - Route through residential proxy — pass a
proxyConfiggroup at run time to get past anti-bot blocking.
Resources
- Crawlee Quick Start
- Actor Deployment
- Input Schema Spec
- Full implementation walkthrough — complete Actor source, Dockerfile, and retrieval code
- Worked examples — three end-to-end run scenarios
Next Steps
Once your Actor is deployed and producing data, move on to dataset and key-value
store management — pagination over large datasets, deduplication, exporting to
external stores, and scheduling recurring runs — covered in apify-core-workflow-b.
Related Skills
Agent-Reach
86.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.1kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.6k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
84.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
