apify-core-workflow-b
Manage Apify datasets, key-value stores, and request queues programmatically, and orchestrate multi-Actor pipelines. Use when you need to read or write Apify datasets, export scraped data to CSV/JSON/XLSX, store config or binary artifacts in a key-value store, manage a resumable request queue, chain…
Install / Use
npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-bInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of apify-core-workflow-b
apify-core-workflow-b scores 93/100 on our quality scale, 697th of 3,055 Automation skills we index (top 23%).
Its SKILL.md is 6.9 KB long, well organised into 15 sections with 6 code examples: a thorough specification that gives an agent plenty to work with.
With 2,785 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 8 days ago, so apify-core-workflow-b is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
apify-core-workflow-b compared with similar skills
All 4 of these similar skills score higher than apify-core-workflow-b; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| apify-core-workflow-b (this skill)by jeremylongshore | 93 | 2.8k | 8d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 88.1k | 17d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.7k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.2k | 1d ago | MCP Server |
Frequently asked questions
- How do I install apify-core-workflow-b?
- Run
npx skills add jeremylongshore/tons-of-skills-marketplace --skill apify-core-workflow-b. The install tabs above show the steps for each supported agent. - Which AI agents does apify-core-workflow-b work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is apify-core-workflow-b safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is apify-core-workflow-b still maintained?
- The repository was last updated 8 days ago, so apify-core-workflow-b is actively maintained.
Skill content
View source on GitHubname: apify-core-workflow-b description: | Manage Apify datasets, key-value stores, and request queues programmatically, and orchestrate multi-Actor pipelines.
Use when you need to read or write Apify datasets, export scraped data to CSV/JSON/XLSX, store config or binary artifacts in a key-value store, manage a resumable request queue, chain Actors into a scrape → transform → export pipeline, or monitor Actor run status and cost.
Trigger with "apify dataset", "apify key-value store", "apify storage", "export apify data", "apify pipeline", "apify request queue". allowed-tools: Read, Write, Edit, Bash(npm:), Bash(npx:), Grep version: 1.5.0 license: MIT author: Jeremy Longshore jeremy@intentsolutions.io tags:
- saas
- scraping
- automation
- apify compatibility: Designed for Claude Code
Apify Core Workflow B — Storage & Pipelines
Overview
Manage Apify's three storage types (datasets, key-value stores, request queues)
and orchestrate multi-Actor pipelines using the apify-client JS SDK. Covers
CRUD operations, data export, automatic pagination, and chaining Actors
together (scrape → transform → export).
This SKILL.md gives you the high-level workflow plus the essential first example for each storage type. Drill into the reference files for the complete, copy-ready code:
- Storage operations — full reference — every dataset, key-value store, and request queue operation with pagination, format export, and binary records.
- Pipelines & run monitoring — full reference — the multi-Actor pipeline function and Actor-run status/cost/abort monitoring.
Prerequisites
- Node.js with
apify-clientinstalled (npm install apify-client). - An Apify account token exported as
APIFY_TOKEN(see Authentication below). - Familiarity with
apify-core-workflow-a(Actor invocation and run lifecycle), since pipelines chain Actor runs and read their default storages.
Authentication
All operations authenticate with an Apify API token. Never hard-code it — read it from the environment and construct the client once:
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
Generate a token at Apify Console → Settings → Integrations, then export it
(export APIFY_TOKEN=apify_api_...) or load it from your secrets manager.
Storage Types at a Glance
| Storage | Best For | Analogy | Retention | |---------|----------|---------|-----------| | Dataset | Lists of similar items (products, pages) | Append-only table | 7 days (unnamed) | | Key-Value Store | Config, screenshots, summaries, any file | S3 bucket | 7 days (unnamed) | | Request Queue | URLs to crawl (managed by Crawlee) | Job queue | 7 days (unnamed) |
Named storages persist indefinitely. Unnamed (default run) storages expire after 7 days.
Instructions
Pick the storage type you need, use the skeleton below to get started, then open the linked reference for the full operation set.
Datasets — append-only item lists
getOrCreate a named dataset, push items, and list them (pagination is manual):
const dataset = await client.datasets().getOrCreate('product-catalog');
const dsClient = client.dataset(dataset.id);
await dsClient.pushItems([{ sku: 'ABC123', name: 'Widget', price: 9.99 }]);
const { items, total } = await dsClient.listItems({ limit: 100, offset: 0 });
Full auto-pagination loop, CSV/JSON/XLSX export, and field filtering: storage-operations.md, Step 1.
Key-value stores — config, files, and Actor OUTPUT
Store JSON or binary records by key, then retrieve them:
const store = await client.keyValueStores().getOrCreate('scraper-config');
const kvClient = client.keyValueStore(store.id);
await kvClient.setRecord({ key: 'settings', value: { maxRetries: 3 }, contentType: 'application/json' });
const record = await kvClient.getRecord('settings');
Binary records, key listing, and reading a run's default OUTPUT:
storage-operations.md, Step 2.
Request queues — resumable crawl URLs
Create a named queue and add requests (deduplicated by uniqueKey):
const queue = await client.requestQueues().getOrCreate('my-crawl-queue');
const rqClient = client.requestQueue(queue.id);
await rqClient.addRequest({ url: 'https://example.com/page1', uniqueKey: 'page1' });
Batch adds and queue stats: storage-operations.md, Step 3.
Multi-Actor pipelines & monitoring
Chain Actors (scrape → transform → export) and monitor run status and cost.
Full runPipeline() function and run-monitoring code:
pipelines.md.
Output
- Datasets return
{ items, total, count, offset, limit }fromlistItems();downloadItems(format)returns aBufferincsv/json/xlsx. - Key-value stores return
{ key, value, contentType }fromgetRecord()and{ items }(each{ key, size }) fromlistKeys(). - Request queues return
{ pendingRequestCount, handledRequestCount, ... }fromget(). - Pipelines return the named export dataset id; run monitoring yields
{ status, statusMessage, stats, usage, usageTotalUsd }per run.
Error Handling
| Error | Cause | Solution |
|-------|-------|----------|
| Dataset not found | Expired (unnamed, >7 days) | Use named datasets for persistence |
| Record too large | KV store 9MB record limit | Split into multiple records |
| Push failed | Dataset items >9MB batch | Push in smaller batches |
| Request already exists | Duplicate uniqueKey | Expected behavior, queue deduplicates |
Examples
Export a named dataset to CSV — get the client, download the buffer, write it:
const csvBuffer = await client.dataset('product-catalog').downloadItems('csv');
require('fs').writeFileSync('products.csv', csvBuffer);
Read an Actor run's OUTPUT record — after a run completes:
const run = await client.actor('apify/web-scraper').call(input);
const output = await client.keyValueStore(run.defaultKeyValueStoreId).getRecord('OUTPUT');
Longer end-to-end examples — the full pagination loop, binary record storage,
and the three-stage runPipeline() — live in the reference files:
storage-operations.md and
pipelines.md.
Resources
- Dataset Documentation
- Key-Value Store Documentation
- Request Queue Documentation
- JS Client API Reference
Next Steps
For common errors and their fixes across the Apify pack, see the
apify-common-errors skill. For Actor invocation and run lifecycle basics that
pipelines build on, see apify-core-workflow-a.
Related Skills
Agent-Reach
88.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.7k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
85.2k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
