Apify Core Workflow B — Storage & Pipelines

SkillDatabases & data

Lets your agent manage Apify datasets, key-value stores, request queues, and run multi-step scraping pipelines.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Apify Core Workflow B skill

About this skill

Manage Apify datasets, key-value stores, and request queues programmatically, and orchestrate multi-Actor pipelines.

What this skill tells your AI

The instructions your AI receives, as published by jeremylongshore/tons-of-skills-marketplace in skills/.curated/apify-core-workflow-b/SKILL.md and read by ahel’s review.

Overview

Manage Apify's three storage types (datasets, key-value stores, request queues) and orchestrate multi-Actor pipelines using the apify-client JS SDK. Covers CRUD operations, data export, automatic pagination, and chaining Actors together (scrape → transform → export).

This SKILL.md gives you the high-level workflow plus the essential first example for each storage type. Drill into the reference files for the complete, copy-ready code:

Prerequisites

  • Node.js with apify-client installed (npm install apify-client).
  • An Apify account token exported as APIFY_TOKEN (see Authentication below).
  • Familiarity with apify-core-workflow-a (Actor invocation and run lifecycle), since pipelines chain Actor runs and read their default storages.

Authentication

All operations authenticate with an Apify API token. Never hard-code it — read it from the environment and construct the client once:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });

Generate a token at Apify Console → Settings → Integrations, then export it (export APIFY_TOKEN=apify_api_...) or load it from your secrets manager.

Storage Types at a Glance

StorageBest ForAnalogyRetention
DatasetLists of similar items (products, pages)Append-only table7 days (unnamed)
Key-Value StoreConfig, screenshots, summaries, any fileS3 bucket7 days (unnamed)
Request QueueURLs to crawl (managed by Crawlee)Job queue7 days (unnamed)

Named storages persist indefinitely. Unnamed (default run) storages expire after 7 days.

Instructions

Pick the storage type you need, use the skeleton below to get started, then open the linked reference for the full operation set.

Datasets — append-only item lists

getOrCreate a named dataset, push items, and list them (pagination is manual):

const dataset = await client.datasets().getOrCreate('product-catalog');
const dsClient = client.dataset(dataset.id);
await dsClient.pushItems([{ sku: 'ABC123', name: 'Widget', price: 9.99 }]);
const { items, total } = await dsClient.listItems({ limit: 100, offset: 0 });

Full auto-pagination loop, CSV/JSON/XLSX export, and field filtering: storage-operations.md, Step 1.

Key-value stores — config, files, and Actor OUTPUT

Store JSON or binary records by key, then retrieve them:

const store = await client.keyValueStores().getOrCreate('scraper-config');
const kvClient = client.keyValueStore(store.id);
await kvClient.setRecord({ key: 'settings', value: { maxRetries: 3 }, contentType: 'application/json' });
const record = await kvClient.getRecord('settings');

Binary records, key listing, and reading a run's default OUTPUT: storage-operations.md, Step 2.

Request queues — resumable crawl URLs

Create a named queue and add requests (deduplicated by uniqueKey):

const queue = await client.requestQueues().getOrCreate('my-crawl-queue');
const rqClient = client.requestQueue(queue.id);
await rqClient.addRequest({ url: 'https://example.com/page1', uniqueKey: 'page1' });

Batch adds and queue stats: storage-operations.md, Step 3.

Multi-Actor pipelines & monitoring

Chain Actors (scrape → transform → export) and monitor run status and cost. Full runPipeline() function and run-monitoring code: pipelines.md.

Output

  • Datasets return { items, total, count, offset, limit } from listItems(); downloadItems(format) returns a Buffer in csv / json / xlsx.
  • Key-value stores return { key, value, contentType } from getRecord() and { items } (each { key, size }) from listKeys().
  • Request queues return { pendingRequestCount, handledRequestCount, ... } from get().
  • Pipelines return the named export dataset id; run monitoring yields { status, statusMessage, stats, usage, usageTotalUsd } per run.

Error Handling

ErrorCauseSolution
Dataset not foundExpired (unnamed, >7 days)Use named datasets for persistence
Record too largeKV store 9MB record limitSplit into multiple records
Push failedDataset items >9MB batchPush in smaller batches
Request already existsDuplicate uniqueKeyExpected behavior, queue deduplicates

Examples

Export a named dataset to CSV — get the client, download the buffer, write it:

const csvBuffer = await client.dataset('product-catalog').downloadItems('csv');
require('fs').writeFileSync('products.csv', csvBuffer);

Read an Actor run's OUTPUT record — after a run completes:

const run = await client.actor('apify/web-scraper').call(input);
const output = await client.keyValueStore(run.defaultKeyValueStoreId).getRecord('OUTPUT');

Longer end-to-end examples — the full pagination loop, binary record storage, and the three-stage runPipeline() — live in the reference files: storage-operations.md and pipelines.md.

Resources

Next Steps

For common errors and their fixes across the Apify pack, see the apify-common-errors skill. For Actor invocation and run lifecycle basics that pipelines build on, see apify-core-workflow-a.

Signals

GitHub stars
3k
Forks
408
Last commit
Sep 2026

ahel review

  • K1binfo
    installs-packages

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
apify-core-workflow-b
Source
github.com/jeremylongshore/tons-of-skills-marketplace