Skip to content

SDKs

TypeScript / JavaScript — @justcrawl/sdk

Section titled “TypeScript / JavaScript — @justcrawl/sdk”

The official client. Zero runtime dependencies, Node 20.3+, ESM and CommonJS, Apache-2.0, published with npm provenance.

Terminal window
npm install @justcrawl/sdk

Its types are generated from the same OpenAPI spec that powers this site, and a CI gate fails our build if the two ever disagree — so the SDK cannot drift from the API.

import { JustCrawl } from '@justcrawl/sdk';
const jc = new JustCrawl({
apiKey: process.env.JUSTCRAWL_API_KEY!, // sr_live_…
baseUrl: 'https://api.justcrawl.io', // default
timeoutMs: 30_000, // default
maxRetries: 2, // default; safe reads + idempotent BI submit
});

Your API key is stored non-enumerably and masked in toJSON() and console.log, so it does not end up in logs, crash reports, or serialized state.

One group per API area:

Group Covers
jobs submit, list, poll, batch status, results
urls URL items, batch create/delete, per-URL jobs and extractions
workflows create, validate, publish, clone, versions, routing
smartWorkflows auto-optimization mode and routing suggestions
schedules recurring crawls, runs, toggle, manual trigger
extraction results, schemas, custom attributes, XPath testing, backfill
analytics overview, timeseries, providers, domains, matrix
benchmarks provider benchmark runs and results
plans plan status, credit ledger, recharge requests
integrations SQS inputs, S3/webhook outputs, storage config
webhooks public URL ingestion
bi SQL console: schema, queries, saved queries, exports

Anything without a wrapper — authentication, account, and audit-log endpoints included — is reachable through the typed escape hatch:

const me = await jc.request('get', '/api/v1/auth/me');
const job = await jc.request('get', '/api/v1/jobs/{id}', { path: { id: jobId } });

request() is typed against the generated spec, so an unknown path, an unsupported method, or a missing path parameter is a compile error. Paths are the spec’s own templates and the values go in path; interpolating them into the string yourself produces a plain string, which nothing can check.

For an endpoint published after your installed version was cut, opt out of the checking explicitly:

const preview = await jc.requestUnchecked<{ ok: boolean }>('get', '/api/v1/preview/thing');

New structured analysis sends a bounded aggregate definition, never SQL:

const query = await jc.bi.runStructuredQuery({
definition: {
version: 1,
table: 'flat__shop__product',
dimensions: [{ kind: 'time_bucket', column: 'captured_at', granularity: 'day' }],
measures: [{ function: 'average', column: 'price' }],
limit: 100,
},
});
const saved = await jc.bi.createSavedQuery({
name: 'Daily shop price',
sourceQueryId: query.jobId,
});
const replay = await jc.bi.runSavedQuery(saved.id);

The source save copies the completed execution’s content, immutable org/agent scope, dialect, and provenance on the server. Do not send scope fields or copy sqlPreview back as executable input. runSavedQuery likewise sends only the saved id. To edit a structured saved object in SQL, call convertSavedQueryToSql(id); conversion does not mutate it, and a later updateSavedQuery preserves scope and the server-derived PostgreSQL dialect pin.

Export a completed result without rerunning it. Create once, retain the export handle, and poll it; each ready response contains a newly signed URL:

const { export: started } = await jc.bi.createExport(query.jobId, {
format: 'parquet',
});
const { export: current } = await jc.bi.getExport(started.id);
if (current.status === 'ready' && current.downloadUrl) {
const response = await fetch(current.downloadUrl); // no Authorization header
const bytes = await response.arrayBuffer();
}

Creation requires bi:write; polling requires bi:read. The default admission limits are 100,000 rows and 256 MiB of stored source JSONL. Recovery has a finite attempt budget, and terminal failures retain both the query and export IDs. Fetch the storage URL directly without the JustCrawl API key. CSV exports neutralize formula-like headers and cells with a leading apostrophe before CSV escaping; choose Parquet when exact values must round-trip without that safety prefix. Result pages expose the server-derived exportEligibility decision, including caller permission and both admission-cap checks.

const job = await jc.jobs.submitAndWait(
{ urlItemId },
{ maxWaitMs: 120_000, signal: controller.signal, onPoll: (j) => console.log(j.status) },
);

Polling in chunks with exponential backoff — there is no streaming interface, because the API exposes none. submitAndWait resolves on both completed and failed; check job.status.

The SDK exposes the published request shapes directly. The single-URL handler currently consumes url and optional priority; do not send tags on this endpoint. Schedule creation can use a published workflow or Automatic routing, and enablement is explicit desired state:

const target = await jc.urls.create({
url: 'https://example.com/products/123',
priority: 10,
});
const schedule = await jc.schedules.create({
name: 'Six-hour product sweep',
workflowId: null, // Automatic routing per URL
frequency: 'every_6_hours',
timezone: 'UTC',
tagFilters: [], // every enabled URL in the organization
isEnabled: false, // configure first; enable when ready
});
await jc.schedules.toggle(schedule.id, { isEnabled: true });

URL creation requires urls:write. Schedule creation and enable/disable require schedules:write plus a verified account email. Every lookup and mutation is scoped to the API key’s organization. Repeating either write is not treated as idempotent: ordinary POST and PATCH calls are never retried by the SDK.

const result = await jc.jobs.fetchResult(jobId);
switch (result.kind) {
case 'url': return result.data; // fetched from a presigned URL for you
case 'blob': return result.blobKey; // your bucket — fetch with your own creds
case 'expired': return result.expiredAt; // aged out of your retention window
case 'refunded': return result.message; // no provider delivered a page; the job was free
}

expired and refunded are outcomes, not exceptions — neither is a fault on your side, so neither throws. They mean different things: expired is a page that existed and was swept by your retention policy, refunded is a page that was never delivered at all, so the job’s credit came back and no retry of the fetch will ever produce a body. The job resource says the same thing in its creditOutcome / creditNote fields.

The presigned URL is fetched with no Authorization header — your key is never sent to the storage host.

import { JustCrawlError } from '@justcrawl/sdk';
try {
await jc.jobs.submit({ urlItemId });
} catch (err) {
if (err instanceof JustCrawlError) {
err.status; // HTTP status
err.code; // machine-readable code
err.inferredCode; // true when the SDK derived the code from the status
err.requestId; // quote this in support requests
err.raw; // the untouched response body
}
}

inferredCode is worth understanding: not every endpoint sends a machine code yet. When one doesn’t, the SDK derives a sensible code from the HTTP status and sets this flag — it never presents a guess as something we said. Branch on err.code only after checking it, or branch on err.status.

Keep your own UUID in the Idempotency-Key header when you need to recover from a lost response after all SDK retries. Then use jc.bi.listQueries({ submissionKey }) to find the retained handle without submitting again. This exact lookup is tenant-scoped and remains available after newer queries fill recent history; a missing key returns an empty list.

GET requests retry on network errors, 5xx, and 429, honoring Retry-After when present and falling back to exponential backoff with jitter when it isn’t.

Ordinary writes are never retried. Submitting a job twice charges two credits. bi.runQuery(), bi.runStructuredQuery(), and bi.runSavedQuery() are the endpoint-specific exceptions: each generates an Idempotency-Key UUID (or preserves the caller’s header) and reuses that key for every attempt, so a lost response resolves to the original BI query rather than a duplicate. A crawl that fails at every provider has its credit returned automatically, but a retried submit that succeeds is a second real charge, so the refund is no substitute for the dedup pattern. 402 Insufficient credits is never retried either — it needs a human to add credits, not another request.

There is no first-class SDK for other languages yet. The API is plain REST with bearer auth, and our OpenAPI spec is published — most generators produce a workable client from it directly.

If you’re writing your own client, three rules are worth copying from ours:

  1. Do not retry ordinary writes. A retried POST /jobs charges twice. BI query submission and POST /api/v1/bi/saved-queries/{id}/run are the exceptions, and only when every attempt reuses the same UUID Idempotency-Key.
  2. Retry 429 and honor Retry-After on safe methods.
  3. Never attach your API key when fetching a presigned resultUrl — it leaks the key into storage access logs, and the request is rejected anyway.

Model Context Protocol — @justcrawl/mcp-server

Section titled “Model Context Protocol — @justcrawl/mcp-server”

If you want an agent to use justcrawl rather than your code, reach for the MCP server instead of this SDK. It is a local stdio process your agent host spawns, and it exposes twenty-nine tools over the same API — a guided scrape that builds the routing workflow for you, raw job submission and inspection, browsing workflows, creating and browsing schedules and URLs, discovering the authorized BI catalog, running safe structured or saved BI queries, creating and polling CSV/Parquet exports, and searching these docs.

{
"mcpServers": {
"justcrawl": {
"command": "npx",
"args": ["-y", "@justcrawl/mcp-server"],
"env": { "JUSTCRAWL_API_KEY": "sr_live_your_key_here" }
}
}
}

It is built on this SDK, so the three rules above hold there too; the server inherits them rather than reimplementing them. One rule is its own: it rejects any sql argument on any tool. jc_bi_query accepts only structured aggregate intent, jc_bi_save_query can save a completed structured run, and new jc_bi_run_saved calls take the id of either an MCP-created or dashboard-created saved query. Its optional queryId continuation alias remains only until the next MCP npm major; jc_bi_get_results is the supported continuation tool and resumes either returned query handle without resubmitting it. jc_bi_create_export and jc_bi_get_export preserve one export handle through finite recovery and return the fresh storage URL without fetching it. A prompt injection hidden in a scraped page therefore cannot turn into a query over your data or receive your API credential through an object-store request.