SDKs
TypeScript / JavaScript — @justcrawl/sdk
Section titled “TypeScript / JavaScript — @justcrawl/sdk”The official client. Zero runtime dependencies, Node 20.3+, ESM and CommonJS, Apache-2.0, published with npm provenance.
npm install @justcrawl/sdk- npm:
@justcrawl/sdk - Source: github.com/justcrawl-io/justcrawl-sdk-js
Its types are generated from the same OpenAPI spec that powers this site, and a CI gate fails our build if the two ever disagree — so the SDK cannot drift from the API.
Client
Section titled “Client”import { JustCrawl } from '@justcrawl/sdk';
const jc = new JustCrawl({ apiKey: process.env.JUSTCRAWL_API_KEY!, // sr_live_… baseUrl: 'https://api.justcrawl.io', // default timeoutMs: 30_000, // default maxRetries: 2, // default; safe reads + idempotent BI submit});Your API key is stored non-enumerably and masked in toJSON() and
console.log, so it does not end up in logs, crash reports, or serialized
state.
Resources
Section titled “Resources”One group per API area:
| Group | Covers |
|---|---|
jobs |
submit, list, poll, batch status, results |
urls |
URL items, batch create/delete, per-URL jobs and extractions |
workflows |
create, validate, publish, clone, versions, routing |
smartWorkflows |
auto-optimization mode and routing suggestions |
schedules |
recurring crawls, runs, toggle, manual trigger |
extraction |
results, schemas, custom attributes, XPath testing, backfill |
analytics |
overview, timeseries, providers, domains, matrix |
benchmarks |
provider benchmark runs and results |
plans |
plan status, credit ledger, recharge requests |
integrations |
SQS inputs, S3/webhook outputs, storage config |
webhooks |
public URL ingestion |
bi |
SQL console: schema, queries, saved queries, exports |
Anything without a wrapper — authentication, account, and audit-log endpoints included — is reachable through the typed escape hatch:
const me = await jc.request('get', '/api/v1/auth/me');const job = await jc.request('get', '/api/v1/jobs/{id}', { path: { id: jobId } });request() is typed against the generated spec, so an unknown path, an
unsupported method, or a missing path parameter is a compile error. Paths are
the spec’s own templates and the values go in path; interpolating them into
the string yourself produces a plain string, which nothing can check.
For an endpoint published after your installed version was cut, opt out of the checking explicitly:
const preview = await jc.requestUnchecked<{ ok: boolean }>('get', '/api/v1/preview/thing');Structured and saved BI queries
Section titled “Structured and saved BI queries”New structured analysis sends a bounded aggregate definition, never SQL:
const query = await jc.bi.runStructuredQuery({ definition: { version: 1, table: 'flat__shop__product', dimensions: [{ kind: 'time_bucket', column: 'captured_at', granularity: 'day' }], measures: [{ function: 'average', column: 'price' }], limit: 100, },});
const saved = await jc.bi.createSavedQuery({ name: 'Daily shop price', sourceQueryId: query.jobId,});const replay = await jc.bi.runSavedQuery(saved.id);The source save copies the completed execution’s content, immutable org/agent
scope, dialect, and provenance on the server. Do not send scope fields or copy
sqlPreview back as executable input. runSavedQuery likewise sends only the
saved id. To edit a structured saved object in SQL, call
convertSavedQueryToSql(id); conversion does not mutate it, and a later
updateSavedQuery preserves scope and the server-derived PostgreSQL dialect
pin.
BI result exports
Section titled “BI result exports”Export a completed result without rerunning it. Create once, retain the export handle, and poll it; each ready response contains a newly signed URL:
const { export: started } = await jc.bi.createExport(query.jobId, { format: 'parquet',});const { export: current } = await jc.bi.getExport(started.id);
if (current.status === 'ready' && current.downloadUrl) { const response = await fetch(current.downloadUrl); // no Authorization header const bytes = await response.arrayBuffer();}Creation requires bi:write; polling requires bi:read. The default admission
limits are 100,000 rows and 256 MiB of stored source JSONL. Recovery has a
finite attempt budget, and terminal failures retain both the query and export
IDs. Fetch the storage URL directly without the JustCrawl API key. CSV exports
neutralize formula-like headers and cells with a leading apostrophe before CSV
escaping; choose Parquet when exact values must round-trip without that safety
prefix. Result pages expose the server-derived exportEligibility decision,
including caller permission and both admission-cap checks.
Waiting for jobs
Section titled “Waiting for jobs”const job = await jc.jobs.submitAndWait( { urlItemId }, { maxWaitMs: 120_000, signal: controller.signal, onPoll: (j) => console.log(j.status) },);Polling in chunks with exponential backoff — there is no streaming interface,
because the API exposes none. submitAndWait resolves on both completed
and failed; check job.status.
Creating URLs and schedules
Section titled “Creating URLs and schedules”The SDK exposes the published request shapes directly. The single-URL handler
currently consumes url and optional priority; do not send tags on this
endpoint. Schedule creation can use a published workflow or Automatic routing,
and enablement is explicit desired state:
const target = await jc.urls.create({ url: 'https://example.com/products/123', priority: 10,});
const schedule = await jc.schedules.create({ name: 'Six-hour product sweep', workflowId: null, // Automatic routing per URL frequency: 'every_6_hours', timezone: 'UTC', tagFilters: [], // every enabled URL in the organization isEnabled: false, // configure first; enable when ready});
await jc.schedules.toggle(schedule.id, { isEnabled: true });URL creation requires urls:write. Schedule creation and enable/disable require
schedules:write plus a verified account email. Every lookup and mutation is
scoped to the API key’s organization. Repeating either write is not treated as
idempotent: ordinary POST and PATCH calls are never retried by the SDK.
Results
Section titled “Results”const result = await jc.jobs.fetchResult(jobId);
switch (result.kind) { case 'url': return result.data; // fetched from a presigned URL for you case 'blob': return result.blobKey; // your bucket — fetch with your own creds case 'expired': return result.expiredAt; // aged out of your retention window case 'refunded': return result.message; // no provider delivered a page; the job was free}expired and refunded are outcomes, not exceptions — neither is a fault on
your side, so neither throws. They mean different things: expired is a page
that existed and was swept by your retention policy, refunded is a page that
was never delivered at all, so the job’s credit came back and no retry of the
fetch will ever produce a body. The job resource says the same thing in its
creditOutcome / creditNote fields.
The presigned URL is fetched with no Authorization header — your key is
never sent to the storage host.
Errors
Section titled “Errors”import { JustCrawlError } from '@justcrawl/sdk';
try { await jc.jobs.submit({ urlItemId });} catch (err) { if (err instanceof JustCrawlError) { err.status; // HTTP status err.code; // machine-readable code err.inferredCode; // true when the SDK derived the code from the status err.requestId; // quote this in support requests err.raw; // the untouched response body }}inferredCode is worth understanding: not every endpoint sends a machine code
yet. When one doesn’t, the SDK derives a sensible code from the HTTP status and
sets this flag — it never presents a guess as something we said. Branch on
err.code only after checking it, or branch on err.status.
Recover a lost BI submission
Section titled “Recover a lost BI submission”Keep your own UUID in the Idempotency-Key header when you need to recover
from a lost response after all SDK retries. Then use
jc.bi.listQueries({ submissionKey }) to find the retained handle without
submitting again. This exact lookup is tenant-scoped and remains available
after newer queries fill recent history; a missing key returns an empty list.
Retries
Section titled “Retries”GET requests retry on network errors, 5xx, and 429, honoring Retry-After
when present and falling back to exponential backoff with jitter when it isn’t.
Ordinary writes are never retried. Submitting a job twice charges two
credits. bi.runQuery(), bi.runStructuredQuery(), and
bi.runSavedQuery() are the endpoint-specific exceptions: each generates an
Idempotency-Key UUID (or preserves the caller’s header) and reuses that key
for every attempt, so a lost response resolves to the original BI query rather
than a duplicate. A crawl that fails at every provider has its credit returned
automatically, but a retried submit that succeeds is a second real charge, so
the refund is no substitute for the dedup pattern. 402 Insufficient credits
is never retried either — it needs a human to add credits, not another request.
Other languages
Section titled “Other languages”There is no first-class SDK for other languages yet. The API is plain REST with bearer auth, and our OpenAPI spec is published — most generators produce a workable client from it directly.
If you’re writing your own client, three rules are worth copying from ours:
- Do not retry ordinary writes. A retried
POST /jobscharges twice. BI query submission andPOST /api/v1/bi/saved-queries/{id}/runare the exceptions, and only when every attempt reuses the same UUIDIdempotency-Key. - Retry
429and honorRetry-Afteron safe methods. - Never attach your API key when fetching a presigned
resultUrl— it leaks the key into storage access logs, and the request is rejected anyway.
Model Context Protocol — @justcrawl/mcp-server
Section titled “Model Context Protocol — @justcrawl/mcp-server”If you want an agent to use justcrawl rather than your code, reach for the MCP server instead of this SDK. It is a local stdio process your agent host spawns, and it exposes twenty-nine tools over the same API — a guided scrape that builds the routing workflow for you, raw job submission and inspection, browsing workflows, creating and browsing schedules and URLs, discovering the authorized BI catalog, running safe structured or saved BI queries, creating and polling CSV/Parquet exports, and searching these docs.
{ "mcpServers": { "justcrawl": { "command": "npx", "args": ["-y", "@justcrawl/mcp-server"], "env": { "JUSTCRAWL_API_KEY": "sr_live_your_key_here" } } }}- npm:
@justcrawl/mcp-server - Setup, per host: Agents & MCP — Claude Desktop, Cursor, Codex CLI
- Capability policy: every SDK resource method mapped or explicitly omitted
It is built on this SDK, so the three rules above hold there too; the server
inherits them rather than reimplementing them. One rule is its own: it rejects
any sql argument on any tool. jc_bi_query accepts only structured aggregate
intent, jc_bi_save_query can save a completed structured run, and new
jc_bi_run_saved calls take the id of either an MCP-created or dashboard-created
saved query. Its optional queryId continuation alias remains only until the
next MCP npm major; jc_bi_get_results is the supported continuation tool and resumes either returned query
handle without resubmitting it. jc_bi_create_export and jc_bi_get_export
preserve one export handle through finite recovery and return the fresh storage
URL without fetching it. A prompt injection hidden in a scraped page therefore
cannot turn into a query over your data or receive your API credential through
an object-store request.