API Quickstart
From API key to scraped HTML in five minutes. The TypeScript SDK is the fastest
path; the raw curl flow underneath it is fully supported and documented below.
Step 1: Get your API key
Section titled “Step 1: Get your API key”Go to Settings > API Keys in the dashboard and create a new key. Copy the
token — it’s only shown once. Keys look like sr_live_….
Step 2: Install the SDK
Section titled “Step 2: Install the SDK”npm install @justcrawl/sdkZero runtime dependencies, Node 20.3+, ESM and CommonJS, fully typed from our OpenAPI spec. Source and license: Apache-2.0 on GitHub.
Step 3: Scrape something
Section titled “Step 3: Scrape something”import { JustCrawl } from '@justcrawl/sdk';
const jc = new JustCrawl({ apiKey: process.env.JUSTCRAWL_API_KEY! });
// Jobs run against a URL item, so register the target first.const { id: urlItemId } = await jc.urls.create({ url: 'https://example.com' });
// Submit and poll to completion in one call.const job = await jc.jobs.submitAndWait({ urlItemId }, { maxWaitMs: 120_000 });
if (job.status === 'completed') { const result = await jc.jobs.fetchResult(String(job.id)); if (result.kind === 'url') console.log(result.data); // the scraped HTML}That’s the whole flow. submitAndWait polls with backoff, honors an
AbortSignal, and gives up at maxWaitMs with the last status it saw.
Note job.status is checked rather than assumed: a job that reached failed is
a result, not an exception, so submitAndWait resolves for both. Statuses
are pending, running, waiting_retry, extracting, extraction_done,
completed, failed — only the last two are terminal.
Errors and retries
Section titled “Errors and retries”Every failure throws a JustCrawlError carrying status, code, and
requestId (quote that in support requests):
import { JustCrawlError } from '@justcrawl/sdk';
try { await jc.jobs.submit({ urlItemId });} catch (err) { if (err instanceof JustCrawlError && err.code === 'quota_exceeded') { console.error(`Out of credits: ${err.remainingCredits} left`); }}The SDK retries safe reads on network errors, 5xx, and 429, honoring
Retry-After when the server sends it. Ordinary writes are not retried:
submitting a crawl job twice charges twice. The one exception is
POST /api/v1/bi/queries: the SDK generates an Idempotency-Key UUID and
reuses it for every attempt, so a lost response resolves to the original query.
Hand-rolled clients may retry that BI endpoint only when every attempt carries
the same UUID key. (A crawl that fails at every provider is refunded
automatically — see Billing & Plans — but a duplicate that
succeeds is a second charge.)
The same flow with curl
Section titled “The same flow with curl”Everything above is plain REST. Use this if you’re not on Node, or want to see the wire format.
Submit a job
Section titled “Submit a job”# First register the URL as a URL item.curl -X POST https://api.justcrawl.io/api/v1/urls \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"url": "https://example.com"}'
# Then submit a job against it. workflowId is optional — omit it and the# gateway resolves the workflow from the URL's domain route.curl -X POST https://api.justcrawl.io/api/v1/jobs \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"urlItemId": "URL_ITEM_UUID"}'Response (201):
{ "jobId": "job-uuid", "status": "pending" }Poll for completion
Section titled “Poll for completion”curl https://api.justcrawl.io/api/v1/jobs/JOB_ID \ -H "Authorization: Bearer YOUR_API_KEY"Poll until status is completed or failed. Back off between polls — a
one-second fixed interval on a long job is wasted requests on both sides.
To watch many jobs at once, batch it:
curl -X POST https://api.justcrawl.io/api/v1/jobs/batch-status \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"jobIds": ["job-1", "job-2", "job-3"]}'Get the result
Section titled “Get the result”curl https://api.justcrawl.io/api/v1/jobs/JOB_ID/result \ -H "Authorization: Bearer YOUR_API_KEY"This returns JSON metadata, not the HTML itself — exactly one of:
{ "resultUrl": "https://s3.../presigned?X-Amz-Signature=...", "providerId": "brightdata", "statusCode": 200, "bodySize": 51234, "latencyMs": 1250}resultUrl— a presigned URL (~15 minute TTL) on platform storage. Fetch it with noAuthorizationheader: the signature is the credential, and sending your API key would leak it into storage access logs.blobKey— a raw object key, when your org delivers to its own S3 bucket. Fetch it with your own credentials.
A 410 Gone here means the result aged out of your organization’s retention
window; the body carries expiredAt.
What happens under the hood
Section titled “What happens under the hood”- justcrawl resolves which workflow to use for your URL (based on domain routing)
- The job enters the SQS queue with per-org fairness sharding
- The worker executes the workflow DAG: tries the primary provider, falls back on failure
- Results are stored and optionally delivered to your S3 bucket or webhook
Next steps
Section titled “Next steps”- TypeScript SDK for the full client reference
- Workflows to understand how DAGs route requests
- Schedules to automate recurring scrapes
- API Reference for the full endpoint list