Skip to content

API Quickstart

From API key to scraped HTML in five minutes. The TypeScript SDK is the fastest path; the raw curl flow underneath it is fully supported and documented below.

Go to Settings > API Keys in the dashboard and create a new key. Copy the token — it’s only shown once. Keys look like sr_live_….

Terminal window
npm install @justcrawl/sdk

Zero runtime dependencies, Node 20.3+, ESM and CommonJS, fully typed from our OpenAPI spec. Source and license: Apache-2.0 on GitHub.

import { JustCrawl } from '@justcrawl/sdk';
const jc = new JustCrawl({ apiKey: process.env.JUSTCRAWL_API_KEY! });
// Jobs run against a URL item, so register the target first.
const { id: urlItemId } = await jc.urls.create({ url: 'https://example.com' });
// Submit and poll to completion in one call.
const job = await jc.jobs.submitAndWait({ urlItemId }, { maxWaitMs: 120_000 });
if (job.status === 'completed') {
const result = await jc.jobs.fetchResult(String(job.id));
if (result.kind === 'url') console.log(result.data); // the scraped HTML
}

That’s the whole flow. submitAndWait polls with backoff, honors an AbortSignal, and gives up at maxWaitMs with the last status it saw.

Note job.status is checked rather than assumed: a job that reached failed is a result, not an exception, so submitAndWait resolves for both. Statuses are pending, running, waiting_retry, extracting, extraction_done, completed, failed — only the last two are terminal.

Every failure throws a JustCrawlError carrying status, code, and requestId (quote that in support requests):

import { JustCrawlError } from '@justcrawl/sdk';
try {
await jc.jobs.submit({ urlItemId });
} catch (err) {
if (err instanceof JustCrawlError && err.code === 'quota_exceeded') {
console.error(`Out of credits: ${err.remainingCredits} left`);
}
}

The SDK retries safe reads on network errors, 5xx, and 429, honoring Retry-After when the server sends it. Ordinary writes are not retried: submitting a crawl job twice charges twice. The one exception is POST /api/v1/bi/queries: the SDK generates an Idempotency-Key UUID and reuses it for every attempt, so a lost response resolves to the original query. Hand-rolled clients may retry that BI endpoint only when every attempt carries the same UUID key. (A crawl that fails at every provider is refunded automatically — see Billing & Plans — but a duplicate that succeeds is a second charge.)


Everything above is plain REST. Use this if you’re not on Node, or want to see the wire format.

Terminal window
# First register the URL as a URL item.
curl -X POST https://api.justcrawl.io/api/v1/urls \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'
# Then submit a job against it. workflowId is optional — omit it and the
# gateway resolves the workflow from the URL's domain route.
curl -X POST https://api.justcrawl.io/api/v1/jobs \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"urlItemId": "URL_ITEM_UUID"}'

Response (201):

{ "jobId": "job-uuid", "status": "pending" }
Terminal window
curl https://api.justcrawl.io/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_API_KEY"

Poll until status is completed or failed. Back off between polls — a one-second fixed interval on a long job is wasted requests on both sides.

To watch many jobs at once, batch it:

Terminal window
curl -X POST https://api.justcrawl.io/api/v1/jobs/batch-status \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"jobIds": ["job-1", "job-2", "job-3"]}'
Terminal window
curl https://api.justcrawl.io/api/v1/jobs/JOB_ID/result \
-H "Authorization: Bearer YOUR_API_KEY"

This returns JSON metadata, not the HTML itself — exactly one of:

{
"resultUrl": "https://s3.../presigned?X-Amz-Signature=...",
"providerId": "brightdata",
"statusCode": 200,
"bodySize": 51234,
"latencyMs": 1250
}
  • resultUrl — a presigned URL (~15 minute TTL) on platform storage. Fetch it with no Authorization header: the signature is the credential, and sending your API key would leak it into storage access logs.
  • blobKey — a raw object key, when your org delivers to its own S3 bucket. Fetch it with your own credentials.

A 410 Gone here means the result aged out of your organization’s retention window; the body carries expiredAt.

  1. justcrawl resolves which workflow to use for your URL (based on domain routing)
  2. The job enters the SQS queue with per-org fairness sharding
  3. The worker executes the workflow DAG: tries the primary provider, falls back on failure
  4. Results are stored and optionally delivered to your S3 bucket or webhook