BI / SQL
The BI endpoints let you run read-only SQL over your own scraped data, save queries under a name, and export results as CSV or Parquet. This page covers what you need before your first call; the API Reference carries the full request and response shapes.
Before you start: BI is entitled per organization
Section titled “Before you start: BI is entitled per organization”The published surface
Section titled “The published surface”| Method | Path | What it does |
|---|---|---|
GET |
/api/v1/bi/schema |
Tables you can query |
GET |
/api/v1/bi/tables/{name} |
Columns of one table |
POST |
/api/v1/bi/queries |
Submit SQL |
GET |
/api/v1/bi/queries |
Query history |
GET |
/api/v1/bi/queries/{id} |
Query status |
GET |
/api/v1/bi/queries/{id}/result-manifest |
Row count, page count, data freshness |
GET |
/api/v1/bi/queries/{id}/results |
One page of rows |
POST |
/api/v1/bi/queries/{id}/cancel |
Cancel a running query |
GET |
/api/v1/bi/saved-queries |
List saved queries |
POST |
/api/v1/bi/saved-queries |
Save a query |
GET |
/api/v1/bi/saved-queries/{id} |
Load one saved query |
PATCH |
/api/v1/bi/saved-queries/{id} |
Update a saved query |
DELETE |
/api/v1/bi/saved-queries/{id} |
Delete a saved query |
POST |
/api/v1/bi/saved-queries/{id}/run |
Replay stored content and scope |
POST |
/api/v1/bi/saved-queries/{id}/convert-to-sql |
Preview editable PostgreSQL for a structured saved query |
POST |
/api/v1/bi/queries/{id}/exports |
Request a CSV or Parquet export |
GET |
/api/v1/bi/exports/{exportId} |
Poll an export, get its download link |
Each successful /results page includes exportEligibility. The gateway
derives it from the current caller’s bi:write permission and the materialized
result’s recorded row and byte caps, so clients do not need to predict whether
the export endpoint will accept that query.
CSV exports are safe by default for spreadsheet use: headers and cells starting
with =, +, -, @, a tab, or a carriage return receive a leading
apostrophe before CSV escaping. Choose Parquet when those exact values must
round-trip without the safety prefix.
Two BI surfaces are deliberately not part of the API contract: schedule management
(/api/v1/bi/query-schedules) and share-link mint/revoke
(/api/v1/bi/saved-queries/{id}/share). They power the dashboard, but they are not published, so
do not build against them — they can change without notice.
Saving and replaying without losing scope
Section titled “Saving and replaying without losing scope”Create a saved query with exactly one source: literal sql, or the id of a completed query as
sourceQueryId. The source form copies the execution’s server-owned content, org/agent scope,
dialect and provenance. Scope fields are never accepted from the request.
Replay through POST /api/v1/bi/saved-queries/{id}/run. The server resolves the stored object
and copies its immutable scope into the new execution; clients must not reconstruct a run from
sql or the display-only sqlPreview. Agent-scoped raw SQL and structured definitions return
sql: null, so list, get, MCP and share consumers cannot mistake preview text for a portable,
org-scoped execution input.
The dashboard can request an editable PostgreSQL rendering of a structured saved query through
POST /api/v1/bi/saved-queries/{id}/convert-to-sql. Conversion itself does not mutate the saved
object. A later update can persist the rendered SQL, but keeps the original scope and pins the
saved object to PostgreSQL so threshold routing cannot reinterpret its dialect.
Submitting a query
Section titled “Submitting a query”A submit waits briefly for the query to finish. If it completes in that window, the first page of
rows comes back in the same response; otherwise you get a jobId to poll.
curl -X POST 'https://api.justcrawl.io/api/v1/bi/queries' \ -H 'Authorization: Bearer $JUSTCRAWL_API_KEY' \ -H 'Idempotency-Key: 123e4567-e89b-42d3-a456-426614174000' \ -H 'Content-Type: application/json' \ -d '{"sql":"SELECT status, count(*) FROM jobs GROUP BY 1"}'Use one UUID idempotency key per intended run. If the connection drops after
the gateway accepted the query, repeat the same body and key: the response
returns the original jobId with replayed: true, without another queue entry.
The official TypeScript SDK creates and reuses this key automatically inside
each bi.runQuery() call. Query-history rows include submissionKey, label,
and submittedAt so a caller can correlate a recovered handle.
Poll GET /api/v1/bi/queries/{id} until status is success, failed, or canceled, then read
pages from GET /api/v1/bi/queries/{id}/results.
One shape difference is worth knowing before it bites you: the page object comes back bare
from /results, but nested under results in the submit response. Unwrapping one as the
other silently yields undefined.
Errors
Section titled “Errors”Unlike the rest of the API — where the structured error shape is opt-in via an Accept header —
BI always emits the structured envelope, on every error:
{ "error": { "code": "not_ready", "message": "Query is running", "docs_url": "https://docs.justcrawl.io/guides/errors/bi", "request_id": "req_01HF3JKB9Q2W4X" }, "status": "running"}Match on code, never on message. Note that some errors carry extra fields beside error
rather than inside it — status on not_ready, rowCount/maxRowCount on result_too_large —
while has_active_schedules puts its schedules list inside.
Every code, its status, and what to do about it: BI error codes.
Two conventions apply across the whole surface:
- 404, not 403, for anything you do not own. A resource id belonging to another organization returns 404, identical to one that never existed, so cross-tenant ids cannot be probed.
feature_disabledbefore everything. The entitlement check runs ahead of every handler, so it is the first thing an unentitled integration sees on any route.
Rate limits
Section titled “Rate limits”Your organization has a cap on concurrent and queued BI queries. Exceeding it returns 429 with
code too_many_queries.
Two values to treat as credentials
Section titled “Two values to treat as credentials”Both of these grant access without any authentication at all. Anyone who obtains one can use it.
shareToken — returned by every single-row saved-query response when a share link exists:
GET /api/v1/bi/saved-queries/{id}, POST /api/v1/bi/saved-queries (create), and PATCH /api/v1/bi/saved-queries/{id} (update) all echo the live token back, not just the fields you
changed — a PATCH that only renames a query still returns it. It resolves that query’s results
publicly, outside authentication and outside the BI feature flag. Never log it, never put it in
an agent transcript, and never fan it out. Revoked via DELETE /api/v1/bi/saved-queries/{id}/share — a working route, deliberately unpublished — never
dashboard-only.
This is exactly why the list endpoint does not return it: GET /api/v1/bi/saved-queries gives
each row an isShared boolean instead, so listing your saved queries can never dump every share
token your organization holds. Fetch the single query when you genuinely need the link.
downloadUrl — returned by GET /api/v1/bi/exports/{exportId} once the export is ready. It
is a presigned URL to the export bytes, valid for downloadUrlExpiresInSec. It is re-issued on
every poll, so fetch it when you need it rather than storing or forwarding it.
Exporting results
Section titled “Exporting results”curl -X POST 'https://api.justcrawl.io/api/v1/bi/queries/$JOB_ID/exports' \ -H 'Authorization: Bearer $JUSTCRAWL_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"format":"csv"}'The status code tells you whether work started: 202 means a new export was enqueued
(reused: false), 200 means an existing queued, running, or ready export for the same query
and format came back instead (reused: true). Repeated requests therefore collapse onto one job
rather than fanning out. A previously failed export is never reused, so asking again after a
failure starts a fresh attempt.
Exports are capped by row count; a larger result returns 400 result_too_large with rowCount
and maxRowCount so you can see by how much. Narrow the query and re-run it before exporting.