Skip to content

BI / SQL

The BI endpoints let you run read-only SQL over your own scraped data, save queries under a name, and export results as CSV or Parquet. This page covers what you need before your first call; the API Reference carries the full request and response shapes.

Before you start: BI is entitled per organization

Section titled “Before you start: BI is entitled per organization”
Method Path What it does
GET /api/v1/bi/schema Tables you can query
GET /api/v1/bi/tables/{name} Columns of one table
POST /api/v1/bi/queries Submit SQL
GET /api/v1/bi/queries Query history
GET /api/v1/bi/queries/{id} Query status
GET /api/v1/bi/queries/{id}/result-manifest Row count, page count, data freshness
GET /api/v1/bi/queries/{id}/results One page of rows
POST /api/v1/bi/queries/{id}/cancel Cancel a running query
GET /api/v1/bi/saved-queries List saved queries
POST /api/v1/bi/saved-queries Save a query
GET /api/v1/bi/saved-queries/{id} Load one saved query
PATCH /api/v1/bi/saved-queries/{id} Update a saved query
DELETE /api/v1/bi/saved-queries/{id} Delete a saved query
POST /api/v1/bi/saved-queries/{id}/run Replay stored content and scope
POST /api/v1/bi/saved-queries/{id}/convert-to-sql Preview editable PostgreSQL for a structured saved query
POST /api/v1/bi/queries/{id}/exports Request a CSV or Parquet export
GET /api/v1/bi/exports/{exportId} Poll an export, get its download link

Each successful /results page includes exportEligibility. The gateway derives it from the current caller’s bi:write permission and the materialized result’s recorded row and byte caps, so clients do not need to predict whether the export endpoint will accept that query.

CSV exports are safe by default for spreadsheet use: headers and cells starting with =, +, -, @, a tab, or a carriage return receive a leading apostrophe before CSV escaping. Choose Parquet when those exact values must round-trip without the safety prefix.

Two BI surfaces are deliberately not part of the API contract: schedule management (/api/v1/bi/query-schedules) and share-link mint/revoke (/api/v1/bi/saved-queries/{id}/share). They power the dashboard, but they are not published, so do not build against them — they can change without notice.

Create a saved query with exactly one source: literal sql, or the id of a completed query as sourceQueryId. The source form copies the execution’s server-owned content, org/agent scope, dialect and provenance. Scope fields are never accepted from the request.

Replay through POST /api/v1/bi/saved-queries/{id}/run. The server resolves the stored object and copies its immutable scope into the new execution; clients must not reconstruct a run from sql or the display-only sqlPreview. Agent-scoped raw SQL and structured definitions return sql: null, so list, get, MCP and share consumers cannot mistake preview text for a portable, org-scoped execution input.

The dashboard can request an editable PostgreSQL rendering of a structured saved query through POST /api/v1/bi/saved-queries/{id}/convert-to-sql. Conversion itself does not mutate the saved object. A later update can persist the rendered SQL, but keeps the original scope and pins the saved object to PostgreSQL so threshold routing cannot reinterpret its dialect.

A submit waits briefly for the query to finish. If it completes in that window, the first page of rows comes back in the same response; otherwise you get a jobId to poll.

Terminal window
curl -X POST 'https://api.justcrawl.io/api/v1/bi/queries' \
-H 'Authorization: Bearer $JUSTCRAWL_API_KEY' \
-H 'Idempotency-Key: 123e4567-e89b-42d3-a456-426614174000' \
-H 'Content-Type: application/json' \
-d '{"sql":"SELECT status, count(*) FROM jobs GROUP BY 1"}'

Use one UUID idempotency key per intended run. If the connection drops after the gateway accepted the query, repeat the same body and key: the response returns the original jobId with replayed: true, without another queue entry. The official TypeScript SDK creates and reuses this key automatically inside each bi.runQuery() call. Query-history rows include submissionKey, label, and submittedAt so a caller can correlate a recovered handle.

Poll GET /api/v1/bi/queries/{id} until status is success, failed, or canceled, then read pages from GET /api/v1/bi/queries/{id}/results.

One shape difference is worth knowing before it bites you: the page object comes back bare from /results, but nested under results in the submit response. Unwrapping one as the other silently yields undefined.

Unlike the rest of the API — where the structured error shape is opt-in via an Accept header — BI always emits the structured envelope, on every error:

{
"error": {
"code": "not_ready",
"message": "Query is running",
"docs_url": "https://docs.justcrawl.io/guides/errors/bi",
"request_id": "req_01HF3JKB9Q2W4X"
},
"status": "running"
}

Match on code, never on message. Note that some errors carry extra fields beside error rather than inside it — status on not_ready, rowCount/maxRowCount on result_too_large — while has_active_schedules puts its schedules list inside.

Every code, its status, and what to do about it: BI error codes.

Two conventions apply across the whole surface:

  • 404, not 403, for anything you do not own. A resource id belonging to another organization returns 404, identical to one that never existed, so cross-tenant ids cannot be probed.
  • feature_disabled before everything. The entitlement check runs ahead of every handler, so it is the first thing an unentitled integration sees on any route.

Your organization has a cap on concurrent and queued BI queries. Exceeding it returns 429 with code too_many_queries.

Both of these grant access without any authentication at all. Anyone who obtains one can use it.

shareToken — returned by every single-row saved-query response when a share link exists: GET /api/v1/bi/saved-queries/{id}, POST /api/v1/bi/saved-queries (create), and PATCH /api/v1/bi/saved-queries/{id} (update) all echo the live token back, not just the fields you changed — a PATCH that only renames a query still returns it. It resolves that query’s results publicly, outside authentication and outside the BI feature flag. Never log it, never put it in an agent transcript, and never fan it out. Revoked via DELETE /api/v1/bi/saved-queries/{id}/share — a working route, deliberately unpublished — never dashboard-only.

This is exactly why the list endpoint does not return it: GET /api/v1/bi/saved-queries gives each row an isShared boolean instead, so listing your saved queries can never dump every share token your organization holds. Fetch the single query when you genuinely need the link.

downloadUrl — returned by GET /api/v1/bi/exports/{exportId} once the export is ready. It is a presigned URL to the export bytes, valid for downloadUrlExpiresInSec. It is re-issued on every poll, so fetch it when you need it rather than storing or forwarding it.

Terminal window
curl -X POST 'https://api.justcrawl.io/api/v1/bi/queries/$JOB_ID/exports' \
-H 'Authorization: Bearer $JUSTCRAWL_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"format":"csv"}'

The status code tells you whether work started: 202 means a new export was enqueued (reused: false), 200 means an existing queued, running, or ready export for the same query and format came back instead (reused: true). Repeated requests therefore collapse onto one job rather than fanning out. A previously failed export is never reused, so asking again after a failure starts a fresh attempt.

Exports are capped by row count; a larger result returns 400 result_too_large with rowCount and maxRowCount so you can see by how much. Narrow the query and re-run it before exporting.