Skip to content

Get access to an extraction's raw HTML

GET
/api/v1/extraction/results/{id}/raw
Code sample: cURL
curl -X GET 'https://api.justcrawl.io/api/v1/extraction/results/00000000-0000-0000-0000-000000000001/raw' \
-H 'Authorization: Bearer $JUSTCRAWL_API_KEY'

Returns a way to reach the raw HTML the extraction was derived from — useful for debugging a selector that stopped matching, or for re-parsing without re-scraping.

The response shape depends on where the blob lives, so branch on kind:

  • presigned — the blob is in JustCrawl’s bucket; url is a signed S3 link valid for one hour. Fetch it directly, and re-request this endpoint rather than caching the link.
  • key — the blob is in the org’s own bucket; blobKey is the object key to read with your own credentials.
id
required
string format: uuid

Extraction result ID.

Either a signed URL or an object key, depending on which bucket holds the blob.

Media typeapplication/json
One of:
RawBlobPresigned
object
kind
required
string
Allowed values: presigned
url
required

Signed S3 URL, valid for one hour.

string format: uri
Example
{
"kind": "presigned"
}

Missing or invalid authentication token

Media typeapplication/json
object
error
string
Example
{
"error": "Missing or invalid authentication token"
}

Insufficient permissions for this operation

Media typeapplication/json
object
error
string
Example
{
"error": "No organization. Complete onboarding first."
}

Extraction not found, or it has no stored raw blob.

Unexpected server error. Logs and PostHog $exception capture

Media typeapplication/json
object
error
string
Example
{
"error": "Something went wrong"
}