Agents & MCP
The justcrawl MCP server gives any LLM-powered agent (Claude Desktop, Cursor, Codex CLI) the ability to submit scrape jobs, list workflows, run safe structured or saved BI queries, and search docs — all from inside the chat.
Which host are you on?
Section titled “Which host are you on?”Each MCP host has a different config file and process model. Pick yours:
What you can do once it’s wired
Section titled “What you can do once it’s wired”The server exposes these tools to the agent:
| Tool | What it does |
|---|---|
jc_scrape |
Scrape a URL: builds or reuses a multi-vendor workflow for its domain, starts the job, and returns the workflow diagram |
jc_scrape_result |
Wait for a jc_scrape job and return the page content, plus extracted fields once the workflow has them |
jc_scrape_extract |
Name fields — price, title, rating — to pull from this domain’s pages from now on, and get them for the page just scraped |
jc_jobs_submit |
Submit a single scrape job for a URL |
jc_jobs_submit_and_wait |
Submit a job and wait for its result in one call — the un-guided path for scripted or batch work |
jc_jobs_list |
Paginated job list, filter by status or workflow |
jc_jobs_get |
Fetch a job + its result (auto-resolves the presigned URL when done) |
jc_workflows_list / jc_workflows_get |
Browse the workflows in your org |
jc_providers_list |
List the providers the platform supports and what each can do |
jc_urls_create |
Add one URL to your library with an optional dispatch priority |
jc_urls_list |
List URLs in your library, filter by tag or search |
jc_bi_get_schema |
Discover only the BI tables authorized for your organization, with a versioned continuation |
jc_bi_get_table |
Describe one discovered table’s columns, types, and available org-scoped samples |
jc_bi_query |
Run a bounded structured aggregate over discovered tables and return resumable results plus display-only SQL |
jc_bi_list_saved_queries |
List BI saved queries available to your organization |
jc_bi_save_query |
Save a completed jc_bi_query execution by query id while preserving its server-owned scope |
jc_bi_run_saved |
Run a saved query by id and return its resumable query handle |
jc_bi_get_results |
Read or continue one materialized query without running it again |
jc_bi_create_export |
Create or reuse a CSV or Parquet export for one completed query |
jc_bi_get_export |
Poll an export handle and return a fresh presigned download URL when ready |
jc_bi_cancel_query |
Request cancellation and wait briefly for the terminal state |
jc_bi_list_queries |
List recent handles or recover one by exact submissionKey |
jc_schedules_create |
Create a recurring schedule, optionally disabled or using Automatic routing |
jc_schedules_list / jc_schedules_trigger |
List schedules and trigger a one-off run |
jc_schedules_set_enabled |
Set a schedule explicitly enabled or disabled |
jc_docs_search / jc_docs_get |
Search and read justcrawl.io docs |
Twelve tools do something other than read: jc_scrape, jc_scrape_extract, jc_jobs_submit, jc_jobs_submit_and_wait, jc_urls_create, jc_schedules_create, jc_schedules_set_enabled, jc_schedules_trigger, jc_bi_save_query (it creates a reusable saved object from a completed query), jc_bi_run_saved (it starts a run of a query you already saved, spending a BI concurrency slot), jc_bi_create_export (it starts or reuses bounded export work), and jc_bi_cancel_query. Everything else is read-only, and nothing here can delete a workflow, a URL, or a schedule.
jc_urls_create accepts only the URL and optional priority; the single-create
handler does not support tags. An enabled jc_schedules_create schedule keeps
queuing real jobs on its cadence and spending credits. Empty or omitted tag
filters select every enabled URL in your organization, while a null or omitted
workflow uses Automatic routing. jc_schedules_set_enabled sends the desired
boolean state rather than flipping blindly. The API still enforces
urls:write / schedules:write, organization isolation, and verified email for
schedule writes even if a host ignores tool annotations.
The package’s capability policy records the MCP mapping or omission reason for every public SDK resource method.
Use jc_bi_get_schema before jc_bi_get_table or jc_bi_query; do not guess relation names.
Both tools return bounded continuations and a catalogVersion. If the version
changes mid-traversal, restartRequired tells the host to restart discovery so
one answer never mixes two catalog snapshots.
Two of those write more than a job. jc_scrape leaves a real workflow behind: the first scrape of a host creates a published workflow routed to that domain in your organization, visible in the dashboard, and every later scrape of the host — from the agent, the API, or the UI — runs through it. If you already have an unpublished workflow routed to that domain, that one is rebuilt from the standard template and published rather than left alone, so edits you had not published yet are replaced; the tool says so in its answer. jc_scrape_extract then attaches an extractor to that same workflow — and if the domain has no extractor yet, it rebuilds the workflow from the standard template, so any hand-edits you made to it in the dashboard are replaced. This persistence is the point rather than a side effect: an agent scrape leaves you with something you can open, edit, and schedule. If you would rather it did not, use jc_jobs_submit_and_wait, which submits against your existing routing and creates nothing — though extracting fields afterwards still builds the domain workflow, whichever tool ran the scrape.
Continue BI results without rerunning them
Section titled “Continue BI results without rerunning them”jc_bi_query and jc_bi_run_saved return a queryId whenever the submit response arrives,
whether the query completes immediately, finishes after polling, fails, is
canceled, times out, or succeeds while its stored result is temporarily
unreadable. Pass that handle to jc_bi_get_results; submitting the saved query
again starts a second execution. The underlying REST submit uses one
idempotency key across transport retries. If every response attempt is lost,
the tool returns submissionStatus: "indeterminate", the exact
submissionKey, and a jc_bi_list_queries next step instead of guessing a
handle. The next step passes submissionKey for an exact lookup, so more
than ten newer queries cannot hide the accepted run. Without that filter,
history returns up to ten recent handles. Neither mode resubmits work. Query history also includes each handle’s label and timestamp
while deliberately omitting SQL text.
Result pages are zero-based. jc_bi_get_results fetches 50 rows by default
(maximum 200) and may stop partway through that logical page to stay within the
agent response budget. Follow the returned continuation exactly: when only
rowOffset changes, keep the same page. truncatedCells and omittedRows
make partial data explicit. Results above 1,000 rows direct the agent to
jc_bi_create_export only when the server confirms the caller has bi:write
and the source is within both export caps. Otherwise the next step remains a
callable jc_bi_get_results continuation, or static refine-and-rerun guidance
after the last page.
Export large BI results
Section titled “Export large BI results”Call jc_bi_create_export once with a completed queryId and a csv or
parquet format. Keep the returned exportId in every state and poll that
exact handle with jc_bi_get_export; do not rerun the query or create another
export while it is queued or running. Recovery is finite, so a terminal failure
retains both IDs and a stable failure code rather than polling forever.
CSV exports neutralize formula-like headers and cells with a leading apostrophe before CSV escaping. Choose Parquet when exact values must round-trip without that safety prefix.
When the export is ready, jc_bi_get_export returns a newly signed object-store
URL and its lifetime. The MCP server does not fetch it and never forwards the
JustCrawl API key to storage. Download it directly without an Authorization
header before it expires.
Save an Agent Playground analysis
Section titled “Save an Agent Playground analysis”The in-site Agent Playground can turn a completed or resumed analysis into the same saved-query object used by the SQL console, schedules, public API, SDK and MCP. These are the tools involved inside the Agent conversation:
| Agent tool | What it does |
|---|---|
queryAgentResults |
Runs one agent-scoped analysis or resumes its durable queryId without rerunning it |
offerSaveQuery |
Verifies that completed query belongs to the current Agent conversation, then shows an immutable proposal with its name, description and provenance |
saveAgentQuery |
After a genuine Save-button click, reloads the trusted offer and creates the saved query; its only input is the offer card id |
resolveCard |
Records a typed decline or replacement-name proposal; it refuses typed acceptance of a save offer |
A typical flow is:
- Ask the Agent to analyze the data, or to continue an earlier query by its
queryId. - Ask to retain a useful result. The Agent shows a Save-query card with the proposed name and original query/Agent provenance.
- Click Save. Typing “yes”, quoting an old acceptance, or asking the model to call the save tool does not authorize it. A completion wake or resumed background turn cannot click for you.
- Follow Open saved query. The link restores the correct organization, source result, saved object and Agent scope in the SQL console.
If the proposed name is already used, the failed proposal stays unchanged. Reply with a new name; the Agent creates a new offer, and you click Save again. Declining ends that offer. Reconnects and duplicate delivery do not create another saved object, and deleting the saved query does not make its old offer reusable.
This differs from MCP’s jc_bi_save_query: MCP sends a completed query id
directly under API-key authorization, while the in-site Agent adds an explicit
human confirmation card. Both paths copy the source query’s server-owned
content and immutable organization/Agent scope; neither accepts scope or
protected SQL from the model.
What you’ll need before you start
Section titled “What you’ll need before you start”- A justcrawl.io API key. Settings → API Keys in the dashboard. Copy the token (it starts with
sr_live_); it’s only shown once. - Node.js 20.3 or newer on the machine running the agent — the MCP server runs as a local stdio process via
npx. (node --versionto check. On 20.0–20.2 the package installs and then fails at runtime, so the minor version matters.) - The host you’re configuring (Claude Desktop / Cursor / Codex CLI) installed and running.
Pick a host card above. Every page is one config snippet, one restart, and a quick sanity check.