Skip to content

Agents & MCP

The justcrawl MCP server gives any LLM-powered agent (Claude Desktop, Cursor, Codex CLI) the ability to submit scrape jobs, list workflows, run safe structured or saved BI queries, and search docs — all from inside the chat.

Each MCP host has a different config file and process model. Pick yours:

The server exposes these tools to the agent:

Tool What it does
jc_scrape Scrape a URL: builds or reuses a multi-vendor workflow for its domain, starts the job, and returns the workflow diagram
jc_scrape_result Wait for a jc_scrape job and return the page content, plus extracted fields once the workflow has them
jc_scrape_extract Name fields — price, title, rating — to pull from this domain’s pages from now on, and get them for the page just scraped
jc_jobs_submit Submit a single scrape job for a URL
jc_jobs_submit_and_wait Submit a job and wait for its result in one call — the un-guided path for scripted or batch work
jc_jobs_list Paginated job list, filter by status or workflow
jc_jobs_get Fetch a job + its result (auto-resolves the presigned URL when done)
jc_workflows_list / jc_workflows_get Browse the workflows in your org
jc_providers_list List the providers the platform supports and what each can do
jc_urls_create Add one URL to your library with an optional dispatch priority
jc_urls_list List URLs in your library, filter by tag or search
jc_bi_get_schema Discover only the BI tables authorized for your organization, with a versioned continuation
jc_bi_get_table Describe one discovered table’s columns, types, and available org-scoped samples
jc_bi_query Run a bounded structured aggregate over discovered tables and return resumable results plus display-only SQL
jc_bi_list_saved_queries List BI saved queries available to your organization
jc_bi_save_query Save a completed jc_bi_query execution by query id while preserving its server-owned scope
jc_bi_run_saved Run a saved query by id and return its resumable query handle
jc_bi_get_results Read or continue one materialized query without running it again
jc_bi_create_export Create or reuse a CSV or Parquet export for one completed query
jc_bi_get_export Poll an export handle and return a fresh presigned download URL when ready
jc_bi_cancel_query Request cancellation and wait briefly for the terminal state
jc_bi_list_queries List recent handles or recover one by exact submissionKey
jc_schedules_create Create a recurring schedule, optionally disabled or using Automatic routing
jc_schedules_list / jc_schedules_trigger List schedules and trigger a one-off run
jc_schedules_set_enabled Set a schedule explicitly enabled or disabled
jc_docs_search / jc_docs_get Search and read justcrawl.io docs

Twelve tools do something other than read: jc_scrape, jc_scrape_extract, jc_jobs_submit, jc_jobs_submit_and_wait, jc_urls_create, jc_schedules_create, jc_schedules_set_enabled, jc_schedules_trigger, jc_bi_save_query (it creates a reusable saved object from a completed query), jc_bi_run_saved (it starts a run of a query you already saved, spending a BI concurrency slot), jc_bi_create_export (it starts or reuses bounded export work), and jc_bi_cancel_query. Everything else is read-only, and nothing here can delete a workflow, a URL, or a schedule.

jc_urls_create accepts only the URL and optional priority; the single-create handler does not support tags. An enabled jc_schedules_create schedule keeps queuing real jobs on its cadence and spending credits. Empty or omitted tag filters select every enabled URL in your organization, while a null or omitted workflow uses Automatic routing. jc_schedules_set_enabled sends the desired boolean state rather than flipping blindly. The API still enforces urls:write / schedules:write, organization isolation, and verified email for schedule writes even if a host ignores tool annotations.

The package’s capability policy records the MCP mapping or omission reason for every public SDK resource method.

Use jc_bi_get_schema before jc_bi_get_table or jc_bi_query; do not guess relation names. Both tools return bounded continuations and a catalogVersion. If the version changes mid-traversal, restartRequired tells the host to restart discovery so one answer never mixes two catalog snapshots.

Two of those write more than a job. jc_scrape leaves a real workflow behind: the first scrape of a host creates a published workflow routed to that domain in your organization, visible in the dashboard, and every later scrape of the host — from the agent, the API, or the UI — runs through it. If you already have an unpublished workflow routed to that domain, that one is rebuilt from the standard template and published rather than left alone, so edits you had not published yet are replaced; the tool says so in its answer. jc_scrape_extract then attaches an extractor to that same workflow — and if the domain has no extractor yet, it rebuilds the workflow from the standard template, so any hand-edits you made to it in the dashboard are replaced. This persistence is the point rather than a side effect: an agent scrape leaves you with something you can open, edit, and schedule. If you would rather it did not, use jc_jobs_submit_and_wait, which submits against your existing routing and creates nothing — though extracting fields afterwards still builds the domain workflow, whichever tool ran the scrape.

Continue BI results without rerunning them

Section titled “Continue BI results without rerunning them”

jc_bi_query and jc_bi_run_saved return a queryId whenever the submit response arrives, whether the query completes immediately, finishes after polling, fails, is canceled, times out, or succeeds while its stored result is temporarily unreadable. Pass that handle to jc_bi_get_results; submitting the saved query again starts a second execution. The underlying REST submit uses one idempotency key across transport retries. If every response attempt is lost, the tool returns submissionStatus: "indeterminate", the exact submissionKey, and a jc_bi_list_queries next step instead of guessing a handle. The next step passes submissionKey for an exact lookup, so more than ten newer queries cannot hide the accepted run. Without that filter, history returns up to ten recent handles. Neither mode resubmits work. Query history also includes each handle’s label and timestamp while deliberately omitting SQL text.

Result pages are zero-based. jc_bi_get_results fetches 50 rows by default (maximum 200) and may stop partway through that logical page to stay within the agent response budget. Follow the returned continuation exactly: when only rowOffset changes, keep the same page. truncatedCells and omittedRows make partial data explicit. Results above 1,000 rows direct the agent to jc_bi_create_export only when the server confirms the caller has bi:write and the source is within both export caps. Otherwise the next step remains a callable jc_bi_get_results continuation, or static refine-and-rerun guidance after the last page.

Call jc_bi_create_export once with a completed queryId and a csv or parquet format. Keep the returned exportId in every state and poll that exact handle with jc_bi_get_export; do not rerun the query or create another export while it is queued or running. Recovery is finite, so a terminal failure retains both IDs and a stable failure code rather than polling forever.

CSV exports neutralize formula-like headers and cells with a leading apostrophe before CSV escaping. Choose Parquet when exact values must round-trip without that safety prefix.

When the export is ready, jc_bi_get_export returns a newly signed object-store URL and its lifetime. The MCP server does not fetch it and never forwards the JustCrawl API key to storage. Download it directly without an Authorization header before it expires.

The in-site Agent Playground can turn a completed or resumed analysis into the same saved-query object used by the SQL console, schedules, public API, SDK and MCP. These are the tools involved inside the Agent conversation:

Agent tool What it does
queryAgentResults Runs one agent-scoped analysis or resumes its durable queryId without rerunning it
offerSaveQuery Verifies that completed query belongs to the current Agent conversation, then shows an immutable proposal with its name, description and provenance
saveAgentQuery After a genuine Save-button click, reloads the trusted offer and creates the saved query; its only input is the offer card id
resolveCard Records a typed decline or replacement-name proposal; it refuses typed acceptance of a save offer

A typical flow is:

  1. Ask the Agent to analyze the data, or to continue an earlier query by its queryId.
  2. Ask to retain a useful result. The Agent shows a Save-query card with the proposed name and original query/Agent provenance.
  3. Click Save. Typing “yes”, quoting an old acceptance, or asking the model to call the save tool does not authorize it. A completion wake or resumed background turn cannot click for you.
  4. Follow Open saved query. The link restores the correct organization, source result, saved object and Agent scope in the SQL console.

If the proposed name is already used, the failed proposal stays unchanged. Reply with a new name; the Agent creates a new offer, and you click Save again. Declining ends that offer. Reconnects and duplicate delivery do not create another saved object, and deleting the saved query does not make its old offer reusable.

This differs from MCP’s jc_bi_save_query: MCP sends a completed query id directly under API-key authorization, while the in-site Agent adds an explicit human confirmation card. Both paths copy the source query’s server-owned content and immutable organization/Agent scope; neither accepts scope or protected SQL from the model.

  1. A justcrawl.io API key. Settings → API Keys in the dashboard. Copy the token (it starts with sr_live_); it’s only shown once.
  2. Node.js 20.3 or newer on the machine running the agent — the MCP server runs as a local stdio process via npx. (node --version to check. On 20.0–20.2 the package installs and then fails at runtime, so the minor version matters.)
  3. The host you’re configuring (Claude Desktop / Cursor / Codex CLI) installed and running.

Pick a host card above. Every page is one config snippet, one restart, and a quick sanity check.