scrapling
v0.2.9-rc.6Scrapling as an iii worker — scrapling::* functions run HTTP / anti-bot / browser fetches, CSS/XPath/regex/adaptive extraction, persistent sessions, and a declarative crawl over the iii bus.
exact versions are immutable; binary and bundle artifacts are digest-pinned.
full markdown
/workers/scrapling.md?version=0.2.9-rc.6. paste it into an llm prompt or pipe it through curl from a worker.install
configuration
- adaptive_storage_path: ./data/scrapling/elements.db
defaults:
headless: true
impersonate: chrome
include_html: false
network_idle: false
proxy:
engine_url: ws://127.0.0.1:49134
max_bulk_concurrency: 5
max_sessions: 8
session_idle_timeout_s: 900dependencies
readme
scrapling
Scrapling as an iii worker. It maps
Scrapling's three fetch tiers and its parsing engine to scrapling::* functions
on the iii bus: fast HTTP with TLS impersonation, a Camoufox anti-bot browser, a
full Playwright/Chromium browser, screenshots, and CSS/XPath/regex/adaptive
extraction.
While connected (and unless turned off) it also injects a usage section into
the agent system prompt via the harness pre-generate hook
(scrapling::inject-guidance). inject_guidance in the scrapling
configuration entry is ON by default; turning it off (the console's config
dialog, or configuration::set) hot-applies with no restart, and the
harness's # Granted functions catalog still advertises the scrapling::*
ids. The guidance stays presence-gated: no
scrapling worker, no prompt text. When enabled, the binding relies on the
engine's recoverable triggers (iii #1962,
engine ≥ 0.21.8): bound before the harness is up, it parks as a pending intent
and activates when the harness registers the trigger type. On older engines the
bind is silently dropped.
Install
iii trigger compose::add worker=scraplingThe worker is a deploy: image Python worker. The image build runs
scrapling install, which downloads the Camoufox and Chromium browsers used by
stealthy-fetch, dynamic-fetch, and screenshot.
Functions
| Function | What it does |
|---|---|
scrapling::fetch |
HTTP get/post/put/delete (curl_cffi, TLS impersonation) |
scrapling::stealthy-fetch |
Camoufox stealth browser — Cloudflare/Turnstile bypass, WebRTC/canvas hardening |
scrapling::dynamic-fetch |
Playwright/Chromium — JS render, waits, XHR capture, CDP |
scrapling::screenshot |
Page screenshot (image content blocks) via a browser fetcher |
scrapling::extract |
Parse HTML with a declarative selector list |
scrapling::css |
One CSS query over provided HTML |
scrapling::xpath |
One XPath query over provided HTML |
scrapling::regex |
Regex over the visible text of provided HTML |
scrapling::find-similar |
An example element + structurally similar elements |
Fetch output
Fetch functions return:
{ "status": 200, "url": "...", "headers": {}, "cookies": {}, "encoding": "utf-8",
"extracted": { "...": "..." }, "html": "<...>" }extracted appears only when selectors are given; html only when
include_html: true. Called with a urls array, the response is
{ "results": [ ...one object per url... ] } (a failed URL yields
{ "url", "error" }, so one bad URL doesn't sink the batch).
Selector contract
selectors (on the fetchers and scrapling::extract) is a list of specs:
[
{ "name": "title", "css": "h1" },
{ "name": "links", "css": "a", "attr": "href", "all": true },
{ "name": "price", "regex": "price (\\d+)" },
{ "name": "card_html", "css": ".card", "html": true }
]- one of
css/xpath/regexper spec attrpulls an attribute;htmlpulls inner HTML; otherwise textall: truereturns every match as a list, else the first match (ornull)
Examples
# HTTP fetch + extract in one call
iii trigger scrapling::fetch --payload '{
"url": "https://example.com",
"selectors": [{ "name": "title", "css": "title" }]
}'
# Anti-bot page
iii trigger scrapling::stealthy-fetch --payload '{
"url": "https://nopecha.com/demo/cloudflare",
"solve_cloudflare": true,
"selectors": [{ "name": "body", "css": "#padded_content a", "all": true }]
}'
# Parse HTML you already have
iii trigger scrapling::extract --payload '{
"html": "<ul><li>a</li><li>b</li></ul>",
"selectors": [{ "name": "items", "css": "li", "all": true }]
}'
# Fetch many URLs at once
iii trigger scrapling::fetch --payload '{ "urls": ["https://a.com", "https://b.com"] }'Config
config.yaml holds operator defaults applied when a call omits the field:
defaults:
impersonate: chrome # HTTP fetcher fingerprint
headless: true # browser fetchers
network_idle: false
proxy: "" # "" = none
include_html: false
max_bulk_concurrency: 5timeout is passed per call, not defaulted here — it means seconds for the
HTTP fetcher and milliseconds for the browser fetchers.
How it maps
| Scrapling | iii |
|---|---|
Fetcher.{get,post,put,delete} |
scrapling::fetch |
StealthyFetcher.fetch |
scrapling::stealthy-fetch |
DynamicFetcher.fetch |
scrapling::dynamic-fetch |
browser page.screenshot() |
scrapling::screenshot |
Selector.css / .xpath / .re / .find_similar |
extract / css / xpath / regex / find-similar |
Boundaries
- Non-JSON Scrapling options (Python
page_action/page_setupcallbacks, proxy rotators, persistent sessions, theSpidercrawl layer) are not exposed. Pass a singleproxystring. - The fetch functions are not agent-callable without human approval (outbound
requests to arbitrary URLs); the pure parsers are. See
iii-permissions.yaml.
api reference (json)
{
"functions": [
{
"description": "BFS-crawl from start_urls (follow same-domain links), extract per page, stream items.",
"metadata": {},
"name": "scrapling::crawl",
"request_schema": {
"properties": {
"allowed_domains": {
"description": "only follow links on these hosts",
"items": {
"type": "string"
},
"type": "array"
},
"concurrency": {
"type": "integer"
},
"css_selector": {
"description": "scope the render to this CSS subtree (e.g. a page's content div)",
"type": "string"
},
"download_delay": {
"description": "seconds to wait between crawl rounds",
"type": "number"
},
"fetcher": {
"enum": [
"http",
"stealthy",
"dynamic"
],
"type": "string"
},
"format": {
"description": "render page body to this format",
"enum": [
"markdown",
"text"
],
"type": "string"
},
"impersonate": {
"type": "string"
},
"include_html": {
"type": "boolean"
},
"main_content_only": {
"description": "strip nav/scripts/hidden before rendering",
"type": "boolean"
},
"max_depth": {
"type": "integer"
},
"max_pages": {
"type": "integer"
},
"same_domain": {
"description": "follow only same-host links (default true)",
"type": "boolean"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"start_urls": {
"items": {
"type": "string"
},
"type": "array"
},
"stream_name": {
"description": "stream to emit items on (default scrapling::crawl)",
"type": "string"
},
"url": {
"description": "single start URL (alternative to start_urls)",
"type": "string"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"items": {
"description": "a small sample of streamed items",
"items": {
"type": "object"
},
"type": "array"
},
"stats": {
"properties": {
"crawled": {
"type": "integer"
},
"errors": {
"type": "integer"
},
"items": {
"type": "integer"
},
"stopped": {
"type": "string"
}
},
"type": "object"
},
"stream": {
"description": "read the full item stream via stream::on with this name + group_id",
"properties": {
"group_id": {
"type": "string"
},
"name": {
"type": "string"
}
},
"type": "object"
}
},
"type": "object"
}
},
{
"description": "One CSS query over HTML; first-or-all; `attr` pulls an attribute else text.",
"metadata": {},
"name": "scrapling::css",
"request_schema": {
"properties": {
"adaptive": {
"description": "relocate elements after a site change via saved identities",
"type": "boolean"
},
"adaptive_domain": {
"description": "page URL/domain that keys saved identities",
"type": "string"
},
"attr": {
"type": "string"
},
"auto_save": {
"description": "save matched identities (defaults on when adaptive)",
"type": "boolean"
},
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"identifier": {
"description": "stable key for the saved element",
"type": "string"
},
"query": {
"type": "string"
}
},
"required": [
"html",
"query"
],
"type": "object"
},
"response_schema": {
"properties": {
"result": {
"items": {
"type": [
"string",
"null"
]
},
"type": [
"array",
"string",
"null"
]
}
},
"type": "object"
}
},
{
"description": "Describe the first css/xpath match: attrs, generated selectors, class list, DOM context.",
"metadata": {},
"name": "scrapling::describe",
"request_schema": {
"properties": {
"html": {
"type": "string"
},
"kind": {
"enum": [
"css",
"xpath"
],
"type": "string"
},
"query": {
"type": "string"
}
},
"required": [
"html",
"query"
],
"type": "object"
},
"response_schema": {
"properties": {
"element": {
"properties": {
"attrs": {
"type": "object"
},
"children": {
"type": "integer"
},
"classes": {
"items": {
"type": "string"
},
"type": "array"
},
"css": {
"type": "string"
},
"full_css": {
"type": "string"
},
"full_xpath": {
"type": "string"
},
"html": {
"type": "string"
},
"parent_tag": {
"type": [
"string",
"null"
]
},
"siblings": {
"type": "integer"
},
"tag": {
"type": "string"
},
"text": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"type": "object"
},
"found": {
"type": "boolean"
}
},
"type": "object"
}
},
{
"description": "Playwright/Chromium fetch: JS render, waits, XHR capture, CDP; extraction + bulk.",
"metadata": {},
"name": "scrapling::dynamic-fetch",
"request_schema": {
"properties": {
"block_ads": {
"type": "boolean"
},
"blocked_domains": {
"items": {
"type": "string"
},
"type": "array"
},
"capture_xhr": {
"type": "string"
},
"cdp_url": {
"type": "string"
},
"cookies": {
"type": "object"
},
"css_selector": {
"description": "scope the render to this CSS subtree (e.g. a page's content div)",
"type": "string"
},
"disable_resources": {
"type": "boolean"
},
"dns_over_https": {
"type": "boolean"
},
"extra_flags": {
"items": {
"type": "string"
},
"type": "array"
},
"extra_headers": {
"type": "object"
},
"format": {
"description": "render page body to this format",
"enum": [
"markdown",
"text"
],
"type": "string"
},
"google_search": {
"type": "boolean"
},
"headless": {
"type": "boolean"
},
"include_html": {
"type": "boolean"
},
"load_dom": {
"type": "boolean"
},
"locale": {
"type": "string"
},
"main_content_only": {
"description": "strip nav/scripts/hidden before rendering",
"type": "boolean"
},
"max_pages": {
"type": "integer"
},
"network_idle": {
"type": "boolean"
},
"proxy": {
"type": "string"
},
"real_chrome": {
"type": "boolean"
},
"retries": {
"type": "integer"
},
"retry_delay": {
"type": "number"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"timeout": {
"description": "milliseconds (browser fetcher)",
"type": "number"
},
"timezone_id": {
"type": "string"
},
"url": {
"type": "string"
},
"urls": {
"items": {
"type": "string"
},
"type": "array"
},
"useragent": {
"type": "string"
},
"wait": {
"description": "extra ms to wait after load",
"type": "number"
},
"wait_selector": {
"type": "string"
},
"wait_selector_state": {
"enum": [
"attached",
"detached",
"visible",
"hidden"
],
"type": "string"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"captured_xhr": {
"items": {
"type": "object"
},
"type": "array"
},
"content": {
"description": "markdown/text render when `format` requested",
"type": "string"
},
"cookies": {
"type": "object"
},
"encoding": {
"type": [
"string",
"null"
]
},
"error": {
"type": "string"
},
"extracted": {
"type": "object"
},
"format": {
"type": "string"
},
"headers": {
"type": "object"
},
"html": {
"type": "string"
},
"results": {
"items": {
"type": "object"
},
"type": "array"
},
"status": {
"type": [
"integer",
"null"
]
},
"url": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "Parse HTML with a selector list (css/xpath/regex, text/attr/html, all-or-first).",
"metadata": {},
"name": "scrapling::extract",
"request_schema": {
"properties": {
"adaptive": {
"description": "relocate elements after a site change via saved identities",
"type": "boolean"
},
"adaptive_domain": {
"description": "page URL/domain that keys saved identities",
"type": "string"
},
"auto_save": {
"description": "save matched identities (defaults on when adaptive)",
"type": "boolean"
},
"html": {
"type": "string"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
}
},
"required": [
"html",
"selectors"
],
"type": "object"
},
"response_schema": {
"properties": {
"extracted": {
"type": "object"
}
},
"type": "object"
}
},
{
"description": "Fast HTTP fetch, TLS impersonation: get/post/put/delete, inline extraction, bulk `urls`.",
"metadata": {},
"name": "scrapling::fetch",
"request_schema": {
"properties": {
"cookies": {
"type": "object"
},
"css_selector": {
"description": "scope the render to this CSS subtree (e.g. a page's content div)",
"type": "string"
},
"data": {
"type": "object"
},
"follow_redirects": {
"type": "boolean"
},
"format": {
"description": "render page body to this format",
"enum": [
"markdown",
"text"
],
"type": "string"
},
"headers": {
"type": "object"
},
"http3": {
"type": "boolean"
},
"impersonate": {
"description": "TLS/UA fingerprint, e.g. 'chrome'",
"type": "string"
},
"include_html": {
"type": "boolean"
},
"json": {
"type": "object"
},
"main_content_only": {
"description": "strip nav/scripts/hidden before rendering",
"type": "boolean"
},
"max_redirects": {
"type": "integer"
},
"method": {
"enum": [
"get",
"post",
"put",
"delete"
],
"type": "string"
},
"params": {
"type": "object"
},
"proxies": {
"description": "per-scheme proxies, e.g. {\"https\": \"http://...\"}",
"type": "object"
},
"proxy": {
"type": "string"
},
"proxy_auth": {
"description": "[user, password]",
"items": {
"type": "string"
},
"type": "array"
},
"retries": {
"type": "integer"
},
"retry_delay": {
"type": "number"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"stealthy_headers": {
"type": "boolean"
},
"timeout": {
"description": "seconds (HTTP fetcher)",
"type": "number"
},
"url": {
"type": "string"
},
"urls": {
"items": {
"type": "string"
},
"type": "array"
},
"verify": {
"type": "boolean"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"captured_xhr": {
"items": {
"type": "object"
},
"type": "array"
},
"content": {
"description": "markdown/text render when `format` requested",
"type": "string"
},
"cookies": {
"type": "object"
},
"encoding": {
"type": [
"string",
"null"
]
},
"error": {
"type": "string"
},
"extracted": {
"type": "object"
},
"format": {
"type": "string"
},
"headers": {
"type": "object"
},
"html": {
"type": "string"
},
"results": {
"items": {
"type": "object"
},
"type": "array"
},
"status": {
"type": [
"integer",
"null"
]
},
"url": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "Find elements by tag/attribute filters (+ optional text regex); BeautifulSoup-style.",
"metadata": {},
"name": "scrapling::find",
"request_schema": {
"properties": {
"attrs": {
"description": "attribute filters, e.g. {\"class\": \"card\"}",
"type": "object"
},
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"limit": {
"type": "integer"
},
"tag": {
"description": "tag name or list of tag names",
"items": {
"type": "string"
},
"type": [
"string",
"array"
]
},
"text_regex": {
"description": "keep only elements whose text matches this regex",
"type": "string"
}
},
"required": [
"html"
],
"type": "object"
},
"response_schema": {
"properties": {
"count": {
"type": "integer"
},
"items": {
"items": {
"properties": {
"attrs": {
"type": "object"
},
"css": {
"type": "string"
},
"html": {
"type": "string"
},
"tag": {
"type": "string"
},
"text": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"description": "Find elements whose visible text matches a regex pattern.",
"metadata": {},
"name": "scrapling::find-by-regex",
"request_schema": {
"properties": {
"case_sensitive": {
"type": "boolean"
},
"clean_match": {
"type": "boolean"
},
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"limit": {
"type": "integer"
},
"pattern": {
"type": "string"
}
},
"required": [
"html",
"pattern"
],
"type": "object"
},
"response_schema": {
"properties": {
"count": {
"type": "integer"
},
"items": {
"items": {
"properties": {
"attrs": {
"type": "object"
},
"css": {
"type": "string"
},
"html": {
"type": "string"
},
"tag": {
"type": "string"
},
"text": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"description": "Find elements whose visible text matches a string (exact or `partial`).",
"metadata": {},
"name": "scrapling::find-by-text",
"request_schema": {
"properties": {
"case_sensitive": {
"type": "boolean"
},
"clean_match": {
"description": "ignore surrounding/collapsing whitespace",
"type": "boolean"
},
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"limit": {
"type": "integer"
},
"partial": {
"description": "match elements that contain the text",
"type": "boolean"
},
"text": {
"type": "string"
}
},
"required": [
"html",
"text"
],
"type": "object"
},
"response_schema": {
"properties": {
"count": {
"type": "integer"
},
"items": {
"items": {
"properties": {
"attrs": {
"type": "object"
},
"css": {
"type": "string"
},
"html": {
"type": "string"
},
"tag": {
"type": "string"
},
"text": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"description": "Structural auto-match: given one example element, return it plus similar elements.",
"metadata": {},
"name": "scrapling::find-similar",
"request_schema": {
"properties": {
"anchor": {
"description": "CSS selector to one example element",
"type": "string"
},
"html": {
"type": "string"
},
"match_text": {
"type": "boolean"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"similarity_threshold": {
"type": "number"
}
},
"required": [
"html",
"anchor"
],
"type": "object"
},
"response_schema": {
"properties": {
"count": {
"type": "integer"
},
"items": {
"items": {
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"description": "Internal pre_generate hook: appends scrapling usage guidance to the agent system prompt. Bound to harness::hook::pre-generate while inject_guidance is on; not called directly.",
"metadata": {
"internal": true
},
"name": "scrapling::inject-guidance",
"request_schema": {
"properties": {
"generate": {
"properties": {
"system_prompt": {
"description": "system prompt assembled so far",
"type": "string"
}
},
"type": "object"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"mutations": {
"properties": {
"system_prompt": {
"description": "full replacement prompt (base + guidance)",
"type": "string"
}
},
"type": "object"
}
},
"required": [
"mutations"
],
"type": "object"
}
},
{
"description": "Internal: reload scrapling settings from the authoritative configuration on change.",
"metadata": {
"internal": true
},
"name": "scrapling::on-config-change",
"request_schema": {
"properties": {},
"type": "object"
},
"response_schema": {
"properties": {
"ok": {
"type": "boolean"
}
},
"required": [
"ok"
],
"type": "object"
}
},
{
"description": "Run a regex over the visible text of provided HTML; `first` returns the first match, else all.",
"metadata": {},
"name": "scrapling::regex",
"request_schema": {
"properties": {
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"pattern": {
"type": "string"
}
},
"required": [
"html",
"pattern"
],
"type": "object"
},
"response_schema": {
"properties": {
"result": {
"items": {
"type": [
"string",
"null"
]
},
"type": [
"array",
"string",
"null"
]
}
},
"type": "object"
}
},
{
"description": "Capture a page screenshot as image content blocks via a browser fetcher (dynamic or stealthy).",
"metadata": {},
"name": "scrapling::screenshot",
"request_schema": {
"properties": {
"fetcher": {
"enum": [
"dynamic",
"stealthy"
],
"type": "string"
},
"format": {
"enum": [
"png",
"jpeg"
],
"type": "string"
},
"full_page": {
"type": "boolean"
},
"headless": {
"type": "boolean"
},
"network_idle": {
"type": "boolean"
},
"proxy": {
"type": "string"
},
"timeout": {
"type": "number"
},
"url": {
"type": "string"
},
"wait_selector": {
"type": "string"
}
},
"required": [
"url"
],
"type": "object"
},
"response_schema": {
"properties": {
"content": {
"description": "image blocks (one per tile, width<=1024/height<=1536) + a text caption",
"items": {
"properties": {
"data": {
"description": "base64 image bytes (image blocks)",
"type": "string"
},
"mime": {
"type": "string"
},
"text": {
"type": "string"
},
"type": {
"enum": [
"image",
"text"
],
"type": "string"
}
},
"required": [
"type"
],
"type": "object"
},
"type": "array"
},
"mime": {
"type": "string"
},
"url": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "Close a session and free its browser/connection.",
"metadata": {},
"name": "scrapling::session-close",
"request_schema": {
"properties": {
"session_id": {
"type": "string"
}
},
"required": [
"session_id"
],
"type": "object"
},
"response_schema": {
"properties": {
"closed": {
"type": "boolean"
}
},
"type": "object"
}
},
{
"description": "Fetch a URL on an open scraping session (reuses its cookies/state); returns page content, shows nothing. For a page the user should see, use browser::sessions::start + browser::navigate.",
"metadata": {},
"name": "scrapling::session-fetch",
"request_schema": {
"properties": {
"css_selector": {
"description": "scope the render to this CSS subtree (e.g. a page's content div)",
"type": "string"
},
"data": {
"type": "object"
},
"format": {
"description": "render page body to this format",
"enum": [
"markdown",
"text"
],
"type": "string"
},
"headers": {
"type": "object"
},
"include_html": {
"type": "boolean"
},
"json": {
"type": "object"
},
"main_content_only": {
"description": "strip nav/scripts/hidden before rendering",
"type": "boolean"
},
"method": {
"enum": [
"get",
"post",
"put",
"delete"
],
"type": "string"
},
"params": {
"type": "object"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"session_id": {
"type": "string"
},
"url": {
"type": "string"
},
"wait_selector": {
"type": "string"
}
},
"required": [
"session_id",
"url"
],
"type": "object"
},
"response_schema": {
"properties": {
"captured_xhr": {
"items": {
"type": "object"
},
"type": "array"
},
"content": {
"description": "markdown/text render when `format` requested",
"type": "string"
},
"cookies": {
"type": "object"
},
"encoding": {
"type": [
"string",
"null"
]
},
"error": {
"type": "string"
},
"extracted": {
"type": "object"
},
"format": {
"type": "string"
},
"headers": {
"type": "object"
},
"html": {
"type": "string"
},
"results": {
"items": {
"type": "object"
},
"type": "array"
},
"status": {
"type": [
"integer",
"null"
]
},
"url": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "List open sessions with their type and idle time.",
"metadata": {},
"name": "scrapling::session-list",
"request_schema": {
"properties": {
"type": {
"description": "filter by type",
"enum": [
"http",
"dynamic",
"stealthy"
],
"type": "string"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"sessions": {
"items": {
"properties": {
"created_at": {
"type": "number"
},
"idle_s": {
"type": "number"
},
"last_used": {
"type": "number"
},
"session_id": {
"type": "string"
},
"type": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"description": "Open a persistent scraping session (HTTP or headless fetcher) whose session_id reuses cookies and state across fetches. Renders nothing the user can see; to open a page in the visible browser use browser::sessions::start.",
"metadata": {},
"name": "scrapling::session-open",
"request_schema": {
"properties": {
"capture_xhr": {
"description": "regex; capture matching XHRs (browser sessions)",
"type": "string"
},
"headers": {
"type": "object"
},
"headless": {
"type": "boolean"
},
"impersonate": {
"type": "string"
},
"proxies": {
"type": "object"
},
"proxy": {
"type": "string"
},
"real_chrome": {
"type": "boolean"
},
"solve_cloudflare": {
"type": "boolean"
},
"timeout": {
"type": "number"
},
"type": {
"description": "session engine",
"enum": [
"http",
"dynamic",
"stealthy"
],
"type": "string"
},
"useragent": {
"type": "string"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"session_id": {
"type": "string"
},
"type": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "Camoufox stealth browser: solves Cloudflare, hardens WebRTC/canvas; extraction + bulk.",
"metadata": {},
"name": "scrapling::stealthy-fetch",
"request_schema": {
"properties": {
"allow_webgl": {
"type": "boolean"
},
"block_ads": {
"type": "boolean"
},
"block_webrtc": {
"type": "boolean"
},
"blocked_domains": {
"items": {
"type": "string"
},
"type": "array"
},
"capture_xhr": {
"type": "string"
},
"cookies": {
"type": "object"
},
"css_selector": {
"description": "scope the render to this CSS subtree (e.g. a page's content div)",
"type": "string"
},
"disable_resources": {
"type": "boolean"
},
"dns_over_https": {
"type": "boolean"
},
"extra_flags": {
"items": {
"type": "string"
},
"type": "array"
},
"extra_headers": {
"type": "object"
},
"format": {
"description": "render page body to this format",
"enum": [
"markdown",
"text"
],
"type": "string"
},
"google_search": {
"type": "boolean"
},
"headless": {
"type": "boolean"
},
"hide_canvas": {
"type": "boolean"
},
"include_html": {
"type": "boolean"
},
"load_dom": {
"type": "boolean"
},
"locale": {
"type": "string"
},
"main_content_only": {
"description": "strip nav/scripts/hidden before rendering",
"type": "boolean"
},
"max_pages": {
"type": "integer"
},
"network_idle": {
"type": "boolean"
},
"proxy": {
"type": "string"
},
"retries": {
"type": "integer"
},
"retry_delay": {
"type": "number"
},
"selectors": {
"items": {
"properties": {
"all": {
"description": "return every match as a list",
"type": "boolean"
},
"attr": {
"description": "extract this attribute instead of text",
"type": "string"
},
"css": {
"type": "string"
},
"html": {
"description": "extract inner HTML instead of text",
"type": "boolean"
},
"name": {
"type": "string"
},
"regex": {
"type": "string"
},
"xpath": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"type": "array"
},
"solve_cloudflare": {
"type": "boolean"
},
"timeout": {
"description": "milliseconds (browser fetcher)",
"type": "number"
},
"timezone_id": {
"type": "string"
},
"url": {
"type": "string"
},
"urls": {
"items": {
"type": "string"
},
"type": "array"
},
"useragent": {
"type": "string"
},
"wait": {
"description": "extra ms to wait after load",
"type": "number"
},
"wait_selector": {
"type": "string"
},
"wait_selector_state": {
"enum": [
"attached",
"detached",
"visible",
"hidden"
],
"type": "string"
}
},
"type": "object"
},
"response_schema": {
"properties": {
"captured_xhr": {
"items": {
"type": "object"
},
"type": "array"
},
"content": {
"description": "markdown/text render when `format` requested",
"type": "string"
},
"cookies": {
"type": "object"
},
"encoding": {
"type": [
"string",
"null"
]
},
"error": {
"type": "string"
},
"extracted": {
"type": "object"
},
"format": {
"type": "string"
},
"headers": {
"type": "object"
},
"html": {
"type": "string"
},
"results": {
"items": {
"type": "object"
},
"type": "array"
},
"status": {
"type": [
"integer",
"null"
]
},
"url": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "Convert HTML to compact Markdown (or text/html); optional CSS scope + main-content clean.",
"metadata": {},
"name": "scrapling::to-markdown",
"request_schema": {
"properties": {
"css_selector": {
"description": "convert only the subtree matching this CSS selector",
"type": "string"
},
"format": {
"enum": [
"markdown",
"text",
"html"
],
"type": "string"
},
"html": {
"type": "string"
},
"main_content_only": {
"description": "strip nav/scripts/hidden nodes first",
"type": "boolean"
}
},
"required": [
"html"
],
"type": "object"
},
"response_schema": {
"properties": {
"content": {
"type": "string"
},
"format": {
"type": "string"
}
},
"type": "object"
}
},
{
"description": "One XPath query over HTML; first-or-all; `attr` pulls an attribute else text.",
"metadata": {},
"name": "scrapling::xpath",
"request_schema": {
"properties": {
"adaptive": {
"description": "relocate elements after a site change via saved identities",
"type": "boolean"
},
"adaptive_domain": {
"description": "page URL/domain that keys saved identities",
"type": "string"
},
"attr": {
"type": "string"
},
"auto_save": {
"description": "save matched identities (defaults on when adaptive)",
"type": "boolean"
},
"first": {
"type": "boolean"
},
"html": {
"type": "string"
},
"identifier": {
"description": "stable key for the saved element",
"type": "string"
},
"query": {
"type": "string"
}
},
"required": [
"html",
"query"
],
"type": "object"
},
"response_schema": {
"properties": {
"result": {
"items": {
"type": [
"string",
"null"
]
},
"type": [
"array",
"string",
"null"
]
}
},
"type": "object"
}
}
],
"triggers": [
{
"description": "Worker lifecycle events emitted by every worker::* op. Subscribe with `operations` / `stages` / `workers` filters.",
"invocation_schema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"definitions": {
"WorkerOperation": {
"description": "Every worker lifecycle operation a subscriber can watch.",
"enum": [
"add",
"remove",
"update",
"start",
"stop",
"clear"
],
"type": "string"
},
"WorkerStage": {
"description": "Lifecycle stage attached to every event. Each [`WorkerOperation`] emits a fixed sequence of these stages.",
"enum": [
"started",
"downloading",
"downloaded",
"removing",
"updating",
"starting",
"stopping",
"clearing",
"done",
"failed"
],
"type": "string"
}
},
"description": "Subscriber filters (`trigger_request_format`). All fields are optional; semantics are AND across fields, OR within a vector. `Some(vec![])` is treated identically to `None`.",
"properties": {
"operations": {
"description": "Subset of operations to subscribe to. `None` = all operations.",
"items": {
"$ref": "#/definitions/WorkerOperation"
},
"type": [
"array",
"null"
]
},
"stages": {
"description": "Subset of stages to subscribe to. `None` = all stages.",
"items": {
"$ref": "#/definitions/WorkerStage"
},
"type": [
"array",
"null"
]
},
"workers": {
"description": "Subset of worker names (exact match). `None` = all workers.",
"items": {
"type": "string"
},
"type": [
"array",
"null"
]
}
},
"title": "WorkerTriggerConfig",
"type": "object"
},
"metadata": {},
"name": "worker",
"return_schema": {
"$schema": "http://json-schema.org/draft-07/schema#",
"definitions": {
"CallerModeWire": {
"description": "Wire-format mirror of [`CallerMode`] so `WorkerCallRequest` can carry it without forcing serde derives onto the internal control-flow enum in `core::add`.",
"enum": [
"cli",
"trigger"
],
"type": "string"
},
"WorkerErrorPayload": {
"description": "Structured failure information, populated on the `failed` stage.",
"properties": {
"code": {
"description": "`Wxxx` code lifted from `WorkerOpError` when available, otherwise `W900` for an opaque internal failure.",
"type": "string"
},
"message": {
"type": "string"
}
},
"required": [
"code",
"message"
],
"type": "object"
},
"WorkerOperation": {
"description": "Every worker lifecycle operation a subscriber can watch.",
"enum": [
"add",
"remove",
"update",
"start",
"stop",
"clear"
],
"type": "string"
},
"WorkerSourceInfo": {
"description": "Where a worker came from. Only populated on `add`/`update` events.",
"properties": {
"kind": {
"$ref": "#/definitions/WorkerSourceKind"
},
"ref": {
"description": "Source-shaped identifier: registry slug (optionally `name@version`), full OCI reference, or local filesystem path.",
"type": "string"
}
},
"required": [
"kind",
"ref"
],
"type": "object"
},
"WorkerSourceKind": {
"description": "Tag for the source variant a `worker::add`/`update` operation came from.",
"enum": [
"registry",
"oci",
"local"
],
"type": "string"
},
"WorkerStage": {
"description": "Lifecycle stage attached to every event. Each [`WorkerOperation`] emits a fixed sequence of these stages.",
"enum": [
"started",
"downloading",
"downloaded",
"removing",
"updating",
"starting",
"stopping",
"clearing",
"done",
"failed"
],
"type": "string"
}
},
"description": "Event payload subscribers receive (`call_request_format`).",
"properties": {
"caller_mode": {
"$ref": "#/definitions/CallerModeWire"
},
"error": {
"anyOf": [
{
"$ref": "#/definitions/WorkerErrorPayload"
},
{
"type": "null"
}
],
"description": "Populated when `stage == failed`."
},
"operation": {
"$ref": "#/definitions/WorkerOperation"
},
"progress": {
"description": "Pull progress (0.0–1.0) for the `downloading` stage.",
"format": "double",
"type": [
"number",
"null"
]
},
"source": {
"anyOf": [
{
"$ref": "#/definitions/WorkerSourceInfo"
},
{
"type": "null"
}
]
},
"stage": {
"$ref": "#/definitions/WorkerStage"
},
"status": {
"description": "`installed` / `already_current` / `repaired` / `replaced` for `add`/`update`; absent otherwise.",
"type": [
"string",
"null"
]
},
"timestamp_ms": {
"description": "Unix timestamp in milliseconds.",
"format": "int64",
"type": "integer"
},
"version": {
"description": "Worker version. Populated on terminal `add`/`update` stages.",
"type": [
"string",
"null"
]
},
"worker": {
"description": "Canonical worker name (or the source label until a name is resolved).",
"type": "string"
}
},
"required": [
"caller_mode",
"operation",
"stage",
"timestamp_ms",
"worker"
],
"title": "WorkerCallRequest",
"type": "object"
}
}
]
}