web
v1.2.3Outbound HTTP client on the iii bus (web::fetch).
- macOS: arm64 · x64
- Linux: arm64 · armv7 · x64
- Windows: arm64 · x64 · x86
skill doc
web
indexCallable id:
web::fetch— pass toagent_trigger { function: "web::fetch" }(NOT theiii://...skill path; that's docs, not a function id). Use this instead ofshell::execwithcurl/wgetfor any HTTP request: you get a parsed{ ok, status, headers, body }envelope, enforced size/timeout caps, and server-side SSRF protection a shellcurldoesn't have.
Decide fast
| You want to… | Call |
|---|---|
| Read a web page (docs, articles) | { "url": "https://…", "format": "markdown" } |
| GET a page/API | { "url": "https://…" } |
| Parse a JSON API response | { "url": "https://…", "response_format": "json" } |
| POST/PUT JSON | { "url": "…", "method": "post", "json": { … } } |
| Send a raw body / form | { "url": "…", "method": "post", "body": "…", "headers": { "content-type": "…" } } |
| Download binary | { "url": "…", "response_format": "base64" } |
Run a shell curl |
❌ stop — use web::fetch |
Minimal call
// agent_trigger { function: "web::fetch", payload: { "url": "https://api.github.com/zen" } }
// → { "ok": true, "status": 200, "body": "Design for failure.", ... }That's the whole T0 path: a URL is the only required field. Everything else has a default.
Read the result correctly (the #1 trap)
ok tells you if the request COMPLETED, not whether the server was happy.
- A 404 or 500 is a successful fetch →
ok: true,status: 404. Branch onstatusfor HTTP outcomes. ok: falsemeans the request never produced a response (bad URL, blocked host, timeout, transport failure). Branch onerrorfor those.
if (!r.ok) → fetch failed; look at r.error (see table below)
else if (r.status>=400) → server returned an HTTP error; look at r.status / r.body
else → success; use r.body or r.jsonDo not treat ok: true as "2xx". Always check status too.
One exception: in page-reading mode (format set), an image/* response returns the image itself plus a one-line text summary instead of this envelope — there is no top-level ok/status to branch on (see "Images in page-reading mode" below).
Request fields
| Field | Default | Notes |
|---|---|---|
url (required) |
— | absolute http(s):// |
method |
GET |
case-insensitive ("post" works); GET HEAD POST PUT PATCH DELETE OPTIONS |
headers |
{} |
object; keys as you write them |
json |
— | structured payload; auto-stringified + sets content-type: application/json. Wins over body. |
body |
— | raw string body; use for non-JSON. Ignored on GET/HEAD. |
response_format |
"text" |
"text" | "base64" (binary) | "json" (also parses into json) |
format |
— | page-reading mode: "markdown" (HTML→Markdown) | "text" (HTML→plain text) | "html" (raw). Sends a browser UA + matching Accept; retries once with the honest UA on a Cloudflare challenge; images come back viewable. Forces text transport (response_format ignored). |
timeout_ms |
30000 | clamped DOWN to the worker ceiling (120000 by default); can't raise past it |
max_bytes |
5 MiB (256 KiB in format mode) |
raw fetches default to the 5 MiB ceiling; page-reading mode (format set) uses a context-safe 256 KiB. Pass an explicit value to override (up to the 5 MiB ceiling). Over-cap body is truncated, not errored |
follow_redirects |
true |
each hop re-checked against the SSRF blocklist |
Content pipeline
Four request controls for page-reading mode (format set):
| Field | Effect |
|---|---|
content_filter |
{ type: "pruning"|"bm25", query?, threshold?, threshold_type?, min_word_threshold? } — filters body to signal content. Pruning default threshold 0.48; BM25 default 1.0. query falls back to page / if omitted. |
target_elements |
Restrict body to matching regions. Selector subset: tag, .class, #id — other forms are silently ignored. |
excluded_tags |
Drop matching elements before rendering (e.g. ["nav", "footer"]). |
include_links |
true → adds links: { internal, external } arrays (absolute URLs, classified by host). |
include_media |
true → adds media: { images, videos, audios } arrays (absolute URLs). |
When content_filter is set, body is the filtered output — no separate field. Feed it straight to a model.
Response
Success (ok: true) — returned for any completed response, 2xx through 5xx:
{
"ok": true,
"status": 200,
"status_text": "OK",
"headers": { "content-type": "application/json", ... }, // keys LOWER-CASED; set-cookie joined with ", "
"body": "<utf8 text | base64 when response_format=base64>",
"json": { ... }, // only when response_format="json" AND parse succeeded AND not truncated
"parse_error": "…", // only when response_format="json" AND JSON.parse failed (body still set)
"response_format": "json",
"bytes_truncated": false, // true when body hit max_bytes (NOT an error)
"redirect_chain": ["https://…/a"], // omitted when no redirects
"content_type": "text/html", // only in page-reading mode (format set)
"transformed": "markdown" // only when an HTML transform actually ran
}Failure (ok: false) — the fetch never produced a response:
{ "ok": false, "error": "blocked_host", "message": "…" } // no status fielderror → cause → fix
error |
Cause | Fix |
|---|---|---|
invalid_payload |
Payload failed schema (bad method, wrong types). message lists the bad fields. |
Correct the named fields. |
invalid_url |
url isn't a parseable absolute http(s):// URL. |
Pass a full absolute URL incl. scheme. |
blocked_host |
Target resolves to a private / link-local / cloud-metadata IP (SSRF guard). | Don't target internal/metadata hosts. For loopback in dev, the operator sets web.allow_loopback. |
timeout |
Slower than timeout_ms (30 s default; 120 s ceiling). |
Raise timeout_ms (up to ceiling) or shrink work via max_bytes. |
too_many_redirects |
More than max_redirects (5) hops. |
Use the final URL directly, or set follow_redirects: false and read the location header. |
transport_error |
Connection refused/reset, TLS failure, DNS failure, or a redirect Location that won't parse. |
Check host/port/cert; retry if transient. |
There is no too_large error — oversize responses come back ok: true with bytes_truncated: true. Branch on the flag, not on an error.
Page reading vs API fetch
Two orthogonal knobs — pick ONE:
response_format= transport encoding for APIs/binaries (text/base64/json). The body is returned untouched.format= page-reading mode (markdown/text/html). The request goes out with a browser User-Agent + format-matchedAcceptheader, andtext/htmlresponses are transformed (markdownis the right default for reading pages — far fewer tokens than raw HTML). Non-HTML bodies pass through unchanged. If Cloudflare answers403with a challenge, the worker retries once with its honest UA (beats the UA-fingerprint rule only, not full JS challenges). For large pages, lowermax_bytes— conversion runs on the capped body.
Don't combine them: when format is set, response_format is ignored (treated as "text").
Images in page-reading mode: a viewable image/* response returns the actual image (plus a one-line text summary like Image fetched (image/png, 8123 bytes)) instead of the JSON envelope — providers that support tool-result images (Anthropic) show it to the model; others see the text line. "Viewable" means jpeg/png/gif/webp, 2xx status, complete (not truncated), and non-empty — anything else (svg, error pages served as images, truncated bytes) comes back as the normal envelope with response_format: "base64" so a hostile image can't fail the provider request. Without format, use response_format: "base64" as before.
Transform bounds: the HTML→markdown/text conversion runs only on bodies ≤ web.max_transform_bytes (1 MiB default) — larger pages come back raw with transformed unset (lower max_bytes to read huge pages). If the body was truncated at max_bytes, the transformed text ends with a visible [Content truncated at max_bytes — …] line.
Rules that save a turn
jsonvsbody: set exactly one.jsonwins if both are present and forcescontent-type: application/json. Usebody+ your owncontent-typefor form/text/XML.- GET/HEAD ignore the body. Putting
json/bodyon a GET sends nothing — switchmethod. response_format: "json"can succeed withparse_error. Non-JSON (or truncated) bodies returnok: true, nojson,parse_errorset; the raw text is still inbody. Check forjsonbefore reading it.- Caps only go down.
timeout_ms/max_bytesabove the worker ceiling are clamped; you can't request more than the operator allows. - Response header keys are lower-cased. Index
headers["content-type"], never"Content-Type". - Auth headers don't survive a cross-origin redirect (see below) — for an authed endpoint, hit the final URL directly.
SSRF guard (server-side, not optional)
DNS is resolved once, every resolved address is checked, then the request is dialed at the validated IP (pinned lookup + TLS servername) — so the IP that passed the check is the IP connected to (no DNS-rebind window).
Blocked unconditionally → error: "blocked_host":
- Private RFC1918 (
10/8,172.16/12,192.168/16) - Link-local incl. cloud metadata
169.254.169.254 - IPv6 ULA (
fc00::/7), link-local (fe80::/10), multicast ::ffff:-mapped IPv4 in both dotted (::ffff:169.254.169.254) and hex (::ffff:a9fe:a9fe) forms
Loopback (127.0.0.0/8, ::1, localhost) is allowed by default (dev convenience); operators flip web.allow_loopback: false for prod, after which loopback returns blocked_host.
On each 3xx the Location is re-resolved and re-validated before following, and Authorization / Cookie / Proxy-Authorization are stripped when the redirect changes host or downgrades https → http.
Examples
// Read a documentation page as Markdown
{ "url": "https://docs.example.com/guide", "format": "markdown" }
// → { ok:true, status:200, body:"# Guide\n\n…", transformed:"markdown", content_type:"text/html" }
// Parse a JSON API
{ "url": "https://api.example.com/status", "response_format": "json" }
// → { ok:true, status:200, json:{ healthy:true } }
// POST JSON (content-type set for you)
{ "url": "https://api.example.com/things", "method": "post",
"json": { "name": "x" }, "response_format": "json" }
// Authenticated GET (auth stays only because there's no cross-host redirect)
{ "url": "https://api.example.com/me",
"headers": { "authorization": "Bearer TOKEN" }, "response_format": "json" }
// Bounded download of a maybe-large file
{ "url": "https://example.com/big.log", "max_bytes": 65536 }
// → { ok:true, body:"<first 64 KiB>", bytes_truncated:true }Authoritative schema
This page covers behavior the schema can't. For exact field types, call
engine::functions::info { function_id: "web::fetch" } — the live schema wins if it ever disagrees with this page.
Related
shell/index— local filesystem + process ops;web::fetchis the network counterpart (use it, notshell::exec curl).sandbox/index— code inside a sandbox reaches the host engine via the boot-timeIII_ENGINE_URLrewrite, notweb::fetch.