# harness

> Thin durable turn loop that wires session-manager, context-manager, and llm-router into an agent loop; spawns sub-agents as child sessions.

| field | value |
|-------|-------|
| version | 1.8.9 |
| type | binary |
| license | Apache-2.0 |
| repo | https://github.com/iii-hq/workers |
| supported_targets | x86_64-apple-darwin, aarch64-apple-darwin, x86_64-unknown-linux-gnu, aarch64-unknown-linux-gnu, x86_64-unknown-linux-musl, armv7-unknown-linux-gnueabihf |
| author | iii |

## installation

```sh
iii trigger compose::add worker=harness@1.8.9
```

## dependencies

- `configuration` @ `0.x`
- `console` @ `^1.9.11`
- `context-manager` @ `^1.1.3`
- `cron` @ `^0.21.9`
- `iii-directory` @ `^1.2.3`
- `iii-observability` @ `0.x`
- `iii-stream` @ `0.x`
- `llm-router` @ `^1.4.12`
- `provider-anthropic` @ `^1.2.8`
- `provider-openai` @ `^1.2.7`
- `provider-openai-codex` @ `^0.4.4`
- `queue` @ `^0.21.5`
- `session-manager` @ `^1.0.13`
- `shell` @ `^0.11.9`
- `state` @ `^0.22.2`

## readme

<div align="center">

# harness

**A thin, durable turn loop that turns a model plus a few iii workers into an agent.**

<p>
  <a href="#install"><img alt="Install: iii trigger compose::add worker=harness" src="https://img.shields.io/badge/install-iii%20trigger%20compose%3A%3Aadd%20worker%3Dharness-0a84ff?style=flat-square"></a>
  <a href="../LICENSE"><img alt="License: Apache 2.0" src="https://img.shields.io/badge/license-Apache%202.0-blue.svg?style=flat-square"></a>
  <a href="https://www.rust-lang.org"><img alt="Built with Rust" src="https://img.shields.io/badge/built%20with-rust-orange?style=flat-square&logo=rust&logoColor=white"></a>
  <a href="https://workers.iii.dev/workers/harness"><img alt="harness on the workers registry" src="https://workers.iii.dev/workers/harness/badge.svg"></a>
  <a href="https://workers.iii.dev/workers/console"><img alt="console on the workers registry" src="https://workers.iii.dev/workers/console/badge.svg"></a>
</p>

</div>

`harness` is the thin, durable turn loop that turns a model plus a few iii
workers into an agent. It takes an incoming message, persists it, assembles a
context, streams a completion, runs any function calls the model requests, and
repeats until the turn stops — all as durable, resumable steps so a crash or
restart picks up mid-turn. It wires [`session-manager`](https://github.com/iii-hq/workers/tree/main/session-manager)
(transcript), [`context-manager`](https://github.com/iii-hq/workers/tree/main/context-manager) (token budgeting, soft
dependency), and [`llm-router`](https://github.com/iii-hq/workers/tree/main/llm-router) (generation); install those
alongside it for the full loop.

## Quickstart

Install the engine, export the Anthropic credential in the terminal that will
run the engine, initialize a project, and start it:

```bash
curl -fsSL https://install.iii.dev/iii/main/install.sh | sh
export ANTHROPIC_API_KEY='<your-anthropic-api-key>'
export OPENAI_API_KEY='<your-openai-api-key>'
iii project init iii-app && cd iii-app
iii compose --up
```

```bash
# New terminal, same folder. `compose::add` updates this Compose project.
cd iii-app
iii trigger compose::add worker=harness worker=console
```

```bash
open http://localhost:3113
```

Create a session, select **Anthropic → Claude Sonnet 5**, and send your first
message. Then select **OpenAI → GPT-5.6 Luna** in the same chat and send another
message. Create a new chat and send one more message with GPT-5.6 Luna. The
providers read `ANTHROPIC_API_KEY` and `OPENAI_API_KEY` from the engine
environment, so credentials do not need to be pasted into or stored by the
Console.

`iii trigger compose::add worker=harness` installs every worker the loop needs (see the badges
above); you do not add them one by one. During bootstrap, the harness asks the
`queue` worker to define a dedicated `harness-turn` queue before it reports
ready. That queue is FIFO within each `session_id` and processes separate
sessions concurrently; startup fails if the queue cannot be ensured.

Every turn, sub-agent spawn, and provider call is one correlated trace: the
harness turn waterfall in the console. Failed descendants stamp the whole trace
as failed and carry standard error attributes. The session transcript keeps the
same recovery, partial-output, and blocked-reaction explanation after refresh.

<p align="center">
  <img src="https://raw.githubusercontent.com/iii-hq/workers/main/harness/docs/images/console-traces.webp" alt="Harness turn waterfall in the iii console" width="100%">
</p>

The agent-facing function surface is deny-by-default: with no `functions.allow`
globs, every model-requested call is refused and the harness is a plain chat
loop. Allow functions in per-send (`options.functions.allow`) and gate them
with the optional [`approval-gate`](https://github.com/iii-hq/workers/tree/main/approval-gate) sibling.

The full function reference (every `harness::*` id and its request/response
schema) lives in the code and `iii worker info harness`.

Building a consumer — a chat UI, a Telegram/WhatsApp bridge, a cron worker, or
any event-driven loop on top of the harness? Start with the integration
contract in
[`architecture/integration.md`](architecture/integration.md): the functions to
trigger, the triggers to bind, and the canonical consumer patterns.

## Local development

See [`DEVELOPMENT.md`](DEVELOPMENT.md) to run the Harness and its required
workers from the local source tree with `iii compose`.

## Working with iii

iii is a language agnostic runtime where services, agents, and tools are
composed of the same things: workers, triggers, and functions. One engine
holds a live registry of every connected worker, their functions, and the
triggers bound to them. Calls route worker to engine to worker, so the
language, runtime, and location of a worker are invisible; the function id is
the only contract.

**1. Discover what is already there (the engine is the source of truth)**
- `engine::functions::list` — every function across all workers (filter with `prefix` / `search` / `worker`)
- `engine::functions::info { function_id }` — the request/response schema for ONE function (this is your API reference)
- `engine::workers::list` / `engine::workers::info { name }` — connected workers and their surface
- `engine::triggers::list` / `engine::triggers::info { id }` — legal trigger types and their config schemas
- `engine::registered-triggers::list` — every trigger instance already bound

**2. Call a function.** Use `agent_trigger` with `{ function: "<worker>::<fn>", description: "<short user-facing action>", payload: { ... } }`. The description is shown as the agent's activity in chat; keep it concise and in the user's language. The payload is a JSON object (never a stringified one), and you fetch the contract via `engine::functions::info` before the first call.

**3. Need a capability that is not registered?**
- `directory::registry::workers::list { search: "<capability>" }`
- `directory::registry::workers::info { name }` to judge fit
- `compose::add { worker: "<name>" }` to declare and start it
- confirm with `engine::functions::list { prefix: "<worker>::" }` and fetch each contract

**4. Worker lifecycle.** `compose::status`, `compose::add`, `compose::up`, `compose::down`, `compose::restart`, `compose::update`, and `compose::remove`. Fetch their contracts with `compose::schema { function_id: "compose::<operation>" }`. The harness routes `compose::*` to its supervising daemon and scopes each call to its own compose file.

**5. Triggers, not polling.** To react to events (HTTP, schedule, webhook, file change), bind a trigger instead of polling. Discover the type with `engine::triggers::list`, copy config from its schema, and confirm the binding fires with a real call (e.g. `web::fetch` to its local URL).

**6. Handy workers.**
- `web::fetch` — all HTTP(S); pass `format: "markdown"` to read docs without flooding context
- `coder::*` — file ops for any code task (read/search/create/update/move/delete)
- `slack::*` — post to Slack

**7. Authoring a worker.** Read the SDK reference for your language first (Node / Python / Rust / Browser / Engine WS) at https://iii.dev/docs/reference/. Use the SDK's `registerWorker(...)` and call `iii.registerFunction` / `iii.registerTrigger` / `iii.trigger` on the returned value; they are methods, not top-level exports. Always declare `description`, `request_format`, and `response_format` so the next caller gets a real contract.

TL;DR: list, info, call. The engine tells you the truth; trust it over memory.

## Configuration

The `harness` configuration entry is owned by the `configuration` worker; every
field hot-reloads (no restart). The fields a deployment is most likely to tune:

```yaml
default_max_turns: 16            # per-turn generate-step cap when a send omits it
default_pending_timeout_ms: 1800000  # legacy parked-call (hold / pre-deploy child) wait guard
max_depth: 3                     # sub-agent depth budget
max_children: 8                  # sub-agent spawns-per-turn budget
max_transient_resumes: 1         # recovery generations after a partial stream failure
sweep_expression: "0 * * * * *"  # cron for the pending-call expiry sweep
```

Other keys (RPC timeouts, stream coalescing, idempotency TTL, validation
retries) and their defaults live in [`src/config.rs`](src/config.rs).

## System prompt

The identity prompt is assembled once at send/spawn time. EVERY agent —
top-level turns (`harness::send`) and spawned children alike — is seeded with
the same single identity
([`prompts/default.txt`](prompts/default.txt)): a deliberately minimal prompt
carrying only the basic engine functions and the discovery loop (list, info,
call). A `default` entry in the directory's system-prompt store
(`<skills_folder>/system-prompts/default.md`, served by
`directory::system-prompts::get`) overrides the embedded prompt for every new
composition — edit or delete it and the next send picks that up, no restart;
any store failure (directory absent, entry missing, blank body) falls back to
the embedded prompt. What makes a child a leaf is its POLICY, not its prompt: children are
capability-walled out of the orchestration surface (`harness::spawn`,
`harness::send`, trigger registration) unless spawned with
`options: { orchestrator: true }`; spawn `options.system_prompt` remains the
identity escape hatch.

A spawn may also give its child a display-only identity with
`display: { name, icon?, color? }`. `name` is trimmed, limited to 48 characters,
and becomes the title of a newly created child session; `icon` and `color` are
closed semantic tokens recorded with the child linkage in
`metadata.subagent_display`. These fields never affect routing or execution,
and reusing an existing `session_id` retains that session's original title and
metadata. Icons are `agent`, `code`, `search`, `terminal`, `database`, `test`,
`review`, `docs`, or `design`; colors are `neutral`, `blue`, `purple`, `teal`,
`green`, `amber`, or `rose`.

No prompt prescribes an orchestration process — identity prompts carry tool
guidance only, enforced repo-wide by [`tests/prompts.rs`](tests/prompts.rs).
The opt-in fan-out playbook (parent-owned control plane: pick a medium, arm
notifications, spawn leaves directly, define completion per medium) lives in
[`skills/orchestration.md`](skills/orchestration.md) — paste it into a task
prompt or pass it via `options.system_prompt`.

An optional `mode` (`ask` | `agent`) prepends a short operating-mode
paragraph; `ask` is also enforced structurally — the dispatch policy of an
ask-mode send (a steer's inherited one, and an ask-mode spawned child's
resolved one) is capped at the configured default policy (`default_functions`).
The cap applies to a NEW turn; a steer folded into an already-running turn
keeps that turn's frozen policy until it finalises.

A non-empty `options.system_prompt` is combined with the built-in
prompt per `options.system_prompt_strategy`: `enrich` (default) appends it to
the built-in prompt, while `override` uses it verbatim. Assembly is tested in
[`src/prompt/tests.rs`](src/prompt/tests.rs).

The resolved prompt is STICKY per session, like `model`/`provider` and the
dispatch policy: a send to an existing session that names neither
`system_prompt` nor `system_prompt_strategy` inherits the prior turn's
resolved prompt verbatim (a prior `disabled` turn's absent prompt inherits
too). Naming either field resolves fresh — an explicit bare
`system_prompt_strategy` (e.g. `"enrich"`) is the reset-to-default escape
hatch. Because the inherited string is frozen at its original resolution,
changing `mode` on a later send without prompt fields keeps the old
operating-mode paragraph — resend the prompt fields to re-resolve.

### Agent profiles

`options.agent` on a session-creating `harness::send` names a directory agent
profile (`directory::agents::*`, one markdown file per profile). The harness
resolves it ONCE via `directory::agents::get` and freezes the result onto the
turn: the profile's RESOLVED system prompt — the directory composes `extends`
chains root-first, so `tech-lead` extending the bundled `iii` base arrives as
the full iii doctrine followed by the tech-lead body — IS the session
identity. Nothing built-in sits underneath it and no prefix is added; only the
per-send `mode` paragraph goes in front, then the usual per-step runtime
context (session id, working directory, policy aid, skills index, hook
injections). A profile whose `extends` chain does not resolve is refused as an
invalid request with the directory's D415 text. The profile's skill filter
becomes the session's skill selection (an explicit `options.skills` wins), its
`model` and optional provider-native `reasoning_effort` are authoritative for
the session, and — when the send also omits
`options.functions` — the dispatch policy defaults to the configured
`default_functions` baseline instead of deny-all (an identity picked to DO
something must be able to dispatch; the ask-mode cap still applies). The
The frozen name/icon/color/model/effort snapshot is also written to session
metadata for clients that render established sessions. The frozen identity
travels with the prompt-stickiness rule: bare later sends
inherit it, an explicit prompt field sheds it. Refused on an existing
session or combined with either prompt field. Directory edits after
resolution never reach a live session — start a new one to pick them up.

`harness::spawn` takes the same id as a top-level `agent` field: the profile's
resolved prompt is the child's whole identity, its skills/model/effort slot in the same way
(model precedence profile → explicit `model` → parent, without dragging the parent's
provider onto a foreign model), and its name and icon become the display
defaults. Which agent profile a spawn names is the prompt's decision — the profile
body steers it, nothing gates it. Spawning
with `agent` into an already RUNNING session of the caller's own tree merges
the task like any reuse and does not re-apply the profile.

New sessions also freeze a names-and-descriptions-only skill index into the
system-prompt prefix. `options.skills` on `harness::send` or `harness::spawn`
accepts exact skill ids. For a fresh session, omitted or empty means all
model-invocable skills. On an existing new-format session, omission inherits
the previous filter while an explicit empty list resets to all; explicit
changes are rejected while its turn is active.
This is curation, not authorization: the turn's function policy must still
allow `directory::skills::get`, and the function must exist in the live
registry. Skill bodies enter context only when the model calls that function.
Catalog changes are appended as durable user-role corrections, leaving the
frozen prefix unchanged. Legacy sessions keep their already-frozen prompt;
start a new session to apply an id filter to one.

Trusted console surfaces can preview the built-in, selected, frozen skill,
runtime-context, registry-notice, and declarative worker-injection layers with
`harness::system-prompt::get`, without making a model request. When the caller
passes no `selected_prompt` and the session has a turn record, the preview
reports the record's RESOLVED prompt (labeled `session (frozen at send)`) —
the truth for what ran and what the next send inherits — instead of
rebuilding the built-in. Set `default_only: true` to read the exact embedded
Harness default without consulting session or runtime state. Static
`pre_generate` hooks publish their exact contribution as trigger metadata
`inject_prompt`; request-dependent hook functions and compaction are not run
by the read-only preview and may change content when the prompt is sent.

## Custom trigger types

The harness emits two async orchestration trigger types siblings and consumers
bind to, and registers five synchronous hook points operator-trusted siblings
plug into in-path. Bind with the standard two-step pattern.

| Trigger type | Kind | Fires / runs |
|---|---|---|
| `harness::turn-started` | async event | A turn began executing (first loop step). Worker-bindable via direct engine registration only — the agent path (`engine::register_trigger`) refuses harness-internal types in every shape. |
| `harness::turn-completed` | async event | A turn reached a terminal status (`completed` / `cancelled` / `failed`), carrying the result and `terminal: bool` — `false` while the session still owns an armed wake (a one-shot notify), meaning a later turn carries the run's real outcome; consumers finalize a logical exchange only on `terminal: true`. Worker-bindable only, same as above. |
| `harness::hook::pre-turn` | sync hook | First step of a turn, before any model spend. May veto. |
| `harness::hook::pre-generate` | sync hook | After context assembly, before generation. May extend the system prompt, append messages, or veto. A static-only hook may publish its exact contribution as trigger metadata `inject_prompt`; the harness appends it directly and skips the compatibility handler. |
| `harness::hook::post-generate` | sync hook | After the final assistant message. Observe only. |
| `harness::hook::pre-trigger` | sync hook | After the allow/deny policy passes, before the target runs. May deny, hold, or rewrite arguments. |
| `harness::hook::post-trigger` | sync hook | After the target returns, before the result is persisted. May rewrite the result. |

Event configs accept `{ session_id?, parent_session_id? }`; hook configs accept
`{ functions?, priority?, timeout_ms?, on_error? }`. See the spec at
[`tech-specs/2026-06-agentic/harness.md`](https://github.com/iii-hq/workers/blob/main/tech-specs/2026-06-agentic/harness.md)
for the hook contract and chain semantics.
