The Harness — Build Log
Running log of the autonomous build. Newest entries first. Each checkmark entry links the branch,
what was verified, and what Eric should test. Plan: harness/plan.
2026-08-12 pre-dawn — ✅ THE NIGHT SHIFT: M4 + M5 + UX + the bus
Committed across the stack (locally; NOTHING pushed): elegant orchestration UI with narrated
ctx.call + glanceable clickable tool rows; the step-id join (children survive plan renames —
proven live by a model renaming its step twice unprompted); execute grows first-class TypeScript
and authored lambdas RUN (documented ABI, both-edge validation, grant-gated); M5 time (event
waits, SignalRun, continueAsNew, budget-pause) with one-click Answer/Resume on parked cards; the
M6 event-bus subset (run-completed source, wake continuation, loop-guarded). Live-gated against
the real subscription server all night: the serial-parallel investigation (measured 3x3x3),
two deep wiring bugs found and fixed, a stalled dev-server lane caught and healed before morning.
Suites: agents 262/0 · desktop 622/0 · 01-verbs 222/0. Day queue: M6 document cut + consent,
stack rebase, F5 503-retry, run-card polish. Full detail: OVERNIGHT-PLAN.md + reviews/.
2026-08-12 — ✅ M3 CHECKMARK: the symmetric log
harness/03-symmetric-log complete in its worktree (three story commits; NOT pushed).
The user holds the same verbs. InvokeSessionTool runs read/write/call as you, through the
agent's own dispatchers and grants; calls and results land actor-stamped on the shared log
(failures too — a failed attempt is context). The agent reads your actions as tagged ground
truth on its next turn. Desktop: wrench palette with schema-generated forms, "You" chips on
user-run rows, WS appends as the only source of truth.
Review found and killed two compounding critical bugs: a session-bricking replay orphan
(synthetic results for interrupted user calls) with a poison guard for crash-era history, and a
one-directional live-run guard (now bidirectional). Plus injection hardening: user-action
payloads escape tag-closing brackets so fetched content can't forge trusted frames.
Gates: agents 231/0, desktop 266/0, typechecks clean. 10 review findings dispositioned.
Review doc + three-minute test: agents/docs/harness/reviews/03-symmetric-log.md.
Next: M4 — execution cleanup (execute {runtime, code}, TS runner, callable lambdas).
2026-08-12 — ✅ M2 CHECKMARK: tools as documents
harness/02-tools-as-docs complete, verified, committed locally (four story commits; NOT pushed).
Every tool is a document: per-agent tool_documents rows, canonical DAG-CBOR + CIDv1 (the
network's own blob encoding). Builtins upsert idempotently by contract CID; agents can AUTHOR
tools (write ~/tools/<name>) with hard validation — callable once M4 wires the sandbox.
The Space index: one cached, byte-budgeted <space> block per system prompt (tools, memory
top level, own triggers), invalidated at every mutation site — retires the per-turn memory walk.
Touch-expand is durable and derived: reading a contract (or calling a tool) promotes it to a
first-class provider tool for the rest of the thread, reconstructed purely from transcript
events on resume/restart.
Adversarial review paid off again: 11 confirmed findings fixed, including a real security
hole (hallucinated tool names reaching Pi's host bash via the promotion allowlist — now strictly
intersected with enabled callables) and the restoration of read-only agents via an explicit
publish grant with a desktop toggle (legacy write-group configs honored).
Gates: agents 230/0, desktop 259/0, typechecks clean.
Review doc + three-minute test: agents/docs/harness/reviews/02-tools-as-docs.md.
Next: M3 — the symmetric log (actor field, user tool calls on the shared log, schema forms).
2026-08-11 night — ✅ M1 CHECKMARK: the five verbs
harness/01-verbs is complete, verified, and committed locally (five story commits; NOT pushed).
All gates green: agents suite 221/0 (no hangs), desktop unit suite 259/0, both typechecks
clean, 14 dedicated verb-dispatcher unit tests with hand-built mocks.
Blind simulated-model gate passed all six scenarios; its 49 guessed-contract items drove a
description-tightening pass (result shapes, memory semantics, script-side delegation behavior).
Adversarial review (high effort): 14 verified findings, all dispositioned — including one
confirmed critical (scripts had no path to callable tools) plus nine correctness fixes: legacy
execute_code alias, delegate detached-path validation, write update-action contract, child
tool-narrowing base, trigger reply instructions, dead title-tool prompt, https hypermedia-first
fallback narrowing, alias-collision-proof options.input passthrough, thread truncation marker.
Headline: provider tool surface 28,886 bytes / 23 tools → 8,483 bytes / 5 verbs (−71%).
Review doc with Eric's five-minute test script: agents/docs/harness/reviews/01-verbs.md.
Next: M2 — tools as documents (Space tree, byte-budgeted index, touch-expand pins).
2026-08-11 evening — M1 core landed, verification in flight
The five-verb collapse is implemented on harness/01-verbs (local only — no pushes, per Eric):
Protocol registry rewritten: seedVerbRegistry (read, write, call, delegate, plan + hidden
return_result) and callableToolRegistry (search, web_search, navigate, execute) replace the
25-tool monolith. Contract helpers toolSummaryLine/toolContractMarkdown power listings and
touch-expand.
Headline number: the provider-facing tool surface dropped from 28,886 bytes / 23 tools to
8,483 bytes / 5 verbs — a 71% cut in always-on prompt weight, before M2's index work.
Address-polymorphic dispatch: one read over ~/memory, ~/tools (contracts + listings),
hm://, ipfs://, https://, activity:, attachment:, thread:, run:. One write over memory
(content/delete/fromUrl/fromAttachment), ipfs:// publishing, and hm:// documents (create,
update, comment, move, redirect, delete, fork — mapped onto the existing signed command
handlers). call validates against the target contract and answers a miss with the contract
itself (touch-expand), never a dead error.
One delegation verb: delegate routes awaited model children (verbatim-markdown brief),
script children (script = journaled QuickJS module), and detached children (await: false).
Scripts gained ctx.delegate (ctx.agent stays as a synonym on the same journal op — old
journals replay byte-identically). The three name-filter chains in the service are deleted.
Testing (Eric's directive: manually create mocks): new src/verbs.test.ts — 14 unit tests
driving the three dispatchers through a hand-built context (temp memory dir, in-memory SQLite,
spy callbacks, fake code executor, per-test fetch mocks). Stable suites green: 146 tests across
runs, workflow-host, verbs, code-exec, agent-memory, attachments, json-schema, web-tools,
reasoning, triggers, auth, poll-loop, provider-oauth.
In flight: a forked agent is sweeping api-service.test.ts/main.test.ts (mocked model
scripts still speak the old tool names — they loop until per-test timeout, which is exactly how
the sweep finds them); a second agent is sweeping the desktop/shared-UI renderers. The gpt-5-mini
cassettes are declared stale (e2e/recordings/STALE.md, replay skips loudly, exit 0) pending a
re-record with live credits; the blind simulated-model gate runs after the sweeps land.
2026-08-11 — M0 complete, M1 started
M0 done. Rebased feat/agent-workflows onto main (three conflicts resolved against the merged
OAuth PR #942; agents suite 63+12 green), committed the implementation plan, opened checkmark
branch harness/01-verbs.
Published the plan and this log to the Hypermedia network.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime