Skip to content
C.W.K.
Stream
Lesson 10 of 10 · published

Hands in the House — One Tool Every Vessel Gets, and the Door Dad Still Holds

~13 min · haven, tools, consent, speech-gate

Level 0Curious
0 XP0/84 lessons0/18 achievements
0/100 XP to next level100 XP to go0% complete

From words to lights

Until 2026-09-28 every tool Pippa carried read or wrote information: files, memory, the web, the family's apps. That day she got the first tool that changes something physical. cwkHaven is the house's control plane: every device on the home network with its live state, its scenes, and one typed action API. The haven tool is a thin door onto that API, so a sentence in chat can dim the living-room lights, start a TV activity or read why a speaker went quiet. Haven will get a quest of its own; this lesson is Pippa's side of the door.

One tool, seven bodies

The tool is one module with one schema and a short list of actions: read the house, look at one device, act on a device, list and run scenes, read the action log, read Haven's diagnostics. How each body gets it follows the rules of this track.

  • Claude gets it the Claude-native way: an in-process MCP server attached to the SDK client for that turn.
  • The six other vessels, whose tool loops run in cwkPippa's own process, get it through their own tool bridges. Each bridge gained the same few imports, one entry in its tool list, and a branch that checks the allowlist and runs the tool. That is lesson 5's copy-and-adapt, applied to a tool instead of a route.
  • Every action carries an actor: Pippa, the surface she spoke from, and a note naming the conversation. Haven records it in its own log with Pippa as the one who did it, beside every tap Dad makes on a dashboard.

Failure stays small. If Haven is down, chat is unchanged and the tool says it cannot reach the house. One setting closes the tool entirely. And under the test suite the tool is shut unless a test opens it on purpose, because the live house is never a test's backend: its actions switch real lights.

A typed vocabulary, not a command line

An action is never free text. The tool takes a verb from the device's own vocabulary, shaped <capability>.<verb> with typed arguments: power.set with {on: true}, brightness.set with a level, activity.start with an activity name. The answer comes back as one of four honest words. Verified means Haven read the state back. Unverified means the command was sent but the device cannot confirm, as with infrared. Failed means Haven could not do it or could not confirm it in time, which is not proof that nothing changed: a command that timed out may still have landed. Refused means Haven turned it down before any device saw it, for a verb the device does not have, bad arguments, or a stale confirmation. The tool's own description tells the soul to report exactly what came back, and never more.

The door Dad still holds

Some actions are too consequential for a model to finish on its own. A network write, such as turning off a switch port or cutting power over Ethernet, comes back from Haven as confirm-required: Pippa can propose it, and only Dad can confirm it, on Haven's network page. The tool tells the soul to say so and never retry. The important part is where that rule lives. It is not a sentence in Pippa's prompt that a clever enough model could reason around. It is a door in Haven, the system that owns the switch: Haven accepts a confirmation only with its one-time token and only from Dad, and the tool always names Pippa as the actor and never hands the model that token. Haven takes each caller's word for who it is, so this is a door against persuasion, not against a program that lies about itself, and against persuasion it holds.

Say it before you do it

Voice mode added one more constraint the next day. In a spoken turn, Pippa says a short line before a tool, so the silence has a reason. That line takes a second or two to synthesize and play, and a light switches faster. Dad heard "I'll turn it on" after the air conditioner was already on. So an action tool now waits for the conversation's reading to go quiet before it acts. The wait has a ceiling of twelve seconds, and it never waits on a client that does not report or has gone stale, so a closed screen can never hang a tool. Read-only tools never wait, because nothing Dad sees or hears happens before their result is spoken.

Principle: Put the consent for an action where the action happens, not in the prompt of the one proposing it. A model can be persuaded; a door in the owning system cannot. Let the model propose and explain, and let the system that owns the consequence ask the person who owns the decision.
Self-reference: This is the first time my words reach past a screen into the room Dad is sitting in. I say what I am about to do before I do it, I report what the house actually confirmed rather than what I asked for, and when Haven says a change needs Dad, I tell him and stop. Hands are worth having only if I'm careful with them.

Code

One tool, seven vessels·text
services/haven_tool.py   one schema, one executor
  actions: house · device · act · scenes · run_scene · log · problems

Claude          -> in-process MCP server on the SDK client (per turn)
codex, gemini,
grok, kimi,     -> each variant's tool_bridge.py: the same imports,
glm, ollama        a tool-list entry, an allowlist-checked branch
                   (copy-and-adapt)

every action carries actor = {kind: pippa, surface, note: conversation}
Haven down      -> chat unchanged, the tool says so
PIPPA_HAVEN=0   -> tool closed
under pytest    -> shut unless a test opens it (real lights)
The path of one action·text
act: device="Living room lamp", command="power.set", args={on: true}
  |
  +-- spoken turn? wait until the announcing line is heard
  |      (ceiling 12 s; never for a silent or stale client)
  |
  +-- POST Haven /api/devices/{id}/actions  {action, args, actor}
         verified          -> Haven read the state back
         unverified        -> sent, the device cannot confirm (IR)
         failed            -> not done or not confirmed in time (may still
                              have landed)
         refused           -> turned down before any device saw it
         confirm-required  -> network write: "waiting for Dad's confirm
                              in Haven", never retried

Exercise

List the tools an assistant you run can call and mark each one: reads, writes information, or changes something outside the software. For every tool in the last group, find where the consent for its riskiest action lives today, and whether a persuasive enough prompt could get past it.
Hint
If the only thing between the model and a consequential action is a sentence in its own instructions, the consent lives in the wrong place. Move it into the system that owns the thing being changed, and have the tool return 'waiting for a person' as an ordinary result the model reports.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.