"Come to think of it, it has to be every UI where you can talk with Pippa or any other soul. Sidekicks included." — Dad, 2026-09-25
Two Dozen Doors
Pippa is not one chat window. She is the main WebUI, a sidekick panel docked into twenty-odd sibling apps (the art studio, the music engine, the prose editor, the travel journal and the rest), a native phone app, a watch, a Chrome side panel, and a menu-bar door on every Mac. Dad's ruling made the scope all of them. The naive way to meet that ruling is a voice feature per app. The house has seen what that becomes: twenty copies of the rules, each drifting in its own direction within a week.
Rules in One Place
Every one of those doors already reaches the same chat routes in cwkPippa. Web siblings embed a sidekick that is an iframe of cwkPippa's own /embed/<app> page. Native apps talk through a small shared door type, or through a web view of the same embed. So one field on the wire gives every surface the whole behavior: the spoken instruction, the dictated instruction, the fast posture, the recording rules. The shared kit that siblings vendor carries only the names of the two fields, taken from cwkPippa's contract. It never carries a copy of what they mean.
One Hook for Every Chat Panel
On the web, the voice logic lives in a single React hook: the per-turn toggle restored from the last reply, spoken units played in order, send or queue or composer routing, hands-free listening, the face screen. The main app mounts it, and so does every panel that draws the chat view, each passing its own send function so the voice facts ride beside that panel's host context. A test finds every file that renders the chat view and fails if any of them lacks the hook, the controls, the voice send or the face. A new panel cannot quietly ship without a voice.
The Frame Has to Allow It
An iframe gets no microphone unless its parent grants one. The kit's default permission list for the sidekick frame grew from clipboard-write to clipboard-write; microphone; autoplay, and every sibling that passed its own list was widened too. Even then the browser asks once per origin, and only in a secure context: localhost, or HTTPS. An HTTPS frame inside a plain-HTTP page is not a secure context, so each sibling opens through its HTTPS twin. When the context is not secure, the panel says why instead of showing a dead microphone.
Parts in the Kit, Loops in the App
Native went the other way round on purpose. The voice session was built first inside the phone app, behind a boundary shaped like a kit: it depended only on the kit's core and speech types and a door to the brain. Only when a second consumer existed, Firekeeper on the Mac, did it lift into the kit as a voice product for iOS and macOS. The kit holds the parts: the listener, the PCM sources, the microphone owner, the barge-in detector, the spoken-reply gate. Each app keeps its own turn loop. Why not kit-first? Because a dictation feature for the health app had needed three rounds on a real phone and surfaced twelve traps no simulator showed. Proving voice once in the app closest to the brain was cheaper than proving an abstraction.