"Defer Telegram and Council, and just leave them on record. We might never do them." — Dad, 2026-09-26
The Web and Every Sidekick
The WebUI's main chat was the first surface, and the shared hook from Track 1 carried voice mode into every sidekick panel the same day: the prose editor, the music engine, the art studio and the rest all talk like the main chat, because each one renders the same chat view through the same embed and mounts the same hook. A browser still asks for the microphone once per origin, and only over HTTPS or localhost.
The Phone
The native app got the most work, and most of Tracks 4 through 7 happened there first: its own streaming decoder, one owner for the audio session so the talk survives a dark screen, a face screen, a Live Activity that follows the conversation, echo cancellation with the reply on the microphone's own engine. Every change went out through TestFlight to Dad's own phone, where many of the bugs in this quest were found.
The Mac: Firekeeper
Firekeeper, the family's dictation app, became the Mac's door into voice mode. A global hotkey opens a small floating panel for a hands-free talk from any app; its menu can start a talk with any soul, continue a recent voice chat from any surface, or open a private talk that lives only in memory. It doesn't use its own on-device speech engine for this: it listens through cwkPippa's realtime session, so the partial words, the end-of-speech window and the keyterms are exactly the web's and the phone's. The turn rides the brain's own chat route with reply_modality: spoken, on a model slot reserved for voice chat rather than the main one, and a talk is an ordinary conversation that shows up in the sidebar like any other. As Track 7 told, it cuts in by key rather than by voice, and it lets each Mac choose its own microphone and speaker, because remote-desktop software kept taking over the default devices.
The Wrist
The watch has no speech recognizer an app can drive outside its own dictation screen, and it isn't on the family's private network, so its door works through the phone. A Pippa Voice complication takes a dictated message on the wrist, the phone turns it into a dictated, spoken turn, then fetches the finished take in the soul's voice and hands the MP3 back to the watch as a file, not a stream. It was proven on Dad's watch through the speaker and through earbuds, with the crown as the volume knob. A fully hands-free wrist loop exists too: the watch records, decides the end of speech with its own loudness rule, and the phone transcribes; which parts of it are proven on a real wrist and which aren't is written down next to it.
Handing a Thread Across
The interpreter app, Spark, is a case of saying no to a door. Its web sidekick talks like every other. But its phone app has no conversation screen, and Dad ruled it shouldn't grow a voice loop of its own. So Spark's "Talk with Pippa" buttons open the moment's thread in the phone app through a link that names the conversation and the language, and the phone app opens it in voice mode, already listening.
The Doors Not Built
Telegram voice notes and multi-soul Council turns have complete designs in the doc. Dad deferred both and said they may never happen, so the design stays as a record, not a plan. A written-down "not now" is part of the architecture too.