Skip to content
C.W.K.
Stream
Lesson 04 of 04 · published

One Capability, Every Surface

~13 min · architecture, shared-kit, iframe, secure-context

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Come to think of it, it has to be every UI where you can talk with Pippa or any other soul. Sidekicks included." — Dad, 2026-09-25

Two Dozen Doors

Pippa is not one chat window. She is the main WebUI, a sidekick panel docked into twenty-odd sibling apps (the art studio, the music engine, the prose editor, the travel journal and the rest), a native phone app, a watch, a Chrome side panel, and a menu-bar door on every Mac. Dad's ruling made the scope all of them. The naive way to meet that ruling is a voice feature per app. The house has seen what that becomes: twenty copies of the rules, each drifting in its own direction within a week.

Rules in One Place

Every one of those doors already reaches the same chat routes in cwkPippa. Web siblings embed a sidekick that is an iframe of cwkPippa's own /embed/<app> page. Native apps talk through a small shared door type, or through a web view of the same embed. So one field on the wire gives every surface the whole behavior: the spoken instruction, the dictated instruction, the fast posture, the recording rules. The shared kit that siblings vendor carries only the names of the two fields, taken from cwkPippa's contract. It never carries a copy of what they mean.

One Hook for Every Chat Panel

On the web, the voice logic lives in a single React hook: the per-turn toggle restored from the last reply, spoken units played in order, send or queue or composer routing, hands-free listening, the face screen. The main app mounts it, and so does every panel that draws the chat view, each passing its own send function so the voice facts ride beside that panel's host context. A test finds every file that renders the chat view and fails if any of them lacks the hook, the controls, the voice send or the face. A new panel cannot quietly ship without a voice.

The Frame Has to Allow It

An iframe gets no microphone unless its parent grants one. The kit's default permission list for the sidekick frame grew from clipboard-write to clipboard-write; microphone; autoplay, and every sibling that passed its own list was widened too. Even then the browser asks once per origin, and only in a secure context: localhost, or HTTPS. An HTTPS frame inside a plain-HTTP page is not a secure context, so each sibling opens through its HTTPS twin. When the context is not secure, the panel says why instead of showing a dead microphone.

Parts in the Kit, Loops in the App

Native went the other way round on purpose. The voice session was built first inside the phone app, behind a boundary shaped like a kit: it depended only on the kit's core and speech types and a door to the brain. Only when a second consumer existed, Firekeeper on the Mac, did it lift into the kit as a voice product for iOS and macOS. The kit holds the parts: the listener, the PCM sources, the microphone owner, the barge-in detector, the spoken-reply gate. Each app keeps its own turn loop. Why not kit-first? Because a dictation feature for the health app had needed three rounds on a real phone and surfaced twelve traps no simulator showed. Proving voice once in the app closest to the brain was cheaper than proving an abstraction.

Code

A test that no chat panel can ship without voice·typescript
import { readdirSync, readFileSync, statSync } from 'node:fs';
import { join } from 'node:path';
import { describe, expect, it } from 'vitest';

const SRC = join(__dirname, '..');

function walk(dir: string): string[] {
  return readdirSync(dir).flatMap((name) => {
    const path = join(dir, name);
    if (name === 'node_modules' || name === '__tests__') return [];
    return statSync(path).isDirectory() ? walk(path) : [path];
  });
}

// Every file that draws the chat view is a surface Dad can talk to.
const surfaces = walk(SRC).filter(
  (path) => path.endsWith('.tsx') && readFileSync(path, 'utf8').includes('<ChatView'),
);

describe('every chat surface talks', () => {
  it('finds the surfaces at all', () => {
    expect(surfaces.length).toBeGreaterThan(0);
  });

  it.each(surfaces)('%s mounts the shared voice hook and its parts', (path) => {
    const source = readFileSync(path, 'utf8');
    expect(source).toContain('useVoiceMode(');   // the one home of the rules
    expect(source).toContain('<VoiceControls');  // KO / EN mics, Voice, Confirm
    expect(source).toContain('<VoiceFace');      // the face screen
    expect(source).toContain('onSend={voiceSession.sendWithVoice}'); // voice facts ride the send
  });
});

External links

Exercise

List every place in a product of yours where a user can talk to the same assistant. For each, write which layer it would use: the shared server rule, a shared client part, or its own loop. Then write the surface test for your stack: find every file that renders the chat view and assert it mounts the voice hook.
Hint
If two surfaces would need to agree on a rule, that rule belongs on the server. If two surfaces would need the same audio plumbing, it belongs in the shared client kit. The loop is the only piece that should differ per app, because each app decides differently when a turn starts.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.