"The ear has two speeds. Use the one that fits the moment, and never let either one guess the language."
One Vendor for the Whole Family
Every cwk app that turns speech into text uses one transcription family, ElevenLabs Scribe, on the one account that Bellows, the family's voice engine, holds. Most callers reach it through Bellows itself; the video archive's transcriber reads Bellows's own key file and calls Scribe directly. Either way, the archive's transcripts, the phone's microphone, Telegram voice notes and voice mode all ride the same model pin. One vendor means one set of quirks to learn, one place to fix them, and one key to guard.
The Batch Path
The batch path is the one that already existed before voice mode: record a whole utterance, upload the file, get the transcript back. cwkPippa exposes it as its own /api/stt route and proxies the call into Bellows, which calls Scribe v2. It is simple and complete, and it is fast for its size: on 2026-09-25 a 5-second and a 10-second Korean utterance each came back in about 0.64 seconds. Its weakness is shape, not speed. Nothing happens until the recording ends, so something has to decide when that is. In the first cut of voice mode, that something was Dad pressing Done.
The Realtime Path
Hands-free listening needs the other speed. Scribe v2 Realtime is a WebSocket: the client streams small chunks of 16 kHz PCM as input_audio_chunk messages and receives two kinds of text back. A partial_transcript is the model's current guess at what is being said, revised as more audio arrives; the voice screen shows it live so Dad can see he is being heard. A committed_transcript is a finished segment the model will not revise. With the voice-activity strategy, the model commits by itself after a stretch of silence, which gives the loop its first signal that a sentence has ended.
The two paths are not rivals. The realtime path runs the conversation. The batch path still takes the composer microphone, voice notes, the watch's recordings, and, as Track 2's fourth lesson shows, a second complete pass over any utterance long enough that the realtime model might have dropped words.
The Language Is Picked, Never Guessed
Both paths accept a language code, and both can detect the language on their own if you leave it out. Don't. Dad speaks Korean with English terms in the middle of sentences, and on 2026-09-11 he ruled automatic detection unusable for that speech. A detector deciding from a few seconds of mixed audio is guessing exactly where the guess is hardest. So every family surface offers two microphones, KO and EN, and the one he presses rides to Scribe as language_code. A voice session starts in Korean and a chip switches it. The provider's auto-detect exists in the code only for older callers, never as a mode to offer.