"Keep the audio. When it piles up, I'll delete it." — Dad, 2026-09-25
Two Things to Keep
A spoken turn produces two artifacts: the words the transcriber heard, and the sound itself. Both are kept, and both are kept as they were, with one honest exception for the sound, below.
The Words, Unedited
It is tempting to clean a transcript before the model sees it: fix the spacing, correct the obvious homophones, maybe run a small model over it to make it read well. cwkPippa does none of that. The soul holds the conversation, the roster of apps, the names of the family, and is the best corrector there is. A cleanup pass in front of her would add latency to every turn and change meaning in a place nobody can see, so when it guesses wrong, the error is invisible and permanent. The raw transcript is Dad's stored message, exactly as the transcriber produced it. How the soul reads it is the next track's business.
The Sound, Filed Beside the Words
The audio of every dictated utterance is kept as an ordinary chat upload. On the batch path, the transcription route takes a keep_audio flag and, only after a successful transcription, stores the recording and returns its upload id. A failed transcription stores nothing, so the client still holds its own copy and can retry. On the hands-free path, the client turns the utterance's PCM into a WAV file and uploads it; for a long utterance the same batch call that re-transcribes it also keeps it. If that upload fails, the turn doesn't wait on it: the words go out without the recording, and the loop says on screen that the recording was not kept. A turn held hostage by storage would be the worse failure, but it means 'every recording is kept' is true only as long as the upload is. Either way the upload id rides the turn as input_audio_ids and lands in the user row's voice_meta, next to the language he picked.
The model never sees these recordings. They are a record, not an input. Nothing prunes them automatically; by Dad's ruling he deletes them when they pile up. And they pay for themselves quickly: the forced-commit fix in the previous lesson was proven by replaying four days of these recordings through the new rule, and the missing twenty seconds of counting were proven lost by the transcriber, not the microphone, because the recording still had them.
The Scope Travels With the Store
Voice created stores the chat never had: uploaded recordings, the voice engine's audio cache, its job history. Every one of them has to inherit the conversation's scoping from its first line. A soul's conversations belong to that soul, and a private conversation keeps nothing: its recordings go into that conversation's own temporary folder and disappear when it ends. This is not theoretical caution. The day before voice mode was designed, the diary feature taught the house this exact lesson, and its whole cause was one missing soul filter on one query. A new store is exactly where that filter gets forgotten.