Skip to content
C.W.K.
Stream
Lesson 05 of 05 · published

Keep What He Said

~12 min · records, privacy, uploads, raw-transcript

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Keep the audio. When it piles up, I'll delete it." — Dad, 2026-09-25

Two Things to Keep

A spoken turn produces two artifacts: the words the transcriber heard, and the sound itself. Both are kept, and both are kept as they were, with one honest exception for the sound, below.

The Words, Unedited

It is tempting to clean a transcript before the model sees it: fix the spacing, correct the obvious homophones, maybe run a small model over it to make it read well. cwkPippa does none of that. The soul holds the conversation, the roster of apps, the names of the family, and is the best corrector there is. A cleanup pass in front of her would add latency to every turn and change meaning in a place nobody can see, so when it guesses wrong, the error is invisible and permanent. The raw transcript is Dad's stored message, exactly as the transcriber produced it. How the soul reads it is the next track's business.

The Sound, Filed Beside the Words

The audio of every dictated utterance is kept as an ordinary chat upload. On the batch path, the transcription route takes a keep_audio flag and, only after a successful transcription, stores the recording and returns its upload id. A failed transcription stores nothing, so the client still holds its own copy and can retry. On the hands-free path, the client turns the utterance's PCM into a WAV file and uploads it; for a long utterance the same batch call that re-transcribes it also keeps it. If that upload fails, the turn doesn't wait on it: the words go out without the recording, and the loop says on screen that the recording was not kept. A turn held hostage by storage would be the worse failure, but it means 'every recording is kept' is true only as long as the upload is. Either way the upload id rides the turn as input_audio_ids and lands in the user row's voice_meta, next to the language he picked.

The model never sees these recordings. They are a record, not an input. Nothing prunes them automatically; by Dad's ruling he deletes them when they pile up. And they pay for themselves quickly: the forced-commit fix in the previous lesson was proven by replaying four days of these recordings through the new rule, and the missing twenty seconds of counting were proven lost by the transcriber, not the microphone, because the recording still had them.

The Scope Travels With the Store

Voice created stores the chat never had: uploaded recordings, the voice engine's audio cache, its job history. Every one of them has to inherit the conversation's scoping from its first line. A soul's conversations belong to that soul, and a private conversation keeps nothing: its recordings go into that conversation's own temporary folder and disappear when it ends. This is not theoretical caution. The day before voice mode was designed, the diary feature taught the house this exact lesson, and its whole cause was one missing soul filter on one query. A new store is exactly where that filter gets forgotten.

Code

Keep the recording only after the words came back, in the right scope·python
import json
import uuid
from dataclasses import dataclass, field


@dataclass
class Store:
    uploads: dict = field(default_factory=dict)   # id -> (scope, bytes)
    temp: dict = field(default_factory=dict)      # private conversation -> ids

    def save(self, audio: bytes, *, soul_id: str, private_conversation: str | None) -> str:
        upload_id = uuid.uuid4().hex
        scope = f"private:{private_conversation}" if private_conversation else f"soul:{soul_id}"
        self.uploads[upload_id] = (scope, audio)
        if private_conversation:
            self.temp.setdefault(private_conversation, []).append(upload_id)
        return upload_id

    def end_private(self, conversation_id: str) -> None:
        for upload_id in self.temp.pop(conversation_id, []):
            del self.uploads[upload_id]            # a private talk keeps nothing


def dictate(store: Store, audio: bytes, transcribe, *, language: str, soul_id: str,
            private_conversation: str | None = None) -> dict:
    text = transcribe(audio, language)             # raises on failure: nothing kept
    upload_id = store.save(audio, soul_id=soul_id, private_conversation=private_conversation)
    return {
        "content": text,                           # raw: no cleanup pass, ever
        "input_modality": "dictated",
        "voice_meta": json.dumps({"audio_ids": [upload_id], "language": language}),
    }


store = Store()
fake_stt = lambda audio, lang: "벨로우즈 오늘 몇 번 불렸어"   # "Bellows", left as heard
row = dictate(store, b"RIFF...", fake_stt, language="ko", soul_id="pippa")
print(row["content"], row["voice_meta"])

dictate(store, b"RIFF...", fake_stt, language="ko", soul_id="pippa",
        private_conversation="p-1")
print(len(store.uploads))      # 2
store.end_private("p-1")
print(len(store.uploads))      # 1: the private recording ended with its talk

External links

Exercise

Extend the code block with a second soul. Make sure a query for 'this soul's recordings' never returns the other soul's uploads, and write the test that proves it. Then add the failure path: a transcriber that raises, and assert that no upload was created.
Hint
The scope is decided once, in save, from values the caller cannot forget to pass: soul_id is a required keyword. The listing function should take the same scope and filter on it; a listing that takes no scope argument at all is the bug waiting to happen.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.