Skip to content
C.W.K.
Stream
Lesson 01 of 05 · published

Batch and Realtime

~13 min · speech-to-text, scribe, batch, realtime, language

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The ear has two speeds. Use the one that fits the moment, and never let either one guess the language."

One Vendor for the Whole Family

Every cwk app that turns speech into text uses one transcription family, ElevenLabs Scribe, on the one account that Bellows, the family's voice engine, holds. Most callers reach it through Bellows itself; the video archive's transcriber reads Bellows's own key file and calls Scribe directly. Either way, the archive's transcripts, the phone's microphone, Telegram voice notes and voice mode all ride the same model pin. One vendor means one set of quirks to learn, one place to fix them, and one key to guard.

The Batch Path

The batch path is the one that already existed before voice mode: record a whole utterance, upload the file, get the transcript back. cwkPippa exposes it as its own /api/stt route and proxies the call into Bellows, which calls Scribe v2. It is simple and complete, and it is fast for its size: on 2026-09-25 a 5-second and a 10-second Korean utterance each came back in about 0.64 seconds. Its weakness is shape, not speed. Nothing happens until the recording ends, so something has to decide when that is. In the first cut of voice mode, that something was Dad pressing Done.

The Realtime Path

Hands-free listening needs the other speed. Scribe v2 Realtime is a WebSocket: the client streams small chunks of 16 kHz PCM as input_audio_chunk messages and receives two kinds of text back. A partial_transcript is the model's current guess at what is being said, revised as more audio arrives; the voice screen shows it live so Dad can see he is being heard. A committed_transcript is a finished segment the model will not revise. With the voice-activity strategy, the model commits by itself after a stretch of silence, which gives the loop its first signal that a sentence has ended.

The two paths are not rivals. The realtime path runs the conversation. The batch path still takes the composer microphone, voice notes, the watch's recordings, and, as Track 2's fourth lesson shows, a second complete pass over any utterance long enough that the realtime model might have dropped words.

The Language Is Picked, Never Guessed

Both paths accept a language code, and both can detect the language on their own if you leave it out. Don't. Dad speaks Korean with English terms in the middle of sentences, and on 2026-09-11 he ruled automatic detection unusable for that speech. A detector deciding from a few seconds of mixed audio is guessing exactly where the guess is hardest. So every family surface offers two microphones, KO and EN, and the one he presses rides to Scribe as language_code. A voice session starts in Korean and a chip switches it. The provider's auto-detect exists in the code only for older callers, never as a mode to offer.

Code

Batch transcription, and what the realtime socket sends back·python
import os

import httpx

STT_URL = "https://api.elevenlabs.io/v1/speech-to-text"


def transcribe(path: str, language: str) -> str:
    """The batch path: one whole recording in, one transcript out.

    `language` is the mic Dad pressed ("ko" or "en"). Never omit it.
    """
    if language not in {"ko", "en"}:
        raise ValueError("pick KO or EN; auto-detect is not offered")
    with open(path, "rb") as audio:
        response = httpx.post(
            STT_URL,
            headers={"xi-api-key": os.environ["ELEVENLABS_API_KEY"]},
            data={"model_id": "scribe_v2", "language_code": language},
            files={"file": (os.path.basename(path), audio, "audio/wav")},
            timeout=60,
        )
    response.raise_for_status()
    return response.json()["text"].strip()


def on_realtime_message(message: dict, segments: list[str]) -> str | None:
    """The realtime path, one server message at a time.

    Returns the live caption to show, or None when nothing changed on screen.
    """
    kind = message.get("message_type")
    words = (message.get("text") or "").strip()
    if kind == "partial_transcript":
        return " ".join([*segments, words]).strip()   # a guess, revised later
    if kind == "committed_transcript" and words:
        segments.append(words)                         # final for this segment
        return " ".join(segments)
    if kind in {"session_started", "committed_transcript_with_timestamps", "warning"}:
        return None
    if message.get("error"):
        raise RuntimeError(f"Scribe refused: {kind}: {message['error']}")
    return None


if __name__ == "__main__":
    segments: list[str] = []
    for msg in (
        {"message_type": "partial_transcript", "text": "내일 일산"},
        {"message_type": "partial_transcript", "text": "내일 일산 날씨"},
        {"message_type": "committed_transcript", "text": "내일 일산 날씨 어때?"},
    ):
        print(on_realtime_message(msg, segments))

External links

Exercise

Record the same 20-second sentence twice, once in one language and once mixing in five English technical terms. Transcribe each with the language set explicitly, then with it omitted. Compare all four transcripts word by word and note where auto-detection went wrong.
Hint
Watch the English terms inside the mixed sentence. Whatever you find, write down which version you could have acted on without re-reading it twice. If pinning the language ever produced the only usable transcript, that one case is the argument for two microphones, because a voice loop cannot ask you to re-read anything.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.