Skip to content
C.W.K.
Stream
Lesson 02 of 05 · published

The Key Stays in the Engine

~14 min · security, single-use-token, websocket, keyterms

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The client may hold a token for one socket. It never holds the key to the account."

The Relay You Don't Want

Realtime transcription has an awkward shape. The audio must stream continuously from the microphone, and the account key must never reach a browser or a phone. The obvious answer is a relay: the client streams to your server, your server holds the key and streams on to the provider. It works, and it puts every second of Dad's voice through one more hop, one more process that can stall, and one more place that sees raw audio. It also doubles your bandwidth for no benefit.

Single-Use Tokens

ElevenLabs offers a cleaner shape, and recommends it for clients: a single-use token. A server that holds the key asks the provider for a token of type realtime_scribe. The token opens exactly one realtime socket, is consumed on first use, and expires after fifteen minutes. The client connects to the provider directly with that token in the URL, so the audio passes through no relay and the key never leaves the server. Minting a token is free; the socket bills for the audio it hears.

Three Layers, Three Owners

In the family this splits across two services, and the split is the lesson:

  • Bellows owns the provider. It holds the account key and exposes one route that mints a token and returns it with the socket URL and model name. That is all it knows about listening.
  • cwkPippa owns the conversation's policy. Its /api/stt/realtime-session route asks Bellows for a token, then builds the complete socket URL around it: the model, the token, pcm_16000 audio, the vad commit strategy with its silence threshold, the language Dad picked, and the keyterms. It returns that URL together with the sample rate, the silence window and the token's lifetime.
  • The client just connects. The web page, the phone and Firekeeper on the Mac all receive the same finished URL and open it. None of them decides the silence window or the vocabulary, so none of them can disagree.

Keyterms, Derived Not Written

Scribe accepts keyterms, words it should expect and bias toward. A transcriber has never heard of "Bellows" or a soul named Ttori; without help it writes whatever common word sounds closest. The realtime session passes the display names of the souls the current soul is allowed to see, read from the soul registry at the moment the session is built. It is never a hand-written list, for two reasons. A hand list goes stale the day a new soul is born. And a hand list would leak: a soul another soul must not see stays out of every listing in the house, and the keyterm list follows the same rule because it is built by the same function.

Code

Mint a token in the engine, build the socket URL in the brain·python
import os
from urllib.parse import urlencode

import httpx

API = "https://api.elevenlabs.io/v1"
REALTIME_URL = "wss://api.elevenlabs.io/v1/speech-to-text/realtime"
PROVIDER_SILENCE_MAX = 3.0   # Scribe's VAD refuses more than 3.0 s


# --- the engine: the only process that ever sees the key -------------------
def mint_realtime_token() -> dict:
    response = httpx.post(
        f"{API}/single-use-token/realtime_scribe",
        headers={"xi-api-key": os.environ["ELEVENLABS_API_KEY"]},
        timeout=10,
    )
    response.raise_for_status()
    return {"token": response.json()["token"], "websocket_url": REALTIME_URL,
            "model_id": "scribe_v2_realtime", "expires_in_seconds": 15 * 60}


# --- the brain: conversation policy, decided once for every client -----------
def build_session(grant: dict, *, language: str, silence: float,
                  keyterms: list[str]) -> dict:
    query = [
        ("model_id", grant["model_id"]),
        ("token", grant["token"]),
        ("audio_format", "pcm_16000"),
        ("commit_strategy", "vad"),
        ("vad_silence_threshold_secs", str(min(silence, PROVIDER_SILENCE_MAX))),
        ("language_code", language),
    ]
    query += [("keyterms", term) for term in sorted(set(keyterms))]  # repeated key
    return {
        "websocket_url": f"{grant['websocket_url']}?{urlencode(query)}",
        "sample_rate": 16000,
        "silence_seconds": silence,
        "provider_silence_seconds": min(silence, PROVIDER_SILENCE_MAX),
        "expires_in_seconds": grant["expires_in_seconds"],
    }


if __name__ == "__main__":
    fake_grant = {"token": "sut_example", "websocket_url": REALTIME_URL,
                  "model_id": "scribe_v2_realtime", "expires_in_seconds": 900}
    visible_souls = ["Pippa", "Ttori", "Feynman"]   # from the registry, not typed here
    session = build_session(fake_grant, language="ko", silence=3.0,
                            keyterms=visible_souls)
    print(session["websocket_url"])
    print(session["provider_silence_seconds"], session["expires_in_seconds"])

External links

Exercise

Run the code block and read the URL it prints. Then change it so keyterms are passed as a dictionary instead of a list of pairs, run it again, and see which terms survived. Finally, write the route that sits between them: it takes the language and an optional silence window from the client, validates both, mints a token from the engine and returns the session.
Hint
A dictionary keeps one value per key, so only the last keyterm survives, and the transcriber quietly stops expecting the others. In the route, validate before you mint. Minting is free, but every minted token is a live credential for fifteen minutes, and a request with a bad language code should fail without creating one.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.