Skip to content
C.W.K.
Stream
Lesson 03 of 04 · published

Play From the First Chunk

~15 min · streaming-audio, ios, timeouts, measurement

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Success. It was never a problem in the native app. From now on, iOS native apps should stream." — Dad, 2026-09-25

Whole Units, Streamed

The previous lesson refused to split a unit. That doesn't mean waiting for the whole take before playing it. The provider streams audio for a whole unit as it synthesizes, and the listener can start hearing the first chunk while the rest is still being made. On the desktop, a v3 unit's first chunk arrives in about 0.8 to 1.4 seconds. A finished take is another story: about 3 seconds for a 43-character sentence and 10 to 13 seconds for 178 characters. Playing from the first chunk is the difference between a pause and a wait.

Why the Phone Waited

The phone didn't stream at first, for a reason the family had learned the hard way. Earlier that summer, the fountain-pen app's web version had played a live MP3 in Safari, which started at once and then cut off roughly four-fifths of the way through. So Bellows' rule said: iOS plays the finished file. The design doc then wrote down a trigger in advance: if the finished-file wait exceeded the desktop's first chunk by more than about 1.5 seconds, build a native decoder. The measurement blew past it.

The Timeout That Left No Trace

Before the decoder, a subtler bug made spoken replies silent on the phone. The request that asked the engine to prepare a take rode the kit's quick transport, which gives up after 5 to 8 seconds. But preparing a take means synthesizing all of it before answering: 24 seconds for a 414-character reply on the bench. So every reply longer than a sentence or two timed out on the phone. The server finished anyway and cached the take, which is why Dad's later tap on the speaker played instantly and made the bug look like flakiness. And the server's access log writes its line only when a response starts, so the client that gave up left no line at all. Finding it took one more build whose only job was to make a failed reading keep its reason on screen instead of clearing itself; the next silent reply named its own cause. The fix was to move the call to the transport built for speech, with minutes of patience instead of seconds.

A Decoder of Its Own

Dad then remembered the real cause of the old Safari bug: it belongs to players that guess a stream's length up front, like Safari's and AVPlayer's, not to iOS. So the phone got a player that never guesses. It receives the body through a URL session delegate, splits it into audio packets with Audio File Stream Services, feeds them to an Audio Queue, and ends only when the server closes the body. On the bench, a 1,000-character reply produced its first sound 1.21 seconds after play and ran 66.27 seconds to the end. The engine's cached finished take for the same text was 1,060,407 bytes and 66.2726 seconds; the stream carried 1,060,407 bytes and 66.2727 seconds. Nothing was cut. The finished file stays as a second path, used only when the stream fails before any sound. Only a browser on iOS still plays finished files now.

Code

Stream one whole unit and start playing on the first chunk·python
import os
import shutil
import subprocess
import time

import httpx

API = "https://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream"


def speak_streaming(voice_id: str, text: str, model_id: str = "eleven_v3") -> None:
    """One unit, one request; audio starts as soon as the first bytes land."""
    if shutil.which("ffplay") is None:
        raise SystemExit("install ffmpeg (for ffplay) to hear the stream")
    player = subprocess.Popen(
        ["ffplay", "-nodisp", "-autoexit", "-loglevel", "quiet", "-i", "pipe:0"],
        stdin=subprocess.PIPE,
    )
    started = time.perf_counter()
    first_chunk = None
    total = 0
    with httpx.stream(
        "POST",
        API.format(voice_id=voice_id),
        params={"output_format": "mp3_44100_128"},
        headers={"xi-api-key": os.environ["ELEVENLABS_API_KEY"]},
        json={"text": text, "model_id": model_id},
        timeout=httpx.Timeout(10.0, read=600.0),   # patience per chunk, sized for speech
    ) as response:
        response.raise_for_status()
        for chunk in response.iter_bytes():
            if first_chunk is None:
                first_chunk = time.perf_counter() - started
            total += len(chunk)
            player.stdin.write(chunk)                 # play while the rest arrives
            player.stdin.flush()                      # now, not when a buffer fills
    player.stdin.close()
    player.wait()
    print(f"first chunk {first_chunk:.2f}s, {total:,} bytes, "
          f"done {time.perf_counter() - started:.2f}s")


if __name__ == "__main__":
    speak_streaming(os.environ["VOICE_ID"], "This whole paragraph is one unit. "
                    "It is synthesized once, and you hear it from the first chunk.")

External links

Exercise

Run the code with your own voice and key, twice: once with a one-sentence text and once with a five-sentence paragraph. Record the first-chunk time and the total time for each. Then give the long one a total deadline of 5 seconds (stop reading and close the response once 5 seconds have passed since the request started) and run it again. Explain what the server did with the request you abandoned.
Hint
The first-chunk time barely moves with length, while the total time grows with it; that gap is the whole case for streaming. With a 5-second deadline the long request is abandoned on your side, but the provider may still finish and bill the synthesis. Note that lowering httpx's read timeout would not do this: it limits the wait for each next chunk, not the whole request, and a healthy stream never goes five seconds between chunks. The phone's call was the other kind, one that answered only when the whole take was ready, so its first wait was the entire synthesis. Your code saw an error; the account saw a completed job. That asymmetry is exactly what hid the phone's bug.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.