"Text becomes audible fast enough to feel immediate — or a deterministic asset with inspectable lineage. Never a coin flip."
Two Shapes of Output
Bellows produces exactly two kinds of result, and it's worth naming them because they set different bars. One is immediate: you press Speak and the sentence lands in the room fast enough that it feels like the machine just said it. The other is a deterministic Studio asset: a rendered file whose every input is recorded, so the same project produces the same audio and you can trace exactly where each second came from. Bellows owns that mechanism. It does not own personality, conversation history, or model judgment — only the guarantee that the same request yields the same sound.
Immediacy Is the Cache, Not a Faster Model
The instinct is to chase immediacy with speed — a lighter model, a closer region. Bellows chases it with memory. The first time a given sentence is spoken at a given identity, it costs a paid request. Every identical time after, it's a content-addressed cache hit: no provider call, no credits, and the audio is already on disk. The felt-immediacy of a repeat isn't a fast synthesis; it's no synthesis at all.
The Proof Is a Second Request
The engine's first real canary was exactly this: synthesize one short sentence, validate and cache it, play it — then fire the identical request again and watch it come back with zero provider cost. That second request is the whole cache thesis in one line. If it isn't free, the cache key is wrong, and Track 4 is where that gets fixed.