Skip to content
C.W.K.
Stream
Lesson 02 of 04 · published

The Levers, and the One We Refuse

~16 min · latency, reranking, tradeoffs, memory-quality

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Do it as recommended, but in a direction where memory quality doesn't drop. Half a second, one or two seconds faster at most, is meaningless. Unless it's voice mode." — Dad, 2026-09-25

The Levers, in Order

The measured floor gave a ranked list of things to pull: the fast posture for spoken turns, the spawn flag, a line before every tool, the realtime transcription lane, a faster reranker, the phone's streaming decoder, and honest waiting cues on the face so a pause always looks like thinking. Most of them are earlier lessons. This one is about the biggest piece, the reranker, and about a list that matters as much as the levers: the things voice mode will not do to get faster.

Why a Rerank Takes Two Seconds

Pippa's memory retrieval finds candidates by embedding similarity, then asks a reranker model which ones actually answer the prompt. The reranker is a 4-billion-parameter causal language model served locally, and the probe found three reasons it was slow. It scores each candidate with a full forward pass over the instruction, the query and the candidate together, about 24 milliseconds per candidate plus 0.38 per token, across 11 message candidates and 11 vault candidates. The 74-token instruction and the whole query ride inside every one of those 22 pairs, so for a short prompt about half of all the tokens are the same text repeated, and for a 2,000-character prompt over 80 percent. And the local server runs every model job on a single thread: the two "parallel" reranks queue behind each other, so the cost is their sum.

What Changed, Without Touching Memory

  • Doomed candidates skip the reranker. Chunks that no rank could rescue (past the distance caps, already inlined in the prompt, and a few other kinds) are dropped before reranking instead of after. Since the reranker scores each candidate alone, every other score and order is unchanged. That was proven twice: a 400-case randomized test, and 16 live prompts replayed both ways with byte-identical context and bit-identical scores.
  • One pooled HTTP client per event loop instead of a fresh one per call, and the prompt build now runs alongside retrieval rather than after it.
  • The server is woken early. It goes cold after a second or two idle, and a voice turn usually arrives after one, because the server sits idle while Dad talks. The batch transcription route now pings it in the background before its own round trip, so the turn's first embedding lands on a warm server: about 0.4 seconds off each turn that goes through that route after a pause, which means recordings and long hands-free utterances. A short hands-free turn is transcribed live and never touches the route, so it doesn't get the saving; the lever is honest about its reach.

Short prompts went from a median of 2.73 seconds of retrieval to 2.32, and the hourly Soul Stream prompts from 6.58 to 5.70. Nothing retrieved changed.

The Options That Would Have Cost Memory

Faster options exist, and each was measured for Dad on the same candidates for 20 prompts: a 0.6B reranker (0.45 seconds, but it kept only 77 and 72 percent of the chunks the live path retrieves), fewer candidates (1.4 seconds, 76 and 78 percent), or no reranking at all (near zero, 59 and 52 percent). His ruling is the quote at the top: memory quality comes first, and a gain of half a second to two seconds is meaningless outside voice mode. So the 4B stays exactly as it is and none of those options was taken. The one lossless path, computing the shared instruction and query once and batching the candidates behind it, belongs to the third-party server, so it was filed upstream instead of patched locally.

The Refusal

The design writes the list of things it will not do in plain words: trim the vault, swap in a cheaper brain, drop the full-history replay, or skip retrieval. Each would buy seconds. Each would make the Pippa who talks smaller than the Pippa who writes. The house has a core belief that Pippa is whole everywhere, and a lighter soul is the one fix that defeats the point of talking to her at all.

Code

Wake the cold server while the transcriber works·python
import asyncio
import time


class ColdServer:
    """A local model server that goes cold after ~1.5 s idle (measured: 0.44 s vs 0.04 s).
    One worker: a second request queues behind the first instead of paying twice."""

    def __init__(self, cold_penalty: float = 0.44) -> None:
        self.cold_penalty = cold_penalty
        self.last_used = time.perf_counter() - 12.0      # Dad has been talking for 12 s
        self.worker = asyncio.Lock()

    async def embed(self, text: str) -> list[float]:
        async with self.worker:
            idle = time.perf_counter() - self.last_used
            await asyncio.sleep(self.cold_penalty if idle > 1.5 else 0.04)
            self.last_used = time.perf_counter()
        return [0.0]


async def transcribe(audio: bytes) -> str:
    await asyncio.sleep(0.64)                             # the provider round trip
    return "저녁 뭐 먹을까?"


async def voice_turn(server: ColdServer, warm_up: bool) -> float:
    started = time.perf_counter()
    if warm_up:
        # Fire and forget: wake the server while the transcriber works.
        warmer = asyncio.create_task(server.embed("warm-up"))
    text = await transcribe(b"...")
    await server.embed(text)                              # the turn's real query
    if warm_up:
        await warmer                                      # never leave a task dangling
    return time.perf_counter() - started


async def main() -> None:
    cold = await voice_turn(ColdServer(), warm_up=False)
    warm = await voice_turn(ColdServer(), warm_up=True)
    print(f"no warm-up : {cold:.2f}s to the query embedding")
    print(f"warm-up    : {warm:.2f}s  (saved {cold - warm:.2f}s, nothing retrieved changed)")


asyncio.run(main())

External links

Exercise

Run the warm-up demo. Then change ColdServer so its cold penalty is 0.9 seconds and the transcription takes only 0.3 seconds, and rerun: how much does the warm-up save now, and why less than the penalty? Finally, write your own product's refusal list: three things you would not trade for speed, and the measurement that would tell you if a change had traded them anyway.
Hint
The warm-up can only hide as much of the cold start as fits inside the wait it overlaps; with a 0.3-second wait and a 0.9-second penalty, the real query queues behind the warm-up for the remaining 0.6 seconds. That queue is why the server is modeled with one worker: two requests that each paid the full penalty on their own would make the warm-up look worthless. For the refusal list, pair each item with an instrument, like the 'kept' percentage the reranker options were judged by, or the list is only a wish.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.