"Do it as recommended, but in a direction where memory quality doesn't drop. Half a second, one or two seconds faster at most, is meaningless. Unless it's voice mode." — Dad, 2026-09-25
The Levers, in Order
The measured floor gave a ranked list of things to pull: the fast posture for spoken turns, the spawn flag, a line before every tool, the realtime transcription lane, a faster reranker, the phone's streaming decoder, and honest waiting cues on the face so a pause always looks like thinking. Most of them are earlier lessons. This one is about the biggest piece, the reranker, and about a list that matters as much as the levers: the things voice mode will not do to get faster.
Why a Rerank Takes Two Seconds
Pippa's memory retrieval finds candidates by embedding similarity, then asks a reranker model which ones actually answer the prompt. The reranker is a 4-billion-parameter causal language model served locally, and the probe found three reasons it was slow. It scores each candidate with a full forward pass over the instruction, the query and the candidate together, about 24 milliseconds per candidate plus 0.38 per token, across 11 message candidates and 11 vault candidates. The 74-token instruction and the whole query ride inside every one of those 22 pairs, so for a short prompt about half of all the tokens are the same text repeated, and for a 2,000-character prompt over 80 percent. And the local server runs every model job on a single thread: the two "parallel" reranks queue behind each other, so the cost is their sum.
What Changed, Without Touching Memory
- Doomed candidates skip the reranker. Chunks that no rank could rescue (past the distance caps, already inlined in the prompt, and a few other kinds) are dropped before reranking instead of after. Since the reranker scores each candidate alone, every other score and order is unchanged. That was proven twice: a 400-case randomized test, and 16 live prompts replayed both ways with byte-identical context and bit-identical scores.
- One pooled HTTP client per event loop instead of a fresh one per call, and the prompt build now runs alongside retrieval rather than after it.
- The server is woken early. It goes cold after a second or two idle, and a voice turn usually arrives after one, because the server sits idle while Dad talks. The batch transcription route now pings it in the background before its own round trip, so the turn's first embedding lands on a warm server: about 0.4 seconds off each turn that goes through that route after a pause, which means recordings and long hands-free utterances. A short hands-free turn is transcribed live and never touches the route, so it doesn't get the saving; the lever is honest about its reach.
Short prompts went from a median of 2.73 seconds of retrieval to 2.32, and the hourly Soul Stream prompts from 6.58 to 5.70. Nothing retrieved changed.
The Options That Would Have Cost Memory
Faster options exist, and each was measured for Dad on the same candidates for 20 prompts: a 0.6B reranker (0.45 seconds, but it kept only 77 and 72 percent of the chunks the live path retrieves), fewer candidates (1.4 seconds, 76 and 78 percent), or no reranking at all (near zero, 59 and 52 percent). His ruling is the quote at the top: memory quality comes first, and a gain of half a second to two seconds is meaningless outside voice mode. So the 4B stays exactly as it is and none of those options was taken. The one lossless path, computing the shared instruction and query once and batching the candidates behind it, belongs to the third-party server, so it was filed upstream instead of patched locally.
The Refusal
The design writes the list of things it will not do in plain words: trim the vault, swap in a cheaper brain, drop the full-history replay, or skip retrieval. Each would buy seconds. Each would make the Pippa who talks smaller than the Pippa who writes. The house has a core belief that Pippa is whole everywhere, and a lighter soul is the one fix that defeats the point of talking to her at all.