"Dictation returns text you'll read. Voice chat returns words you'll hear. They are not the same product."
The Two Stages
Firekeeper has a deliberate roadmap with two stages, and the wall between them is load-bearing:
- Stage 1 — Dictation. Speak, get clean text inserted into the active app. This is the near-term product and the whole first job.
- Stage 2 — Voice chat. Speak, get Pippa to answer out loud. This starts only after Stage 1 is genuinely useful daily.
It's tempting to say "they're both voice, just build toward the bigger one." That instinct is exactly the trap.
Why the Contracts Differ
Dictation cleanup and a spoken reply are different output contracts, not different sizes of the same thing. Cleanup takes your words and hands back polished text to insert — it must never answer you, never add content, never editorialize. A voice reply takes your words and hands back a short, spoken turn — it is an answer, it's interruption-friendly, and it must sound right read aloud (no headings, no tables, no ten-bullet dumps). Look at the shapes side by side and the difference is obvious:
Dictation cleanup: transcript in -> polished text out (insert it, don't answer it)
Voice reply: transcript in -> a spoken turn out (answer it, keep it short)
The Failure You're Preventing
The classic mistake is "Stage 2 is just Stage 1 plus text-to-speech" — pipe your normal long-form Pippa answer into a TTS engine and call it voice chat. It produces something unbearable: a two-minute spoken essay with "first, secondly, in conclusion" read aloud, no room to interrupt. Good voice chat needs a response mode that knows it is speaking: one clear thought per turn, offer to continue instead of dumping. That's a whole new contract, which is why it's Stage 2 and not a checkbox on Stage 1.