"The soul's avatar is the default screen, with its emotions. If something needs to be shown, show it as a card." — Dad, 2026-09-25
Some Things Must Be Seen
A spoken reply can't contain a table, a code snippet, a URL or a list of twelve flight times, and it shouldn't try. But Dad will still ask for those things in the middle of a talk. The answer is not to refuse them and not to read them aloud; it is to put them somewhere else. That somewhere is a card: a fenced block whose info string is card. Everything inside is ordinary Markdown and renders on the screen as a card. None of it is ever spoken. The instruction tells the soul to put what must be seen in a card, say one sentence about it, and keep talking.
One Syntax, Every Screen
Choosing a fence was deliberate. Every Markdown renderer already treats a fenced block as a unit, every model already knows how to write one, and the voice engine already drops fenced blocks before synthesis. So the new code was small: a card renderer in the shared message component. The phone's answer cell is that same packaged component, so the phone showed cards the same day the web did. On the voice screen, cards appear under the soul's face; the phone draws a Markdown table inside a card as a real grid.
The Four-Backtick Problem
Then the edge case that makes this lesson worth its own page. A card that needs to contain a code block can't use three backticks, because the inner fence would close the card early. Markdown's answer is a longer fence: open and close the card with four backticks and the three-backtick block inside is just content. The instruction says exactly that. But a generic "drop code blocks" pass that only knows three-backtick fences pairs the wrong fences and leaves the middle behind, and the middle is the code. Run the example and listen to what the naive pass would hand the voice: the word "python" and a line of code.
So cwkPippa strips cards itself, with a pattern that remembers how long the opening fence was and closes only on the same fence, and every TTS route it owns applies that before any text reaches the voice engine. No client has to know. The web client splits a reply into text and card segments with the same pattern, mirrored and tested on both sides.
URLs Go in Cards
A link is the most common thing a spoken turn wants to share. Read aloud it is noise; dropped it is lost. Inside a card it becomes a link on the screen, and the response-parts machinery that already renders links in written replies can still render it, because a card is plain Markdown.