Skip to content
C.W.K.
Stream
Lesson 03 of 05 · published

Cards: Shown, Never Spoken

~13 min · cards, markdown, rendering, tts-safety

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The soul's avatar is the default screen, with its emotions. If something needs to be shown, show it as a card." — Dad, 2026-09-25

Some Things Must Be Seen

A spoken reply can't contain a table, a code snippet, a URL or a list of twelve flight times, and it shouldn't try. But Dad will still ask for those things in the middle of a talk. The answer is not to refuse them and not to read them aloud; it is to put them somewhere else. That somewhere is a card: a fenced block whose info string is card. Everything inside is ordinary Markdown and renders on the screen as a card. None of it is ever spoken. The instruction tells the soul to put what must be seen in a card, say one sentence about it, and keep talking.

One Syntax, Every Screen

Choosing a fence was deliberate. Every Markdown renderer already treats a fenced block as a unit, every model already knows how to write one, and the voice engine already drops fenced blocks before synthesis. So the new code was small: a card renderer in the shared message component. The phone's answer cell is that same packaged component, so the phone showed cards the same day the web did. On the voice screen, cards appear under the soul's face; the phone draws a Markdown table inside a card as a real grid.

The web voice screen while Pippa speaks: her laughing face in a violet ring, the line 'Speaking. Talk or click the face to cut in.', a Pause button, and under it a card titled '보이스 모드는 이렇게 움직여요' listing three steps: listening, thinking, speaking.
A card under her face while she speaks. The voice never reads it; it exists only on the screen.

The Four-Backtick Problem

Then the edge case that makes this lesson worth its own page. A card that needs to contain a code block can't use three backticks, because the inner fence would close the card early. Markdown's answer is a longer fence: open and close the card with four backticks and the three-backtick block inside is just content. The instruction says exactly that. But a generic "drop code blocks" pass that only knows three-backtick fences pairs the wrong fences and leaves the middle behind, and the middle is the code. Run the example and listen to what the naive pass would hand the voice: the word "python" and a line of code.

So cwkPippa strips cards itself, with a pattern that remembers how long the opening fence was and closes only on the same fence, and every TTS route it owns applies that before any text reaches the voice engine. No client has to know. The web client splits a reply into text and card segments with the same pattern, mirrored and tested on both sides.

URLs Go in Cards

A link is the most common thing a spoken turn wants to share. Read aloud it is noise; dropped it is lost. Inside a card it becomes a link on the screen, and the response-parts machinery that already renders links in written replies can still render it, because a card is plain Markdown.

Code

A card-aware stripper next to the common naive one·python
import re

CARD = re.compile(
    r"^(?P<fence>`{3,})card[ \t]*\n.*?\n(?P=fence)[ \t]*$\n?",
    re.MULTILINE | re.DOTALL,
)
# The common "drop code blocks" regex, which knows only three-backtick fences:
ANY_FENCE = re.compile(r"```.*?```", re.DOTALL)

reply = """Here's the fix, it's one line. I put it on your screen.

````card
```python
retries = 3
```
````

Want me to run the tests after?
[emotion:happy]"""


def speakable(text: str) -> str:
    return CARD.sub("", text)


def splits(text: str) -> list[tuple[str, str]]:
    """What a screen shows: text segments and card segments, in order."""
    out, last = [], 0
    for match in CARD.finditer(text):
        if match.start() > last:
            out.append(("text", text[last:match.start()]))
        body = match.group(0).split("\n", 1)[1].rsplit(match.group("fence"), 1)[0]
        out.append(("card", body.strip()))
        last = match.end()
    if last < len(text):
        out.append(("text", text[last:]))
    return out


print("--- generic fence stripper would speak:")
print(ANY_FENCE.sub("", reply))
print("--- card-aware stripper speaks:")
print(speakable(reply))
print("--- the screen shows:")
for kind, body in splits(reply):
    print(kind, repr(body[:40]))

External links

Exercise

Run the example and compare the three outputs. Then add a second card that holds a table and a URL, and a normal three-backtick code block outside any card. Decide what should happen to the uncarded code block in a spoken turn, and adjust speakable() to match your decision.
Hint
A code block outside a card in a spoken turn means the model broke the instruction. The voice engine already drops ordinary fenced blocks, so it won't be read; the question is whether the screen should still show it. Showing it is kinder, and the fix belongs in the prompt, not in a stripper that silently hides content.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.