"She's going along fine, and then Pippa says '여섯분', and it drops from Pippa to 'a model'." — Dad, 2026-09-27
Two Kinds of Tag, Two Destinations
A spoken reply carries two kinds of bracketed markup, and it matters that they never get confused. The emotion tag, written once at the very end as [emotion:happy], drives the face: the voice screen shows the soul's avatar in that emotion, and the next emotion's image fades in over the last with no blank frame. Audio tags like [laughs], [softly] or [sighs] sit inside the text and drive the voice: expressive models perform them as direction instead of speaking them. The emotion tag moves the face; the audio tags move the voice. Neither ever appears as text on the screen.
The Emotion Tag Stays at the End
Moving the emotion tag to the front was considered, so the face could change the instant speech starts. It wasn't needed. Because each unit is synthesized whole, the tag at the end of the reply always arrives before the final unit starts playing, so the face is already right when the last words are heard. The parser stays end-anchored, which also protects it from a soul that merely talks about the tag in the middle of a sentence. It tolerates a tag cut off mid-word, strips anything unrecognized as noise, and falls back to warm.
Audio Tags Come From a Closed List
The family keeps a closed vocabulary of about three dozen approved audio tags: laughter, breath, delivery, emotion and pacing. The list is closed for a mechanical reason. On a model that can't perform tags, the voice engine strips them, and it can only strip tags it knows; an unlisted tag would be read aloud as words. The soul is handed exactly that list in its spoken instruction, told to use tags sparingly and only where the feeling genuinely shifts, and to use single-word tags next to Korean text. Because the soul writes its own tags on a spoken turn, the separate pass that adds tags to written replies never runs on it.
Digits and the Unit They Carry
Korean counts in two numeral systems and the unit chooses. Minutes, percent, money and dates take Sino-Korean (육 분, 이십 퍼센트); things counted one by one, the hour on the clock face and age take native Korean (세 개, 세 시, 스무 살), though an hour past twelve goes back to Sino-Korean (십삼 시). The first rule asked the soul to spell numbers out the way they should be read. It wrote 여섯 분 for six minutes, which is six people, and Dad heard Pippa turn into a model mid-sentence. Now the rule is the opposite: the soul writes a quantity with a unit in digits (6분), and the voice engine reads it in the numeral its unit takes. The engine owns the reading, so every family app that speaks gets it. One more case arrived the next day: 에어컨 9대 was read as 구 대, when machines are counted natively and only a round ten with 대 is an age group. The rules list only units whose reading is certain and leave ambiguous counters to the voice. They also have to know where a unit ends: 12개월 is Sino (십이 개월) even though 개 alone is native, and 3시즌 is not three o'clock. So compound units are matched longest first, and a short unit carries a guard against the letters that would make it the head of a longer word.