Skip to content
C.W.K.
Stream
Lesson 04 of 04 · published

The Face and the Voice

~14 min · emotion-tag, audio-tags, korean-numerals, avatar

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"She's going along fine, and then Pippa says '여섯분', and it drops from Pippa to 'a model'." — Dad, 2026-09-27

Two Kinds of Tag, Two Destinations

A spoken reply carries two kinds of bracketed markup, and it matters that they never get confused. The emotion tag, written once at the very end as [emotion:happy], drives the face: the voice screen shows the soul's avatar in that emotion, and the next emotion's image fades in over the last with no blank frame. Audio tags like [laughs], [softly] or [sighs] sit inside the text and drive the voice: expressive models perform them as direction instead of speaking them. The emotion tag moves the face; the audio tags move the voice. Neither ever appears as text on the screen.

The Emotion Tag Stays at the End

Moving the emotion tag to the front was considered, so the face could change the instant speech starts. It wasn't needed. Because each unit is synthesized whole, the tag at the end of the reply always arrives before the final unit starts playing, so the face is already right when the last words are heard. The parser stays end-anchored, which also protects it from a soul that merely talks about the tag in the middle of a sentence. It tolerates a tag cut off mid-word, strips anything unrecognized as noise, and falls back to warm.

Audio Tags Come From a Closed List

The family keeps a closed vocabulary of about three dozen approved audio tags: laughter, breath, delivery, emotion and pacing. The list is closed for a mechanical reason. On a model that can't perform tags, the voice engine strips them, and it can only strip tags it knows; an unlisted tag would be read aloud as words. The soul is handed exactly that list in its spoken instruction, told to use tags sparingly and only where the feeling genuinely shifts, and to use single-word tags next to Korean text. Because the soul writes its own tags on a spoken turn, the separate pass that adds tags to written replies never runs on it.

Digits and the Unit They Carry

Korean counts in two numeral systems and the unit chooses. Minutes, percent, money and dates take Sino-Korean (육 분, 이십 퍼센트); things counted one by one, the hour on the clock face and age take native Korean (세 개, 세 시, 스무 살), though an hour past twelve goes back to Sino-Korean (십삼 시). The first rule asked the soul to spell numbers out the way they should be read. It wrote 여섯 분 for six minutes, which is six people, and Dad heard Pippa turn into a model mid-sentence. Now the rule is the opposite: the soul writes a quantity with a unit in digits (6분), and the voice engine reads it in the numeral its unit takes. The engine owns the reading, so every family app that speaks gets it. One more case arrived the next day: 에어컨 9대 was read as 구 대, when machines are counted natively and only a round ten with 대 is an age group. The rules list only units whose reading is certain and leave ambiguous counters to the voice. They also have to know where a unit ends: 12개월 is Sino (십이 개월) even though 개 alone is native, and 3시즌 is not three o'clock. So compound units are matched longest first, and a short unit carries a guard against the letters that would make it the head of a longer word.

Code

The face's tag, the screen's text, and the voice's numerals·python
import re

VALID_EMOTIONS = {"warm", "happy", "excited", "playful", "serious", "concerned", "surprised", "sad"}
TRAILING_EMOTION = re.compile(r"\[emotion:(\w*)\]?\s*$", re.IGNORECASE)  # end-anchored
AUDIO_TAG = re.compile(r"\[(?:[a-z]+(?: [a-z]+){0,2})\](?!\()[ \t]?")  # not [x](url)


def split_reply(reply: str) -> tuple[str, str]:
    """(body, emotion). The face reads the emotion; the tag never shows."""
    match = TRAILING_EMOTION.search(reply)
    if not match:
        return reply.strip(), "warm"
    emotion = match.group(1).lower()
    return reply[: match.start()].rstrip(), emotion if emotion in VALID_EMOTIONS else "warm"


def for_screen(body: str) -> str:
    """Audio tags are performed by the voice, never shown as text."""
    return AUDIO_TAG.sub("", body)


SINO = "_일이삼사오육칠팔구"
NATIVE_ONES = ["", "한", "두", "세", "네", "다섯", "여섯", "일곱", "여덟", "아홉"]
NATIVE_TENS = ["", "열", "스무", "서른", "마흔", "쉰", "예순", "일흔", "여든", "아흔"]
UNIT_SYSTEM = {"분": "sino", "초": "sino", "달러": "sino", "퍼센트": "sino",
               "개월": "sino", "개국": "sino", "시간": "native",
               "개": "native", "명": "native", "시": "hour", "살": "native"}
# A unit must not be the head of a longer word: 3시즌 is not three o'clock.
GUARD = {"개": "(?![국년월소발념])", "명": "(?![령단함소])", "살": "(?![짝림균])",
         "시": "(?![즌리점기각절대청장민내외험작공])"}


def sino(n: int) -> str:
    tens, ones = divmod(n, 10)
    head = "" if tens == 0 else ("십" if tens == 1 else SINO[tens] + "십")
    return head + ("" if ones == 0 else SINO[ones])


def native(n: int) -> str:
    tens, ones = divmod(n, 10)
    ten = NATIVE_TENS[tens]
    if tens == 2 and ones:
        ten = "스물"                       # 스무 only stands alone before a counter
    return ten + NATIVE_ONES[ones]


def read_numerals(text: str) -> str:
    """Digits with a unit, read in the numeral the unit takes (1-99 here)."""
    def spell(match: re.Match) -> str:
        n, unit = int(match.group(1)), match.group(2)
        if not 0 < n < 100:
            return match.group(0)
        system = UNIT_SYSTEM[unit]
        if system == "hour":                    # the clock face is native, 1-12 only
            system = "native" if n <= 12 else "sino"
        word = sino(n) if system == "sino" else native(n)
        return f"{word} {unit}"
    units = "|".join(re.escape(u) + GUARD.get(u, "")        # longest first: 개월 before 개
                     for u in sorted(UNIT_SYSTEM, key=len, reverse=True))
    return re.sub(rf"(\d+)\s?({units})", spell, text)


reply = "[happy] 좋아, 3시에 시작하면 6분이면 끝나. 20퍼센트 할인도 있어!\n[emotion:happy]"
body, emotion = split_reply(reply)
print("face:  ", emotion)
print("screen:", for_screen(body))
print("voice: ", read_numerals(body))
print("voice: ", read_numerals("12개월 동안 3개국, 3시즌. 13시 회의는 2시간, 발표는 13시간 뒤."))
print(split_reply("알았어. [emotion:exc"))              # truncated tag -> stripped, warm

External links

Exercise

Run the code, then add three units to UNIT_SYSTEM: one Sino-Korean (for example 미터), one native (for example 마리), and one you think is ambiguous. Decide what the reader should do with the ambiguous one and justify leaving it to the voice or not.
Hint
번 is the classic ambiguous counter: 세 번 is three times, but 일 번 is number one. Any rule you write for it will be wrong half the time. The engine's own rule is to list only units whose reading is certain and let the voice handle the rest, which is also why authors write 6분 in digits: 분 is certain, but a spelled-out 여섯 분 is not.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.