"Make the strength of the motion a setting, and default to barely noticeable." — Dad, 2026-09-29
A Portrait That Looked Frozen
The previous lesson moves the face between replies: the emotion tag at the end of each one picks the expression. But most of a talk is not her replying. It is Dad talking and her listening, and for all of that time the face was a still portrait. On a live screen a still portrait doesn't read as calm; it reads as frozen, the same problem Track 6 solved for the microphone with a moving meter. After a brainstorm on the options, from new animated assets to reacting to his words, Dad ruled for the cheapest thing that could work: motion that needs no new assets first, its strength a setting, barely noticeable by default. New assets later, and only if use asks for them.
Signals In, a Pose Out
One small model turns signals the screen already has into a pose: a scale, a sideways and a vertical offset, and a rotation. The screen applies the pose to the whole face, ring included, so nothing is ever cropped and the frame never shows an edge. The signals are the loop's phase (listening, waiting, speaking, or held when paused or stopped), Dad's microphone level, whether words are arriving, and a count of the transcript pieces committed so far. What she does with them:
- Breathes and drifts in every live phase: a 4.2-second breath of under one percent of scale, and a slow wander built from three sines whose periods never line up, so it never visibly repeats and is identical on every surface.
- Leans in with his voice: up to 1.5 percent larger at his loudest, following his level with a time constant of 0.12 seconds.
- Nods once, over 0.7 seconds, each time a piece of his words is committed.
- Tilts her head once he has been talking for 4 seconds, fully by 7.
- Looks up and a little aside while she prepares a reply.
- Settles when paused or stopped: the drift stops and the breath halves.
The tilt, the look up and the settling each ease toward their targets with a 0.35-second time constant (the lean follows his level faster, at 0.12 seconds), so a change of phase is a glide, never a jump. The nod is the one motion with a shape of its own: a half sine that starts and ends at rest. A piece that lands while a nod is still playing restarts it from rest, a small jump the model accepts rather than hides. And a long frame, a hidden tab or a stall, counts as at most a tenth of a second, so coming back to the tab never snaps her into place.
Which Face While She Listens
The expression follows the same care. When listening starts, the last reply's emotion is held for 1.2 seconds, so a laugh doesn't vanish the instant she stops talking; then she listens with warm. She keeps warm while she prepares a reply, and while she speaks until the reply's own emotion is known. A new emotion fades in over the last in 350 milliseconds, on both surfaces.
Motion Is a Setting
Every movement is multiplied by one number: 0 is still, 1 is barely noticeable (the default Dad chose), 2 is clear. The web reads it from Admin; the phone keeps its own in its voice settings. A system that asks for reduced motion always gets the still face, whatever the setting says; accessibility outranks the house default. All three surfaces follow that request live, so a change in the middle of a talk takes hold at once.
One Model, Two Surfaces, Then the Kit
The web and the phone held the same numbers and passed the same test cases, one file mirroring the other, with a standing rule to change a number on both sides or on neither. When Firekeeper became the model's second native user, the phone's half moved into the shared voice kit, and the phone now keeps only how its own phases map to a presence mode. Two things were deliberately not built: new assets, and reactions to what Dad says. A keyword list would be hardcoding, and a model call per transcript piece would be latency and cost on every sentence. Both wait for use. On Air's broadcast face, in Track 8, is built on this same presence.