Skip to content
C.W.K.
Stream
Lesson 05 of 05 · published

A Face That Listens

~14 min · avatar, presence, animation, accessibility

Level 0Muted
0 XP0/42 lessons0/13 achievements
0/100 XP to next level100 XP to go0% complete
"Make the strength of the motion a setting, and default to barely noticeable." — Dad, 2026-09-29

A Portrait That Looked Frozen

The previous lesson moves the face between replies: the emotion tag at the end of each one picks the expression. But most of a talk is not her replying. It is Dad talking and her listening, and for all of that time the face was a still portrait. On a live screen a still portrait doesn't read as calm; it reads as frozen, the same problem Track 6 solved for the microphone with a moving meter. After a brainstorm on the options, from new animated assets to reacting to his words, Dad ruled for the cheapest thing that could work: motion that needs no new assets first, its strength a setting, barely noticeable by default. New assets later, and only if use asks for them.

Signals In, a Pose Out

One small model turns signals the screen already has into a pose: a scale, a sideways and a vertical offset, and a rotation. The screen applies the pose to the whole face, ring included, so nothing is ever cropped and the frame never shows an edge. The signals are the loop's phase (listening, waiting, speaking, or held when paused or stopped), Dad's microphone level, whether words are arriving, and a count of the transcript pieces committed so far. What she does with them:

  • Breathes and drifts in every live phase: a 4.2-second breath of under one percent of scale, and a slow wander built from three sines whose periods never line up, so it never visibly repeats and is identical on every surface.
  • Leans in with his voice: up to 1.5 percent larger at his loudest, following his level with a time constant of 0.12 seconds.
  • Nods once, over 0.7 seconds, each time a piece of his words is committed.
  • Tilts her head once he has been talking for 4 seconds, fully by 7.
  • Looks up and a little aside while she prepares a reply.
  • Settles when paused or stopped: the drift stops and the breath halves.

The tilt, the look up and the settling each ease toward their targets with a 0.35-second time constant (the lean follows his level faster, at 0.12 seconds), so a change of phase is a glide, never a jump. The nod is the one motion with a shape of its own: a half sine that starts and ends at rest. A piece that lands while a nod is still playing restarts it from rest, a small jump the model accepts rather than hides. And a long frame, a hidden tab or a stall, counts as at most a tenth of a second, so coming back to the tab never snaps her into place.

Which Face While She Listens

The expression follows the same care. When listening starts, the last reply's emotion is held for 1.2 seconds, so a laugh doesn't vanish the instant she stops talking; then she listens with warm. She keeps warm while she prepares a reply, and while she speaks until the reply's own emotion is known. A new emotion fades in over the last in 350 milliseconds, on both surfaces.

Motion Is a Setting

Every movement is multiplied by one number: 0 is still, 1 is barely noticeable (the default Dad chose), 2 is clear. The web reads it from Admin; the phone keeps its own in its voice settings. A system that asks for reduced motion always gets the still face, whatever the setting says; accessibility outranks the house default. All three surfaces follow that request live, so a change in the middle of a talk takes hold at once.

One Model, Two Surfaces, Then the Kit

The web and the phone held the same numbers and passed the same test cases, one file mirroring the other, with a standing rule to change a number on both sides or on neither. When Firekeeper became the model's second native user, the phone's half moved into the shared voice kit, and the phone now keeps only how its own phases map to a presence mode. Two things were deliberately not built: new assets, and reactions to what Dad says. A keyword list would be hardcoding, and a model call per transcript piece would be latency and cost on every sentence. Both wait for use. On Air's broadcast face, in Track 8, is built on this same presence.

Code

A presence model: breath, lean, nod, tilt and a glide between moods·typescript
// The face's presence: signals in, a pose out. The same numbers on every surface.
type Mode = 'listening' | 'waiting' | 'speaking' | 'held';
interface Signals { mode: Mode; level: number; hearingWords: boolean; heardPieces: number }
interface Pose { scale: number; x: number; y: number; rotate: number }

const P = {
  breathSeconds: 4.2, breathScale: 0.008, heldBreath: 0.5,
  swayX: 0.008, swayY: 0.006, swayRotate: 0.6, swaySpeed: 0.5,
  leanScale: 0.015, nodSeconds: 0.7, nodY: 0.02,
  tiltAfter: 4, tiltRamp: 3, tiltRotate: 1.5, thinkY: -0.01, thinkRotate: -1.2,
  settle: 0.35, levelSeconds: 0.12, maxStep: 0.1,
};

/** Three sines whose periods never line up: a wander that doesn't visibly repeat. */
const wander = (t: number, phase: number) =>
  (Math.sin(t + phase) + 0.6 * Math.sin(1.618 * t + 1.3 + phase) + 0.3 * Math.sin(2.718 * t + 0.7 + phase)) / 1.9;

/** Ease toward a target: a change of mode is a glide, never a jump. */
const approach = (now: number, target: number, dt: number, tau: number) =>
  now + (target - now) * (1 - Math.exp(-dt / tau));

export class Presence {
  private t = 0; private level = 0; private live = 1; private think = 0; private tilt = 0;
  private pieces: number | null = null;
  private nodAt: number | null = null;
  private wordsSince: number | null = null;

  step(s: Signals, dt: number, motion: number, reducedMotion = false): Pose {
    const step = Math.max(0, Math.min(dt, P.maxStep));       // a hidden tab is not a jump
    this.t += step;
    const listening = s.mode === 'listening';
    this.level = approach(this.level, listening ? s.level : 0, step, P.levelSeconds);
    if (listening && this.pieces !== null && s.heardPieces > this.pieces) this.nodAt = this.t;
    this.pieces = s.heardPieces;
    this.wordsSince = listening && s.hearingWords ? this.wordsSince ?? this.t : null;
    const heardFor = this.wordsSince === null ? 0 : this.t - this.wordsSince;
    const tiltTarget = Math.max(0, Math.min(1, (heardFor - P.tiltAfter) / P.tiltRamp));
    this.tilt = approach(this.tilt, tiltTarget, step, P.settle);
    this.think = approach(this.think, s.mode === 'waiting' ? 1 : 0, step, P.settle);
    this.live = approach(this.live, s.mode === 'held' ? 0 : 1, step, P.settle);
    let nod = 0;
    if (this.nodAt !== null) {
      const progress = (this.t - this.nodAt) / P.nodSeconds;
      if (progress >= 1) this.nodAt = null; else nod = Math.sin(Math.PI * progress);
    }
    const amount = reducedMotion ? 0 : Math.max(0, Math.min(2, motion));
    if (amount === 0) return { scale: 1, x: 0, y: 0, rotate: 0 };
    const w = this.t * P.swaySpeed;
    const breath = Math.sin((2 * Math.PI * this.t) / P.breathSeconds)
      * (P.heldBreath + (1 - P.heldBreath) * this.live);
    return {
      scale: 1 + amount * (P.breathScale * breath + P.leanScale * this.level),
      x: amount * P.swayX * this.live * wander(w, 0),
      y: amount * (P.swayY * this.live * wander(w, 2.1) + P.nodY * nod + P.thinkY * this.think),
      rotate: amount * (P.swayRotate * this.live * wander(w, 4.2) + P.tiltRotate * this.tilt
        + P.thinkRotate * this.think),
    };
  }
}

const face = new Presence();
const marks = new Set([0, 10, 33, 50, 63, 80, 95, 110]);   // tenths of a second worth printing
for (let i = 0; i <= 110; i += 1) {
  const t = i / 10;
  const mode: Mode = t < 9 ? 'listening' : 'waiting';
  const talking = t >= 1 && t < 8.5;
  const pieces = t >= 6 ? 2 : t >= 3 ? 1 : 0;                // two pieces of his words land
  const pose = face.step({ mode, level: talking ? 0.7 : 0.05, hearingWords: talking,
    heardPieces: pieces }, 0.1, 1);
  if (marks.has(i)) {
    console.log(`${t.toFixed(1).padStart(4)}s ${mode.padEnd(9)} scale ${pose.scale.toFixed(4)}`
      + `  y ${(pose.y * 100).toFixed(2)}%  rotate ${pose.rotate.toFixed(2)}°`);
  }
}
console.log('reduced motion:', face.step(
  { mode: 'listening', level: 1, hearingWords: true, heardPieces: 2 }, 0.1, 2, true));

External links

Exercise

Run the model with npx tsx and read the eight lines: find the lean while he talks, the nod after each piece, the tilt after four seconds of words, and the look up once she prepares a reply. Then run it with motion 0 and with motion 2. Finally, add a 'held' stretch (a pause) and confirm the drift fades out and the breath halves rather than stopping dead.
Hint
The nod shows in y within a second of each piece landing, deepest about 0.35 seconds in, because it is a half sine over 0.7 seconds. In a held stretch, the live factor eases toward zero, which fades the wander and scales the breath toward half; nothing snaps, because the live factor eases with the same time constant as the tilt and the look up.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.