"For a face, the engine computes the answer. For a random bottle, it refuses to — and that refusal is the most principled feature it has."
No Canonical Model, No Claimed Orientation
Portrait, pose, and hand each rest on a canonical model — a known reference shape the geometry can solve against. A generic object — a bottle, a book, a car, a coffee cup — has no such model. There is no universal 'canonical bottle' whose landmarks you can match. And without a model, there is nothing to honestly recover an orientation from. So the object track does the disciplined thing: it claims no 3-D orientation at all. It will not infer a turn it cannot justify.
Detector-Seeds, Human-Steers
What the engine can honestly do is help. A 2-D detection seeds a box — a reasonable starting guess for where the object sits and how big it is — and from there the human fully controls the box's 3-D orientation by hand. The engine provides the scaffold and the manipulation tools; the artist provides the spatial judgment the data can't supply. It's the same human-in-the-loop philosophy as the focal and pose overrides, taken to its logical end: when the geometry has nothing to solve, the human owns the orientation entirely, and the engine is honest about handing it over.
Orthographic at Rest
The default representation is orthographic — parallel edges, no convergence, no invented vanishing points. This is deliberate, and it's the honest choice: an orthographic box asserts nothing about the camera or the perspective. It's the least-committal thing the engine can draw, which is exactly right when it has no basis to commit. Only when a human opts in by setting a focal length does the box switch to true converging perspective, complete with vanishing-point guides. Perspective is a claim about the camera, so the engine makes you author that claim rather than fabricating one.
The Discipline of Not Faking It
Step back and this feature carries the whole engine's ethic in miniature. Where a face gives geometry enough to compute an answer, it computes one. Where a random object gives it nothing, it computes nothing and says so — seeding, assisting, and defaulting to the humblest representation available. Knowing when not to claim an answer is as much a part of a trustworthy tool as computing one when it can. An engine that guessed an orientation for every object would be more impressive in a demo and worthless as a reference. Loomis chose worthless-in-the-demo, trustworthy-in-truth.