A model that wins on the benchmark and a model that works for your users are not the same model.
The benchmark is clean, balanced and finished. The world is none of those things. Lighting
changes, the camera moves, the user phrases it differently, the gripper slips, the
distribution drifts, and the number that looked settled in the paper quietly stops holding.
Most teams find this out in production. They have a result they trust, a
demo that impressed the room, and no instrumentation between that and the thing their
customers actually touch. The fix is rarely a bigger model. It is measurement that reflects
the deployed setting, and the engineering discipline that measurement makes possible.
We work on both sides of that gap, because we do not think they are separable.