What the models still can't see
Show a modern multimodal model a photograph of a kitchen and it will tell you almost everything worth knowing. Countertop material, time of day inferred from the light, the make of the stovetop, the fact that someone left a knife out who probably shouldn’t have. Ask it to write the scene as a paragraph and it will out-describe most people I know. This is not a small thing. Five years ago it was science fiction.
But describing a kitchen and knowing a kitchen are different acts, and the gap between them is where I keep getting stuck. I don’t mean this as a gotcha about consciousness — I have no interest in that argument. I mean something narrower and more useful: the model has never burned toast in that kitchen. It has never reached for a mug at 6 a.m. and found the cabinet rearranged. It has no history with the room. What it has is an extremely good statistical portrait of what kitchens, in general, tend to look like and how people, in general, tend to talk about them.
That portrait is built entirely from the outside. Every image a vision model has ever trained on is a frozen slice, stripped of the causal chain that produced it and the one that follows. The model sees the plate but not the meal before it or the sink after. It sees the person mid-stride but not where they were going or why they were late. Humans build our sense of a place by moving through it thousands of times and accumulating consequence — this drawer sticks, that burner runs hot, the light through this window means it’s almost time to leave. None of that accrues to a system that only ever receives isolated frames.
I think this is the real content behind the intuition that these models “don’t understand” the world, an intuition that’s usually stated too vaguely to be useful. It isn’t that the model lacks facts about kitchens — it may have more facts than I do. It’s that its relationship to any single kitchen is stateless. Every image is the first image. There is no thread of experience connecting one observation to the next, and without that thread there’s no way to develop the kind of knowledge that comes from being somewhere repeatedly and being changed by it, even slightly.
Robotics researchers have been running into a version of this for a decade, and it’s instructive that embodiment turned out to be so much harder than perception. A model can caption a hallway perfectly and still fail to walk down it, because walking down it requires representing your own position in a world that pushes back — friction, momentum, the fact that a decision made three steps ago constrains what’s possible now. Perception is a spectator sport. Acting in a place, and living with the results, is not.
None of this means the current systems are shallow or that the trajectory stops here. Give a model persistent memory of its own interactions with a specific environment, let its outputs feed back into what it perceives next, and you start to build something closer to a resident than a visitor. A few labs are already doing versions of this with long-horizon agents that operate in the same digital environment across sessions, and the resulting behavior looks qualitatively different from single-shot description — more like habit, less like commentary.
Until that’s the default rather than the exception, I’d treat fluent description as a separate skill from situated understanding, not a stepping stone to it. They can improve on independent curves. A model can get better at describing kitchens forever without getting one inch closer to knowing one, the same way you can memorize a city’s street map without ever having lived on any of its streets.
What I take from this, practically, is a bias in how I evaluate new multimodal releases. I’ve stopped being impressed by caption quality alone and started asking a different question: does this system’s account of a place change the second time it encounters it, in a way that reflects the first encounter? Right now the honest answer is almost always no. That gap — not raw description quality — is the frontier I’m actually watching.