Spatial data, in three parts

Context is the scarce part.

Three posts from September 2026 on what gets thrown away when we record people moving through the world.

  1. The frame is data too18 Sep 2026
  2. You cannot add that later23 Sep 2026
  3. Scaling the place, not the person25 Sep 2026

Part 1  ·  18 September 2026

The frame is data too

Point at a car and say "the ball is in front of it." If you think in English, you mean the side facing you. A Hausa speaker means the far side.

Neither of you is wrong. Linguists call this a frame of reference. English reverses the front/back axis and leaves left and right alone, so you treat the car as facing you. Tamil rotates the whole frame. Hausa applies your own frame directly, so the car faces where you face. Same room, same ball, three geometries.

Diagram. A ball has no front, so each language lends it one. In English and Tamil, in front of the ball is the side facing you; in Hausa it is the far side. English and Hausa agree on left; Tamil flips it. No two languages agree on both.

Machines have picked one. A benchmark called COMFORT tested nine vision-language models across 109 languages. Ask GPT-4o in Tamil and it gives the English answer. Ask in Hausa, same. Speakers of both prefer something else. Prompt a model to use a different frame explicitly and accuracy falls to roughly chance. Telling it to think in someone else's terms is the same as breaking it.

That is a benchmark, so it is easy to file under interesting.

Then you find it shipping. GuideDog is a dataset of 22,000 real street scenes from 183 cities in 46 countries, built to teach models to guide blind and low-vision people through traffic. Every hazard gets a clock-face direction and a distance in steps.

One step is 0.7 meters. It is 0.7 meters in all 46 countries, for every user, whatever their height or gait or mobility aid. The number comes from a paper published in 1997.

The authors know. Their limitations section has a heading called Cross-Cultural Spatial Language, and under it they write that the annotations rely on universal spatial references and do not model cross-cultural variation. They mark it as future work and ship.

That is two careful teams, working honestly, arriving at the same gap from opposite directions and both leaving it open. Meanwhile we are capturing human point-of-view video at enormous scale and labeling it as though position and orientation were the whole signal.

They are not. The frame is data too. Nobody is recording it, so the machine inherits one person's defaults and files them as measurement.

Math is universal and close to solved. Convention is local, unlabeled, and being standardized by accident. If you are building anything that tells a person where to go, that is not a research problem. It is a product decision someone on your team is making right now without knowing they are making it.

I trained as an archaeologist before any of this. You learn early that an object pulled from the ground without its depth, position and orientation recorded is mostly worthless, and that digging destroys the evidence, so the record is all that survives. Twenty-five years later I am reading papers about spatial data and finding the same lesson, unlearned.


Part 2  ·  23 September 2026

You cannot add that later

Ego4D is 3,670 hours of first-person video from 931 people in 9 countries. The most geographically diverse egocentric dataset anyone has built, and it will still teach a machine one culture's sense of direction.

Last week: "in front of the car" means a different place in English, Tamil and Hausa. Asked in all three, GPT-4o gives the English answer. (Disclosure: I advise Biel Glasses, which builds a device for people with low vision.)

The fix everyone reaches for is more footage from more places. But look at what actually gets written down. Position, orientation, object, action. Never whose frame the annotator was standing in.

Venn diagram. Ego4D records position, orientation, object and action. GuideDog records obstacle type, clock direction, and distance in steps at 0.7 meters per step. The overlap, whose frame, is recorded by neither.

GuideDog is the cleaner example. 22,084 street scenes, 183 cities, 46 countries, annotated to guide blind and low-vision people. Every obstacle gets a clock direction and a distance in steps. One step is 0.7 meters, from a paper published in 1997.

Nobody measured those steps. A depth model estimates the distance, and the estimate gets divided by 0.7. The images come from YouTube walking videos, so the person wearing the camera is someone filming a walking tour, and nothing about them is recorded either.

Video is about to be everywhere. Glasses are shipping and capture is nearly free. Context is the scarce part: who wore the camera, how they think about direction, what one step means for the person actually walking it.

You cannot add that later. It exists only at the moment of capture, which is the moment it gets thrown away.

Anyone building a dataset this week is deciding whether in five years they own something nobody can rebuild, or another pile of footage that looks like everyone else's. Very few of them are treating it as a decision.


Part 3  ·  25 September 2026

Scaling the place, not the person

Niantic Spatial shipped 100 real environments for robots to train in this month. World Labs, a billion dollars in the bank, can now generate one from a photograph.

Both are geometry. Neither has a person in it.

(Disclosure: I advise a company whose product generates the missing signal as a by-product. I am not neutral here.)

That gap matters more than it sounds, because of a result from June. HumanScale matched egocentric human video against teleoperated real-robot trajectories at 5,000 hours each. The human video won: 24 percent lower validation loss on real-robot action prediction, 52.5 percent higher success in distribution, and 90 percent higher out of it.

So the field now knows that recordings of people moving through the world are the better pretraining source. And the field is spending its money on the world, not on the people.

Look at the supply side. Every outdoor, naturalistic, sensor-equipped egocentric navigation dataset I can find, added together, comes to 67 hours. EgoWalk is 50 of them. EgoTraj is 10.7. EgoCogNav is 6. One hundred people wearing a camera for an hour a day would produce all of it in a single day.

Chart, log scale. Hours of egocentric video by corpus: Egocentric-1M 1,000,000; Egocentric-10K 10,000; Ego4D 3,670; Ego-Exo4D 1,286; EgoDex 829; EPIC-KITCHENS 100. Outdoor human navigation: EgoWalk 50, EgoTraj 10.7, EgoCogNav 6. Combined, 67 hours.

EgoCogNav is the one to read. They set out to capture what they call perceived path uncertainty: the hesitation, the visual scanning, the backtracking a person produces when they are not sure the ground ahead is safe. They name it as the human factor every existing navigation dataset leaves out. To measure it, they recruited 17 sighted volunteers, walked them along predetermined waypoints in eye-tracking glasses, and had them report their own uncertainty on a handheld joystick.

Six hours. That is the state of the art for recording how a person behaves when they are not sure.

We are about to have photorealistic, physically plausible worlds and almost no record of how a human actually moves through one. Simulation scales the stage. It does not scale the actor.

If you are training anything that has to move through a real place: what fraction of your data is the place, and what fraction is the person?


Talk to me

If you are building a dataset, or a product that tells a person where to go, book time directly.

Book a call