In October 2019 a delivery robot stopped at a curb cut on the University of Pittsburgh campus. To the robot, nothing was wrong. It was waiting out of the road, on pavement its map said was clear. To Emily Ackerman, a doctoral student who uses a wheelchair, it was blocking the only accessible entrance to the sidewalk, and she was stuck in the street with traffic coming. Starship disputed her account after reviewing the footage, then updated its mapping at that intersection.
The fix was a change to the place. Nobody changed the robot's idea of who uses a curb cut.
One model, every body
Odyssey-3 launched on 8 October as a research preview. Odyssey is a Palo Alto lab that raised $310 million in June at a $1.45 billion valuation, with Amazon and AMD Ventures among its investors. One base world model sits underneath policies for robot arms, Flexion's humanoids, a car on real roads in India and indoor drones. CEO Oliver Cameron calls it "a big step toward a single intelligence that can understand and operate in the world around us."
One model, many bodies, is where this is going. That makes the important question a simple one: whose world does it understand?
The rig is fine. The walker is the variable.
Odyssey's real-world capture began with 25-pound backpacks: six cameras, two lidar sensors and an inertial measurement unit, carried on foot through California. That is serious hardware, more than most egocentric datasets ever get. This is not a complaint about sensors.
A world model learns two things from footage: how the world looks, and how the people in it behave. The camera records the first. The person carrying it decides the second.
What a script teaches
Capture walkers walk to a script. Go here, cover this route, keep moving. They are paid, rested, adult and alone. They see the bollard and the scooter dumped across the path and step around both without breaking stride. When they stop, the reason is in the frame.
Thousands of hours of that teach a model a set of rules nobody wrote down:
- People move at a steady speed.
- People move in a line, toward somewhere.
- People see you, and they get out of the way.
- People are one body, not two attached by a hand or a leash.
- When people stop, something visible stopped them.
Each rule is true of the walker on the script. Each is false for a large share of the people on any real sidewalk.
The hundred other walkers
Stand on a corner for ten minutes and count who breaks a rule.
A parent pushing a stroller with a four-year-old on the other hand, who lets go. Children are bad at reading traffic: a Royal Holloway study in Psychological Science found primary school children could not accurately judge the speed of vehicles going faster than 20 mph.
An 80-year-old crossing slower than the signal assumes. US crosswalk timing is built on 3.5 feet per second, and the research traffic engineers cite drops that to 3.0 wherever older walkers are common. By 2030, one person in six will be 60 or over.
A blind man trailing a wall with a cane, who stops when the sound of traffic changes. The Global Burden of Disease estimate for 2020 counted 43.3 million people who are blind and 295 million with moderate or severe vision impairment.
Then the rest: the wheelchair user who needs that one curb cut, the man on crutches, the dog walker with a leash across the path, the courier backing a trolley out of a van, the teenager reading a phone, the tourist who stops dead to photograph a cathedral, the crowd leaving a stadium.
None of them is following a script. Together they are most of the sidewalk. The scripted walker is the edge case.
And the unscripted behavior is the information. A hesitation marks a curb that is hard to read. A detour marks a surface that is hard to cross. A stop marks a sound, a gap or a risk the camera cannot see. A model that never saw any of it files all of that as noise, and files the people producing it as anomalies to route around. Then an arm, a humanoid or a car built on that model shares a sidewalk with them.
Nobody records who wears the camera
Ego4D, the largest academic egocentric corpus, has 931 wearers in 74 locations across 9 countries. It reports gender and age for the wearers who disclosed them, and nothing about disability for any of them. EgoCogNav, built specifically to capture hesitation and backtracking, recruited 17 sighted volunteers and walked them along predetermined waypoints. GuideDog, a dataset for guiding blind and low-vision walkers, samples its frames from YouTube walking videos and has its labels checked by three sighted annotators. Odyssey has not said who carried its backpacks.
What would change it
Not another sensor. Not a bigger simulator either: scaling the place does not scale the person.
The fix is who carries the camera, and whether they carry it on a script. Recruit capture walkers from the people the robots will meet: parents, older people, wheelchair and cane users. Let them walk their own routes, at their own pace, on their own errands. Record, with consent and with enough context to read it, everything a script filters out: the stops, the hesitations, the reroutes. An hour of a real person deciding whether the ground ahead is safe teaches a model something a thousand hours of confident walking never will.
Odyssey has the money, the hardware and one model already driving cars and humanoids. It has not said who carried the backpacks. Neither has anyone else.
Another well-funded stage with nobody on it who cannot see, who walks slowly, or who is holding a four-year-old's hand.
Disclosure: I advise a company working on this problem. I am not neutral here.
Related: Spatial data, in three parts