How BitRobot Crowdsources Real-World Data for Embodied AI, with Jonathan Victor
The usual story of how modern AI arrived is about algorithms, the transformer that finally cracked language. It leaves out the part that made it possible, how thirty years of the internet had produced a training set, and the models that emerged were just reading it back. The raw material is free.
Nothing on the web can teach a machine what it feels like to close a hand around a cup, or to catch its balance when a foot slips. That data has to be made physically, one interaction at a time. A robot can't learn to fold clothes from Wikipedia. Real-world interaction data is the bottleneck now, more than compute or model design, and it leaves embodied AI with a question language never had to answer: who makes all that data, and how?
Jonathan Victor, president of BitRobot, joins Amira Valliani on the latest episode of Bits to Bricks to answer it, and he has bet his career on one answer: don't build the factory, build a network.
How FrodoBots built a real-world robot dataset
BitRobot started as FrodoBots, little sidewalk robots run as a game. Anyone could log in, take control of someone's robot through a browser, and drive it around a real city on a scavenger hunt, picking up virtual items overlaid on the street. Thousands of people across dozens of cities played, and the exhaust of all that play was the data a navigation model needs: how a human handles a dog running up, a puddle, a crowd. FrodoBots open-sourced roughly 2,000 hours of it as FrodoBots-2K, the largest urban robot navigation dataset in the public domain, since used by teams at DeepMind, Meta, and UC Berkeley, where it trained a general navigation model. Before that release, Victor says, the largest public dataset for urban navigation was about 60 hours. "Suddenly overnight your data set has like 30x in size," he says of the researchers who came out of the woodwork. That reaction was the seed of BitRobot.
Why is real-world data the bottleneck for robots?
The hard part about robotics data is hands: a robot picking things up, folding a towel, handling a cup without crushing it. Real-world interaction data like this is a harder problem than compute or model design, and the default way to make it looks like a factory: hire people, put them in teleoperation rigs, and record them driving robots through one task after another. It is clean and controlled, but slow, because a skilled operator produces only a few dozen usable demonstrations a day. What labs actually want, Victor says, is data collected "not in some sanitized setting, but when it's representative of the real world."
How BitRobot's subnets replace teleoperation
Instead of one game for one domain, BitRobot is a market of subnets, each collecting a different kind of robot data (urban navigation, mobile manipulation, dexterity) through a different method (teleoperation, simulation, egocentric video). "Let's create the tools so you can create a market for that domain," Victor says, "and then let the market decide which are the most valuable data sets." The current flagship is the RoboCap, a low-cost egocentric-video device with six cameras that captures a first-person view and the wearer's hands. It is priced around $1,000 for researchers, roughly a seventh of comparable rigs, and cheaper still for network contributors. The plan is to get thousands of them onto people doing ordinary jobs, in bakeries, hotels, and factories. You wear it, upload, and the data is scored, annotated, anonymized, and face-blurred before it is sold to a lab and you are paid retroactively. A centralized company chasing the same coverage, Victor argues, would have to run operations in 200-plus countries.
What training data do AI labs actually pay for?
Raw volume is close to worthless for AI labs. The value sits in the long tail, the 0.1% case almost no one captures. After a hundred hours in one factory, everyone has seen everything useful there. BitRobot scores contributions on entropy, how much new information a clip adds, through what it calls verifiable robotic work, so novel data earns more than the millionth laundry fold.
It’s similar to Tesla's flywheel for autonomous cars—but using a distributed model. The more cars on the road, the more edge cases the fleet catches. In a centralized model, Victor says, "the useful hours of data are subsidizing all of the non-useful hours. The point of a network like BitRobot is to cut that tax so the person who happens to witness the rare task is the one who gets paid. The network leans on fast feedback, a quality score, and incentives to ensure good operators.
How Solana and tokens reward contributors
BitRobot uses Solana for two jobs: accounting for who contributed what work, and paying them for it. "Solana makes that very cheap, very easy, tons of tools and infrastructure, a very strong community," he says. "Solana is the best place for it." Victor also clarifies that tokens alone do not make a network work, which is why BitRobot spends most of its effort lining up commercial contracts with labs before scaling supply.
The larger stake is ownership. If physical-AI data becomes strategic infrastructure, whoever owns the corpus owns the field. A distributed is not obviously better here, as Valliani puts it, only messier, and the real question is whether that broader, (sometimes) messier data set is a structural advantage, or a problem that labs would rather solve with capital and control. BitRobot's bet is that the mess, the diversity and the long tail, is the advantage, and that the data training embodied AI should end up with a network anyone could join rather than a few companies with the biggest teleoperation farms.
This article draws from our conversation with Jonathan Victor, president of BitRobot. For the full discussion, listen to the episode of Bits to Bricks.
