Physical AI Has a Data Problem. India's Kitchens, Ironing Boards, and Textile Floors Are Solving It.
The world's most advanced AI labs have run into a wall no amount of text can solve. Language models learned to talk — they still don't know how to fold a shirt. The raw material for Physical AI isn't text. It's footage of human hands doing ordinary work.
In April 2026, a video of textile workers in a Delhi NCR factory went quietly viral. They were assembling and finishing garments with small cameras strapped to their foreheads. The internet's first assumption was correct: someone was recording them to train robots. The headgear traced back to Egolab.AI, a first-person point-of-view data aggregator founded by two teenagers who, within months of starting the company, had built something global robotics labs were willing to pay seven figures to acquire.
That single clip is a useful snapshot of a much bigger shift. The world's most advanced AI labs — the same ones that spent the last five years scraping every word on the internet to train language models — have run into a wall that no amount of text can solve. Language models learned to talk. They still don't know how to fold a shirt, iron a collar without scorching it, or thread a bobbin without watching a human do it first. That gap has a name now: Physical AI. And the raw material it runs on isn't text. It's footage of human hands doing ordinary work.
From simulation to the real world
For years, robotics labs bet on simulation — building physics engines good enough to teach a robot arm to grasp, twist, and place objects without ever touching the real world. That bet has plateaued. Simulating physics with perfect accuracy has proven extraordinarily difficult, and for robots to fold laundry, wash dishes, or work safely alongside humans, they need to learn from real footage of real people doing real things. The result is a demand curve that has gone vertical: investors put more than $6 billion into humanoid robotics in 2025 alone, and Physical AI companies are now spending upward of $100 million a year buying real-world training data.
The industry's own demand estimates are staggering. Estimates for the data requirement of leading robotics labs run from hundreds of millions to over a billion hours of egocentric data over the next two to three years, with broader industry consensus placing the total closer to a few billion hours of data that has to be created, not scraped. A single task context alone can require anywhere from 100,000 to 1 million hours. Read that last part again. One task. One context — say, picking up a glass in a specific kitchen — can require six figures of footage before a model generalizes reliably.
This is why the leading edge of robotics data collection has moved past sanitized lab teleoperation. Early research efforts in handheld and wearable capture rigs started the movement toward cheaper, human-generated demonstration data. Since then, Physical AI companies have built out exoskeleton-based capture systems, wearable rigs matched to a robot's own joint configuration, and proprietary teleoperation fleets embedded inside live manufacturing deployments, alongside closed-loop "data flywheels" run from their own deployed robot fleets. But the cheapest, fastest-scaling category of all doesn't need a robot in the room at all. It needs a human, a task, and a camera. That's egocentric data — and it has found its largest, lowest-cost supply base in India.
Why the data is being recorded from rural BPO cities in India, not California
India didn't choose this role by accident. Three things made it inevitable.
India's textile industry directly employs nearly 4.5 crore people, much of it in rural India, and holds roughly 3.9% of global market share as the world's sixth-largest exporter of textiles and apparel. That is precisely the kind of repetitive, fine-motor, high-variation manual work — stitching, ironing, folding, quality-checking — that robotics labs are desperate to model and that no simulation has been able to replicate convincingly.
India's gig and BPO ecosystem is already capturing thousands of hours of 4K first-person footage a day at national scale, with major annotation players running comparable volumes against strong internal demand. This is India's BPO scale advantage, redeployed from voice and text to motion and gaze.
Industry estimates put India's AI data annotation industry — the layer that labels, structures, and quality-checks this footage — on a path to exceed $7 billion by 2030 and employ up to a million people. That is not a forecast about robots. It's a forecast about India's role in the physical AI supply chain, whether or not a single humanoid ever ships from an Indian factory floor.
Volume is already being commoditized
Here is what the market has not yet fully absorbed: the egocentric data gold rush is moving faster into commoditization than the last one did. Industry data shows collection rates falling roughly 30% between January and June 2026 alone, as more suppliers entered the market — a compression curve that took the BPO industry the better part of a decade to replicate, and one physical AI data supply is running through in six months.
Market watchers who track this space closely put it plainly: most businesses currently flourishing in raw data collection are not durable ventures. The businesses that will still matter in three years won't be the ones who can strap the most cameras to the most heads fastest. They'll be the ones who treat egocentric capture as the first stage of a governed data pipeline — controlled task protocols, calibrated hardware, multi-stage quality review, and domain-specific structuring — rather than a one-time footage sale. Leading Physical AI companies are already converging on stage, purpose-built capture facilities for kitchens, laundry, and industrial settings, and calibrated-facility approaches are emerging from multiple markets across South Asia pointing to the same conclusion from opposite ends of the industry: infrastructure and quality governance are the moat.
Governance as the moat
This is the exact discipline Nextura.ai has practiced over the years, applied to a new substrate. We have spent years building and governing large, distributed human workforces for language, content and now data — built specifically to extend that governance model into AI data: sourcing, structuring, quality-checking, and delivering data that AI systems can actually learn from, not just consume.
The parallel to physical AI's current moment is direct. Just as our multilingual annotation workforce brings language and cultural context that generic labeling pipelines miss, the same workforce discipline — protocol design, calibrated capture, layered QA, domain-specific task structuring — is exactly what separates a durable physical AI data partner from a gig-economy footage vendor. Household chores, ironing and garment finishing, and textile-floor workflows are not generic motion capture; they carry cultural, material, and technique-specific nuance that only a workforce trained to recognize it can encode correctly. That is the gap between data that trains a model and data that trains a model well, and it is where we believe India's next data export category — and Nextura.ai's own AI Data practice — is headed.
For robotics and LLM labs
The physical AI race will not be won by whoever collects the most hours of footage. Leading Physical AI companies have already logged well over 100,000 hours each, with some delivering several hundred thousand hours across hundreds of real-world settings. Volume is table stakes now, not a differentiator. The labs that win will be the ones who partner with data operators who understand governance, protocol design, cultural and domain context, and quality control as deeply as they understand collection.
India is going to train a meaningful share of the world's robots — in its factories, its homes, and its ironing rooms — whether or not the world is watching closely. The question worth asking is not whether this data gets created. It's who creates it with the discipline to make it worth using.
