Egocentric video: teaching AI to see through human eyes.
Robots and assistants that act in the physical world need to learn how people actually do things. First-person video is one of the richest ways to show them.
A camera where the eyes are.
Egocentric video is footage recorded from the point of view of the person doing an activity, usually with smart glasses or a head-mounted camera. Instead of watching someone from across the room, the model sees what they see: their hands, the tools they reach for, and where they look.
That view is exactly what physical AI needs: humanoid and industrial robots, AR assistants, and embodied agents that must understand intent, plan steps and manipulate objects in messy, real environments.
Main kinds of egocentric data.
Hand–object interaction.
Close-up footage of hands grasping, using and handing over objects — the core signal for robot manipulation.
Procedural & task video.
Step-by-step recordings of real tasks: cooking, assembly, repair, lab work, warehouse picking.
Navigation & locomotion.
Walking through homes, streets and facilities to teach spatial awareness and path planning.
Social & conversational.
First-person views of people talking and collaborating, with gaze and speech in context.
Ego-exo (multi-view).
The same activity filmed from the wearer's eyes and from external cameras, synchronised.
Multi-sensor capture.
Video paired with eye tracking, IMU motion, depth, audio or hand pose for richer training signals.
Why good first-person data is hard to get.
Motion and blur.
Head movement makes footage shaky, and hands often block the very action you need to see.
Privacy and consent.
Wearable cameras capture bystanders, faces, screens and homes. Every frame must be consented and handled under GDPR.
Diversity.
Models trained in one kitchen fail in another. You need many people, places, cultures, and lighting conditions.
Long, unstructured footage.
Hours of video contain a few seconds of useful action. Finding and segmenting it is costly.
Hard-to-label actions.
Where does 'pick up' end and 'pour' begin? Fine-grained action boundaries need clear rules and trained annotators.
Hardware and sync.
Glasses, head mounts and extra sensors must be calibrated and time-aligned across devices.
Public datasets that shaped the field.
Ego4D
Around 3,670 hours of daily-life first-person video from 900+ participants in 9 countries.
Ego-Exo4D
Synchronised first-person and external views of skilled activities such as cooking, sports and repair.
EPIC-KITCHENS-100
100 hours of unscripted kitchen activity with dense action annotations.
Project Aria
Research glasses and open datasets combining video, eye tracking and motion sensors.
Public datasets are great for research, but most have non-commercial licences and don't cover your tasks, environments or robot. That's where custom capture comes in.
Custom egocentric data, captured and labeled for you.
Global contributor network.
70,000+ vetted experts in many countries, so footage covers real diversity of people, homes and cultures.
Privacy built in.
Informed consent, face and screen blurring, ISO 27001, SOC 2 Type II and GDPR-compliant handling.
Domain experts on camera.
Chefs, technicians, nurses and warehouse staff performing the tasks your robot needs to learn.
Device flexible.
Smart glasses, head or chest mounts, and extra sensors, set up to your specification.
From brief to training-ready footage.
Design.
We define tasks, environments, devices and the label schema with your robotics or research team.
Capture.
Consented contributors record real tasks in homes, workplaces and facilities, on the devices you choose.
Annotate.
Action segments, hand–object labels, object tracking, narrations and step boundaries, by trained annotators.
Verify.
Double-pass review, agreement scoring and privacy redaction before every delivery.
Learn more about our AI Data Foundations services.