Human demonstrations, seen from the inside and the outside.
Sourcebae captures head-mount, egocentric, and exocentric demonstration data - time-synced, calibrated, and annotated - to train robot manipulation policies at scale.
One demonstration, every viewpoint a policy needs
Most manipulation models learn best from complementary views. We run all three modalities — independently or rig-synced on the same take — so you choose the coverage your architecture wants.
Head-mount capture
Lightweight head rigs let demonstrators move and work hands-free — ideal for mobile and bimanual tasks where a fixed camera can't follow.
Forward RGB-D + onboard IMU
Natural, unconstrained motion
Optional eye tracking for gaze
Indoor, outdoor, in-the-wild
Egocentric
The world as the demonstrator sees it: hand–object interaction, occlusion, and intent captured from the point of action.
First-person manipulation frames
2D/3D hand & finger pose
Gaze and attention signals
Aligned action & contact labels
Exocentric
Calibrated multi-camera rigs observe the scene from the outside, giving full-body context and a third-person frame robots can map to.
2–8 synchronized external views
Full extrinsic/intrinsic calibration
Whole-scene & trajectory coverage
Cross-view 3D reconstruction-ready
Ego shows intent. Exo shows context. You need both
First-person grounds the action
Egocentric frames keep the manipulated object and the hand in view through occlusion, so policies learn what the operator actually attends to and does.
Third-person generalizes the viewpoint
Exocentric views expose the same task from arbitrary angles, helping policies transfer to a robot's own camera placement instead of overfitting one POV.
Paired views unlock cross-view learning
Time-synced ego + exo on a single take is the substrate for view-invariant representations, retargeting, and sim-to-real alignment.
Synchronized, calibrated, annotated streams
Every take ships as aligned multi-modal data with the metadata your training pipeline expects — not raw footage you have to clean.
Video
30–60 FPS, ego & exo
Depth
Per-frame, registered
Motion
Accel + gyro, 200 Hz
Hand pose
2D/3D keypoints
Gaze
Fixation + attention
Action labels
Segmented, timestamped
Calibration
Intrinsics + extrinsic
Sync
Shared clock, μs-level
From task spec to policy-ready dataset
Scope & design
We map your tasks, environments, and target schema, then design the modality mix and capture protocol.
Rig & calibrate
Head rigs and external cameras are set up, calibrated, and clock-synced to a shared timeline.
Collect
Trained operators perform demonstrations across the variations you specify — objects, lighting, locations.
Sync & QA
Streams are aligned to a shared clock and reviewed for coverage, calibration drift, and label quality.
Annotate
Hand pose, gaze, action segments, and contact events are labeled and reviewed against your guidelines.
Deliver
Packaged in your format with calibration and metadata, shipped in batches with a documented spec.
A managed collection workforce, not a one-off shoot
We run capture as an operation: trained operators, repeatable rigs, and QA built in - so volume goes up without quality going down.
500+
Trained operators
1k+
Hrs demo data / month
3
Modes | Head · Ego · Exo
48
Hour pilot turnaround
Frequently asked questions
Robot imitation learning data consists of recorded human demonstrations that show how physical tasks are performed. These datasets can include synchronized video, depth, motion, gaze, hand pose, contact events, action segments and task metadata for training and evaluating robot manipulation policies.
Tell us the task. We'll capture it.
Share your manipulation tasks and target schema - we'll come back with a modality plan, a sample spec, and a pilot you can train on.