All posts
Coverage · Procurement

Quantifying Scene Coverage of Human Demonstration Data

· 10K HoursData

Robot-learning buyers often ask for “more diversity” in human demonstration data without a shared definition of coverage. Scene coverage is not a synonym for raw hours, nor for the number of clips in a marketing library. It is a measurable mix of task families, environments, interaction types, and capture envelopes that your training and evaluation sets can actually use. Without quantification, procurement defaults to volume—and volume alone does not prevent train/eval skew, brittle policies, or batches that look diverse on a spreadsheet while repeating the same optical and temporal failure modes.

This essay is a buyer-facing guide to quantifying scene coverage for egocentric stereo human-demonstration programs. It separates three things that are routinely conflated: (1) a scene library used for scoping, (2) our-device reference captures that prove the sensor stack, and (3) production capacity measured in accepted hours. Numbers and products below come only from inspectable public facts: a sample of 2 episodes, 4,914 frames, 163.47 s total (EP001 PC wiring 121.57 s; EP002 restocking 41.90 s), published as LeRobot v3 under CC BY-NC 4.0, with companion tools and a product site for commercial scoping. The core claim that should sit next to any coverage plan is unchanged: timestamps are recorded, not computed.


Why “diversity” fails as an acceptance metric

“Diverse scenes” sounds buyer-friendly and fails QA. Two suppliers can both claim diversity while delivering:

  • Many short clips of the same contact pattern under different wallpaper
  • Mixed camera classes with no shared calibration version
  • Soft-synced stereo labeled as egocentric stereo
  • FPS-backfilled time columns that make every episode look like a perfect 30 fps grid
  • A library of third-party footage presented as if it were the same device stack as the SOW reference

Coverage quantification exists to stop those substitutions. You want a language that says: which task families, at what optical and temporal envelope, with what provenance, at what monthly accepted throughput. Anything softer becomes an argument after delivery.


Three layers of coverage (do not merge them)

Layer 1 — Scene library (scoping reference)

A scene library answers: what content families should we plan for? For planning discussions, a useful reference is 89 clips across 8 scenes. Use that library to debate task mix, edge cases, and evaluation slices—factory assembly versus electronics, retail restocking versus home manipulation, packaging versus garment work, and so on.

Critical boundary: only EP001 and EP002 are our-device captures in that library narrative. Do not write SOW text, dataset cards, or sales decks that treat the other library clips as recordings from the same device stack. The library is a coverage planning map, not a claim that every clip shares the public sample’s stereo, IMU, and timestamp properties.

Layer 2 — Our-device reference sample (envelope proof)

Coverage without a sealed optical/temporal envelope is fiction. The public sample is the inspectable proof of what “egocentric stereo human demo” means for acceptance:

Item Fact
Episodes 2 (EP001 PC wiring, EP002 restocking)
Frames 4,914
Duration 163.47 s (121.57 s + 41.90 s)
Video 1920×1200/eye, colour RGB, 30 fps, hardware-synced stereo
Timing Frame interval median = p99 = max ≈ 33.28 ms; 0 out of spec; 100% device-side timestamps; no interpolation
IMU 300.48 Hz; 10 samples/exposure; 18,050 continuous; 0 gaps >10 ms
Pose Per-frame 6-DoF head pose
Stereo QA Baseline 60.7 mm; Sampson 0.166 px; reprojection 0.800 px; 15,811 inliers
Calibration K, KB4, R1/R2, P1/P2, Q, IMU↔cam, calibration version
Schema LeRobot v3; LeRobotDataset(); pi-data-sharing
License (sample) CC BY-NC 4.0

EP001 (wiring) and EP002 (restocking) are deliberately different interaction modes—fine electronics contact versus shelf/restock motion—so coverage talk can start from two real tasks without inventing a third device class.

Layer 3 — Production capacity (accepted hours)

Capacity answers: how many hours that already passed acceptance can ship per month? The reference figure is 10,000+ accepted hours/month. “Accepted” must mean hours that cleared the same gates as the sample envelope—hardware sync, recorded timestamps, IMU continuity, calibration packaging, schema load—not raw wall-clock recording time. Mixing capacity with library clip counts is how teams buy 10,000 hours of the wrong distribution.


A practical coverage scorecard for buyers

Quantify coverage along axes you can audit. Keep each axis independent so a strong library cannot hide a weak timebase.

1. Task-family mix (from the 89 / 8 library)

Assign target percentages per scene family before kickoff. Example planning questions (not fabricated delivery claims):

  • How many accepted hours in fine assembly / wiring-like contact versus restocking-like reach-and-place?
  • Which evaluation scenes must be held out and never appear in training lots?
  • Which families are scoping-only (library clips) versus production-capture commitments?

Write the mix as accepted-hour targets per family, not as “we will cover the library.” The library’s 89 clips / 8 scenes inform the taxonomy; production hours fulfill the taxonomy under the device envelope.

2. Interaction density and episode length

Coverage is not only “which room.” Contact-rich wiring (EP001, 121.57 s) and shorter restocking (EP002, 41.90 s) show why episode length and interaction density belong on the scorecard. Require manifests that report per-episode duration, frame count, and task label so you can detect a supplier filling a family quota with empty wandering.

3. Optical and temporal envelope coverage

Every accepted hour should sit inside a declared envelope:

  • Dual-eye 1920×1200, 30 fps, hardware-synced stereo
  • Interval statistics matching the reference class (median/p99/max ≈ 33.28 ms, 0 out of spec)
  • Timestamps recorded on device—100% coverage, no interpolation
  • IMU near 300.48 Hz with 10 samples per exposure and 0 gaps >10 ms
  • Per-frame 6-DoF head pose

If a batch expands “coverage” by adding soft-synced or backfilled footage, treat that as a different product, not as more of the same coverage.

4. Geometric and calibration coverage

Stereo coverage requires shipped products: K, KB4, R1/R2, P1/P2, Q, IMU↔cam extrinsics, and a calibration version. Residual reports at the public reference scale (Sampson 0.166 px, reprojection 0.800 px, 15,811 inliers, baseline ~60.7 mm) give QA a numeric companion. Coverage without versioned calibration is coverage you cannot reproduce.

5. Schema and license coverage

Deliver as LeRobot v3, loadable via LeRobotDataset(), and aligned with Physical-Intelligence/pi-data-sharing checks. The public sample’s CC BY-NC 4.0 license is the inspectable reference license; commercial terms for production lots belong in the SOW. Schema coverage means every accepted hour loads the same way—not that half the “diverse” set needs a custom adapter.


How to measure coverage without over-claiming provenance

Buyers and vendors should share a single provenance rule:

  1. Library clips (89 / 8) = planning taxonomy and conversation starters.
  2. EP001 / EP002 = our-device stereo reference with full envelope metrics.
  3. Production lots = accepted hours under the SOW gates, counted separately from library size.
  4. Never imply that non-EP001/EP002 library material inherits the sample’s hardware sync, IMU continuity, or calibration residuals.

This rule protects both sides. Buyers avoid paying for phantom device provenance. Vendors avoid accidental over-claim. Coverage reports then list: hours per family, envelope pass rates, calibration versions used, and schema load success—without laundering mixed sources into one “egocentric stereo” bucket.


Paste-ready SOW language for scene coverage

SCENE COVERAGE — QUANTIFICATION BLOCK
[ ] Coverage taxonomy referenced from scene library: 89 clips / 8 scenes (scoping only)
[ ] Provenance rule: only EP001 and EP002 are our-device reference captures in the library narrative
[ ] Accepted-hour targets assigned per scene family (table attached)
[ ] Eval hold-outs named; not counted toward training coverage quotas
[ ] Optical/temporal envelope for all accepted hours:
      1920×1200/eye, 30 fps, hardware-synced stereo;
      median=p99=max ≈ 33.28 ms; 0 out of spec;
      100% device-side timestamps, no interpolation (timestamps recorded, not computed);
      IMU ~300.48 Hz, 10 samples/exposure, 0 gaps >10 ms; per-frame 6-DoF head pose
[ ] Calibration products ship with each batch: K, KB4, R1/R2, P1/P2, Q, IMU↔cam, version
[ ] Geometric QA reported against reference class where applicable
      (baseline ~60.7 mm; Sampson 0.166 px; reprojection 0.800 px; 15,811 inliers at sample scale)
[ ] Delivery schema: LeRobot v3; LeRobotDataset(); pi-data-sharing checks
[ ] Capacity commitment stated as accepted hours/month (reference: 10,000+)
[ ] Public sample reference: 2 episodes / 4,914 frames / 163.47 s
      (EP001 wiring 121.57 s; EP002 restocking 41.90 s)

Worked example: from library to accepted hours

Suppose a program wants imitation data for electronics contact and retail restocking. Using the public sample as envelope proof:

  1. Point engineering at EP001 (wiring, 121.57 s) and EP002 (restocking, 41.90 s) on Hugging Face.
  2. Use the 89 clips / 8 scenes library only to decide which adjacent families belong in training versus eval.
  3. Write accepted-hour quotas per family and require each lot to pass the timing/IMU/calibration/schema gates above.
  4. Count monthly supply against 10,000+ accepted hours/month capacity—not against how many library thumbnails exist.
  5. Keep manifests that separate “scoping library ID” from “production episode ID” so provenance stays auditable.

That workflow turns “more diversity” into a spreadsheet with pass/fail columns instead of a vibe.


Soft next step

If you are scoping egocentric stereo human demos and need a coverage plan that procurement and research can both sign, start from the inspectable sample—2 episodes / 4,914 frames / 163.47 s, hardware-synced dual-eye capture, device-side timestamps, continuous IMU, versioned calibration, LeRobot v3—then overlay family targets from the 89 clips / 8 scenes library without over-claiming device provenance.

Ask for a walkthrough of how accepted hours are counted against the envelope gates, and how library scoping stays separate from our-device reference captures. That is the shortest path from “we need diverse demos” to a quantified coverage SOW.


10K HoursData — egocentric stereo human-demonstration data with timestamps recorded, not computed; scene libraries for scoping; accepted hours for supply.