Home Robotics data

ROBO · Robotics & teleoperation LiDAR · trajectories · demonstrations

Data for robots that work in real rooms.

Robot learning fails on the messy middle: a half-occluded handle, an operator who recovered from a bad grasp, a segment boundary nobody defined. We label point clouds, trajectories, action segments and teleoperation demonstrations, and we tell you which episodes are worth keeping.

Your tooling · Multi-pass QA · Under NDA

What we label

From raw sensor frames to usable episodes.

Robotics datasets are rarely one task. A single manipulation episode may need 3D geometry, temporal boundaries, contact points and a judgement about whether the demonstration was any good. We handle each as its own labelled deliverable with its own quality target.

3D

LiDAR and point-cloud labeling

3D cuboids, instance and semantic segmentation of point clouds, ground and free-space marking, and consistent object identity carried across every frame of a sequence rather than re-guessed per frame.

FUSE

Sensor fusion and cross-modal alignment

Camera, depth, LiDAR and proprioception reconciled into one labelled view. We flag calibration drift and timestamp misalignment when we find it, because a label that is right in 2D and wrong in 3D is worse than no label.

TRAJ

Trajectories and waypoints

End-effector and base paths, waypoint annotation, object tracks through occlusion, and interpolation boundaries marked explicitly so downstream training knows what was observed and what was inferred.

SEG

Action and skill segmentation

Temporal boundaries for reach, grasp, transport, place and release, with sub-skill labels and language descriptions where your model needs them. Boundary conventions are agreed and written down before anyone starts cutting.

GRASP

Grasp, contact and pose keypoints

Contact points, approach vectors, gripper state transitions, hand and object keypoints, and articulated-object part labels for handles, drawers, lids and hinges.

DEMO

Teleoperation demonstration review

Human review and scoring of collected demonstrations: success and failure, recovery behaviour, jerk and hesitation, operator error against environment difficulty. Curation is often the highest-leverage labelling you can buy.

EGO

Egocentric collection, not just labeling

Wearable-camera capture of first-person task demonstrations, hand-object interactions and trajectories for embodied AI and robot foundation models. When a model keeps failing on a category, collecting the missing data is often cheaper than relabelling what you already have.

Why it is hard

Consistency over time is the whole game.

Image labelling is judged frame by frame. Robotics data is judged across a sequence, and the failure mode is drift: object 14 becomes object 22 after an occlusion, one annotator cuts the grasp at first contact while another cuts it at gripper close, and the model learns a boundary that does not exist. Most of our effort goes into stopping that, not into drawing the first box.

So we fix conventions before volume. A calibration round on a small sample surfaces the disagreements while they are still cheap, the rules that come out of it are written down, and every batch afterwards is checked against them with multi-pass review and consensus. You get an inter-annotator agreement figure per task type, an edge-case ledger, and the identity chains you can actually train on.

What we watch for

  • Identity drift across occlusion and re-entry.
  • Segment boundaries defined differently by different annotators.
  • Sparse returns at range mislabelled as free space.
  • Timestamp and calibration drift between sensors.
  • Reflective, transparent and thin-structure objects.
  • Successful-looking episodes that contain a silent recovery.
Curating which episodes to keep changes a model more than adding another thousand.
How a project runs

Train. Annotate. Review. Deliver.

01

Train

Annotators learn your spec, your sensor rig and your boundary conventions on a calibration set before production begins.

02

Annotate

Sequence-first labelling inside your tooling, with identity carried across frames and hard cases flagged rather than forced.

03

Review

Multi-pass review and consensus on every batch, with agreement and geometric-overlap checks reported per task type.

04

Deliver

Model-ready episodes in your format and schema, on schedule, with the ledger and QA report attached.

Formats and tooling

Point clouds as PCD or LAS; episodes as ROS bag and MCAP sidecars, HDF5 or LeRobot-style parquet; 2D as COCO, YOLO or Pascal VOC; segment tables as JSONL or CSV. We work in your annotation stack, off-the-shelf or internal, and write to your schema.

How your data is handled

Robotics captures are sensitive: they show real facilities, real products and sometimes real people. Projects run under signed confidentiality agreements with access-controlled, managed workflows from intake to delivery, and annotators see only what their task requires.

Who does the work

Patience is a technical qualification.

Sequence labelling rewards people who will check frame 480 as carefully as frame 4. Our annotators are trained talent from communities the job market overlooks in Hong Kong: single parents, people with disabilities, and people managing long-term illness, supported by partner NGOs and social workers. They work flexible hours that fit around care and treatment, on remote-friendly tasks, and they are selected for exactly the care and focus this work needs. The result is a double bottom line: data your robots can learn from, and dignified income for people who were shut out of the tech economy.

Read the case study →

Robotics data, answered.

Keep reading
For organizations

Have data that needs labeling?

Tell us your formats, volume and timeline. We'll scope a pilot and quote with clear quality targets.

Request a quote
For supporters

Want to back the mission?

Join the program, partner with us, or help us reach more people across Hong Kong.

Get involved