Home Image & video labeling

VIS · Image & video boxes → masks → tracks

Every pixel accounted for.

Boxes, polygons, masks, keypoints, attributes and multi-frame tracks, drawn to your spec and checked against it. The difference between a cheap vendor and a good one is not the first pass; it is what happens to the 8% of frames where the spec runs out.

COCO · YOLO · VOC · your schema · Under NDA

What we label

The full computer-vision toolkit.

Tell us the label type and the acceptance criteria and we will quote against them. Where you are unsure which annotation type your model actually needs, we will say so during scoping rather than sell you the most expensive one.

BOX

Bounding boxes and 2D detection

Tight boxes with agreed rules for occlusion, truncation at the frame edge, crowds and minimum object size. Those four conventions cause most of the disagreement in detection datasets, so we settle them first.

POLY

Polygons and instance segmentation

Per-instance outlines at the vertex density your model needs, with holes, disconnected parts and overlapping instances handled consistently instead of being approximated away.

SEM

Semantic and panoptic segmentation

Dense per-pixel class masks with an explicit policy for boundaries, shadows, reflections and unlabelled regions, delivered as PNG indices or run-length encoding.

KEY

Keypoints, pose and landmarks

Human pose, hand and facial landmarks, and product or machinery keypoints, each with visibility flags so an occluded joint is recorded as occluded rather than quietly invented.

CLS

Classification and attributes

Single-label, multi-label and hierarchical taxonomies, plus per-object attributes such as colour, material, state, damage and quality grade. Confusable pairs get a written decision rule, not a coin flip.

OCR

OCR and text in the wild

Text detection and transcription on signage, menus, notices, labels and packaging, including Traditional Chinese set vertically or in stylised script and mixed freely with English. A Hong Kong team reads these every day.

TRACK

Video tracking and event tagging

Multi-object tracking with stable identities through occlusion and re-entry, interpolated versus observed frames marked explicitly, plus temporal event and action boundaries where you need them.

Quality

Anyone can draw a box. The spec is the product.

The first ninety per cent of any image dataset is straightforward, and every vendor gets it right. Models are made or broken by the rest: the object cut in half by the frame, the two instances that overlap, the sign that is half-legible, the category that only appears eleven times. A pipeline that forces an answer on those cases injects noise you cannot see until training.

We run a calibration round on a small sample first, so the disagreements surface while they are still cheap to fix. Whatever your spec does not cover gets a written decision and an entry in an edge-case ledger that ships with the batch. Every delivery carries its numbers: gold-set accuracy, IoU against reference geometry, and inter-annotator agreement per class rather than one flattering average.

What every delivery includes

  • Gold-set scoring on every batch, not just the first.
  • IoU thresholds agreed per label type before production.
  • Multi-pass review with consensus on disputed items.
  • Per-class agreement figures, not a single blended score.
  • An edge-case ledger that sharpens your guidelines.
  • A named reviewer accountable for each batch.
Quality you can verify, not quality you have to trust.
How a project runs

Train. Annotate. Review. Deliver.

01

Train

Annotators work through your guidelines and a calibration batch, and we return the questions your spec did not answer.

02

Annotate

Careful labelling at a pace that protects accuracy, in your tool or ours, with hard frames flagged rather than forced.

03

Review

Multi-pass review and consensus on every batch, scored against the gold set and the thresholds we agreed.

04

Deliver

Model-ready data in your format, on schedule, with the QA report and edge-case ledger attached.

Formats we deliver

COCO JSON, YOLO text, Pascal VOC XML, PNG or RLE masks, CVAT XML, and CSV or JSONL for classification and attributes. If you have an internal schema we write to it directly rather than making you convert.

How your data is handled

Projects run under signed confidentiality agreements with access-controlled, managed workflows from intake to delivery. Where imagery contains faces, plates or documents, we agree a handling and redaction policy before a single file moves.

Who does the work

Skilled annotators, not anonymous throughput.

Image labelling is where Label Less started, and it is where the model proved itself: around thirty single mothers in Hong Kong trained into paid annotation work, fitting two to three focused hours around childcare. Today the same trained team spans single parents, people with disabilities and people managing long-term illness, supported by partner NGOs and social workers. They are not interchangeable capacity. They are the reason the quality holds, and the reason your project also produces dignified, flexible income for people the tech economy overlooked.

Read the case study →

Image labeling, answered.

Keep reading
For organizations

Have data that needs labeling?

Tell us your formats, volume and timeline. We'll scope a pilot and quote with clear quality targets.

Request a quote
For supporters

Want to back the mission?

Join the program, partner with us, or help us reach more people across Hong Kong.

Get involved