Home Cantonese annotation
Cantonese, heard the way Hong Kong says it.
Written Chinese is not spoken Cantonese, and neither is Mandarin. We annotate Cantonese and Traditional Chinese with the local knowledge that decides whether a label is right: code-switching, sentence-final particles, estate and MTR names, forum slang, and the gap between 書面語 and 口語.
Native Hong Kong annotators · Multi-pass QA · Under NDA
Six task types, one local team.
Most language work arrives as a mix. We scope each task type separately so you can buy only what you need, and so quality targets stay honest per task rather than averaged across the batch.
Cantonese speech transcription
Verbatim or clean-read transcription, with speaker labels, word- or segment-level timestamps, and tagged non-speech events. We deliver in spoken Cantonese or standard written Chinese, or both aligned to the same audio.
Named entities with Hong Kong context
Districts, housing estates, MTR stations, government bureaux, local brands and street names, linked across their Chinese, English and romanised forms so one entity does not become three.
Code-switching segmentation
Mixed utterances such as “send 個 email 俾我” are segmented and tagged at token level with the language of each span recorded. Code-switching is treated as signal, not as a corrupted sequence to be discarded.
Sentiment, stance and intent
Rated against your rubric, including the sarcasm, understatement and particle-carried tone that flip a sentence. 啦, 囉, 㗎, 咩 and 呀 do real semantic work, and a model trained on Mandarin data will miss all of it.
Content moderation and safety
Local slang, coded language, forum idiom and harassment patterns, labelled against your policy. Ambiguous cases are escalated with a written rationale instead of being silently guessed.
Script, romanisation and register
Traditional and Simplified handling, Jyutping romanisation, Hong Kong versus Taiwan vocabulary differences, and paired 書面語 ↔ 口語 rewrites for training or evaluation sets.
The hard part is not the language. It is the city.
A remote annotator working from a dictionary can tell you that 搞掂 means “done”. They cannot tell you that a customer writing it in a support ticket is closing the conversation, not asking for help. That distinction is the label, and it is the reason a generic vendor's Chinese data quietly underperforms on Hong Kong users.
Our annotators live here. They read the forums, ride the same lines, and recognise a building name written three different ways. When a spec does not cover a case, they write down what they decided and why, and that entry goes into an edge-case ledger you receive with the batch. Your guidelines get sharper as the project runs, which is usually worth more than the labels themselves.
Cases a generic pipeline gets wrong
- Particles that carry the sentiment: 好啦 (resigned) against 好呀 (willing).
- English embedded mid-sentence with Chinese grammar around it.
- Estate, ward and station names that look like ordinary nouns.
- Numerals and dates written in mixed script within one line.
- Traditional characters carrying Hong Kong, not Taiwan, vocabulary.
- Coded and evasive language that moderation policies were not written for.
Train. Annotate. Review. Deliver.
Train
Annotators work through your written spec and a calibration set before a single label ships. Disagreements in calibration are where the spec gets fixed.
Annotate
Careful labelling at a pace that protects accuracy, inside your tooling or ours, with ambiguous items flagged rather than forced.
Review
Multi-pass review and consensus on every batch, with inter-annotator agreement reported against the targets we agreed up front.
Deliver
Model-ready data in your format, on schedule, with the edge-case ledger and an agreement report attached.
Formats we deliver
JSONL, CSV, CoNLL, TextGrid, SRT and WebVTT, or your internal schema. UTF-8 throughout, with a documented character-encoding and normalisation policy so Traditional characters survive the round trip intact.
How your data is handled
Projects run under signed confidentiality agreements with access-controlled, managed workflows from intake to delivery. Annotators see only what their task requires, and workflows are auditable so you can verify quality rather than take it on trust.
Local knowledge is not a bonus. It is the workforce.
The people annotating your Cantonese data are trained Hong Kong annotators recruited from communities the job market overlooks: single parents working in the hours between childcare, people with disabilities, and people managing long-term illness. They are supported by partner NGOs and social workers, and they bring exactly the lived local fluency this work demands. Choosing Label Less buys you better data and creates dignified, flexible income at the same time.
Read the case study →Cantonese annotation, answered.
Robotics & teleoperation data
LiDAR, trajectories, action segments and teleoperation demonstrations for embodied AI.
VISImage & video labeling
Boxes, polygons, segmentation, keypoints and multi-frame tracking for computer vision.
CASEA career in AI data for single mothers
How roughly 30 single mothers in Hong Kong trained into paid annotation work.
Have data that needs labeling?
Tell us your formats, volume and timeline. We'll scope a pilot and quote with clear quality targets.
Request a quote →Want to back the mission?
Join the program, partner with us, or help us reach more people across Hong Kong.
Get involved →