
Data Annotation Outsourcing: How to Choose a Vendor (2026 Guide)
Most teams don’t set out to hire an annotation vendor. They start by labeling data themselves — a few hundred images, a spreadsheet of rules, an intern or two — and it works. Then the dataset needs to be ten times larger, the edge cases start piling up, and two of your best engineers are spending their week drawing boxes instead of training models.
That’s the point where outsourcing starts to make sense. This guide covers what actually drives annotation cost, how the three sourcing models differ, seven criteria for evaluating a vendor, and what to prepare before you request a quote.
Table of contents
- When outsourcing makes sense — and when it doesn’t
- In-house vs. crowdsourcing vs. specialist vendor
- What data types change your choice of vendor
- What actually drives annotation cost
- 7 criteria for choosing an annotation vendor
- A pre-order checklist
- Data residency, NDAs and time zones
- Frequently asked questions
- Why teams choose ANOSUPO
When outsourcing makes sense — and when it doesn’t
Outsourcing is a good fit when the annotation work is well-defined enough to be written down, and large enough that doing it internally costs you something you care about — usually engineering time.
Outsourcing tends to work when:
- The volume is beyond what your team can absorb alongside model work
- The task can be specified in writing, including the awkward edge cases
- You need consistent output over weeks or months, not a one-off batch
- You need to scale up and down without hiring
Outsourcing tends not to work when:
- The labeling rules are still changing daily because the problem definition isn’t settled
- The task requires domain expertise that only exists inside your team and can’t be transferred through a spec
- The dataset is small enough that the coordination overhead exceeds the labeling effort
If you’re in the second group, the honest advice is to label a few hundred items yourself first — the process is what forces the spec to become concrete. Our guides on building YOLO training data and semantic segmentation annotation walk through how to do that in-house. Once the spec stabilizes, outsourcing becomes a volume question rather than a judgment question.
In-house vs. crowdsourcing vs. specialist vendor
There are three realistic ways to get labeled data, and they fail in different ways.
| In-house | Crowdsourcing | Specialist vendor | |
|---|---|---|---|
| Quality consistency | High at first, drifts as the team grows | Varies by worker; hard to hold steady on difficult tasks | Held by a defined inspection process |
| Ramp-up speed | Slow — hiring and training | Fast | Days to weeks, depending on spec complexity |
| Scale ceiling | Limited by headcount | Very high | High, with a managed team |
| Confidential data | Fully controlled | Difficult — many hands, little traceability | Depends on the vendor’s certification and handling policy |
| Cost visibility | Hidden in salaries | Low unit price, plus rework | Unit price, if the vendor publishes one |
| Best-fit phase | Problem definition, first PoC | Simple, high-volume, non-sensitive tasks | Production datasets and ongoing pipelines |
Each model has a characteristic failure. In-house fails quietly: nobody logs the hours, so the cost shows up as delayed model releases rather than as a line item. Crowdsourcing fails on difficulty — as soon as a task requires judgment rather than pattern matching, inter-annotator agreement drops and you spend your savings on review. This is the exact reason Kyoto University came to us for evaluation data on a conversational AI project: the task was too subjective for a crowd platform to hold quality on, and it needed a selected, trained group of workers instead. Specialist vendors fail when the spec was never nailed down — a vendor will faithfully produce whatever you asked for, including the wrong thing.
You can see how these played out across different projects on our case studies page.
What data types change your choice of vendor
Vendor shortlists narrow fast once you name your data type. Most annotation companies handle 2D images competently. Fewer handle the rest.
- Images and video — bounding boxes, segmentation, keypoints, object tracking across frames. This is the entry point for most teams and the broadest market. See image & video annotation.
- 3D point cloud and LiDAR — 3D cuboids, 3D segmentation, and sensor fusion where LiDAR and camera data must be labeled consistently across calibrated sensors. See 3D point cloud & LiDAR annotation, or our 3D point cloud annotation guide for what the work involves.
- Text, speech and LLM data — RLHF preference pairs, DPO, SFT instruction-response sets, multi-level LLM evaluation, transcription. See LLM, text & speech annotation.
- Data collection and preprocessing — sourcing or shooting the raw data, cleaning, formatting, anonymization. See data collection & preprocessing.
Two thresholds sharply reduce the number of viable vendors. The first is sensor fusion: labeling a 3D cuboid is one skill, keeping it consistent with the corresponding camera view across a calibrated rig is another. The second is tasks that need a custom annotation interface — when the labeling schema doesn’t fit any off-the-shelf tool. For OMRON SINIC X we built a dedicated UI for a complex, non-standard task rather than forcing the work into a generic tool. If your task falls into either category, ask about it in the first conversation, not after signing.
Video deserves a separate note here. It is often treated as an image project with more files, but the vendor requirements are different: you need someone who can maintain object identity across frames, not just draw accurate boxes. Our guide to video annotation and building a tracking dataset covers what changes and how it affects your estimate.
What actually drives annotation cost
Unit price is not one number. Six factors move it:
- Annotation type — a polygon takes several times longer than a box, and price follows time.
- Object density — 40 objects in a frame costs more than 3, because pricing is usually per object rather than per image.
- Number of classes — more classes means more decisions per object and more chances to disagree.
- Spec ambiguity — if “partially occluded” isn’t defined, someone has to decide, and inconsistency is expensive to fix later.
- Inspection regime — sampling-based QA is cheaper per unit than full-volume inspection, and the difference shows up in your error rate.
- Security requirements — isolated teams, restricted environments and non-retention policies all carry operational cost.
Many annotation vendors in the English-speaking market don’t publish rates at all — you enter a sales process before you can estimate anything. That’s a real cost in itself, because you can’t size a project or compare options without a number. We publish ours:
| Annotation type | Unit price (excl. tax) |
|---|---|
| Bounding box (object detection) | from $0.036 / label |
| Keypoint (pose estimation) | from $0.021 / point |
| Segmentation | from $0.152 / region |
| 3D cuboid (LiDAR) | from $0.121 / cuboid |
| LLM multi-level evaluation | from $0.121 / evaluation |
There is no setup fee and no management fee — billing is fully usage-based, so you pay for what’s annotated. Full rate details are on our pricing page.
One thing a unit price can’t tell you: whether the work will need redoing. A low rate with sampled QA and a high rate with full-volume inspection are not the same product, and the gap only becomes visible when you train on the data. That’s why we recommend running a small paid or free pilot before committing volume — it converts an unknown into a measured error rate.
7 criteria for choosing an annotation vendor
Seven things to evaluate, with the question to ask for each.
- Quality assurance method. Sampling and full-volume inspection produce different error profiles. Ask: do you inspect every item, or a sample — and what percentage?
- Security and data handling. Certification, retention policy, worker NDAs, whether your data is reused for the vendor’s own model training. Ask: where is our data stored, who can see it, and what happens to it after delivery?
- Data type coverage. Whether the vendor can follow you from 2D images into 3D, video or LLM data as your project grows. Ask: which of these have you delivered in the last year?
- Pricing transparency. Whether you can estimate before entering a sales cycle, and whether setup or management fees exist. Ask: what’s the unit price, and what else appears on the invoice?
- Minimum order and pilot options. A high minimum forces you to bet before you have evidence. Ask: what’s the smallest batch you’ll take, and can we trial on our own data first?
- Spec-building support. Most first specs are incomplete. A good vendor surfaces the edge cases before production, not after. Ask: who writes the annotation guidelines, and how are ambiguous cases escalated?
- Post-delivery correction. What happens when delivered data doesn’t match the agreed spec. Ask: what’s covered, for how long, and at what cost?
For reference, our answers: full-volume inspection rather than sampling, with 99.7% quality consistency; ISO/IEC 27001 (ISMS) certification with a non-retention, cloud-only policy detailed on our security page; all four data categories above; published unit prices with no setup or management fee; no minimum order, with image annotation starting from PoC batches of 50–100 items; joint spec development including edge-case definition; and free correction for one year on anything that doesn’t match the agreed spec and is attributable to us.
A pre-order checklist
Settle these internally before you request quotes. Every item you can answer makes the quotes you get more accurate and more comparable.
- ☐ Annotation definition, including edge cases — occlusion, truncation at frame borders, minimum object size, ambiguous classes
- ☐ Rough volume — how many items for the first milestone, and what the full dataset looks like
- ☐ Output format — COCO, YOLO, Pascal VOC, or something custom your pipeline expects
- ☐ Evaluation metric — how you’ll judge whether the delivered data is good enough
- ☐ Deadline — including whether it’s one delivery or a recurring pipeline
- ☐ Confidentiality classification — what the data contains and what handling it requires
- ☐ Sample data — a representative subset, including the hard cases, not just clean examples
- ☐ Internal owner — who answers the vendor’s spec questions, and how fast
The last one matters more than it looks. Annotation projects stall on unanswered questions far more often than on labeling capacity.
Data residency, NDAs and time zones
If you’re outsourcing across borders, three questions come up in procurement before anything technical does.
Where the data is processed. We work cloud-only and don’t retain data locally. Data is separated by project, teams are isolated per engagement, and everything is physically deleted after completion. Your data is never used to train our own models.
The confidentiality framework. All staff work under NDA, and we hold ISO/IEC 27001 (ISMS) certification. In many Western procurement processes, certification is a pass/fail gate before capability is even discussed — so it’s worth confirming early rather than late.
Time zones. We’re based in Fukuoka, Japan. In practice this is a workflow design question rather than a drawback: asynchronous handoff means work progresses while your team is offline, provided the spec is clear enough that the annotation team isn’t blocked waiting for answers. We structure projects around a written spec and a single communication channel — for Pirika, a Slack-centered setup let us minimize meetings entirely while scaling throughput.
Frequently asked questions
How much does data annotation outsourcing cost?
It depends on annotation type and object density. Our published rates start from $0.036 per bounding box, $0.021 per keypoint, $0.152 per segmentation region, and $0.121 per 3D cuboid (excl. tax). There’s no setup fee and no management fee — billing is fully usage-based. For a project-specific figure, send us your data type and volume and we’ll quote against it.
Is outsourcing secure for confidential data?
It can be, if the vendor’s handling policy is explicit. We hold ISO/IEC 27001 (ISMS) certification, work cloud-only without local retention, separate teams per project, and have all staff under NDA. Your data is not used to train our models and is physically deleted after the project ends.
What’s the minimum volume I can order?
There’s no minimum order. Image annotation can start from a PoC batch of roughly 50–100 items, which is usually enough to check whether the spec holds up before you commit to volume.
Can I check quality before placing a real order?
Yes. Our free trial annotates roughly 10–50 items of your actual data using the same team and process as a production order. You review the output and decide afterwards.
Can you handle 3D point cloud and LiDAR data?
Yes — 3D cuboids, 3D segmentation, and sensor fusion across LiDAR and camera data. This is a narrower capability than 2D image annotation, so it’s worth confirming with any vendor you evaluate rather than assuming it.
What happens if the annotations don’t match our spec?
We inspect every item rather than a sample, so spec mismatches should be caught before delivery. If something still doesn’t match the agreed spec and the cause is on our side, we correct it free of charge for one year after delivery. This covers errors against the agreed spec, not changes to the spec itself.
Why teams choose ANOSUPO
- Published unit prices — you can estimate before talking to sales
- No minimum order, no setup or management fee — fully usage-based, so a pilot costs pilot money
- Full-volume inspection, not sampling — 99.7% quality consistency
- ISO/IEC 27001 certified, cloud-only, non-retention, per-project team isolation
- All four data categories — images and video, 3D point cloud and LiDAR, LLM/text/speech, and data collection and preprocessing
- One year of free correction against the agreed spec
- 1,000+ projects delivered for companies, universities and research institutions since 2021
ANOSUPO is operated by Borderless Japan and based in Fukuoka. Recognition includes Forbes 30 Under 30 Asia 2023, the Good Design Award 2023, and selection for Japan Innovation Campus, the Japanese government’s startup hub in the United States.
If you already know your data type and volume, request a quote and we’ll price it directly. If you’d rather see the output before deciding, start with the free trial below.
Try our quality on your own data — for free.

