
Data Labeling Services: What You’re Actually Buying
Send the same dataset description to four vendors and the quotes will often land three times apart. The instinct is to read that spread as a quality signal — the expensive one must be doing something the cheap one skips. Sometimes that is true. More often, the four vendors have quietly assumed four different scopes. Data labeling services are priced against scope, not against a standard unit of work, and the parts that move the number are usually invisible in the quote itself: who writes the labeling rules, how much of the output gets inspected, and who pays when the specification turns out to be wrong on the ten-thousandth image.
This guide covers what data labeling services actually include, how scope and unit price get decided before anyone sends you a number, and how to verify quality on your own data before committing to a volume. It is written from the buyer’s side of the table.
Table of contents
- What data labeling services actually cover
- Three delivery models, and what actually differs
- How scope gets defined before you get a quote
- What moves the unit price
- How to evaluate quality before you commit
- Security and how your data is handled
- Starting small
- Frequently asked questions
- What to take away
What data labeling services actually cover
First, the terminology. “Data labeling” and “data annotation” describe the same work in commercial practice, and providers use one or the other largely out of habit — US machine-learning teams lean toward labeling, academic and Japanese sources lean toward annotation. There is no capability difference hiding behind the word choice, and a vendor is not narrower because their homepage says one rather than the other. Filtering your shortlist on vocabulary will only shorten it arbitrarily.
The distinction worth making is between data types. Most teams arrive with one type in mind and discover a second one halfway through, so it is worth knowing at the outset which of these can sit inside the same engagement.
Image and video
The largest category by volume: bounding boxes for object detection, polygons and pixel-level regions for segmentation, keypoints for pose, and identity-consistent tracking across frames. Video is not simply “more images” — the moment you need the same object to keep the same ID across a sequence, the work and the pricing change shape. Teams building perception, inspection, or behavior-analysis models usually start here. Our image and video annotation service covers this range.
3D point cloud and LiDAR
Cuboids in 3D space, 3D segmentation, and sensor fusion where each point-cloud object is linked to its counterpart in a synchronized camera image. This is the most specification-sensitive category we handle: occlusion handling, ground-plane convention, and minimum point thresholds all have to be agreed before anyone starts, because they are expensive to retrofit. If your program combines cameras and LiDAR, the two can be scoped as one order rather than split across providers — see 3D point cloud and LiDAR annotation.
Text, LLM, and audio
Preference data for RLHF and DPO, instruction–response pairs for SFT, multi-level model evaluation, transcription, and classical NLP tagging. The judgment here is subjective in a way that box-drawing is not, which makes the guideline and the reviewer selection matter more than throughput. Language coverage is a real constraint, and native-speaker work is not interchangeable with translated work. This sits under LLM, text, and audio annotation.
Collection and preprocessing
The step most first-time buyers forget to scope. Raw footage often needs deduplication, format conversion, frame extraction, face or plate anonymization, or simply to be collected in the first place because the scenario you need does not exist in your archive. Labeling a poorly assembled dataset is the most reliable way to spend a budget and learn nothing. Data collection and preprocessing can be quoted alongside the labeling itself.
Three delivery models, and what actually differs
Providers fall into roughly three shapes: a managed service that takes the work and returns finished data, a platform that licenses you tooling and leaves the workforce to you, and a crowdsourcing marketplace that distributes tasks to an open pool. Each is a legitimate answer to a different question, and which one fits depends less on your budget than on how much specification work you want to keep. If you are still deciding between these — and between doing it in-house at all — our guide to data annotation outsourcing covers vendor selection in depth; this section stays on the one thing that is easy to miss when comparing quotes.
That thing is responsibility. Two providers can quote the same rate for the same boxes and be selling substantially different products, because the boundary of what they are accountable for sits in a different place.
Who writes the labeling guidelines
Someone has to decide what counts as an instance, how to treat an object cut off at the frame edge, whether a reflection in a window is a car, and what to do with the case nobody anticipated. If the provider expects a finished specification from you, that labor stays on your engineering team — and it is not a small amount of labor. If the provider builds the guideline with you and takes responsibility for its internal consistency, that work has moved into the engagement. Ask directly which of these you are buying, because both are often described as “labeling services”.
Who absorbs the rework
Specifications change after the first training run. That is normal — the failure cases teach you something the planning document could not. The question is what happens commercially when they do. Some arrangements treat every revision as new billable volume; others distinguish between a defect against the agreed spec and a change to the spec itself, and absorb the first. At ANOSUPO, defects traceable to us against the agreed specification are corrected at no charge for one year after delivery; a change to the specification is a different conversation, and any honest provider will tell you the same. We operate as a managed service, which is why these two clauses sit where they do.
How scope gets defined before you get a quote
By the time a number reaches you, most of the cost has already been fixed by decisions made in the scoping conversation. Three of them do the heavy lifting.
Volume, class count, and edge-case rules
Volume is the obvious lever and the least interesting one. Class count matters more than it looks: every additional class multiplies the number of boundary decisions an annotator has to make, and the marginal class is usually the ambiguous one. Edge-case rules are where estimates break. A dataset of clean, well-lit, unoccluded objects and a dataset drawn from real operating conditions can differ by a factor of two in labor while looking identical in a spreadsheet. Bring your difficult data to the scoping call, not your representative data.
Guideline design is a deliverable, not a freebie
A labeling guideline is a piece of engineering. It has to be specific enough that two annotators reach the same answer independently, and it only earns that property by surviving contact with real examples. Expect at least one revision after the first batch, and treat a provider who produces the guideline with you — and shows you the disagreements it resolved — as delivering something, not as doing sales support. Where this work is unpriced, it is usually also unowned.
Acceptance criteria — sampling vs. full inspection
This is the single largest hidden variable in any quote. Sampled spot checks and full-volume inspection are different products at different costs, and a rate quoted under one is not comparable to a rate quoted under the other. Sampling tells you the error rate; it does not remove the errors from the data you are about to train on. We inspect every deliverable rather than a sample, with a measured 99.7% quality consistency. When you compare providers, hold the inspection standard constant first — otherwise you are comparing the price of two different things.
What moves the unit price
Our published rates start at from $0.036/label for bounding boxes, and every project we quote moves from that starting point according to the same handful of factors. Object density is usually the largest: billing is per object, so a crowded intersection scene costs many times a product photograph even though both count as one image. Class count and ambiguity come next, because slower, rule-checked decisions are slower work. Then the expected revision cycles, and finally turnaround — compressing a schedule means running more annotators in parallel against the same guideline, which raises the coordination cost of keeping them consistent.
What matters as much as the rate is what the rate contains. The comparison that actually predicts your invoice looks like this.
| Question to ask | Why it changes the total |
|---|---|
| Is there a setup or onboarding fee? | A fixed cost that does not scale down for a small first batch |
| Is project management billed separately? | Often a monthly line item independent of volume delivered |
| Is there a monthly minimum? | Turns a variable cost into a fixed one during quiet months |
| Is rework billable? | Determines who pays for the spec revision you have not had yet |
| Is inspection sampled or full-volume? | Changes what the unit rate is buying, not just its size |
ANOSUPO charges no setup fee and no management fee, bills purely on volume created, and applies no minimum-order requirement. Full rate details for every data type are on our pricing page, and a quote against your actual specification is free.
How to evaluate quality before you commit
Quality claims are easy to make and hard to compare, so it is worth knowing which questions actually separate providers. The most informative one is how they handle guideline design: a provider who asks you about occlusion, class boundaries, and what to do with the ambiguous case has done this work before, while one who only asks for a volume and a deadline has not. The second is the inspection standard, asked as a number with a period attached rather than as a target.
After that, ask who does the work. For subjective tasks — evaluation, preference data, anything requiring domain reading — whether annotators are selected for the task or drawn from a general pool changes the output more than any tooling difference. Then ask what the correction flow looks like when you reject a batch: how it is reported, who reviews it, how long it takes, and whether it costs anything.
Finally, look at work the provider has actually delivered rather than at capability lists. Our case studies describe the constraints each project ran into and how the specification was built, which is more useful for calibration than a logo wall. A provider who cannot describe a project’s difficulties in specific terms has usually not been close to one.
Security and how your data is handled
For most buyers this is a gating question rather than a differentiator: either the arrangement satisfies your legal and compliance requirements or the rest of the conversation does not happen. Certification is the baseline evidence — ANOSUPO is ISO/IEC 27001 (ISMS) certified — but the operational details are what your security reviewer will ask about.
Ours are as follows. Data is not retained and not processed locally; work happens in the cloud environment agreed for the project. All staff sign NDAs, and teams are separated per project so that access does not accumulate across clients. We never repurpose client data to train our own models, and data is physically deleted on completion. If you have specific requirements — VPN connection, a named tool, data residency constraints — raise them in the first conversation rather than at contract stage, since some of them change how the work is set up. Details are on our security page.
Starting small
Almost every expensive labeling failure traces back to a specification that read as complete in a document and fell apart on real data. The defense is not a longer document. It is committing a small amount of data first, in two distinct stages that answer two different questions.
The first is a free trial: we annotate roughly 10 to 50 items of your real data with our production team, so you can assess actual quality before placing any order. This is not a demo run by a specialist — it is the workflow you would receive. It answers the question can this provider do the work.
The second is a PoC. There is no minimum-order requirement at ANOSUPO, so this is a planning decision rather than a contractual threshold; for image annotation we take on PoCs from around 50 to 100 items, which is typically enough to train a baseline and see where it fails. It answers a different question: is my specification correct. Keep the two separate in your planning — collapsing them into one batch means you learn one of those answers and assume the other. You can move from the trial into a PoC without renegotiating anything, and scale from there when the guideline has survived the difficult cases. Rates for both are the same published rates on our pricing page.
Frequently asked questions
What’s the difference between data labeling and data annotation?
In commercial practice, none. The two terms are used interchangeably for the same work, with US machine-learning teams tending toward “labeling” and research and Japanese sources tending toward “annotation”. Vendor capability does not follow the vocabulary, so it is not a useful filter when building a shortlist.
How much do data labeling services cost?
Our published rates start at from $0.036/label for bounding boxes, and the total moves with object density per item, class count, how ambiguous the judgment calls are, expected revision cycles, and turnaround. What the rate includes matters as much as its size — setup fees, management fees, monthly minimums, and billable rework all sit outside the headline number at some providers. Full rates for every data type are on our pricing page, and quotes against your specification are free.
Can I test the quality before placing a large order?
Yes. We run a free trial on roughly 10 to 50 items of your real data, annotated by the production team that would handle your project rather than by a sales specialist. You assess the output before placing any order, and there is no obligation to continue.
Do you handle 3D point cloud and LLM data as well as images?
Yes — four data families, quotable within a single engagement: image and video, 3D point cloud and LiDAR, LLM, text, and audio, and data collection and preprocessing. Programs that combine sensors do not need to be split across providers.
How is our data protected?
ANOSUPO is ISO/IEC 27001 (ISMS) certified. Data is not retained and not processed locally, all staff sign NDAs, teams are separated per project, client data is never used to train our own models, and it is physically deleted on completion. Further detail is on our security page.
What to take away
Quotes for data labeling services diverge because scope diverges, not because quality does. Before comparing numbers, hold three things constant across providers: who writes the guideline, whether inspection is sampled or full-volume, and who pays for rework when the specification changes. Once those are fixed, the unit rates become comparable — and usually much closer together than they first appeared. Then commit a small batch of your real data rather than a large batch of your assumptions. We have delivered 1,000+ projects, and the ones that went well almost always started that way.
Try our quality on your own data — for free.

