Surge builds evaluation sets, expert demonstrations, preference data, rubrics and custom RL environments using doctors, lawyers, engineers and finance professionals rather than a general crowd. Its enterprise motion is consultative: it works with your subject-matter experts to define target capabilities and failure modes, then supplies the workforce and infrastructure to execute SFT, RLHF, preference optimization or RL. It also publishes its own benchmarks (the Tuesday Work Index, Chartography), which is a useful proxy for the kind of work it does. Unlike Labelbox or Encord it sells no software you can run yourself — you buy delivered data.
| Delivery model | Managed expert-data service (sales-led) |
|---|
| Workforce | Vetted domain professionals; size not published |
|---|
| Annotator locations | unknown |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: No published pricing, no self-serve entry and no platform you can operate yourself — every engagement starts with a sales conversation, and the company is not set up for small one-off datasets. Quality control and annotator selection are a black box you have to take on trust, and there is no way to audit the pipeline the way you can with software you host.
unknown — no pricing published; sales-led engagements only
Scale's Data Engine spans generation of prompt-response pairs, RLHF preference collection, red teaming, model evaluation, and classic annotation across text, image, video and LiDAR sensor fusion. Its contributor supply runs through Outlier, which states 900K+ onboarded contributors across 50 countries and over $500M paid out. Scale also sells GenAI Portfolio and Donovan, a defense/public-sector product line that distinguishes it from every other vendor here. Francis deSouza is now CEO following Alexandr Wang's 2025 departure to Meta.
| Delivery model | Managed service + annotation platform (sales-led) |
|---|
| Workforce | Outlier: 900K+ contributors onboarded, 50 countries |
|---|
| Annotator locations | Global; US-cleared staff available for public-sector work |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | Dedicated US public-sector/defense offering (Donovan) |
|---|
Watch out: Meta's 2025 investment made Scale a conflict-of-interest problem for competing frontier labs, several of which reportedly wound down work — diligence this if you are training a competing model. Large enterprise minimums make it a poor fit for small eval sets, and parts of the public product site read as stale relative to how fast the category moves.
unknown — no public pricing; enterprise contracts via sales
Mercor matches physicians, lawyers, engineers and consultants to AI labs for evaluation, data generation and RL work, and reports paying out over $4M per day. It also publishes APEX, a benchmark of frontier models on economically valuable professional tasks, and sells an enterprise line for deploying agents against a company's own workflow data. The model is closer to specialist staffing than to a data factory: you get access to vetted people rather than a finished, QA'd dataset.
| Delivery model | Expert marketplace (you run the pipeline) |
|---|
| Workforce | 30k+ experts across 300+ fields |
|---|
| Annotator locations | Global; not broken out publicly |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: You are buying labour, not a delivered dataset — rubric design, calibration and inter-annotator agreement remain your problem, and there is no annotation tooling included. At roughly $113/hr it is among the most expensive options here, which makes it the wrong call for anything a careful generalist could do. Compliance posture and expert data-handling controls are not published.
Average contracted expert rate stated as $113/hr; role listings range roughly $75–$100/hr. Platform markup and contract minimums are not published.
Handshake repurposed its university careers platform — the career-services system for 92% of the top 500 US universities, with 1600+ university partnerships — into an AI training and evaluation workforce. It supplies verified students, alumni and graduate-degree holders for SFT, RLHF, complex reasoning and evaluation tasks, with its own pre-task training and certification layer. Credential verification through the university relationship is the structural differentiator: the PhD claim is checkable in a way a self-declared crowd profile is not.
| Delivery model | Expert network / fellowship (managed tasking) |
|---|
| Workforce | 20M+ network; 3M+ grad degrees, 500k+ PhDs |
|---|
| Annotator locations | Primarily US universities and alumni |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | Identity verification at onboarding; certs not published |
|---|
Watch out: The newest entrant here, with a short buyer-side track record and no published client references, so you are taking delivery risk on a young operation. The network is heavily US-university-centric, which makes it a poor choice for non-English languages or region-specific data. Academic credentials also do not equal practitioner experience — a philosophy PhD is not a practising clinician.
Contributor pay published at $40–$125/hr depending on role; $100M+ paid out to date. Buyer-side pricing not published.
Invisible sells AI training data, RL environments, evaluations and back-office automation through five branded modules: Neuron (data platform), Atomic (process builder), Meridial (the 24,000+ expert network), Synapse (evaluations) and Axon (agents). Named customers include Cohere, Microsoft, AWS, ElevenLabs and Thomson Reuters. It is more of a hybrid BPO than a pure data vendor — the same operation that supplies RLHF annotators will also run your contact-centre or forecasting workflow. Note the domain moved from invisible.co to invisibletech.ai.
| Delivery model | Managed service + process automation platform |
|---|
| Workforce | 24,000+ vetted domain experts (Meridial) |
|---|
| Annotator locations | Global; not broken out publicly |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: Labeling and RLHF are one line among many in a broad services portfolio, so you may not get a team that specialises in your problem. The five-module branding is marketing scaffolding rather than five separable products, which makes scoping and comparing contracts harder than with a single-purpose vendor. No public pricing and no self-serve entry.
unknown — no public pricing
Toloka has repositioned from its Yandex-era self-serve crowdsourcing marketplace into an expert data supplier: agentic RL environments, coding data, preference and reasoning-chain data, red teaming, creative-media evaluation and robotics data. It states 6,000+ active contributors across 90+ domains and 100+ countries, with more than 70% holding advanced degrees. The multilingual coverage — 40+ languages — is stronger than most US-centric expert networks, which is the main reason to shortlist it over Surge or Mercor.
| Delivery model | Managed expert-data service (sales-led) |
|---|
| Workforce | 6,000+ active contributors, 70%+ advanced degrees |
|---|
| Annotator locations | 100+ countries, 40+ languages |
|---|
| Self-serve signup | No (crowd self-serve de-emphasised) |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: A 6,000-contributor pool is orders of magnitude smaller than a crowd platform, so this is not the vendor for high-volume commodity labeling. The old self-serve crowdsourcing product is no longer the front door, so if you used Toloka for cheap microtasks that route appears gone. Its Yandex origins can trigger extended vendor security and geopolitical review in regulated or government procurement.
unknown — no public pricing page found
Snorkel has pivoted from selling Snorkel Flow, its programmatic weak-supervision labeling platform, to a data-services business: the Snorkel Data Series of ready-made expert datasets, custom dataset and RL-environment development, and specialised agents. It runs rubric-guided pipelines with calibrated experts across 1,000+ expert-level domains, and meta-evaluates its own reviewers. It also maintains public benchmarks including Terminal-Bench 4.0, Senior SWE-bench, OSWorld 2.0 and Harvey BigLaw Bench, which are the clearest evidence of what its evaluation work looks like.
| Delivery model | Managed data-development service (sales-led) |
|---|
| Workforce | Calibrated experts across 1,000+ domains |
|---|
| Annotator locations | unknown |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: If you came looking for Snorkel Flow or self-hosted programmatic labeling, that product is no longer marketed — this is now a sales-led services engagement with no software you can run. Off-the-shelf Data Series datasets may not match your task distribution, and custom development carries the usual enterprise minimum and lead time.
unknown — no public pricing
Appen runs the ADAP platform and six data lines: frontier alignment (chain-of-thought, RLHF, red teaming), agentic AI trajectories and RL environments, speech and audio across 500+ locales, multimodal, physical AI/LiDAR, and model integrity work like hallucination benchmarking and bias audits. Its speech and dialectal-language depth is genuinely hard to replicate — that catalogue predates the LLM era by two decades. It is the only vendor here with audited public financials, which cuts both ways.
| Delivery model | Managed service + ADAP platform (sales-led) |
|---|
| Workforce | Global crowd; size not published; 500+ speech locales |
|---|
| Annotator locations | Global crowd across 100+ countries |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | unknown; ASX-listed with public disclosure |
|---|
Watch out: Appen lost its largest customer contract in early 2024 and its revenue has fallen sharply since; as a small-cap listed company it carries more counterparty risk on a multi-year contract than a well-funded private vendor. Its heritage is general crowd work, and on PhD-level reasoning or specialist judgment it is competing against networks purpose-built for that. Neither crowd size nor pricing is disclosed publicly.
unknown — no public pricing
Sama is a managed annotation provider built on an impact-sourcing model: full-time, fair-wage employment for people historically excluded from the digital economy, with B Corp certification and a published impact report. Its strength is computer vision — object detection, segmentation, video object tracking, LiDAR and radar point clouds — plus text instruction-following and preference ranking. Named customers include Microsoft, Walmart, eBay and NASA. Unlike gig-based networks, its annotators are employees, which shows up in consistency on long-running projects.
| Delivery model | Managed annotation service + platform (sales-led) |
|---|
| Workforce | 15,000+ full-time associates (employees, not gig) |
|---|
| Annotator locations | Largely East Africa and India (delivery centres) |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | B Corp certified; security certs not published on site |
|---|
Watch out: Much weaker fit for frontier LLM post-training or PhD-level reasoning data than the expert networks — this is an annotation workforce, not a specialist bench. Delivery is largely from East Africa and India, which can conflict outright with data-residency or citizenship requirements, so confirm before scoping. The company has previously faced public scrutiny over conditions in content-moderation work, which some buyers' ethics reviews will surface.
unknown — no public pricing; custom quotes only
iMerit combines a managed workforce of 25,000+ domain experts with Ango Hub, its own annotation and model-evaluation platform supporting image, video, LiDAR, DICOM, text, PDF and audio. iMerit Scholars embeds credentialed specialists into the development loop, and its Deep Reasoning Lab provides interfaces for chain-of-thought and RLHF workflows. Its compliance posture is the most explicitly published of any managed vendor here — SOC 2, ISO 27001, GDPR, HIPAA and TISAX — which makes it the default shortlist entry for medical imaging and automotive.
| Delivery model | Managed service + Ango Hub platform (sales-led) |
|---|
| Workforce | 25,000+ domain experts, 60+ countries |
|---|
| Annotator locations | India-centric with secure global delivery offices |
|---|
| Self-serve signup | No |
|---|
| Compliance / deployment | SOC 2, ISO 27001, GDPR, HIPAA, TISAX |
|---|
Watch out: Delivery is India-centric, so the certifications solve the security question but not necessarily a data-residency or annotator-nationality requirement. Ango Hub is a bundled tool rather than a competitive standalone product — do not choose iMerit for the platform if you are not also buying the workforce. No public pricing and no self-serve trial.
unknown — no public pricing
Prolific is the only vendor here with fully published pricing and true self-serve access: you define a population using 300+ prescreeners, host the task on your own infrastructure or any tool that produces a URL, and results start arriving within hours. It supports model evaluation, preference and RLHF data, safety testing and SFT across 80+ languages, and maintains HUMAINE, a public human-centred LLM evaluation benchmark. A Human Feedback API and an MCP integration let you drive collection programmatically. Quality control runs through Protocol, which applies 40+ identity and behavioural checks.
| Delivery model | Self-serve participant marketplace (API + UI) |
|---|
| Workforce | 300,000+ active verified participants, 80+ languages |
|---|
| Annotator locations | Global, skewed UK/US/Europe; targetable by demographics |
|---|
| Self-serve signup | Yes |
|---|
| Compliance / deployment | Tasks hosted on your own infrastructure; certs not published |
|---|
Watch out: The default pool is a representative research panel, not a bench of credentialed specialists — for volume PhD-level STEM or code annotation you will fight the prescreeners and pay more per usable label than a dedicated expert network. The 42.8% corporate platform fee is a large markup on already-rising participant rates. You must supply your own task interface; Prolific provides people and routing, not annotation tooling.
Platform fee 42.8% for corporate customers, 33.3% for academic/non-profit, charged on top of participant pay. Participant minimum £6.00/$8.00 per hour; recommended at least £9.00/$12.00 per hour. Managed research services quoted separately.
Labelbox now sells four things: Horizon (RL environments and preference signal for post-training and evals), Terra (robotics foundation-model data collected with purpose-built hardware), Alignerr (a network of 2.6M+ knowledge experts supplying human evaluation signal), and Recursion (an enterprise RL platform for building and continuously improving specialist agents). It cites work with over 90% of leading US AI labs, including producing the 820-problem expert-authored dataset behind Meta's GIM benchmark. The classic annotation tool still exists behind a free signup, but it is no longer what the company leads with.
| Delivery model | Platform + managed expert network (Alignerr) |
|---|
| Workforce | Alignerr: 2.6M+ knowledge experts |
|---|
| Annotator locations | Global; not broken out publicly |
|---|
| Self-serve signup | Yes (free tier); paid pricing not published |
|---|
| Compliance / deployment | unknown |
|---|
Watch out: The pivot has real cost for existing users: the annotation platform is no longer the headline product and the public pricing page is gone, so a tool you may already depend on is now a moving target with opaque cost. Alignerr's 2.6M-strong pool is a scale claim, not a quality claim — variance across such a network is high and QA is on you or on Labelbox's process, unaudited. Recursion is a new product with a short production track record.
unknown — free signup exists but the public pricing page has been removed (returns 404); paid tiers quoted by sales
SuperAnnotate's distinguishing feature is a fully customisable multimodal editor — you construct the exact annotation interface your task needs across image, video, text and audio, rather than adapting to a fixed tool. It sells three things together: the software platform, AI Data Services, and an Expert Talent Network of vetted, professionally managed annotation teams. Tiers differ mainly on Orchestrate compute hours (1K / 2.5K / 10K), SSO, and the level of assigned support. Use cases are framed around SFT, RLHF, RAG, agents and evaluation.
| Delivery model | Annotation platform + optional managed teams |
|---|
| Workforce | Expert Talent Network (managed teams); size not published |
|---|
| Annotator locations | unknown |
|---|
| Self-serve signup | No — demo request required at every tier |
|---|
| Compliance / deployment | SSO on Pro/Enterprise; Trust Center published |
|---|
Watch out: No price is published at any tier, including Starter, and even the entry tier routes through a demo request — so there is no way to budget or trial without engaging sales. The customisable editor is genuinely flexible but that flexibility is setup work; expect real configuration effort before the first label. Weaker positioning for pure text/LLM preference work than the expert-data specialists.
unknown — three tiers (Starter, Pro, Enterprise) published with feature lists but no dollar figures; Orchestrate compute hours 1K/2.5K/10K respectively
Label Studio is the de facto open-source default for multimodal annotation — text, image, audio, video, time series — and the fastest way to stand up a labeling or preference-rating interface without a procurement cycle. HumanSignal sells the commercial editions: Starter Cloud for small teams, and Enterprise which adds SAML/LDAP SSO, LLM-as-a-judge, auto-labeling, bulk labeling and stronger quality workflows. A separate Data Services line supplies project management and domain experts. It is the only entry here where you can read the source, run it locally, and pay nothing.
| Delivery model | Open-source software + managed cloud + optional services |
|---|
| Workforce | None bundled; Data Services quoted separately |
|---|
| Annotator locations | N/A — you supply annotators |
|---|
| Self-serve signup | Yes (free OSS and paid cloud) |
|---|
| Compliance / deployment | Fully self-hostable; SSO/SAML on Enterprise only |
|---|
Watch out: The open-source edition omits most of what makes labeling programs survive at scale: no SSO, no auto-labeling, no LLM-as-a-judge, and thin quality-control and agreement workflows — those are Enterprise-only. Self-hosting past a handful of annotators is real operational work you will own. No workforce is bundled: you still have to find and manage the humans, and the Data Services line is a separate custom engagement.
Community Edition free and open source. Starter Cloud $99/month base plus $49/month per additional user, up to 12 users. Enterprise and Data Services are custom-quoted. Academic program available.
· open source
Encord is the annotation platform of choice for physical AI and medical imaging: alongside image, video, audio and documents it handles DICOM and NIfTI, 3D/LiDAR point clouds, geospatial data and ECG as paid add-ons. It bundles three surfaces — Annotate for labeling, Index for curation across up to 1bn+ assets, and Active for label-error detection, model evaluation and active-learning pipelines — plus data agents for automation. Enterprise adds SSO, multiple workspaces, an SLA, and VPC or on-prem deployment, which is the main reason regulated buyers shortlist it over cloud-only tools.
| Delivery model | Annotation & curation platform (self-serve to enterprise) |
|---|
| Workforce | None bundled; Data-as-a-Service offered separately |
|---|
| Annotator locations | N/A — you supply annotators |
|---|
| Self-serve signup | Yes (Starter, Team); Enterprise is sales-led |
|---|
| Compliance / deployment | Cloud, VPC (add-on), on-prem (add-on), SSO, MFA, enterprise SLA |
|---|
Watch out: The headline tiers are misleading because the modalities most buyers come for — DICOM/NIfTI, LiDAR, geospatial, ECG, LLM evaluations — are all add-ons, so real cost is well above whatever the tier implies, and no figures are published to check that. Index and Active carry hard data-volume caps per tier. On-prem and VPC are also add-ons rather than included. If your work is text-only LLM post-training, most of what you would be paying for is irrelevant.
unknown — Starter, Team and Enterprise tiers are published with detailed feature and volume comparisons but no dollar figures; Enterprise is contact-sales
Argilla is a free, self-hostable collaboration layer where AI engineers and domain experts curate datasets for fine-tuning, RLHF and evaluation, deployable in a Hugging Face Space in minutes. It joined Hugging Face and integrates tightly with the Hub's datasets and models, which makes it the lowest-friction option if your pipeline already lives there. Critically, its maintainers now state the project is stable but will receive no new features — only bug fixes and patches — and are inviting outside maintainers.
| Delivery model | Open-source self-hosted software (or HF Space) |
|---|
| Workforce | None — you supply annotators |
|---|
| Annotator locations | N/A — you supply annotators |
|---|
| Self-serve signup | Yes — install or deploy free |
|---|
| Compliance / deployment | Fully self-hostable; no vendor SLA or support |
|---|
Watch out: Officially feature-frozen: the original authors have moved on, no new features are planned, and support is limited to bug-fix patches — do not build a multi-year annotation program on it. There is no managed workforce, no SLA and no commercial support to escalate to. For anything ongoing, Label Studio is the safer open-source choice.
Free — Apache-2.0; costs are only your own hosting (or a free Hugging Face Space)
· open source · deprecated
NVIDIA acquired Gretel and folded it into NeMo: gretel.ai now redirects to NVIDIA, and the successor is NeMo Data Designer, a compound AI system for generating synthetic datasets. You define columns declaratively — category and subcategory samplers, person samplers, numeric distributions, and LLM-generated text columns driven by templated prompts — using the `nemo-microservices[data-designer]` Python SDK against a hosted endpoint, with Nemotron, GPT-OSS and Llama models available by default. It is the cheapest way to bulk out SFT coverage or produce structured records where no real data exists.
| Delivery model | Synthetic-data SDK + hosted API (self-hostable microservice) |
|---|
| Workforce | None — no humans in the loop |
|---|
| Annotator locations | N/A — fully synthetic |
|---|
| Self-serve signup | Yes — free API key |
|---|
| Compliance / deployment | Self-hostable via NeMo microservices; keeps data in your environment |
|---|
Watch out: Synthetic generation cannot produce genuine human preference or judgment signal — it will not substitute for RLHF raters, and models trained only on it inherit the generator's biases. Gretel's standalone product and its differential-privacy tabular SaaS no longer exist as an independent offering, so prior customers had to migrate. Production use pulls you into the NVIDIA NIM/GPU ecosystem, and pricing beyond the free trial is not published.
Free trial via an NVIDIA API key on build.nvidia.com; production pricing for NeMo microservices / self-hosted deployment is not published on the Data Designer page
· open source
Tonic solves the problem upstream of labeling: getting usable data out of a system that holds PII in the first place. Structural de-identifies structured databases across 12+ database types; Textual redacts and synthesises unstructured content including PDFs, DOCX, images and JSON; Fabricate generates synthetic relational data from scratch. For a labeling program this is the compliance path — you de-identify before the records reach an external annotation workforce, which changes what your DPA has to cover. It is the only entry here with a genuine credit-card self-serve tier at a published price.
| Delivery model | Self-serve SaaS + self-hosted enterprise software |
|---|
| Workforce | None — no humans in the loop |
|---|
| Annotator locations | N/A — de-identification and synthesis only |
|---|
| Self-serve signup | Yes — free tier and $29/mo Plus |
|---|
| Compliance / deployment | Self-hosted option on Textual Enterprise; purpose-built for PII removal |
|---|
Watch out: This is not a labeling or human-evaluation vendor — it supplies no annotators and no preference signal, so it complements rather than replaces anything else in this list. Only Fabricate has published prices; Structural and Textual enterprise tiers are quote-only, and Structural's Professional tier is capped at 2 data sources with Oracle and Db2 gated to Enterprise. Automated de-identification still needs your own validation before you can rely on it for a regulatory claim.
Fabricate: Free $0/month with $5 monthly credits (roughly 9 complex generation sessions); Plus $29/month with $25 monthly credits plus pay-as-you-go; Enterprise custom. Structural: Professional (up to 10TB source data, 10 users, 2 data sources) and Enterprise, both custom-quoted. Textual: pay-as-you-go flat rate per 1,000 words, or Enterprise with self-hosting.