Seminal AI
§6

Data labeling, RLHF and human evaluation vendors

This category covers the companies that supply the human side of model training: annotated datasets, supervised fine-tuning demonstrations, preference and reward signal for RLHF, RL environments, red-teaming, and human evaluation panels. It splits into three structurally different things that buyers routinely conflate — managed expert-data services that deliver finished datasets (Surge, Scale, Snorkel, Toloka, Invisible), talent marketplaces that rent you vetted humans while you run the pipeline (Mercor, Handshake AI, Prolific, Alignerr), and annotation software you staff yourself (Label Studio, Encord, SuperAnnotate, Argilla).

Data checked 2026-09-06

The category has moved hard toward domain expertise since 2024: several vendors that sold labeling tools or crowd work now sell PhD-level expert data and RL environments instead, and at least two — Snorkel and Labelbox — have effectively retired the platform product that made their name. Hourly cost spans roughly $8/hr for general crowd work to $125/hr for a credentialed specialist, which is the single biggest line item in most fine-tuning budgets and the reason procurement, not compute, is usually the bottleneck.

A How to choose

Four axes actually separate these vendors. First, who does the work: a general crowd panel (Prolific's minimum is $8/hr, recommended $12/hr) versus a credentialed specialist network (Handshake AI publishes $40–$125/hr; Mercor's stated average contracted rate is $113/hr). That is the order-of-magnitude gap, and it is only worth paying when the task genuinely requires judgment a smart generalist cannot fake — most preference-rating and instruction-following work does not.

Second, who runs the pipeline: a managed service hands you finished, QA'd data but comes with a six-figure minimum and a multi-week sales cycle; a marketplace hands you people and leaves rubric design, calibration and inter-annotator agreement to you; software hands you neither. Third, where the humans sit and who sees the data — Sama and iMerit deliver largely from East Africa and India, Prolific and Handshake skew US/UK, and if you need annotators inside a jurisdiction or on your own VPC, ask before the demo (Encord and iMerit publish on-prem/VPC and SOC 2 / ISO 27001 / HIPAA / TISAX respectively; most expert networks publish nothing). Fourth, contract shape: only Prolific, Label Studio, Tonic and NVIDIA let you start on a credit card.

Now the anti-defaults. Do not call Scale or Surge for a 500-row domain eval set — the procurement cycle will cost more than the data, and five of your own in-house experts for two days will beat both; use Label Studio or a spreadsheet. Do not use Prolific for PhD-level STEM or code annotation at volume: it is a representative-population research panel, and its 42.8% platform fee on top of participant pay makes expert work there expensive per usable label.

Do not buy an annotation platform seat if all you need is a preference-rating UI. Do not assume synthetic data (NeMo Data Designer, Tonic Fabricate) substitutes for human preference signal — it can bulk out SFT coverage, but the reward model still needs real judgment. Check conflicts of interest: Meta's 2025 stake in Scale caused competing frontier labs to pull work, and if you are training a model that competes with an investor, ask.

Finally, verify the product still exists before you plan around it: Snorkel Flow is no longer marketed, Labelbox's public pricing page is gone and its annotation tool is no longer the headline product, Gretel is now NVIDIA NeMo Data Designer, and Argilla is feature-frozen. Turing, Micro1, CloudFactory, TELUS Digital, Dataloop, Kili and V7 are reasonable additions to a longlist but were cut here for overlap.

B At a glance

Name Delivery modelWorkforceAnnotator locationsSelf-serve signupCompliance / deployment Pricing
Surge AI Managed expert-data service (sales-led)Vetted domain professionals; size not publishedunknownNounknown unknown — no pricing published; sales-led engagements only
Scale AI Managed service + annotation platform (sales-led)Outlier: 900K+ contributors onboarded, 50 countriesGlobal; US-cleared staff available for public-sector workNoDedicated US public-sector/defense offering (Donovan) unknown — no public pricing; enterprise contracts via sales
Mercor Expert marketplace (you run the pipeline)30k+ experts across 300+ fieldsGlobal; not broken out publiclyNounknown Average contracted expert rate stated as $113/hr; role listings range roughly $75–$100/hr. Platform markup and contract minimums are not published.
Handshake AI Expert network / fellowship (managed tasking)20M+ network; 3M+ grad degrees, 500k+ PhDsPrimarily US universities and alumniNoIdentity verification at onboarding; certs not published Contributor pay published at $40–$125/hr depending on role; $100M+ paid out to date. Buyer-side pricing not published.
Invisible Technologies Managed service + process automation platform24,000+ vetted domain experts (Meridial)Global; not broken out publiclyNounknown unknown — no public pricing
Toloka Managed expert-data service (sales-led)6,000+ active contributors, 70%+ advanced degrees100+ countries, 40+ languagesNo (crowd self-serve de-emphasised)unknown unknown — no public pricing page found
Snorkel AI Managed data-development service (sales-led)Calibrated experts across 1,000+ domainsunknownNounknown unknown — no public pricing
Appen Managed service + ADAP platform (sales-led)Global crowd; size not published; 500+ speech localesGlobal crowd across 100+ countriesNounknown; ASX-listed with public disclosure unknown — no public pricing
Sama Managed annotation service + platform (sales-led)15,000+ full-time associates (employees, not gig)Largely East Africa and India (delivery centres)NoB Corp certified; security certs not published on site unknown — no public pricing; custom quotes only
iMerit Managed service + Ango Hub platform (sales-led)25,000+ domain experts, 60+ countriesIndia-centric with secure global delivery officesNoSOC 2, ISO 27001, GDPR, HIPAA, TISAX unknown — no public pricing
Prolific Self-serve participant marketplace (API + UI)300,000+ active verified participants, 80+ languagesGlobal, skewed UK/US/Europe; targetable by demographicsYesTasks hosted on your own infrastructure; certs not published Platform fee 42.8% for corporate customers, 33.3% for academic/non-profit, charged on top of participant pay. Participant minimum £6.00/$8.00 per hour; recommended at least £9.00/$12.00 per hour. Managed research services quoted separately.
Labelbox Platform + managed expert network (Alignerr)Alignerr: 2.6M+ knowledge expertsGlobal; not broken out publiclyYes (free tier); paid pricing not publishedunknown unknown — free signup exists but the public pricing page has been removed (returns 404); paid tiers quoted by sales
SuperAnnotate Annotation platform + optional managed teamsExpert Talent Network (managed teams); size not publishedunknownNo — demo request required at every tierSSO on Pro/Enterprise; Trust Center published unknown — three tiers (Starter, Pro, Enterprise) published with feature lists but no dollar figures; Orchestrate compute hours 1K/2.5K/10K respectively
Label Studio (HumanSignal) Open-source software + managed cloud + optional servicesNone bundled; Data Services quoted separatelyN/A — you supply annotatorsYes (free OSS and paid cloud)Fully self-hostable; SSO/SAML on Enterprise only Community Edition free and open source. Starter Cloud $99/month base plus $49/month per additional user, up to 12 users. Enterprise and Data Services are custom-quoted. Academic program available.
Encord Annotation & curation platform (self-serve to enterprise)None bundled; Data-as-a-Service offered separatelyN/A — you supply annotatorsYes (Starter, Team); Enterprise is sales-ledCloud, VPC (add-on), on-prem (add-on), SSO, MFA, enterprise SLA unknown — Starter, Team and Enterprise tiers are published with detailed feature and volume comparisons but no dollar figures; Enterprise is contact-sales
Argilla Open-source self-hosted software (or HF Space)None — you supply annotatorsN/A — you supply annotatorsYes — install or deploy freeFully self-hostable; no vendor SLA or support Free — Apache-2.0; costs are only your own hosting (or a free Hugging Face Space)
NVIDIA NeMo Data Designer (formerly Gretel) Synthetic-data SDK + hosted API (self-hostable microservice)None — no humans in the loopN/A — fully syntheticYes — free API keySelf-hostable via NeMo microservices; keeps data in your environment Free trial via an NVIDIA API key on build.nvidia.com; production pricing for NeMo microservices / self-hosted deployment is not published on the Data Designer page
Tonic.ai Self-serve SaaS + self-hosted enterprise softwareNone — no humans in the loopN/A — de-identification and synthesis onlyYes — free tier and $29/mo PlusSelf-hosted option on Textual Enterprise; purpose-built for PII removal Fabricate: Free $0/month with $5 monthly credits (roughly 9 complex generation sessions); Plus $29/month with $25 monthly credits plus pay-as-you-go; Enterprise custom. Structural: Professional (up to 10TB source data, 10 users, 2 data sources) and Enterprise, both custom-quoted. Textual: pay-as-you-go flat rate per 1,000 words, or Enterprise with self-hosting.

C Entries

Surge AI

Surge builds evaluation sets, expert demonstrations, preference data, rubrics and custom RL environments using doctors, lawyers, engineers and finance professionals rather than a general crowd. Its enterprise motion is consultative: it works with your subject-matter experts to define target capabilities and failure modes, then supplies the workforce and infrastructure to execute SFT, RLHF, preference optimization or RL. It also publishes its own benchmarks (the Tuesday Work Index, Chartography), which is a useful proxy for the kind of work it does. Unlike Labelbox or Encord it sells no software you can run yourself — you buy delivered data.

Delivery modelManaged expert-data service (sales-led)
WorkforceVetted domain professionals; size not published
Annotator locationsunknown
Self-serve signupNo
Compliance / deploymentunknown

Watch out: No published pricing, no self-serve entry and no platform you can operate yourself — every engagement starts with a sales conversation, and the company is not set up for small one-off datasets. Quality control and annotator selection are a black box you have to take on trust, and there is no way to audit the pipeline the way you can with software you host.

unknown — no pricing published; sales-led engagements only

Scale AI

Scale's Data Engine spans generation of prompt-response pairs, RLHF preference collection, red teaming, model evaluation, and classic annotation across text, image, video and LiDAR sensor fusion. Its contributor supply runs through Outlier, which states 900K+ onboarded contributors across 50 countries and over $500M paid out. Scale also sells GenAI Portfolio and Donovan, a defense/public-sector product line that distinguishes it from every other vendor here. Francis deSouza is now CEO following Alexandr Wang's 2025 departure to Meta.

Delivery modelManaged service + annotation platform (sales-led)
WorkforceOutlier: 900K+ contributors onboarded, 50 countries
Annotator locationsGlobal; US-cleared staff available for public-sector work
Self-serve signupNo
Compliance / deploymentDedicated US public-sector/defense offering (Donovan)

Watch out: Meta's 2025 investment made Scale a conflict-of-interest problem for competing frontier labs, several of which reportedly wound down work — diligence this if you are training a competing model. Large enterprise minimums make it a poor fit for small eval sets, and parts of the public product site read as stale relative to how fast the category moves.

unknown — no public pricing; enterprise contracts via sales

Mercor

Mercor matches physicians, lawyers, engineers and consultants to AI labs for evaluation, data generation and RL work, and reports paying out over $4M per day. It also publishes APEX, a benchmark of frontier models on economically valuable professional tasks, and sells an enterprise line for deploying agents against a company's own workflow data. The model is closer to specialist staffing than to a data factory: you get access to vetted people rather than a finished, QA'd dataset.

Delivery modelExpert marketplace (you run the pipeline)
Workforce30k+ experts across 300+ fields
Annotator locationsGlobal; not broken out publicly
Self-serve signupNo
Compliance / deploymentunknown

Watch out: You are buying labour, not a delivered dataset — rubric design, calibration and inter-annotator agreement remain your problem, and there is no annotation tooling included. At roughly $113/hr it is among the most expensive options here, which makes it the wrong call for anything a careful generalist could do. Compliance posture and expert data-handling controls are not published.

Average contracted expert rate stated as $113/hr; role listings range roughly $75–$100/hr. Platform markup and contract minimums are not published.

Handshake AI

Handshake repurposed its university careers platform — the career-services system for 92% of the top 500 US universities, with 1600+ university partnerships — into an AI training and evaluation workforce. It supplies verified students, alumni and graduate-degree holders for SFT, RLHF, complex reasoning and evaluation tasks, with its own pre-task training and certification layer. Credential verification through the university relationship is the structural differentiator: the PhD claim is checkable in a way a self-declared crowd profile is not.

Delivery modelExpert network / fellowship (managed tasking)
Workforce20M+ network; 3M+ grad degrees, 500k+ PhDs
Annotator locationsPrimarily US universities and alumni
Self-serve signupNo
Compliance / deploymentIdentity verification at onboarding; certs not published

Watch out: The newest entrant here, with a short buyer-side track record and no published client references, so you are taking delivery risk on a young operation. The network is heavily US-university-centric, which makes it a poor choice for non-English languages or region-specific data. Academic credentials also do not equal practitioner experience — a philosophy PhD is not a practising clinician.

Contributor pay published at $40–$125/hr depending on role; $100M+ paid out to date. Buyer-side pricing not published.

Invisible Technologies

Invisible sells AI training data, RL environments, evaluations and back-office automation through five branded modules: Neuron (data platform), Atomic (process builder), Meridial (the 24,000+ expert network), Synapse (evaluations) and Axon (agents). Named customers include Cohere, Microsoft, AWS, ElevenLabs and Thomson Reuters. It is more of a hybrid BPO than a pure data vendor — the same operation that supplies RLHF annotators will also run your contact-centre or forecasting workflow. Note the domain moved from invisible.co to invisibletech.ai.

Delivery modelManaged service + process automation platform
Workforce24,000+ vetted domain experts (Meridial)
Annotator locationsGlobal; not broken out publicly
Self-serve signupNo
Compliance / deploymentunknown

Watch out: Labeling and RLHF are one line among many in a broad services portfolio, so you may not get a team that specialises in your problem. The five-module branding is marketing scaffolding rather than five separable products, which makes scoping and comparing contracts harder than with a single-purpose vendor. No public pricing and no self-serve entry.

unknown — no public pricing

Toloka

Toloka has repositioned from its Yandex-era self-serve crowdsourcing marketplace into an expert data supplier: agentic RL environments, coding data, preference and reasoning-chain data, red teaming, creative-media evaluation and robotics data. It states 6,000+ active contributors across 90+ domains and 100+ countries, with more than 70% holding advanced degrees. The multilingual coverage — 40+ languages — is stronger than most US-centric expert networks, which is the main reason to shortlist it over Surge or Mercor.

Delivery modelManaged expert-data service (sales-led)
Workforce6,000+ active contributors, 70%+ advanced degrees
Annotator locations100+ countries, 40+ languages
Self-serve signupNo (crowd self-serve de-emphasised)
Compliance / deploymentunknown

Watch out: A 6,000-contributor pool is orders of magnitude smaller than a crowd platform, so this is not the vendor for high-volume commodity labeling. The old self-serve crowdsourcing product is no longer the front door, so if you used Toloka for cheap microtasks that route appears gone. Its Yandex origins can trigger extended vendor security and geopolitical review in regulated or government procurement.

unknown — no public pricing page found

Snorkel AI

Snorkel has pivoted from selling Snorkel Flow, its programmatic weak-supervision labeling platform, to a data-services business: the Snorkel Data Series of ready-made expert datasets, custom dataset and RL-environment development, and specialised agents. It runs rubric-guided pipelines with calibrated experts across 1,000+ expert-level domains, and meta-evaluates its own reviewers. It also maintains public benchmarks including Terminal-Bench 4.0, Senior SWE-bench, OSWorld 2.0 and Harvey BigLaw Bench, which are the clearest evidence of what its evaluation work looks like.

Delivery modelManaged data-development service (sales-led)
WorkforceCalibrated experts across 1,000+ domains
Annotator locationsunknown
Self-serve signupNo
Compliance / deploymentunknown

Watch out: If you came looking for Snorkel Flow or self-hosted programmatic labeling, that product is no longer marketed — this is now a sales-led services engagement with no software you can run. Off-the-shelf Data Series datasets may not match your task distribution, and custom development carries the usual enterprise minimum and lead time.

unknown — no public pricing

Appen

Appen runs the ADAP platform and six data lines: frontier alignment (chain-of-thought, RLHF, red teaming), agentic AI trajectories and RL environments, speech and audio across 500+ locales, multimodal, physical AI/LiDAR, and model integrity work like hallucination benchmarking and bias audits. Its speech and dialectal-language depth is genuinely hard to replicate — that catalogue predates the LLM era by two decades. It is the only vendor here with audited public financials, which cuts both ways.

Delivery modelManaged service + ADAP platform (sales-led)
WorkforceGlobal crowd; size not published; 500+ speech locales
Annotator locationsGlobal crowd across 100+ countries
Self-serve signupNo
Compliance / deploymentunknown; ASX-listed with public disclosure

Watch out: Appen lost its largest customer contract in early 2024 and its revenue has fallen sharply since; as a small-cap listed company it carries more counterparty risk on a multi-year contract than a well-funded private vendor. Its heritage is general crowd work, and on PhD-level reasoning or specialist judgment it is competing against networks purpose-built for that. Neither crowd size nor pricing is disclosed publicly.

unknown — no public pricing

Sama

Sama is a managed annotation provider built on an impact-sourcing model: full-time, fair-wage employment for people historically excluded from the digital economy, with B Corp certification and a published impact report. Its strength is computer vision — object detection, segmentation, video object tracking, LiDAR and radar point clouds — plus text instruction-following and preference ranking. Named customers include Microsoft, Walmart, eBay and NASA. Unlike gig-based networks, its annotators are employees, which shows up in consistency on long-running projects.

Delivery modelManaged annotation service + platform (sales-led)
Workforce15,000+ full-time associates (employees, not gig)
Annotator locationsLargely East Africa and India (delivery centres)
Self-serve signupNo
Compliance / deploymentB Corp certified; security certs not published on site

Watch out: Much weaker fit for frontier LLM post-training or PhD-level reasoning data than the expert networks — this is an annotation workforce, not a specialist bench. Delivery is largely from East Africa and India, which can conflict outright with data-residency or citizenship requirements, so confirm before scoping. The company has previously faced public scrutiny over conditions in content-moderation work, which some buyers' ethics reviews will surface.

unknown — no public pricing; custom quotes only

iMerit

iMerit combines a managed workforce of 25,000+ domain experts with Ango Hub, its own annotation and model-evaluation platform supporting image, video, LiDAR, DICOM, text, PDF and audio. iMerit Scholars embeds credentialed specialists into the development loop, and its Deep Reasoning Lab provides interfaces for chain-of-thought and RLHF workflows. Its compliance posture is the most explicitly published of any managed vendor here — SOC 2, ISO 27001, GDPR, HIPAA and TISAX — which makes it the default shortlist entry for medical imaging and automotive.

Delivery modelManaged service + Ango Hub platform (sales-led)
Workforce25,000+ domain experts, 60+ countries
Annotator locationsIndia-centric with secure global delivery offices
Self-serve signupNo
Compliance / deploymentSOC 2, ISO 27001, GDPR, HIPAA, TISAX

Watch out: Delivery is India-centric, so the certifications solve the security question but not necessarily a data-residency or annotator-nationality requirement. Ango Hub is a bundled tool rather than a competitive standalone product — do not choose iMerit for the platform if you are not also buying the workforce. No public pricing and no self-serve trial.

unknown — no public pricing

Prolific

Prolific is the only vendor here with fully published pricing and true self-serve access: you define a population using 300+ prescreeners, host the task on your own infrastructure or any tool that produces a URL, and results start arriving within hours. It supports model evaluation, preference and RLHF data, safety testing and SFT across 80+ languages, and maintains HUMAINE, a public human-centred LLM evaluation benchmark. A Human Feedback API and an MCP integration let you drive collection programmatically. Quality control runs through Protocol, which applies 40+ identity and behavioural checks.

Delivery modelSelf-serve participant marketplace (API + UI)
Workforce300,000+ active verified participants, 80+ languages
Annotator locationsGlobal, skewed UK/US/Europe; targetable by demographics
Self-serve signupYes
Compliance / deploymentTasks hosted on your own infrastructure; certs not published

Watch out: The default pool is a representative research panel, not a bench of credentialed specialists — for volume PhD-level STEM or code annotation you will fight the prescreeners and pay more per usable label than a dedicated expert network. The 42.8% corporate platform fee is a large markup on already-rising participant rates. You must supply your own task interface; Prolific provides people and routing, not annotation tooling.

Platform fee 42.8% for corporate customers, 33.3% for academic/non-profit, charged on top of participant pay. Participant minimum £6.00/$8.00 per hour; recommended at least £9.00/$12.00 per hour. Managed research services quoted separately.

Labelbox

Labelbox now sells four things: Horizon (RL environments and preference signal for post-training and evals), Terra (robotics foundation-model data collected with purpose-built hardware), Alignerr (a network of 2.6M+ knowledge experts supplying human evaluation signal), and Recursion (an enterprise RL platform for building and continuously improving specialist agents). It cites work with over 90% of leading US AI labs, including producing the 820-problem expert-authored dataset behind Meta's GIM benchmark. The classic annotation tool still exists behind a free signup, but it is no longer what the company leads with.

Delivery modelPlatform + managed expert network (Alignerr)
WorkforceAlignerr: 2.6M+ knowledge experts
Annotator locationsGlobal; not broken out publicly
Self-serve signupYes (free tier); paid pricing not published
Compliance / deploymentunknown

Watch out: The pivot has real cost for existing users: the annotation platform is no longer the headline product and the public pricing page is gone, so a tool you may already depend on is now a moving target with opaque cost. Alignerr's 2.6M-strong pool is a scale claim, not a quality claim — variance across such a network is high and QA is on you or on Labelbox's process, unaudited. Recursion is a new product with a short production track record.

unknown — free signup exists but the public pricing page has been removed (returns 404); paid tiers quoted by sales

SuperAnnotate

SuperAnnotate's distinguishing feature is a fully customisable multimodal editor — you construct the exact annotation interface your task needs across image, video, text and audio, rather than adapting to a fixed tool. It sells three things together: the software platform, AI Data Services, and an Expert Talent Network of vetted, professionally managed annotation teams. Tiers differ mainly on Orchestrate compute hours (1K / 2.5K / 10K), SSO, and the level of assigned support. Use cases are framed around SFT, RLHF, RAG, agents and evaluation.

Delivery modelAnnotation platform + optional managed teams
WorkforceExpert Talent Network (managed teams); size not published
Annotator locationsunknown
Self-serve signupNo — demo request required at every tier
Compliance / deploymentSSO on Pro/Enterprise; Trust Center published

Watch out: No price is published at any tier, including Starter, and even the entry tier routes through a demo request — so there is no way to budget or trial without engaging sales. The customisable editor is genuinely flexible but that flexibility is setup work; expect real configuration effort before the first label. Weaker positioning for pure text/LLM preference work than the expert-data specialists.

unknown — three tiers (Starter, Pro, Enterprise) published with feature lists but no dollar figures; Orchestrate compute hours 1K/2.5K/10K respectively

Label Studio (HumanSignal)

Label Studio is the de facto open-source default for multimodal annotation — text, image, audio, video, time series — and the fastest way to stand up a labeling or preference-rating interface without a procurement cycle. HumanSignal sells the commercial editions: Starter Cloud for small teams, and Enterprise which adds SAML/LDAP SSO, LLM-as-a-judge, auto-labeling, bulk labeling and stronger quality workflows. A separate Data Services line supplies project management and domain experts. It is the only entry here where you can read the source, run it locally, and pay nothing.

Delivery modelOpen-source software + managed cloud + optional services
WorkforceNone bundled; Data Services quoted separately
Annotator locationsN/A — you supply annotators
Self-serve signupYes (free OSS and paid cloud)
Compliance / deploymentFully self-hostable; SSO/SAML on Enterprise only

Watch out: The open-source edition omits most of what makes labeling programs survive at scale: no SSO, no auto-labeling, no LLM-as-a-judge, and thin quality-control and agreement workflows — those are Enterprise-only. Self-hosting past a handful of annotators is real operational work you will own. No workforce is bundled: you still have to find and manage the humans, and the Data Services line is a separate custom engagement.

Community Edition free and open source. Starter Cloud $99/month base plus $49/month per additional user, up to 12 users. Enterprise and Data Services are custom-quoted. Academic program available. · open source

Encord

Encord is the annotation platform of choice for physical AI and medical imaging: alongside image, video, audio and documents it handles DICOM and NIfTI, 3D/LiDAR point clouds, geospatial data and ECG as paid add-ons. It bundles three surfaces — Annotate for labeling, Index for curation across up to 1bn+ assets, and Active for label-error detection, model evaluation and active-learning pipelines — plus data agents for automation. Enterprise adds SSO, multiple workspaces, an SLA, and VPC or on-prem deployment, which is the main reason regulated buyers shortlist it over cloud-only tools.

Delivery modelAnnotation & curation platform (self-serve to enterprise)
WorkforceNone bundled; Data-as-a-Service offered separately
Annotator locationsN/A — you supply annotators
Self-serve signupYes (Starter, Team); Enterprise is sales-led
Compliance / deploymentCloud, VPC (add-on), on-prem (add-on), SSO, MFA, enterprise SLA

Watch out: The headline tiers are misleading because the modalities most buyers come for — DICOM/NIfTI, LiDAR, geospatial, ECG, LLM evaluations — are all add-ons, so real cost is well above whatever the tier implies, and no figures are published to check that. Index and Active carry hard data-volume caps per tier. On-prem and VPC are also add-ons rather than included. If your work is text-only LLM post-training, most of what you would be paying for is irrelevant.

unknown — Starter, Team and Enterprise tiers are published with detailed feature and volume comparisons but no dollar figures; Enterprise is contact-sales

Argilla

Argilla is a free, self-hostable collaboration layer where AI engineers and domain experts curate datasets for fine-tuning, RLHF and evaluation, deployable in a Hugging Face Space in minutes. It joined Hugging Face and integrates tightly with the Hub's datasets and models, which makes it the lowest-friction option if your pipeline already lives there. Critically, its maintainers now state the project is stable but will receive no new features — only bug fixes and patches — and are inviting outside maintainers.

Delivery modelOpen-source self-hosted software (or HF Space)
WorkforceNone — you supply annotators
Annotator locationsN/A — you supply annotators
Self-serve signupYes — install or deploy free
Compliance / deploymentFully self-hostable; no vendor SLA or support

Watch out: Officially feature-frozen: the original authors have moved on, no new features are planned, and support is limited to bug-fix patches — do not build a multi-year annotation program on it. There is no managed workforce, no SLA and no commercial support to escalate to. For anything ongoing, Label Studio is the safer open-source choice.

Free — Apache-2.0; costs are only your own hosting (or a free Hugging Face Space) · open source · deprecated

NVIDIA NeMo Data Designer (formerly Gretel)

NVIDIA acquired Gretel and folded it into NeMo: gretel.ai now redirects to NVIDIA, and the successor is NeMo Data Designer, a compound AI system for generating synthetic datasets. You define columns declaratively — category and subcategory samplers, person samplers, numeric distributions, and LLM-generated text columns driven by templated prompts — using the `nemo-microservices[data-designer]` Python SDK against a hosted endpoint, with Nemotron, GPT-OSS and Llama models available by default. It is the cheapest way to bulk out SFT coverage or produce structured records where no real data exists.

Delivery modelSynthetic-data SDK + hosted API (self-hostable microservice)
WorkforceNone — no humans in the loop
Annotator locationsN/A — fully synthetic
Self-serve signupYes — free API key
Compliance / deploymentSelf-hostable via NeMo microservices; keeps data in your environment

Watch out: Synthetic generation cannot produce genuine human preference or judgment signal — it will not substitute for RLHF raters, and models trained only on it inherit the generator's biases. Gretel's standalone product and its differential-privacy tabular SaaS no longer exist as an independent offering, so prior customers had to migrate. Production use pulls you into the NVIDIA NIM/GPU ecosystem, and pricing beyond the free trial is not published.

Free trial via an NVIDIA API key on build.nvidia.com; production pricing for NeMo microservices / self-hosted deployment is not published on the Data Designer page · open source

Tonic.ai

Tonic solves the problem upstream of labeling: getting usable data out of a system that holds PII in the first place. Structural de-identifies structured databases across 12+ database types; Textual redacts and synthesises unstructured content including PDFs, DOCX, images and JSON; Fabricate generates synthetic relational data from scratch. For a labeling program this is the compliance path — you de-identify before the records reach an external annotation workforce, which changes what your DPA has to cover. It is the only entry here with a genuine credit-card self-serve tier at a published price.

Delivery modelSelf-serve SaaS + self-hosted enterprise software
WorkforceNone — no humans in the loop
Annotator locationsN/A — de-identification and synthesis only
Self-serve signupYes — free tier and $29/mo Plus
Compliance / deploymentSelf-hosted option on Textual Enterprise; purpose-built for PII removal

Watch out: This is not a labeling or human-evaluation vendor — it supplies no annotators and no preference signal, so it complements rather than replaces anything else in this list. Only Fabricate has published prices; Structural and Textual enterprise tiers are quote-only, and Structural's Professional tier is capped at 2 data sources with Oracle and Db2 gated to Enterprise. Automated de-identification still needs your own validation before you can rely on it for a regulatory claim.

Fabricate: Free $0/month with $5 monthly credits (roughly 9 complex generation sessions); Plus $29/month with $25 monthly credits plus pay-as-you-go; Enterprise custom. Structural: Professional (up to 10TB source data, 10 users, 2 data sources) and Enterprise, both custom-quoted. Textual: pay-as-you-go flat rate per 1,000 words, or Enterprise with self-hosting.