| Langfuse |
Yes — OTel-native plus Python/JS SDKs | Yes — LLM-judge, custom scores, human annotation queues | Yes — versioned prompt registry with SDK caching | Yes — MIT, free and fully featured for core | 50k units/mo cloud; unlimited self-hosted |
Hobby $0 (50k units/mo, 30-day retention, 2 users); Core $29/mo; Pro $199/mo; Enterprise cloud $2,499/mo — each with 100k units included, then $8/100k units, graduating to $7/$6.50/$6 per 100k at 1M/10M/50M. Self-hosted open source free; self-hosted Enterprise is custom-priced. |
| LangSmith |
Yes — native SDKs plus OTel ingest; LangGraph-aware views | Yes — datasets, experiments, LLM-judge, human review | Yes — prompt hub with versioning | Enterprise only (self-hosted and hybrid) | 5k base traces/mo, 1 seat |
Developer $0 (1 seat, up to 5k base traces/mo then pay-as-you-go); Plus $39/seat/mo (10k base traces/mo included, unlimited seats, includes one small serverless deployment); Enterprise custom. Overages billed as LCU $1.50/unit (compute) and LSU $1.00/unit (storage). Base traces have 14-day retention; 400-day extended traces cost extra. |
| Braintrust |
Yes — spans/logs, though evals are the centre of gravity | Yes — strongest in class: datasets, scorers, experiment diffs, CI | Yes — versioned prompts and playground | Enterprise only — hybrid, data plane in your cloud | Starter: 1 GB and 10k scores/mo |
Starter $0/mo ($10 model credits, 1 GB processed data, 10k scores, 14-day retention, unlimited seats); Pro $249/mo ($100 model credits, 5 GB then $3/GB, 50k scores then $1.50/1k, 30-day retention then $0.50/GB/mo); Enterprise custom. Starter overages: $4/GB and $2.50 per 1k scores. 6–12 months free for qualifying startups. |
| Weights & Biases Weave |
Yes — Python/TS decorators, OTel ingest supported | Yes — Evaluations API, scorers, online monitors | Yes — versioned prompt objects and playground | Enterprise only for commercial use; dedicated cloud available | 1 GB/mo ingestion, 5 seats |
Free $0 (up to 5 Models seats, 5 GB storage, 1 GB/mo Weave ingestion, extra at $0.10/MB); Pro from $60/mo (up to 10 Models seats, 100 GB storage then $0.03/GB, 1.5 GB/mo Weave ingestion, extra at $0.10/MB); Enterprise custom with dedicated cloud or self-managed. Local self-hosted 'Personal' edition is free but non-commercial. |
| Arize AX |
Yes — OpenTelemetry/OpenInference spans | Yes — online and offline evals, LLM-judge, experiments | Yes — prompt playground and versioning | Enterprise only | 25k spans/mo, 15-day retention |
AX Free $0 (25k spans/mo, 1 GB storage, 15-day retention, unlimited seats); AX Pro $50/mo (50k spans/mo, 10 GB, 30-day retention, unlimited seats); AX Enterprise custom (custom volume/retention, SaaS or self-hosted). Startup pricing available. |
| Arize Phoenix |
Yes — OTel/OpenInference, local UI | Yes — eval library with prepackaged judge templates | Yes — playground and prompt versioning | Yes — pip, Docker, Helm; free | Unlimited self-hosted; hosted limits unknown |
Self-hosted: free, no volume limits (infrastructure costs only). Hosted Phoenix Cloud free tier limits are not published on the pricing page — treat as unknown; paid production workloads are directed to Arize AX (Pro $50/mo). |
| Helicone |
Yes — request/session logging via proxy or async | Basic — scores and online evaluators; weakest area | Yes — prompt versioning and experiments | Yes — Apache-2.0, Docker Compose or Helm | 10k requests/mo, 7-day retention |
Hobby $0 (10k requests/mo, 1 GB storage, 1 seat, 7-day retention, 10 logs/min); Pro $79/mo (10k free requests then usage-based, unlimited seats, 1-month retention, 1,000 logs/min); Team $799/mo (5 orgs, 3-month retention, 15,000 logs/min, SOC 2/HIPAA); Enterprise custom (on-prem, SAML SSO). Storage overage roughly $0.97/GB. 50% off first year for startups under 2 years and $5M raised; free for students. |
| Comet Opik |
Yes — OTel-compatible, Python/JS SDKs | Yes — metrics SDK, LLM-judge, pytest-style test suites | Yes — prompt library with versioning, playground | Yes — Apache-2.0, same codebase as hosted | 25k spans/mo cloud; unlimited self-hosted |
Open source self-hosted: free, unlimited spans/users. Free Cloud $0 (25k spans/mo, 60-day retention, up to 10 users); Pro Cloud $19/mo (100k spans/mo, 60-day retention, up to 50 users); Enterprise custom. Overage $5 per 100k spans; extended retention 60→400 days at $29 per 100k spans. |
| Galileo |
Yes — traces and agent-graph views | Yes — Luna evaluator models, custom metrics, runtime guardrails | Yes — prompt versioning and experiments | Enterprise only — VPC or on-prem | 5,000 traces/mo |
Free $0/mo (5,000 traces/mo, unlimited users, unlimited custom evals); Pro $100/mo billed yearly (50,000 traces/mo, standard RBAC, dedicated Slack support), scaling with trace volume; Enterprise custom (unlimited traces, hosted/VPC/on-prem, SSO, guardrails, 24/7 support). Retention is not published per tier. |
| Datadog Agent Observability |
Yes — LLM spans correlated with APM/infra, OTel supported | Yes — datasets, automated evaluators, human review | Limited — experimentation over prompts, not a prompt CMS | No | None (14-day trial) |
Billed annually: first 100k LLM spans $160/mo; each additional 10k LLM spans $3.50/mo. Retention add-ons per 10k spans/mo: 30-day $1.50, 60-day $3, 90-day $4. No free tier beyond Datadog's standard 14-day trial; AI Credits sold in 500-credit bundles at $500/mo or $1.30/credit. |
| Pydantic Logfire |
Yes — OTel-native, SQL over spans, strong Python instrumentation | Partial — via separate Pydantic Evals library, not in-app | No | Enterprise only — Kubernetes/Helm | 10M records/mo, 30-day retention |
Personal $0 (10M records/mo, 30-day retention, 1 admin + 2 read-only guests, ingestion pauses at the limit); Team $49/mo (10M records included, $2 per additional million, 30-day retention, 5 seats, $25/extra seat); Growth $249/mo (10M included, $2/M, up to 90-day retention, unlimited seats); Enterprise custom with Kubernetes/Helm self-hosting. |
| Confident AI (DeepEval) |
Yes — spans, though evals came first | Yes — DeepEval metrics, CI test runs, live-traffic evals, red teaming | Yes — prompt versioning | Enterprise only (framework itself runs anywhere) | 1 GB-month spans, 5 test runs/week, 2 seats |
Free $0 (2 seats, 1 project, 1 GB-month of trace spans, 5 test runs/week); Starter $200/mo (unlimited seats, 5 projects, 5 GB-months); Team $2,000/mo (unlimited projects, 75 GB-months, SSO, custom RBAC, SOC 2); Enterprise custom (red teaming, HIPAA, on-prem). Span overage $1 per GB-month ingested or retained; online eval tokens billed separately at roughly $0.05/M input and $0.40/M output. |
| HoneyHive |
Yes — OTel-based, session and span views | Yes — automated evaluators plus human review workflows | Yes — prompt versioning and playground | Enterprise only — self-hosted, hybrid or single-tenant | 10,000 events/mo, 5 users |
Developer $0 (10,000 events/mo, up to 5 users, 30-day retention, no credit card). Enterprise: custom pricing, not published — custom event volume, unlimited users, custom retention, SSO, SLA, optional self-hosted/hybrid/single-tenant deployment. |
| Laminar |
Yes — OTLP ingest, SQL queries, browser session replay | Yes — evaluations plus LLM-driven 'Signals' analysis | Limited — unknown/secondary to tracing | Yes — open-source repo, self-support | 1 GB, 7-day retention, 1 seat |
Free $0 (1 GB data, $5 Signals credit, 7-day retention, 1 seat, 1 project); Starter $30/mo (3 GB then $2/GB, $15 Signals credit, 30-day retention, unlimited seats/projects); Pro $150/mo (10 GB then $1.50/GB, $50 Signals credit, 6-month retention); Enterprise custom. Signals metered at $0.50/$3 per 1M tokens on Starter, $0.40/$2.50 on Pro. |
| Freeplay |
Yes — observability with cost/latency analytics | Yes — model-graded, code-based and human review, batch testing | Yes — a core strength; prompts, models and params versioned together | unknown — not published | None published |
unknown — no public pricing page or published tiers; contact sales. Deployment options (SaaS versus VPC) are likewise not published. |
| Traceloop (OpenLLMetry) |
Yes — OpenLLMetry OTel instrumentation, exports to 25+ backends | Yes — evaluations and CI/CD checks in the hosted platform | Yes — included on the free tier | Enterprise only for the platform; library runs anywhere | 50k spans/mo, 5 seats, 24-hour retention |
Free Forever $0 (up to 50k spans/mo, up to 5 seats, 24-hour retention, includes monitoring, evaluation, CI/CD and prompt management); Enterprise custom for >50k spans/mo (unlimited seats, custom retention, SOC 2, on-prem/air-gapped, dedicated Slack). 14-day trial; available on AWS, GCP and Azure marketplaces. OpenLLMetry itself is free and Apache-2.0. |
| LangWatch |
Yes — OTel-compatible event/trace capture | Yes — custom evaluators, annotations, scenario-based agent simulation | Yes — versioning and prompt optimisation | Enterprise only for supported deployments; OSS components Apache-2.0 | 50k events/mo, 14-day retention, 2 users |
Developer €0 (50k events/mo, 14-day retention, 2 users, 3 scenarios / 3 simulations / 3 custom evals); Growth €29 per core seat/mo (200k events/mo then €5 per 100k, 30-day retention then €3/GB, unlimited lite users, unlimited simulations and evals, volume discounts above 20 users); Enterprise custom (negotiated volume, custom retention, SSO/RBAC, audit logs, ISO 27001, self-hosted/hybrid/on-prem). |
| PromptLayer |
Partial — request/log-centric, weaker on deep agent span trees | Yes — evaluation pipelines and regression tests, basic depth | Yes — the core product; visual registry, versions, release labels | Enterprise only — GCP/AWS/Azure or single-tenant | 2,500 requests/mo, 5 seats |
Free $0 (2.5k requests/mo, 5 seats, 1 workspace); Pro $49/mo (2.5k+ requests, $0.003/transaction overage, 5 seats, unlimited workspaces); Team $500/mo (100k+ requests, $0.002/transaction overage, 25 seats); Enterprise custom (self-hosted on GCP/AWS/Azure, EU-hosted or single-tenant). Free/Pro/Team are US-hosted only. |
| Openlayer |
Yes — inference tracing and production monitoring | Yes — test-suite framing, dev and production tests with alerts | Limited — version comparison rather than a prompt CMS | Enterprise only — on-premise available | 20,000 inferences/mo, 1 member, 3-month retention |
Basic $0 (20,000 inferences/mo, 1 member, 5 projects, 3-month retention, no on-premise). Enterprise: custom quote based on scale, deployment and support — no dollar figures published. There is no self-serve paid tier between the two. |
| Humanloop |
Was yes — no longer viable | Was yes — no longer viable | Was yes (its main strength) — no longer viable | unknown / moot | Not applicable — sunsetting |
unknown — pricing is no longer meaningful; the platform is being sunset following the Anthropic acquisition. No end-of-service date is published on the announcement page. |