| LlamaParse |
130+ types: PDF, DOCX, PPTX, XLSX, HTML, JPEG, PNG, XML, EPUB | Markdown, plain text, JSON, spatial text | Tier-dependent; agentic tiers target complex tables and charts, Fast tier is text-only | Credits per page (1,000 credits = $1.25) | Async job API; Fast tier is the low-latency option, agentic tiers are markedly slower (no published SLA) |
1,000 credits = $1.25. Free tier 10,000 credits/month. Starter $50/mo (40,000 credits, PAYG up to 400,000 more). Pro $500/mo (400,000 credits, PAYG up to $5,000/mo). Basic parsing 'as low as 1 credit' per page (~$0.00125/page); agentic tiers cost multiple credits per page — the exact per-tier credit table was not retrievable from the docs site at time of writing. |
| Unstructured |
50-65+ types: PDF, DOCX, PPTX, XLSX, HTML, EML, images, spreadsheets | JSON document elements (typed), chunked and optionally embedded; markdown/HTML per element | Good on ordinary documents; table and image enrichment is a paid-platform feature, weaker than VLM parsers on dense tables | Per page ($0.015) | Async pipeline/workflow model, not a low-latency request API; OSS hi_res mode is seconds per page on CPU |
Free: 10,000 pages, no card, all features. Pay-as-you-go: $0.015/page after the free 10,000. Business: custom pricing with multi-user accounts and dedicated instance/VPC deployment. Compliance: HIPAA, SOC 2 Type 2, GDPR, ISO 27001. |
| Reducto |
30+ types: PDF, images, spreadsheets, presentations | Markdown/JSON with bounding boxes, table structure, table summaries; schema-shaped JSON via Extract | Among the strongest on dense/nested tables and multilingual scans; bounding boxes returned for citation and audit | Per 1,000 pages, priced separately per operation | Async API; Deep Extract and Deep Split are markedly slower than Parse (no published SLA) |
Standard PAYG with $150 free credits. Per 1,000 pages: Parse (r-1) $10, Extract $20, Deep Extract $40, Split $20, Deep Split $40, Classify $7.50, Edit $60 ($15 pre-filled). 200 concurrent pages, up to 5 Studio seats. Growth: custom, 350 concurrent, BAA + ZDR + EU/AU endpoints. Enterprise: custom, 500+ concurrent, VPC/on-prem, SSO/SAML. Startup program (under $15M raised / $3M revenue / 50 employees) and up to $5,000 in migration credits. |
| Docling |
PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, LaTeX, EML/MSG, images (PNG/TIFF/JPEG), audio (WAV/MP3), video (MP4/MOV/MKV/WebM) | Markdown, HTML, JSON, DocTags, DocLang, WebVTT | Strong table-structure recognition and OCR for its class; GraniteDocling VLM adds chart-to-table conversion | Free (self-hosted compute only) | Local; fast on digital PDFs, slow on CPU for scans/VLM pipelines — GPU strongly recommended for volume |
Free. MIT license on the codebase (individual model weights carry their own licenses). You pay only for the CPU/GPU you run it on. |
| MinerU |
PDF, DOCX, PPTX, XLSX, images, web pages | Markdown and JSON; formulas as LaTeX, tables as HTML | Strong on formulas and multi-column scientific layouts; 109-language OCR incl. handwriting (PP-OCRv6) | Free (self-hosted compute only) | Local; CPU-only supported but slow — GPU (Volta+ or Apple Silicon) recommended |
Free. MinerU Open Source License (Apache 2.0 derivative) since v3.1.0 — previously AGPLv3. Self-hosted compute only. |
| Mistral Document AI (OCR) |
PDF and images (page-based); document URLs or base64 uploads | Per-page markdown with paragraph bounding boxes, structural block labels, block confidence scores; JSON annotations | Strong general OCR and layout; per-block confidence scores enable triage. Below Reducto on dense nested tables | Per 1,000 pages ($4, $0.40 cached) | Synchronous API, typically seconds per document; no published SLA |
$4.00 per 1,000 pages input; $0.40 per 1,000 pages cached input. Same rate for OCR 4.1 and OCR 4.0. Billed per page, not per token. |
| Datalab (Marker / Surya) |
PDF, images, and common office documents | Markdown, JSON, HTML; schema-shaped JSON via Extraction; filled PDFs via Form Fill | Strong for the price on equations and tables (Surya OCR); accurate tier closes much of the gap to premium parsers | Per 1,000 pages, by processor and speed tier | Fast tier is the low-latency path; accurate tier trades seconds per page for fidelity. Rate limits raised on Team plan |
Free: $20/month allowance for work-email accounts, $10 for personal, then pay-as-you-go. Per 1,000 pages — Convert (fast/balanced) $4, Convert (accurate) $10, Extraction fast $6 / balanced $15+ / accurate $20+, Form fill $10. Add-ons +$3 to +$6 per 1,000 pages. EU processing +25%. Data-retention opt-in -25%. Team plan $400/month. Startup discount 33% for year one; up to $5,000 migration credits. |
| Azure AI Document Intelligence |
PDF, images (JPEG/PNG/BMP/TIFF), and Office formats for selected models | JSON with paragraphs, roles, tables, key/value pairs, selection marks, barcodes, formulas; Layout emits markdown; searchable PDF add-on | Solid, mature OCR with high-resolution add-on for small text; table extraction is reliable rather than best-in-class | Per 1,000 pages, tiered by volume and region | Async analyze-then-poll operation; seconds to tens of seconds per document depending on model and add-ons |
Free F0 tier: 500 pages/month. Pay-as-you-go rates are per 1,000 pages and tiered by volume (0-1M pages vs 1M+) for Read, prebuilt, custom extraction, classification and generative extraction — the specific per-1,000-page dollar figures did not render on the public pricing page at time of writing and are region-specific; verify in the Azure pricing calculator. Custom neural model training is $3/hour after 10 free hours; template model training is free. Monthly commitment tiers available at 20,000 / 100,000 / 500,000 pages. |
| Amazon Textract |
PDF, TIFF, JPEG, PNG | JSON blocks with geometry, relationships, confidence; no native markdown | Very good raw OCR incl. handwriting; table extraction is competent but reading-order reconstruction is left to you | Per page, priced per API and per feature | Synchronous API for single-page documents (sub-second to seconds); async job API required for multi-page PDFs |
US West (Oregon): Detect Document Text $0.0015/page first 1M pages, $0.0006/page after. Analyze Document — Forms $0.05/page, Tables $0.015/page (first 1M), Queries $0.015/page. Analyze Expense $0.01/page first 1M, $0.008/page after. Analyze ID $0.025/page first 100K, $0.01/page after. Analyze Lending $0.07/page first 1M, $0.055/page after. Free tier: 3 months, monthly 1,000 pages Detect Document Text, 100-1,000 pages Analyze Document by feature, 100 pages each Expense and ID. |
| Firecrawl |
URLs (HTML with JS rendering), plus PDFs and DOCX encountered while crawling | Markdown, HTML, structured JSON (schema-guided), screenshots, links | HTML-to-markdown tables are good; document OCR is a secondary capability, not competitive with dedicated parsers | Credits (1 credit = 1 page scraped) | Single scrape typically a few seconds; crawls are async jobs whose duration scales with site size and concurrency limits |
Free: 1,000 credits/month, no card. Hobby $16/mo billed yearly (5,000 credits/mo). Standard $83/mo yearly (100,000). Growth $333/mo yearly (500,000). Scale $599/mo yearly (1,000,000). Enterprise custom. Credit costs: scrape/crawl/map/monitor 1 per page; search 2 per 10 results; interact 2 per browser-minute. PAYG top-ups on paid plans only, $5 per increment (1,000-5,000 credits by plan). Annual discount 16.7% (Hobby/Standard/Growth), 20% (Scale). |
| Jina Reader |
URLs (HTML, with PDF support); search queries via s.jina.ai | Markdown (default), JSON, structured JSON via x-json-schema/x-instruction, streaming mode | Reasonable HTML table conversion; not a document OCR product | Output tokens (search billed from a ~10,000-token floor per request) | Typically a few seconds per URL; stream mode for large pages. Bounded by RPM tier, not by an SLA |
10 million free tokens with each new API key. Billing is per output token for r.jina.ai and a fixed floor of ~10,000 tokens per s.jina.ai search request; tokens are bought via Stripe and failed requests are not charged. Rate limits: no key 20 RPM (reader only, search blocked); free key 500 RPM reader / 100 RPM search; premium 5,000 RPM reader / 1,000 RPM search. The USD-per-token rate is behind the authenticated dashboard and could not be verified at time of writing. |
| Exa |
Natural-language or keyword queries; URLs for contents retrieval | Ranked results with URLs, titles, published dates; full page text, highlights, AI summaries; cited answers | n/a — search and page-text API, not a document parser | Per 1,000 requests (search) and per 1,000 pages (contents) | Search is sub-second to a couple of seconds; Deep Search and Agent runs take substantially longer by design |
Pay-as-you-go, no subscription or minimum. Search $7/1,000 requests (up to 10 results); Deep Search $12-15/1,000; Contents $1/1,000 pages; Answer $5/1,000 requests; Monitors $15/1,000; Agent $0.012-$1.00 per fixed-effort run or usage-metered. Results beyond the first 10 cost $1/1,000 results; AI page summaries $1/1,000 pages on any endpoint. New accounts get $20 in credits (~2,800 searches) plus $10/month on the free tier. Enterprise plans offer volume discounts, custom indexes and SLAs. |
| Tavily |
Natural-language queries; URLs for Extract/Crawl/Map | JSON with ranked snippets, relevance scores, optional LLM answer; markdown page content from Extract/Crawl | n/a — search and HTML extraction; no document OCR | API credits ($0.008 each; 1 credit per basic search) | Basic search typically 1-2 seconds; advanced search and crawl noticeably slower. Rate limits differ by plan |
Free: 1,000 API credits/month, no card. Pay-as-you-go: $0.008/credit. Project plan: adjustable monthly subscription starting at 4,000 credits/month with higher rate limits (base price is set by an on-page slider and was not captured). Enterprise: custom. Credit costs — basic search 1, advanced search 2; basic extract 1 per 5 URLs, advanced extract 2 per 5; map 1 per 10 pages (2 with instructions); crawl = map + extract combined. Credits reset monthly and do not roll over. Free access for students. |
| Brave Search API |
Keyword and natural-language queries; Goggles files for custom reranking | JSON SERP results (web, news, images, video), LLM-context payloads, schema-enriched results, cited streaming answers | n/a — search API, not a document parser | Per 1,000 requests ($5 Search, $4 Answers + $5/M tokens) | Search is fast at 50 QPS sustained; Answers is capped at 2 QPS and streams responses |
Search plan: $5 per 1,000 requests, 50 queries/second, with $5/month in free credits applied automatically. Answers plan: $4 per 1,000 requests plus $5 per million input/output tokens, 2 queries/second, also with $5/month free credits. Enterprise: custom pricing with full-funnel zero data retention, custom agreements and invoicing. Endpoints included: Web, LLM Context, Answers, Image, Video, News, Suggest, Spellcheck. |
| Perplexity Search & Agent API (formerly Sonar) |
Natural-language queries and chat messages; multi-query search supported | Ranked search results JSON (Search API); grounded prose with citations, streaming, tool calls (Agent API); embeddings | n/a — search and answer API, not a document parser | Per 1,000 requests (Search) and per million tokens + per tool invocation (Agent) | Search API is sub-second to seconds; Agent responses take seconds and deep-research runs take minutes |
Search API: $5.00 per 1,000 successful requests. Agent API: no per-request fee — model tokens at each provider's published rate ($0.13-$25 per million in/out), tools $0.0005-$0.005 per invocation (web search, fetch URL, people search, finance search), sandbox sessions $0.03 each. Router API: token-based at each model's published rate, no per-request fee. Legacy Sonar models: sonar $1/$1 per million in/out, sonar-pro $3/$15, sonar-reasoning-pro $2/$8, sonar-deep-research $2/$8, plus per-request search fees of $5-6/1K (low context), $8-10/1K (medium), $12-14/1K (high). |
| ScrapingBee |
URLs (HTML), with optional JS rendering, custom JS scenarios and geotargeting | Raw HTML, screenshots, JSON via CSS/XPath extraction rules; dedicated Google/search endpoints | n/a — returns HTML; no OCR or markdown conversion | API credits (1 plain / 5 JS / 10-25 premium proxy / 75 stealth) | Plain requests are sub-second to seconds; JS rendering adds seconds. Failed URLs are retried for up to 30 seconds |
Hobby $19/mo (75,000 credits, 25 concurrent). Freelance $49/mo (250,000, 50 concurrent). Startup $99/mo (1,000,000, 100 concurrent). Business $249/mo (3,000,000, 200 concurrent). Business+ $599/mo (8,000,000, 400 concurrent). Trial: 1,000 free API credits, no card. Credit costs: 1 default, 5 with render_js, 10 premium proxy without JS, 25 premium proxy with JS, 75 stealth proxy (requires JS). Auto-Mode charges only the tier that succeeded; failed requests cost 0 credits. Prices exclude VAT. |
| Apify |
URLs and site-specific inputs (search terms, profile URLs, categories) defined per Actor | JSON, CSV, XML, Excel datasets via API or dataset export; key-value store for files | n/a — structured web records; no document OCR | Platform credits, drawn by compute units ($0.13-$0.20/CU) plus per-event or per-result Actor fees | Actor runs are async jobs from seconds to hours depending on scope; not suited to synchronous request-time use |
Free $0 (about $5 of platform credits/month). Starter $19/mo ($17 annual, $19 credits). Scale $199/mo ($179 annual, $199 credits). Business $999/mo ($899 annual, $999 credits). Compute units: $0.20/CU on Free and Starter, $0.16/CU on Scale, $0.13/CU on Business, where 1 CU = 1 GB RAM for 1 hour. Actors bill either pay-per-event (fixed price per developer-defined action, usually inclusive of platform usage) or pay-per-usage (compute + data transfer). Paid users are billed for overage; free users are blocked until the next cycle. Unused credits expire monthly with no rollover. |
| Zyte API |
URLs, with HTTP or browser request modes, browser actions and session handling | Raw HTTP body or rendered browser HTML, screenshots, automatic structured extraction (product, article, job posting schemas) | n/a — returns HTML/structured records; no document OCR | Per successful request, priced by site-specific tier 1-5 and request type | HTTP-tier requests are fast; browser-tier and action-heavy requests add seconds. Failed and rate-limited requests are free |
$5 free credit for your first billing month. Standard pay-as-you-go with a $100/month spending limit and no commitment; Standard with commitment carries $200-$2,500/month limits; Enterprise is custom. Per-request cost depends on target site, request type (HTTP or browser) and an automatically assigned tier 1-5, so there is no single headline rate — use the dashboard cost estimator. Add-ons: screenshots $0.002, extraction $0.0004-$0.0016, browser actions billed on CPU/network. Volume discounts: 25% at $100 commitment, 52% at $500. Only successful responses are billed. |
| Bright Data Web Scraper API |
Target URLs, search terms and site-specific inputs for 100+ supported domains | Structured JSON or CSV records with validation; also raw unlocked HTML via companion Web Unlocker product | n/a — structured web records; no document OCR | Per 1,000 records delivered ($1.50 PAYG, $1.30 at scale) | Batch/async collection jobs; unlimited concurrency but delivery is minutes-scale, not request-time |
Free tier: 5,000 records/month, no credit card. Pay-as-you-go: $1.50 per 1,000 records with customisable spend limits. Scale: $499/month including 384,000 records, then $1.30 per 1,000 additional records. Enterprise: custom with volume discounts and a dedicated account manager. All plans include automated proxy management, full browser rendering, CAPTCHA solving, unlimited concurrent requests, batch scheduling and JSON/CSV output. No charge for failed deliveries. |