# DocParse > DocParse is a document parsing API that converts PDFs, scans, images, Office files, HTML, XML, and CSV into clean Markdown and structured DocIR JSON for RAG, search, automation, and AI products. DocParse uses deterministic-first routing where possible and makes model-backed or external OCR an explicit, cost-controlled workspace capability. Public product claims and architecture are documented in the pages below. ## Start here - [Product overview](https://docparse.genedai.me/product): Capabilities and output model. - [API quickstart](https://docparse.genedai.me/developers): Authentication, first request, results, jobs, and production controls. - [Document parser decision library](https://docparse.genedai.me/compare): Direct comparisons, best-tool shortlists, and alternatives. - [Build for RAG](https://docparse.genedai.me/build-for/rag): Structure-aware retrieval ingestion. - [Security and data controls](https://docparse.genedai.me/security): Implemented trust boundaries and limitations. - [OpenAPI 3.1](https://docparse.genedai.me/openapi.json): Machine-readable API contract. ## Product - [A document parsing API built for the ingest layer](https://docparse.genedai.me/product/document-parsing-api): DocParse gives product teams one authenticated API for turning mixed documents into clean Markdown and structured DocIR JSON. It starts with an inline request, then supports jobs, polling, lifecycle events, cancellation, webhooks, scoped keys, quotas, and artifact purge when the workflow moves into production. - [PDF to Markdown that keeps the document useful](https://docparse.genedai.me/product/pdf-to-markdown): DocParse converts PDFs into Markdown for RAG and LLM workflows while also returning a structured representation of pages and blocks. Digital PDFs use a deterministic text-layer path when possible; scans and complex layouts can be routed to OCR or layout-aware processing under explicit workspace controls. - [A structured document model for products that need more than text](https://docparse.genedai.me/product/structured-document-json): DocIR is DocParse’s normalized JSON document representation. It keeps source metadata, pages, ordered typed blocks, optional geometry, assets, parser provenance, and warnings in one versioned record while Markdown remains available as a convenient text view. - [Route each document to the parser it actually needs](https://docparse.genedai.me/product/ocr-routing): DocParse uses source type, size, embedded-text evidence, and requested fidelity to choose a deterministic, OCR, layout-aware, image-vision, Office, or external OCR path. The decision is service controlled, recorded with the job, and bounded by tenant capability and cost policy. ## Build for - [Build a RAG ingest layer on document structure, not cleanup scripts](https://docparse.genedai.me/build-for/rag): For RAG, DocParse converts mixed source files into Markdown plus page-aware DocIR before chunking. The pipeline can split on structural boundaries, attach citations, retain parser provenance, and reprocess only the document cohorts whose quality needs improvement. - [Give AI agents a document tool with a real control plane](https://docparse.genedai.me/build-for/ai-agents): DocParse lets an agent submit a document, observe a job, and consume Markdown or typed DocIR without receiving storage credentials or choosing internal parser infrastructure. Scoped keys, quotas, idempotency, lifecycle events, and purge make the tool safer to expose in an agent workflow. - [Turn mixed files into a maintainable knowledge base](https://docparse.genedai.me/build-for/knowledge-bases): DocParse provides a consistent ingestion boundary for PDFs, Office documents, scans, images, HTML, XML, and CSV. Knowledge-base builders can preserve headings and pages, track source and parser versions, and re-index documents without coupling the product to every file-specific parser. - [Build enterprise search on traceable document structure](https://docparse.genedai.me/build-for/enterprise-search): DocParse normalizes enterprise files into searchable Markdown and DocIR while preserving page and block context. Tenant-scoped keys, outbound host policy, observable jobs, and explicit purge support an ingestion service that can sit behind an existing connector and permission layer. - [Give document automation one reliable intake layer](https://docparse.genedai.me/build-for/document-automation): DocParse turns varied uploads into a normalized document record before classification, field extraction, validation, human review, or downstream routing. The workflow can branch on block type, route, warning, and job outcome instead of embedding parser-specific exceptions throughout the product. ## Comparisons, shortlists, and alternatives - [DocParse vs LlamaParse: choose by workflow boundary](https://docparse.genedai.me/compare/docparse-vs-llamaparse): LlamaParse is a managed parsing product in the LlamaIndex ecosystem with file parsing jobs and text or Markdown expansion. DocParse is a Cloudflare-hosted parsing service that emphasizes deterministic-first routing, Markdown plus versioned DocIR, and an integrated tenant control plane. The right choice depends on ecosystem, output needs, deployment boundary, and the representative files you evaluate. - [DocParse vs Unstructured: managed boundary or parsing toolkit](https://docparse.genedai.me/compare/docparse-vs-unstructured): Unstructured provides open-source partition functions and hosted APIs that break documents into typed elements across many formats and strategies. DocParse provides a Cloudflare-native service with Markdown plus DocIR, deterministic-first routing, and integrated commercial controls. Choose based on how much parser infrastructure you want to own and which output contract your application needs. - [DocParse vs Mistral OCR: orchestration layer or OCR processor](https://docparse.genedai.me/compare/docparse-vs-mistral-ocr): Mistral OCR is a Document AI processor for extracting text and structured content from documents, including layout features documented by Mistral. DocParse is a broader parsing control plane that can use deterministic, Office, OCR, vision, or explicitly permitted external OCR routes and normalize them into Markdown plus DocIR. They are different layers, and DocParse can use Mistral OCR as an optional provider route. - [DocParse vs Docling: managed service or local conversion toolkit](https://docparse.genedai.me/compare/docparse-vs-docling): Docling is an open-source toolkit that converts many document formats into a unified DoclingDocument and exports Markdown, JSON, HTML, text, and chunk formats. DocParse is a managed Cloudflare service with routing, DocIR, API credentials, jobs, quotas, webhooks, and purge. Choose Docling for direct local control; choose DocParse when an operated API boundary is the priority. - [DocParse vs Amazon Textract: ingest layer or AWS OCR service](https://docparse.genedai.me/compare/docparse-vs-aws-textract): Amazon Textract is an AWS document-analysis service centered on text, forms, tables, queries, signatures, and layout blocks. DocParse is a broader document-ingestion control plane that routes mixed file formats and normalizes results into Markdown plus DocIR. Choose by whether you need an AWS-native analysis primitive or a parser-independent product boundary. - [DocParse vs Google Document AI: control plane or Gemini layout parser](https://docparse.genedai.me/compare/docparse-vs-google-document-ai): Google Document AI offers specialized processors, including a layout parser designed to preserve tables, figures, lists, headers, and hierarchy for search and RAG. DocParse sits at a different boundary: it routes mixed formats, can keep clean files on deterministic paths, and returns one Markdown plus DocIR contract with tenant operations. The choice is processor capability versus ingestion-system ownership. - [DocParse vs Azure Document Intelligence for layout-aware ingestion](https://docparse.genedai.me/compare/docparse-vs-azure-document-intelligence): Azure AI Document Intelligence's layout model extracts text, tables, selection marks, figures, sections, and logical roles, and its current API can return Markdown. DocParse adds a product-facing routing and operations layer across deterministic parsers and optional model providers. Choose Azure for a managed Azure analysis model; choose DocParse for a normalized multi-route ingest contract. - [DocParse vs Adobe PDF Extract: PDF specialist or mixed-file ingest layer](https://docparse.genedai.me/compare/docparse-vs-adobe-pdf-extract): Adobe PDF Extract focuses on extracting PDF text, structure, tables, figures, reading order, and renditions into structured outputs; Adobe also documents a PDF-to-Markdown operation. DocParse accepts a wider mixed-file workload and wraps parsing with normalized DocIR, routing policy, tenant controls, lifecycle events, and purge. The key distinction is PDF specialization versus product-wide ingestion. - [DocParse vs Reducto: focused ingest control or document platform](https://docparse.genedai.me/compare/docparse-vs-reducto): Reducto documents a broad agentic document platform spanning parse, extract, classify, split, edit, and reusable pipelines, with cloud and enterprise deployment options. DocParse is deliberately narrower: a Cloudflare-native parsing and normalization layer with deterministic-first routing and bounded provider use. Choose by whether you need a full document-workflow platform or a focused ingest primitive. - [DocParse vs LandingAI ADE for agentic document workflows](https://docparse.genedai.me/compare/docparse-vs-landingai-ade): LandingAI Agentic Document Extraction separates parsing, extraction, splitting, classification, and section operations and returns semantic document chunks. DocParse focuses on format-aware parsing, normalized Markdown and DocIR, and production ingest controls. ADE fits model-led document understanding; DocParse fits teams that want deterministic-first routing and a compact application boundary. - [DocParse vs Nanonets: parsing infrastructure or intelligent automation](https://docparse.genedai.me/compare/docparse-vs-nanonets): Nanonets positions its Document Intelligence API around OCR, field and table extraction, structured data, review workflows, and business-system integrations. DocParse concentrates on converting mixed documents into Markdown and DocIR with deterministic-first routing and explicit lifecycle controls. Nanonets fits extraction automation; DocParse fits a general AI-product ingest layer. - [DocParse vs Marker: managed document API or open-source parser](https://docparse.genedai.me/compare/docparse-vs-marker): Marker is an open-source document converter that produces Markdown, JSON, HTML, or chunks and can run locally on CPU, GPU, or Apple Silicon with optional VLM assistance. DocParse is a hosted multi-tenant API with format routing, normalized DocIR, quotas, jobs, webhooks, and retention controls. Choose by whether you want to operate the parser or consume a service boundary. - [DocParse vs MinerU: self-hosted parsing stack or managed ingest API](https://docparse.genedai.me/compare/docparse-vs-mineru): MinerU is an open-source parsing system that converts complex PDFs and Office documents into Markdown and JSON and now documents router, API, multi-GPU, concurrency, and long-document improvements. DocParse is a Cloudflare-hosted service with normalized DocIR and tenant operations. The decision is infrastructure ownership, output contract, and operational scope—not a universal accuracy ranking. - [DocParse vs PaddleOCR: OCR pipeline or managed ingestion boundary](https://docparse.genedai.me/compare/docparse-vs-paddleocr): PaddleOCR's PP-StructureV3 is an open-source document-parsing pipeline for OCR, layout blocks, tables, formulas, reading order, JSON, and Markdown. DocParse is a hosted orchestration and normalization service that can use deterministic, OCR, vision, or external routes. PaddleOCR fits teams operating models locally; DocParse fits teams consuming a stable product API. - [DocParse vs MarkItDown: conversion library or document API](https://docparse.genedai.me/compare/docparse-vs-microsoft-markitdown): Microsoft MarkItDown is a Python utility for converting common files and Office documents into Markdown, with optional dependencies and plugins. DocParse is a hosted API that also returns structured DocIR and supplies routing, jobs, quotas, webhooks, tenant access, and purge. MarkItDown fits local lightweight conversion; DocParse fits production product ingestion. - [Best document parsing APIs for AI products in 2026](https://docparse.genedai.me/compare/best-document-parsing-apis): The best document parsing API depends on the layer you need. LlamaParse and LandingAI emphasize model-led parsing, cloud platforms provide managed OCR and layout, Reducto spans a broader document lifecycle, and DocParse adds deterministic-first routing plus a normalized product boundary. This shortlist maps fit and trade-offs; it is not a universal accuracy ranking. - [Best PDF-to-Markdown tools for LLM workflows in 2026](https://docparse.genedai.me/compare/best-pdf-to-markdown-tools): A PDF-to-Markdown tool should be chosen by document class and operating model. MarkItDown is lightweight, Marker and MinerU offer deeper local parsing, Docling provides a structured conversion toolkit, Adobe and Mistral provide cloud APIs, and DocParse wraps multiple routes in a production API. No converter is best for every PDF. - [Best OCR APIs for complex documents in 2026](https://docparse.genedai.me/compare/best-ocr-apis): OCR APIs differ in scope. Amazon Textract emphasizes forms and tables, Google and Azure provide layout processors, Mistral offers a focused OCR API, Adobe specializes in PDFs, and Nanonets targets business extraction workflows. DocParse is the orchestration layer when OCR is only one route. Match the service to the output and operating boundary you need. - [Best open-source document parsers for AI in 2026](https://docparse.genedai.me/compare/best-open-source-document-parsers): Open-source document parsers solve different layers. MarkItDown is a lightweight converter, Docling provides a structured toolkit, Unstructured emits typed elements, Marker and MinerU run deeper model pipelines, and PaddleOCR focuses on OCR and layout. The best choice depends on formats, hardware, license obligations, output model, and the operations your team can own. - [Best document ingestion tools for RAG in 2026](https://docparse.genedai.me/compare/best-rag-document-ingestion-tools): RAG ingestion is not one product category. Unstructured and Docling provide conversion primitives, LlamaParse and cloud layout models provide managed parsing, Reducto and LandingAI add broader document operations, and DocParse owns normalization and lifecycle before chunking. Choose the missing layer in your architecture, then test retrieval outcomes end to end. - [LlamaParse alternatives for RAG and document AI in 2026](https://docparse.genedai.me/compare/llamaparse-alternatives): LlamaParse alternatives fall into four groups: managed parsing APIs such as DocParse, cloud layout services from Google, Azure, AWS, Adobe, and Mistral, broad platforms such as Reducto or LandingAI, and self-hosted tools such as Docling, Marker, and Unstructured. The right replacement depends on why you are leaving—not on feature-count alone. - [Unstructured alternatives for document parsing in 2026](https://docparse.genedai.me/compare/unstructured-alternatives): Alternatives to Unstructured depend on which part you use: local partitioning, typed elements, connectors, hosted parsing, or RAG preprocessing. Docling and MarkItDown cover local conversion, Marker and MinerU add model pipelines, managed cloud APIs cover OCR and layout, and DocParse supplies normalized ingestion operations. Map the replacement to the actual dependency. - [Mistral OCR alternatives for document AI in 2026](https://docparse.genedai.me/compare/mistral-ocr-alternatives): Mistral OCR alternatives include AWS, Google, Azure, and Adobe cloud services; managed parsers such as LlamaParse, Reducto, LandingAI, and DocParse; and self-hosted tools such as Marker, MinerU, Docling, and PaddleOCR. Choose based on whether you need focused OCR, layout structure, broader workflows, local control, or a normalized ingest layer. - [Docling alternatives for document parsing in 2026](https://docparse.genedai.me/compare/docling-alternatives): Docling alternatives range from lightweight MarkItDown to typed-element Unstructured, model-backed Marker and MinerU, OCR-focused PaddleOCR, managed parsers such as LlamaParse, and hosted APIs such as DocParse. Select by which Docling capability you need to replace: conversion, structured document modeling, local control, or production service operations. ## Engineering guides - [What a production document parsing API should actually do](https://docparse.genedai.me/blog/document-parsing-api-guide): A production document parsing API should normalize mixed files into readable text and structured document data, expose route and failure state, support synchronous and asynchronous use, and enforce access, cost, retention, and egress controls. Text extraction alone is only one stage of that contract. - [PDF to Markdown for LLMs: preserve structure before tokens](https://docparse.genedai.me/blog/pdf-to-markdown-for-llms): Converting PDF to Markdown for an LLM is useful only when reading order, headings, lists, tables, and page context survive the transformation. Keep a structured companion record beside Markdown so chunks and citations remain traceable to source pages and parser evidence. - [How to build a RAG document ingestion pipeline that can be debugged](https://docparse.genedai.me/blog/build-rag-document-ingestion): A reliable RAG ingestion pipeline records each transformation from source to active index: identify and hash the file, parse it into a structured document, validate output, chunk by structure, embed versioned chunks, switch the active index, and evaluate retrieval with traceable citations. - [How to evaluate a document parsing API without fooling yourself](https://docparse.genedai.me/blog/how-to-evaluate-document-parsing-api): Evaluate a document parsing API on representative document cohorts, label the source evidence you care about, score structure separately from text, record failures and latency, and measure the downstream task. Do not generalize from a single showcase PDF or an unsegmented average. - [Why tables and reading order decide whether parsed text is trustworthy](https://docparse.genedai.me/blog/tables-reading-order-document-parsing): A parser can recover every visible word and still produce a wrong document if it interleaves columns, disconnects captions, or scrambles table cells. Reading order and structural relationships must be evaluated as first-class output, with page geometry retained for audit and correction. - [Deterministic vs AI document parsing is a routing decision](https://docparse.genedai.me/blog/deterministic-vs-ai-document-parsing): Deterministic parsing is fast, repeatable, and effective when a document contains reliable native structure. OCR is necessary for pixels, and vision-language parsing can help with difficult layout or visual semantics. A production system should route by document evidence and policy instead of choosing one method for every file. - [Document parsing security starts before the parser runs](https://docparse.genedai.me/blog/document-parsing-security): A document parsing service accepts untrusted files and often reaches storage, queues, containers, models, URLs, and webhooks. Secure it with strict input bounds, isolated processing, server-owned routing, scoped credentials, outbound host policy, durable cost controls, minimal logs, tenant isolation, and explicit retention and purge. - [DocIR: why document pipelines need a structured intermediate model](https://docparse.genedai.me/blog/docir-structured-document-model): A structured intermediate model decouples downstream applications from parser-specific payloads. DocIR represents source metadata, pages, ordered typed blocks, optional geometry, assets, parser provenance, and warnings under a versioned schema while preserving Markdown as a readable projection. ## Machine-readable resources - [Full content corpus](https://docparse.genedai.me/llms-full.txt): Markdown text for every public content page, including related links and official sources. - [AI content index](https://docparse.genedai.me/ai-index.json): Typed registry with intent, entities, relationships, sources, and canonical URLs. - [XML sitemap](https://docparse.genedai.me/sitemap.xml): Canonical indexable URLs. - [RSS feed](https://docparse.genedai.me/feed.xml): Engineering guides, comparisons, shortlists, and alternative pages. - [Service health](https://docparse.genedai.me/health): Runtime status and parser version. ## Important boundaries - The free workspace includes 20 text-based documents and 200 MB of deterministic parsing; model-backed image and layout modes require an approved workspace. - Comparison pages describe public documentation and do not assert universal benchmark superiority. - Best-tool pages are architecture shortlists, not paid placements or universal rankings. - llms.txt is published as a discovery aid; canonical HTML pages remain the authoritative product content.