opendataloader-pdf

A PDF parser that converts documents into structured data for AI use, such as Markdown and JSON. It also helps automate PDF accessibility.

Share on XLicense: MPL-2.0

Overview

OpenDataLoader PDF is an open-source PDF parser that turns documents into data ready for AI use, and also helps automate PDF accessibility. It extracts Markdown, JSON with bounding boxes, and HTML. A deterministic local mode handles most files, while a hybrid mode adds AI for complex pages, scanned files and OCR. It installs with pip and is used for tasks such as RAG pipelines.

Key features

  • Outputs Markdown, JSON with bounding boxes and HTML
  • Deterministic local mode for standard documents
  • Hybrid AI mode for complex pages, tables and formulas
  • Built-in OCR in hybrid mode for scanned PDFs
  • Automates PDF accessibility work

Best for

Teams preparing PDFs for RAG or other AI pipelines, or handling accessibility requirements. Scanned files and complex layouts need the hybrid mode.

Upstream
opendataloader-project/opendataloader-pdf
Fork on GitHub
Guo-astro/opendataloader-pdf
Upstream stars
29k
Category
Documents, media and content
License
MPL-2.0
Forked
2026-02-26
Sync status
Diverged from upstream