megaparse

A file parser that converts PDFs, Word and PowerPoint files into a format suited for LLMs while keeping information intact. It is meant for feeding documents to AI models.

Share on XLicense: Apache-2.0

Overview

MegaParse is a Python file parser that converts documents into a format suited for LLM ingestion while trying to lose no information. It handles PDF, PowerPoint and Word files, and also lists support for text, Excel and CSV. Content it covers includes tables, table of contents, headers, footers and images. It is open source and installs with pip on Python 3.11 or newer.

Key features

  • Parses PDF, PowerPoint and Word documents
  • Also lists Excel, CSV and plain text support
  • Handles tables, headers, footers and images
  • Aims for no information loss during parsing
  • Installs with pip on Python 3.11 or newer

Best for

Developers feeding business documents into LLM or RAG pipelines who want structure preserved. It requires Python 3.11 or newer.

Upstream
The-Vibe-Company/megaparse
Fork on GitHub
Guo-astro/megaparse
Upstream stars
7.4k
Category
Documents, media and content
License
Apache-2.0
Forked
2026-02-22
Sync status
In syncLast synced 2026-09-29