zpdf

A PDF text extraction library written in Zig that focuses on speed and low memory copying. It is at an early alpha stage and is used to pull text out of PDF files.

Share on XLicense: CC0-1.0

Overview

zpdf is a PDF text extraction library written in Zig. It reads files through memory mapping and avoids copying data where possible, and it extracts text in a streaming way with arena allocation. It handles several compression filters and font encodings, parses cross-reference tables and streams, and can read the structure tree of tagged PDFs. The README labels it an early alpha version.

Key features

  • Memory-mapped, zero-copy reading where possible
  • Supports FlateDecode, ASCII85, ASCIIHex, LZW and RunLength filters
  • Handles WinAnsi, MacRoman and ToUnicode font encodings
  • Structure tree extraction and Markdown export for structured PDFs
  • Strict or permissive error handling

Best for

Zig developers who need fast text extraction from PDFs. It is at an early alpha stage, so expect gaps.

Upstream
Lulzx/zpdf
Fork on GitHub
Guo-astro/zpdf
Upstream stars
922
Category
Documents, media and content
Language
Zig
License
CC0-1.0
Forked
2025-12-31
Sync status
In syncLast synced 2026-09-29