GLM-OCR

A multimodal OCR model for reading complex documents and turning images of text into usable text. It targets accurate and fast document understanding.

Share on XLicense: Apache-2.0

Overview

GLM-OCR is a multimodal OCR model for understanding complex documents. It combines a visual encoder, a small cross-modal connector and a GLM-0.5B language decoder, and uses a two-stage pipeline of layout analysis followed by parallel recognition. It turns document images into text, including formulas and tables, and can be deployed with vLLM, SGLang or Ollama.

Key features

  • Two-stage layout analysis and parallel recognition
  • Handles formulas, tables and information extraction
  • Only 0.9B parameters
  • Deployable with vLLM, SGLang and Ollama
  • Aimed at complex tables, code-heavy documents and seals

Best for

Teams that need to read complex real-world documents into text and want a compact model they can host themselves.

Upstream
zai-org/GLM-OCR
Fork on GitHub
Guo-astro/GLM-OCR
Upstream stars
7.5k
Category
Documents, media and content
Language
Python
License
Apache-2.0
Forked
2026-03-17
Sync status
In syncLast synced 2026-09-29