PaddleOCR
PaddleOCR is a powerful, lightweight OCR and document parsing toolkit that turns PDFs and images into structured, LLM-ready data. It supports 100+ languages, JSON and Markdown output, and helps build RAG and Agentic applications.
Visit PaddlePaddle/PaddleOCROverview
PaddleOCR is an OCR toolkit and document AI engine that converts PDFs and images into structured, LLM-ready data. It provides high-accuracy text recognition, document parsing, and multilingual support for building intelligent RAG and Agentic applications.
Key Features
- Intelligent document parsing converts complex PDFs and images into Markdown or JSON, with structure-aware conversion and fine-grained coordinates.
- PaddleOCR-VL and PP-StructureV3 support text, formula, table, chart, seal, and ancient document recognition.
- Universal text recognition supports 100+ languages, with PP-OCRv6 covering 50 languages in a single model without model switching.
- PP-OCRv6 improves detection by 4.6% and recognition by 5.1% over PP-OCRv5, with 5.2× CPU inference speedup end-to-end.
- The developer ecosystem integrates with Dify, RAGFlow, Pathway, and Cherry Studio and supports building LLM data pipelines.
- One-click deployment supports NVIDIA GPU, Intel CPU, Kunlunxin XPU, and diverse AI accelerators.
Use Cases
- Building RAG pipelines that need accurate parsing of PDFs, images, and complex documents for LLM ingestion.
- Converting scanned documents, tables, formulas, and charts into Markdown or JSON for structured data workflows.
- Multilingual document processing across 100+ languages, including Chinese, English, Japanese, Korean, Arabic, and other scripts.
- Embedding OCR into applications through Python, C++, C#, Java, or browser-based inference.
Getting Started
- Try the online Experience Center and APIs on the official website without setup.
- For local deployment, follow the PP-OCR, PaddleOCR-VL, or PP-StructureV3 documentation based on your needs.
Deployment & Requirements
- Python 3.8–3.12; OS support for Linux, Windows, and macOS; hardware support for CPU, GPU, XPU, and NPU.
- One-click deployment supports NVIDIA GPU, Intel CPU, Kunlunxin XPU, and diverse AI accelerators.
- High-performance inference supports CUDA 12 and can use Paddle Inference or ONNX Runtime backends.
Before You Adopt
- License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
- Basic text recognition requires only minimal core dependencies, while document parsing and information extraction require optional dependencies.
- High-performance inference depends on CUDA 12 and can use Paddle Inference or ONNX Runtime backends.
- Full support is noted for PaddlePaddle framework versions 3.1.0 and 3.1.1.