Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise
Favicon of anydoc

anydoc

Free Listing

anydoc is an open-source Rust library that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF documents into clean GitHub-Flavored Markdown, with Node.js, Python, browser, WebAssembly, and CLI bindings.

Visit firecrawl/anydoc

Overview

anydoc is an open-source Rust library by Firecrawl that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF documents into GitHub-Flavored Markdown. It provides Node.js, Python, browser WebAssembly, and CLI bindings, and its browser demo converts files locally.

Key Features

  • Converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-Flavored Markdown.
  • Uses one shared document model and Markdown serializer so output structure stays consistent across formats.
  • Detects formats from file bytes rather than extensions, so mislabeled files still convert.
  • Runs as pure Rust with no ML models or external services and reports median conversion under 5 ms.
  • Ships bindings for Rust, Node.js, Python, browser WebAssembly, and CLI usage.
  • Preserves headings, lists, tables with merged cells, footnotes, speaker notes, and equations as LaTeX.

Use Cases

  • Converting mixed office document collections into one consistent, structured Markdown output for LLM pipelines.
  • Converting documents locally in the browser so files never leave the user's machine.
  • Letting agents read office documents through the shipped Agent Skill.
  • Extracting text-based PDFs locally without requiring an OCR service.

Getting Started

  • Install anydoc for Rust with cargo add anydoc, for Node.js with npm install @firecrawl/anydoc, for Python with pip install firecrawl-anydoc, or for the browser with npm install @firecrawl/anydoc-wasm.
  • Run the CLI with npx @firecrawl/anydoc report.docx, or use the same conversion API from a path or bytes.
  • Stop at the document model when embedded assets need to be kept.

Deployment & Requirements

  • Pure Rust, no ML models, and no external services for local conversion.
  • Browser WebAssembly demo downloads a few MB of WebAssembly once.
  • Text-based PDFs convert locally through pdf-inspector; scanned or image-only PDFs require opt-in hosted OCR through Firecrawl Parse.
  • Hosted OCR can be used without signup; setting FIRECRAWL_API_KEY provides higher limits.
  • The Rust crate has no OCR option and never makes network calls.

Before You Adopt

  • License: MIT. Review its terms before using, modifying, or distributing the project.
  • Scanned or image-only PDFs are not handled by local conversion and rely on Firecrawl Parse for OCR.
  • When hosted OCR is used, the entire document is sent off the machine because Parse has no page selection.

Comments

Sign in to leave a comment.

More like anydoc

Favicon of PaddleOCR

PaddleOCR

Free ListingStars: 90.2K

Global Leading OCR Toolkit & Document AI Engine

Document Parsing & OCR

PaddleOCR is a powerful, lightweight OCR and document parsing toolkit that turns PDFs and images into structured, LLM-ready data. It supports 100+ languages, JSON and Markdown output, and helps build RAG and Agentic applications.