Code GenerationOpen_source
AnyDoc logo

AnyDoc

Convert 14 document formats to clean Markdown in milliseconds

Overview

About AnyDoc

AnyDoc converts documents into clean GitHub-Flavored Markdown. Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF all go in, and consistent Markdown comes out, which is the format models actually read well. It is written in Rust by the Firecrawl team and released under MIT.

Two design decisions make it worth knowing about. First, there is no machine learning in it at all: it is a parser, so median conversion is under five milliseconds and the output is deterministic rather than a different interpretation each run. That is a sharp contrast with document handling that routes everything through a vision model, which is slower, costs money per page and produces something slightly different every time. Second, every format parses into the same document model and renders through the same serializer, so headings, nested lists, merged table cells and footnotes come out identically whether the source was a .doc from 2003 or yesterday's .pptx. The project claims that of seven converters benchmarked across 100 documents it was the only one to handle all fourteen formats, which is their own benchmark and should be read as such. It also ships as an agent skill, so a coding agent can read any document it encounters. PDF support covers text-based PDFs through a companion tool, which means scanned images are still not solved here.

Pricing

No pricing page found

No free tier

Free and MIT licensed. Available on crates.io, npm and PyPI, as a CLI, as WebAssembly, and as an agent skill.

Checked 2026-09-18

Details

GitHub Stars 18,175
Forks 1,050
Data from: GitHub • Website•Updated: Aug 24, 2026
documentationdocreadmeopen-source