pdfmd
PDF to Markdown, in Rustv0.3.0 · MIT

PDF bytes in.
Markdown out.

The whole stack

pdfmd parses the object graph, inflates streams, decodes fonts, and interprets content operators itself. Then it recovers columns, tables, headings, and lists. 17 pages in 4.4 ms.

Install

One crate, no build scripts. Install the CLI from a checkout, or pip install ./python for the ctypes bindings.

cargo install --path .
Zero dependencies
[package]
name         = "pdfmd"
version      = "0.3.0"
rust-version = "1.70"

[dependencies]
runtime crates →0
page.pdf → page.md4 blocks · real output
Content stream
BT /F1 24 Tf 72 720 Td
(Layout-Aware Extraction) Tj ET
BT /F1 14 Tf 72 684 Td
(1  Two columns, one pass) Tj ET
BT /F1 10 Tf 72 660 Td
(Reading order comes from) Tj
0 -12 Td (glyph positions.) Tj ET
BT /F1 10 Tf 72 620 Td (stage) Tj 80 0 Td (ms) Tj ET
BT /F1 10 Tf 72 606 Td (inflate) Tj 80 0 Td (0.9) Tj ET
Markdown
1
24 pt
# Layout-Aware Extraction
2
14 pt
## Two columns, one pass
3
body
Reading order comes from glyph positions.
4
aligned
| stage | ms | | --- | --- | | inflate | 0.9 |
4 / 4 blockswritten
Font size makes headings. Alignment makes tables. Position makes reading order.Tf · Td · Tj
Usageshell · rust · python

One converter.
Three ways in.

CLI — files, stdin, or a URL
pdfmd input.pdf              # to stdout
pdfmd input.pdf -o out.md
pdfmd input.pdf --page-breaks
pdfmd input.pdf \
  --extract-images figs -o out.md
cat input.pdf | pdfmd -
pdfmd https://example.com/x.pdf

URLs are fetched with curl on PATH. Images pass through as JPEG or JPEG 2000, or are re-encoded as PNG.

Rust — main.rs
use pdfmd::{
    convert_pdf_to_markdown,
    ConvertOptions,
};

let result = convert_pdf_to_markdown(
    &pdf_bytes,
    &ConvertOptions::default(),
)?;
print!("{}", result.markdown);

One error type. result.images is empty unless image_dir is set.

Python — convert.py
import pdfmd

result = pdfmd.convert_file(
    "paper.pdf",
    image_dir="figs",
)
print(result.markdown)

# one thread per core, input order
batch = pdfmd.convert_many(
    ["a.pdf", "b.pdf"],
)

ctypes over the C ABI. No PyO3, no maturin, no runtime packages.

Speedhyperfine · Apple Silicon · release

4.4 milliseconds.
Seventeen pages.

CLI — 1.05 MB arXiv paper
1
mean
4.4 ms ± 0.3
2
min
3.8 ms
3
throughput
~240 MB/s
4
pages/sec
~3,900
5
binary
~700 KB

Pages extract in parallel on std::thread::scope. The tokenizer and DEFLATE decoder borrow straight from the source bytes.

Against published library timings
pagespublishedpdfmdspeedup
12.0 ms0.012 ms167×
2441.0 ms0.118 ms347×
60123.0 ms0.201 ms612×
457777.0 ms0.933 ms833×

Throughput targets against other libraries' published numbers, not a claim on their exact corpora or hardware.

Referencelayout · internals · limits

No crates.
Every layer in-house.

markdownWhat does it recover?

Layout-aware and still best-effort. Reading order comes from glyph positions, not the order bytes appear in the stream.

Columns
Multi-column pages read left to right, then top to bottom in each column.
Tables
GFM tables from ruled path grids and from aligned borderless columns.
Headings
From tagged-PDF roles, font size, bold, numbered sections, and names like Abstract.
Text
Lists, bold and italic from the font name, monospace as fenced code, hyphenated breaks rejoined.
Margins
Repeating running headers and footers are stripped.
src/What is inside?

bytes → xref → objects → inflate → fonts → content ops → layout → heuristics → markdown.

pdf/
xref, objects, streams, DEFLATE, page tree.
extract/
Fonts, CMaps, content operators, layout, tagged PDF, images.
heuristics/
Columns, tables, headings, lists, emphasis.
lib.rs
convert_pdf_to_markdown.
ffi.rs
The C ABI behind the python/ bindings.
ffiCan I call it from another language?

The crate builds a cdylib. Anything that can call C can use it. Buffers are pointer and length pairs with no implied NUL, and every result is freed once.

PdfmdResult *r = pdfmd_convert(bytes, len, false, NULL);
if (r->error) fwrite(r->error, 1, r->error_len, stderr);
else          fwrite(r->markdown, 1, r->markdown_len, stdout);
pdfmd_result_free(r);
limitsWhat won't it do?

It targets academic and prose documents, not forms or invoices.

Layout
Spanned cells, nested tables, and complex math stay best-effort reflow.
Fonts
Fonts without a ToUnicode CMap, a standard encoding, or /Differences drop glyphs.
By design
Encrypted PDFs and LZWDecode streams return an error.