PDF bytes in.
Markdown out.
pdfmd parses the object graph, inflates streams, decodes fonts, and interprets content operators itself. Then it recovers columns, tables, headings, and lists. 17 pages in 4.4 ms.
One crate, no build scripts. Install the CLI from a checkout, or pip install ./python for the ctypes bindings.
cargo install --path .
[package] name = "pdfmd" version = "0.3.0" rust-version = "1.70" [dependencies]
BT /F1 24 Tf 72 720 Td (Layout-Aware Extraction) Tj ET BT /F1 14 Tf 72 684 Td (1 Two columns, one pass) Tj ET BT /F1 10 Tf 72 660 Td (Reading order comes from) Tj 0 -12 Td (glyph positions.) Tj ET BT /F1 10 Tf 72 620 Td (stage) Tj 80 0 Td (ms) Tj ET BT /F1 10 Tf 72 606 Td (inflate) Tj 80 0 Td (0.9) Tj ET
- 24 pt
- # Layout-Aware Extraction
- 14 pt
- ## Two columns, one pass
- body
- Reading order comes from glyph positions.
- aligned
- | stage | ms | | --- | --- | | inflate | 0.9 |
One converter.
Three ways in.
pdfmd input.pdf # to stdout
pdfmd input.pdf -o out.md
pdfmd input.pdf --page-breaks
pdfmd input.pdf \
--extract-images figs -o out.md
cat input.pdf | pdfmd -
pdfmd https://example.com/x.pdf
URLs are fetched with curl on PATH. Images pass through as JPEG or JPEG 2000, or are re-encoded as PNG.
use pdfmd::{
convert_pdf_to_markdown,
ConvertOptions,
};
let result = convert_pdf_to_markdown(
&pdf_bytes,
&ConvertOptions::default(),
)?;
print!("{}", result.markdown);
One error type. result.images is empty unless image_dir is set.
import pdfmd
result = pdfmd.convert_file(
"paper.pdf",
image_dir="figs",
)
print(result.markdown)
# one thread per core, input order
batch = pdfmd.convert_many(
["a.pdf", "b.pdf"],
)
ctypes over the C ABI. No PyO3, no maturin, no runtime packages.
4.4 milliseconds.
Seventeen pages.
- mean
- 4.4 ms ± 0.3
- min
- 3.8 ms
- throughput
- ~240 MB/s
- pages/sec
- ~3,900
- binary
- ~700 KB
Pages extract in parallel on std::thread::scope. The tokenizer and DEFLATE decoder borrow straight from the source bytes.
| pages | published | pdfmd | speedup |
|---|---|---|---|
| 1 | 2.0 ms | 0.012 ms | |
| 24 | 41.0 ms | 0.118 ms | |
| 60 | 123.0 ms | 0.201 ms | |
| 457 | 777.0 ms | 0.933 ms |
Throughput targets against other libraries' published numbers, not a claim on their exact corpora or hardware.
No crates.
Every layer in-house.
markdownWhat does it recover?
Layout-aware and still best-effort. Reading order comes from glyph positions, not the order bytes appear in the stream.
- Columns
- Multi-column pages read left to right, then top to bottom in each column.
- Tables
- GFM tables from ruled path grids and from aligned borderless columns.
- Headings
- From tagged-PDF roles, font size, bold, numbered sections, and names like Abstract.
- Text
- Lists, bold and italic from the font name, monospace as fenced code, hyphenated breaks rejoined.
- Margins
- Repeating running headers and footers are stripped.
src/What is inside?
bytes → xref → objects → inflate → fonts → content ops → layout → heuristics → markdown.
- pdf/
- xref, objects, streams, DEFLATE, page tree.
- extract/
- Fonts, CMaps, content operators, layout, tagged PDF, images.
- heuristics/
- Columns, tables, headings, lists, emphasis.
- lib.rs
- convert_pdf_to_markdown.
- ffi.rs
- The C ABI behind the python/ bindings.
ffiCan I call it from another language?
The crate builds a cdylib. Anything that can call C can use it. Buffers are pointer and length pairs with no implied NUL, and every result is freed once.
PdfmdResult *r = pdfmd_convert(bytes, len, false, NULL);
if (r->error) fwrite(r->error, 1, r->error_len, stderr);
else fwrite(r->markdown, 1, r->markdown_len, stdout);
pdfmd_result_free(r);
limitsWhat won't it do?
It targets academic and prose documents, not forms or invoices.
- Layout
- Spanned cells, nested tables, and complex math stay best-effort reflow.
- Fonts
- Fonts without a ToUnicode CMap, a standard encoding, or /Differences drop glyphs.
- By design
- Encrypted PDFs and LZWDecode streams return an error.