Twitter/X

OpenDataLoader is an open-source PDF-to-Markdown extractor that @DataChaz…

Brief

OpenDataLoader is an open-source PDF-to-Markdown extractor that @DataChaz promoted on 2026-04-08, claiming 100 pages/second throughput. The tweet says it runs "flawlessly on CPU" and can decode tables, complex layouts and nested structures; the author points followers to the repo in the thread and emphasizes it's 100% free and open-source.

Source evidence

title: @DataChaz: 🚨 Extracting data from PDFs just got solved.

Someone open-sourced a tool that turns PDFs into Markd...
author: @DataChaz
contenttype: tweet
publication: Twitter/X
published: 2026-04-08T08:37:46+00:00
source
url: https://x.com/DataChaz/status/2041797637909967236

word_count: 56

🚨 Extracting data from PDFs just got solved.

Someone open-sourced a tool that turns PDFs into Markdown at 100 pages a second 🤯

It’s called OpenDataLoader.

It runs flawlessly on CPU and decodes tables, complex layouts, and nested structures like an absolute pro.

Best part?

100% free and open-source.

Grab the repo link in the 🧵↓