Large scanned PDFs through the OCR queue
A Python script that sends whole PDFs to the flatmark queue, polls each job and saves the Markdown. Docling and OCR run on flatmark's servers, not on your 8 GB machine.
Last updated
| Operations | submit_conversion_job, get_job, get_conversion_result |
|---|---|
| Platform | Code (Python, flatmark SDK from PyPI) |
| Limits | 25 MB, 200 pages and about 2 minutes of conversion per file |
What it does
Docling on a machine with 8 GB of RAM can fail with out of memory errors on large PDFs. Splitting the file into 2-page batches helps, but some batches still fail, and the same limit follows you to a small cloud instance.
This script sends each whole PDF to the queue instead. The queue runs Docling with OCR and table recognition, and your machine only uploads files and downloads Markdown.
Steps
- Download
convert_queue.pyfrom the template folder below. - Set
FLATMARK_API_KEYand runuv run convert_queue.py scans/*.pdf --json. - For each PDF, the script calls
submit_conversion_job, then pollsget_jobafter 5 seconds, doubling the wait up to 30 seconds. On arate_limitedanswer it waits as long asRetry-Aftersays. - When the job has
succeeded,get_conversion_resultdownloads the Markdown to<name>.md, and with--jsonthe DoclingDocument to<name>.json.
A failed job prints its error, and its credits are refunded. Polls and downloads cost nothing.
Get the template
The script and its README are on GitHub: templates/large-scanned-pdfs-ocr-queue.
Get your API key
The queue needs a key. Create an API key and export it as FLATMARK_API_KEY.
FAQ
- Do I still need to split my PDFs into small batches?
- Only above the queue's limits: 25 MB, 200 pages or about 2 minutes of conversion per file. Below them, send the whole PDF in one job.
- Can I get a webhook instead of polling?
- Yes. Pass a public webhook_url with the file, and flatmark POSTs the job's final state there. The script polls because it needs no public endpoint.