Large scanned PDFs through the OCR queue

A Python script that sends whole PDFs to the flatmark queue, polls each job and saves the Markdown. Docling and OCR run on flatmark's servers, not on your 8 GB machine.

Last updated

Operations submit_conversion_job, get_job, get_conversion_result
Platform Code (Python, flatmark SDK from PyPI)
Limits 25 MB, 200 pages and about 2 minutes of conversion per file

What it does

Docling on a machine with 8 GB of RAM can fail with out of memory errors on large PDFs. Splitting the file into 2-page batches helps, but some batches still fail, and the same limit follows you to a small cloud instance.

This script sends each whole PDF to the queue instead. The queue runs Docling with OCR and table recognition, and your machine only uploads files and downloads Markdown.

Steps

  1. Download convert_queue.py from the template folder below.
  2. Set FLATMARK_API_KEY and run uv run convert_queue.py scans/*.pdf --json.
  3. For each PDF, the script calls submit_conversion_job, then polls get_job after 5 seconds, doubling the wait up to 30 seconds. On a rate_limited answer it waits as long as Retry-After says.
  4. When the job has succeeded, get_conversion_result downloads the Markdown to <name>.md, and with --json the DoclingDocument to <name>.json.

A failed job prints its error, and its credits are refunded. Polls and downloads cost nothing.

Get the template

The script and its README are on GitHub: templates/large-scanned-pdfs-ocr-queue.

Get your API key

The queue needs a key. Create an API key and export it as FLATMARK_API_KEY.

FAQ

Do I still need to split my PDFs into small batches?
Only above the queue's limits: 25 MB, 200 pages or about 2 minutes of conversion per file. Below them, send the whole PDF in one job.
Can I get a webhook instead of polling?
Yes. Pass a public webhook_url with the file, and flatmark POSTs the job's final state there. The script polls because it needs no public endpoint.