Documents to a RAG vector store

A Python script that converts a folder of PDF, Word, PowerPoint, Excel and HTML files to Markdown, splits it at headings and hands each chunk to your vector store.

Last updated

Operations convert_document
Platform Code (Python, flatmark SDK from PyPI)
Files PDF, DOCX, PPTX, XLSX, HTML and plain text, up to 8 MB each

What it does

You want to ask questions about a large set of documents with a local model. Ingestion is mostly data preparation: getting every file into one searchable format and cutting it into usable chunks.

The script converts each file with one direct call, splits the Markdown at headings and keeps each chunk's heading path as context. Your embedding model and vector store take the chunks from there.

Steps

  1. Download ingest.py from the template folder below.
  2. Set FLATMARK_API_KEY and run uv run ingest.py corpus/. uv installs the flatmark SDK for the script.
  3. Read chunks.jsonl: one line per chunk, with source, headings and text.
  4. Replace the body of index() with your embedding call and vector store upsert. Nothing else in the script depends on a specific database.

Scans and files over 8 MB are listed on stderr. Send those to the queue with the large scanned PDFs template.

Get the template

The script and its README are on GitHub: templates/documents-to-rag-vector-store.

Get your API key

Create an API key and export it as FLATMARK_API_KEY.

FAQ

Which vector store does the template use?
None by default. It writes the chunks to a JSONL file. Replace the index() function to embed each chunk and store it in the database you already run, local or hosted.
Why are my PDF chunks not split at headings?
The direct call returns PDFs as plain text without headings, so the script splits them by size only. The queue returns PDF headings and tables.