Documents to a RAG vector store
A Python script that converts a folder of PDF, Word, PowerPoint, Excel and HTML files to Markdown, splits it at headings and hands each chunk to your vector store.
Last updated
| Operations | convert_document |
|---|---|
| Platform | Code (Python, flatmark SDK from PyPI) |
| Files | PDF, DOCX, PPTX, XLSX, HTML and plain text, up to 8 MB each |
What it does
You want to ask questions about a large set of documents with a local model. Ingestion is mostly data preparation: getting every file into one searchable format and cutting it into usable chunks.
The script converts each file with one direct call, splits the Markdown at headings and keeps each chunk's heading path as context. Your embedding model and vector store take the chunks from there.
Steps
- Download
ingest.pyfrom the template folder below. - Set
FLATMARK_API_KEYand runuv run ingest.py corpus/. uv installs theflatmarkSDK for the script. - Read
chunks.jsonl: one line per chunk, withsource,headingsandtext. - Replace the body of
index()with your embedding call and vector store upsert. Nothing else in the script depends on a specific database.
Scans and files over 8 MB are listed on stderr. Send those to the queue with the large scanned PDFs template.
Get the template
The script and its README are on GitHub: templates/documents-to-rag-vector-store.
Get your API key
Create an API key and export it as FLATMARK_API_KEY.
FAQ
- Which vector store does the template use?
- None by default. It writes the chunks to a JSONL file. Replace the index() function to embed each chunk and store it in the database you already run, local or hosted.
- Why are my PDF chunks not split at headings?
- The direct call returns PDFs as plain text without headings, so the script splits them by size only. The queue returns PDF headings and tables.