PageIndex searches long PDFs with no vector database

Vectorless RAG skips embeddings completely. PageIndex turns a long PDF into a table-of-contents tree with page ranges and section summaries. A model then searches that tree the way a person flips to a chapter. The project reports 98.7 percent on FinanceBench, a benchmark of questions over financial filings.
Key Takeaways
- PageIndex pulls answers from long documents with no embeddings and no vector database.
- It builds a table of contents with page numbers, then the model walks it.
- Every answer points at a real section and page you can open and check.
- Sections follow the document’s own structure instead of fixed-size chunks.
- Each query spends model calls, so you swap a database bill for a token bill.
What vectorless RAG actually means
You split a document into chunks. You embed each chunk, store the vectors, then embed the question and return the nearest chunks. Nearly every RAG stack in production works this way, even the ones with an agent in the loop .
The vectorless pipeline replaces all of that. It first builds an index of the document’s structure. A model then reads that index and picks the section that answers the question, with no similarity math anywhere in the loop.

The claim underneath it all is that similarity and relevance are different things. Nearest-neighbour search finds text that reads like the question, and that text often has no answer in it.
The PageIndex File System post names two failure modes. Something relevant gets missed because it’s worded differently. Or something close comes back that’s beside the point. In legal and financial writing, two paragraphs can look almost the same to an embedding model while saying opposite things about who is liable.
Going vectorless deletes a familiar list of chores. Chunk size tuning, overlap tuning, and embedding model choice all go away. So does the re-embedding job after a model upgrade, and the vector store itself. In their place you get one tree-building step per document, and a model call on every retrieval.
How the tree is built and searched
Step one builds a tree index. It looks like a table of contents, but it’s written for a model rather than a reader. Each node carries a title, a node id, start and end page indices, a summary, and any nested children.
{
"title": "Financial Stability",
"node_id": "0006",
"start_index": 21,
"end_index": 22,
"summary": "The Federal Reserve ...",
"nodes": [
{
"title": "Monitoring Financial Vulnerabilities",
"node_id": "0007",
"start_index": 22,
"end_index": 28,
"summary": "The Federal Reserve's monitoring ..."
}
]
}The page indices make every answer traceable back to a page a person can open and check, which is something an embedding score cannot do.
Step two is the search itself. The model reads node summaries, decides whether a branch is worth opening for this particular question, and descends into the promising one. It repeats until it reaches the section that holds the answer.
Every one of those calls has to come back as a node id that really exists in the tree. Pinning a model to a fixed set of legal answers is exactly what constrained generation does.
The project compares this to game-playing search, and the comparison holds up. It’s a search over a structured space, with a judgment call at every node. That judgment can fold in chat history and the user’s role, because nothing has been squeezed into a fixed-length vector beforehand.
Sections come from the document’s own structure. A table stays attached to its heading and a clause stays inside its section, never sliced mid-thought at a random token boundary.
A file-level tree layer scales the idea past one document, because a folder hierarchy is already a tree. The same policy that walks a 100-page report can walk a folder tree, then drop into one file.

What the FinanceBench score covers
The headline number is 98.7 percent accuracy on FinanceBench, a benchmark built on financial filings. The full results sit in the Mafin2.5-FinanceBench repository . The score held steady with both GPT-4o and DeepSeek v3 underneath. That points at the retrieval method doing the work, whichever model sits behind it.

FinanceBench asks questions over long, formal filings. That is close to the worst case for chunk-and-embed retrieval, and close to the best case here. Filings have real structure, steady headings, and answers that live in a named section.
The project’s own team ran the benchmark, so treat it as a strong signal rather than an independent result. It also says nothing about loose prose, documents with no usable headings, or answers scattered across a dozen places. Read it as the narrower claim it is: structure-aware retrieval wins on structured documents, and that is still worth having.
What it costs to run
Dropping the vector database swaps one bill for another. Indexing is model work, paid once per document. Cost there tracks document length, and query volume never touches it. Retrieval is model work too, and you pay it on every query.
Tree navigation is several model calls in a row. That’s slower per query than a similarity lookup by a wide margin. PageIndex co-author Mingtian Zhang put the trade-off plainly in the Show HN thread.
When the tree is large, however, this approach may be slower than the vector-based method since it prioritizes accuracy. If you prioritize speed over accuracy, then I guess you should use Vector DB.
Low query volume over long, well-structured documents favours the tree. High query volume over short documents pushes the other way. Contract review and filing analysis land on the tree side of that line, and a public search box lands on the vector side.
| Approach | Index cost | Per-query cost | Traceability |
|---|---|---|---|
| PageIndex | Model calls per document | Several model calls | Page and section references |
| Vector RAG | Embedding pass per chunk | Millisecond lookup | Chunk ids only |
| Agentic RAG over a vector store | Embedding pass per chunk | Lookup plus model calls | Chunk ids only |
In return there’s no vector database to operate and no embedding pipeline to maintain. Nothing needs re-embedding when you swap models.
The project names one limit on itself. The open-source package uses plain PDF parsing, and messy layouts are where the hosted service’s better OCR earns its keep.
Three deployment paths exist: self-host the open repository, reach the cloud service through MCP or an API, or take a private enterprise install. The repository is MIT licensed and written in Python. It sits on a 0.3 development release line with roughly 143 open issues. Treat the self-hosted path as capable but young.
How to build a PageIndex tree from a PDF
Check the structure before you wire it into retrieval.
Install the dependencies
Clone the repository and run pip3 install --upgrade -r requirements.txt.
Add a model key
Create a .env file with your provider key. Many providers work through LiteLLM
, so you’re not tied to one vendor.
Generate the tree
Run python3 run_pageindex.py --pdf_path /path/to/your/document.pdf.
Tune the table-of-contents detection
The --toc-check-pages flag sets how many pages get scanned for a contents page. The --model flag picks the model doing the work. Defaults are 20 pages and GPT-4o.
Read the output before trusting it
The result is JSON with a title, node id, start and end page indices, a summary, and nested child nodes. A bad tree makes everything after it wrong. Open the file and check that the sections match the document.
Try the vectorless retrieval notebook
The repository ships a minimal cookbook that runs retrieval over the tree you just built.
Go agentic if you need multi-step questions
The agentic example wires the tree into an agent loop. The model can then navigate, read, and refine before it answers.
Know when to stop self-hosting
The open-source path uses plain PDF parsing. If your filings arrive as dense scans, move to the hosted OCR instead of tuning flags.
Botmonster Tech