GitHub - VectifyAI/PageIndex: π PageIndex: Document Index for Vectorlessβ¦
Reasoning-based RAG β¦ No Vector DB, No Chunking β¦ Context-Aware Retrieval β¦ Reads Like a Human
Spots Reasoning-based RAG β¦ No Vector DB, No Chunking β¦ Context-Aware Retrieval β¦ Reads Like a Human

[Aug '26] π₯ PageIndex SDK: pip install -U pageindex now ships local mode: index, retrieve, and chat entirely on your machine with your own LLM key, or point the same client at PageIndex Cloud with an API key.
[Aug '26] β‘ PageIndex Flash: fast tree index generation for text-based PDFs, now the default indexing method in PageIndex SDK local mode.
Scale PageIndex to Millions of Documents: PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document. PageIndex App: a human-like document analysis agent for long professional documents.
Are you frustrated with vector database retrieval accuracy for long and complex documents? Vector-based RAG retrieves by semantic similarity. But similarity β relevance β what retrieval actually needs is relevance, and relevance requires reasoning. On professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search misses what is relevant but not similar, and returns what is similar but not relevant.
Inspired by AlphaGo, PageIndex replaces the vector index with a hierarchical tree index and lets an LLM reason its way through it, the way a human expert turns to and reads the right section of a long report. Retrieval happens in two steps: Index: generate a tree-structure index for each document Retrieve: agentically search that tree with LLM reasoning
PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, with no vector DBs or chunking.
It is ideal for financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any other long, complex professional document.
index=: a basic model is sufficient. The tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it, which a basic model does well.
chat=: use the best model you can afford. The chat model searches the tree to retrieve information. See Query cost and accuracy. Configure other models, streaming, multi-document search, citations, and more. Drop PageIndex tools into the OpenAI Agents SDK, the Claude Agent SDK, or any other framework. Local indexing cost and time
Reasoning-based RAG β¦ No Vector DB, No Chunking β¦ Context-Aware Retrieval β¦ Reads Like a Human
