First question
What should AI find?
Document tools differ by ingestion, indexing, retrieval, citations, and how much of the source material they bring into context.
AI Resources
A focused map for retrieval, RAG, document context, knowledge-base, and indexing tools that help AI work with files and source material.
Use this page to narrow what to explore next, then open the original source before relying on setup, privacy, citations, scale, or cost assumptions.
Questions to check
These checks frame the source-linked Resources below. They do not rank products or cover every option.
First question
Document tools differ by ingestion, indexing, retrieval, citations, and how much of the source material they bring into context.
Data path
Look at what data is indexed, where it is stored, which model providers are involved, and who can access the knowledge base.
Trust habit
A useful document workflow should make it easier to return to the original file, citation, note, or record when an answer matters.
Coverage and freshness
These groups are selective starting points, not a complete directory. The date reflects the newest included Resource’s LifeHubber added date, not a recheck of every linked source. Check the original source for current setup, terms, limits, privacy, access, costs, and behaviour.
Fresh in this topic
Recently added Resources from the groups below.
Documents and knowledge bases
Use this group when the job is searching documents, indexing files, building a knowledge base, or giving an agent source material it can retrieve.
infiniflow/ragflow
Document ingestion, chunking, retrieval, citations, and knowledge workflows cover the full RAG pipeline rather than only vector storage.
VectifyAI/PageIndex
A vectorless tree index with traceable reasoning-based retrieval puts document structure and explainable search paths ahead of embedding similarity.
langgenius/dify
Visual RAG pipelines connect knowledge retrieval, agent tools, model providers, and application flows in one configurable workflow layer.
onyx-dot-app/onyx
A self-hostable interface combines RAG with web search, code execution, file creation, and deep research when retrieval must sit inside a broader workbench.
ItzCrazyKns/Vane
Cited search-style answers, file uploads, SearxNG web search, and a Docker route bring private documents and web results into one answering workflow.
nashsu/llm_wiki
Document-to-wiki conversion, source traceability, graph search, and a desktop app show a maintained-knowledge-page route beyond question-by-question retrieval.
colbymchenry/codegraph
Symbol relationships, call graphs, framework-aware routes, and auto-sync give coding agents structural codebase context instead of plain text chunks.
Lum1104/Understand-Anything
A knowledge graph, diff-impact views, tours, and plugins across coding environments show code understanding shared across tools and workflows.
yichuan-w/LEANN
A local vector database with a lower-storage design puts index size and local operation ahead of application features in the comparison.
cocoindex-io/cocoindex
Delta-only processing, lineage, and multiple index targets keep knowledge pipelines current without rebuilding unchanged data.
1Panel-dev/MaxKB
Document and web collection, hybrid retrieval, embedded Q&A, and offline Docker paths span knowledge ingestion, retrieval method, and delivery inside another site.
Also in AI
Keep the thread going with AI Guides for decision habits for messy AI choices, AI Access for free and low-cost ways to compare AI model access, AI Ballot for a clearer view of what readers are leaning toward.