Skip to content
Ayush Gupta

RAG & document intelligence

Retrieval over real documents (scans, tables, long filings) with citations that point at the exact page.

01 · RAG & document intelligence

Answers you can trace to the source

Most RAG breaks on real documents. Tables get flattened, scans lose their layout, and the answer comes back with a page number at best. After the first wrong answer, people stop trusting it.

Talk about this

Getting it right

  • Retrieval that keeps layout (tables, figures, stamps), not only extracted text
  • Citations people can check: the page and the region
  • Retrieval quality measured on real questions, not demo ones
  • Refusing when the evidence isn't there

What I build

  • Question answering over large document collections
  • Search across contracts, filings and reports
  • Field extraction from scanned documents, linked to the source
  • Knowledge assistants grounded in internal documentation

Typical stack

ColPali · pgvector · Neo4j · bge-m3 · cross-encoder rerankers · LangGraph · FastAPI · Gemini / Claude / OpenAI

Questions

Does this work with scanned PDFs and tables?

Yes. Atlas retrieves over page images rather than OCR text, so tables, figures and stamps survive retrieval, and the citation is drawn on the original page.

How does someone verify an answer?

Every claim carries a citation to the page and region it came from. A verifier re-checks each claim against its cited page before the answer is shown.

Can it run without sending documents to a third party?

Yes. Embedding and generation models can be self-hosted. The Atlas architecture notes document the move to a private GPU for confidential documents.

Working on something like this? Tell me about it.

Get in touch