The challenge
A private equity deal comes with a data room full of PDFs, spreadsheets and call transcripts. Analysts need answers they can defend in an investment committee, so every number has to trace back to the page it came from. A generic chatbot can't promise that. It blends sources, and sometimes it makes things up.
What we did
Pages, not chunks
We made the page the unit of retrieval. Every page goes through OCR on its own with Google Vision, Gemini handles charts and graphs, and each page keeps its file name and page number all the way through. Spreadsheets are parsed sheet by sheet. Every page moves through a state machine with retries, so one broken file never stalls a whole data room.
Search that finds the right page
Retrieval blends three methods: vector search, BM25 keyword search and HyDE, where the model drafts a likely answer and searches with that. The results are merged with reciprocal rank fusion, and each question fans out into several searches.
Checking before answering
Before anything is written, a separate model pass scores every retrieved page against the question and drops the ones that don't hold up. The answer is written only from pages that passed, and the citation is just metadata carried along, so it can't cite a page it never saw. When the documents don't cover something, the answer says so.
Agents, evals and security
The agent workflows run on LangGraph and LangChain, with LangSmith for tracing, test datasets and regression runs. The platform also pulls out financial metrics like revenue, EBITDA and margins and reconciles them across documents. Sachin led the platform as CTO, and we set up the Google Cloud infrastructure and led the security work behind SOC 2.
The outcome
Emblem has indexed more than 10 million pages and scored 100% on the Vectara RAG accuracy benchmark across 3,000 queries. It also scored 95.5% on SpreadsheetBench, verified independently by the benchmark team, and completed SOC 2 Type II certification.
Tech stack
- TypeScript, Node.js, Nest.js, Next.js, React
- Python, LangGraph, LangChain, LangSmith
- PostgreSQL, Prisma, TurboPuffer, Redis
- Google Cloud, Gemini, Google Vision, Docker, Grafana
