Draft — not published. This page is noindex and not in the sitemap until it's approved.
Guide · LLM engineering
A RAG Evaluation Checklist Before You Promise Accuracy
Key takeaways
- A RAG demo over five PDFs proves nothing about the real corpus — evaluate on that, at full size.
- Build a question set with known answers before tuning anything; it's the yardstick for every change.
- Most "LLM got it wrong" complaints are retrieval failures — check what was retrieved first.
- Report where retrieval fails, not just an average score. Honest limits are the deliverable.
By Prasanna Patil · updated
Why demos lie
A RAG pipeline over a handful of clean PDFs works impressively. The same pipeline over the real corpus — thousands of scanned pages, duplicated documents, tables, ten-year-old policies that contradict last month's — fails in ways the demo never showed. The gap between the two is where RAG projects die, and it's measurable before anyone signs a contract.
The checklist
- Question set first. 50–100 real questions with known-correct answers, written from the corpus, including questions the corpus can't answer. Every tuning decision gets judged against this.
- Retrieval before generation. For each question, check the retrieved chunks. If the right chunk isn't in the top-k, no prompt will save you — fix chunking, embeddings or search before touching the model.
- Chunking matched to the documents. Fixed-size chunks shred tables and cross-references. Section-aware splitting usually beats cleverness.
- Refusals tested explicitly. The questions with no answer must produce "not in the documents", not a confident hallucination.
- Freshness checked. When two documents contradict (old policy vs new), decide the rule and test it.
What I report
Percentage of the question set answered correctly, percentage refused correctly, and — most useful — the categories where retrieval fails (document types, topics, formats). That's the report I bring to a RAG scoping conversation at Shivohini TechAI or anywhere else: where it works, where it doesn't, and what fixing the gaps would take.
Limitations of this guide
It covers single-corpus, English-language retrieval QA. Agentic RAG, multi-corpus routing, multilingual retrieval and fine-tuning trade-offs each need their own treatment.
Frequently asked questions
What accuracy is "good enough"?
Depends entirely on the cost of a wrong answer. Internal search tolerates misses a compliance answer can't. I report measured performance plus failure categories, and the business decides whether the number is acceptable — that decision shouldn't be made by the team building it.
Do I need a fine-tuned model?
Usually not. Most accuracy problems in RAG are retrieval problems, and fine-tuning doesn't fix retrieval. Exhaust chunking, embeddings and hybrid search first; revisit fine-tuning only with evidence the model's reasoning is the bottleneck.
How big should the evaluation set be?
Fifty well-chosen questions covering document types and known failure modes beats five hundred generated ones. Start small, expand when a failure category surprises you.
RAG demo that won't survive the real corpus?
Tell me about the documents and the questions users actually ask.