Skip to content

Draft — not published. This page is noindex and not in the sitemap until it's approved.

Guide · LLM engineering

A RAG Evaluation Checklist Before You Promise Accuracy

Key takeaways

  • A RAG demo over five PDFs proves nothing about the real corpus — evaluate on that, at full size.
  • Build a question set with known answers before tuning anything; it's the yardstick for every change.
  • Most "LLM got it wrong" complaints are retrieval failures — check what was retrieved first.
  • Report where retrieval fails, not just an average score. Honest limits are the deliverable.

By Prasanna Patil · updated

01

Why demos lie

A RAG pipeline over a handful of clean PDFs works impressively. The same pipeline over the real corpus — thousands of scanned pages, duplicated documents, tables, ten-year-old policies that contradict last month's — fails in ways the demo never showed. The gap between the two is where RAG projects die, and it's measurable before anyone signs a contract.

02

The checklist

  • Question set first. 50–100 real questions with known-correct answers, written from the corpus, including questions the corpus can't answer. Every tuning decision gets judged against this.
  • Retrieval before generation. For each question, check the retrieved chunks. If the right chunk isn't in the top-k, no prompt will save you — fix chunking, embeddings or search before touching the model.
  • Chunking matched to the documents. Fixed-size chunks shred tables and cross-references. Section-aware splitting usually beats cleverness.
  • Refusals tested explicitly. The questions with no answer must produce "not in the documents", not a confident hallucination.
  • Freshness checked. When two documents contradict (old policy vs new), decide the rule and test it.
03

What I report

Percentage of the question set answered correctly, percentage refused correctly, and — most useful — the categories where retrieval fails (document types, topics, formats). That's the report I bring to a RAG scoping conversation at Shivohini TechAI or anywhere else: where it works, where it doesn't, and what fixing the gaps would take.

04

Limitations of this guide

It covers single-corpus, English-language retrieval QA. Agentic RAG, multi-corpus routing, multilingual retrieval and fine-tuning trade-offs each need their own treatment.

05

Frequently asked questions

What accuracy is "good enough"?

Depends entirely on the cost of a wrong answer. Internal search tolerates misses a compliance answer can't. I report measured performance plus failure categories, and the business decides whether the number is acceptable — that decision shouldn't be made by the team building it.

Do I need a fine-tuned model?

Usually not. Most accuracy problems in RAG are retrieval problems, and fine-tuning doesn't fix retrieval. Exhaust chunking, embeddings and hybrid search first; revisit fine-tuning only with evidence the model's reasoning is the bottleneck.

How big should the evaluation set be?

Fifty well-chosen questions covering document types and known failure modes beats five hundred generated ones. Start small, expand when a failure category surprises you.

RAG demo that won't survive the real corpus?

Tell me about the documents and the questions users actually ask.