Ofer Mandelovich discusses challenges of evaluating RAG systems with incomplete or changing ground truth. Key approaches include LLM as a judge with methods like Umbrella and AutoNuggetizer, separating retrieval vs. generation accuracy, and using synthetic data or online sampling for ongoing evaluation.
Summarized by Podsumo
LLM as a judge can evaluate RAG systems by scoring chunk relevance (0-3 scale) with ~80% human correlation.
Retrieval and generation accuracy require separate evaluation; synthetic data generation helps bootstrap when ground truth is scarce.
Agentic RAG improves results by rewriting queries and running multiple iterations, but tool quality (e.g., Notion Search) varies.
Context engineering encompasses RAG as an attention mechanism for selecting relevant data into the LLM's context window.
Key production challenges include role-based permissions, multimodal data, and sovereign AI demands for on-premise deployments.
"You take every chunk in your top 10 and run an LLM as a judge that gives you a score between zero and three... the correlation is very high, I think it was like something like maybe 80%."