AI Engineering Podcast
AI Engineering Podcast

How to Evaluate RAG Systems When Ground Truth Keeps Changing

1h 3min

Ofer Mandelovich discusses challenges of evaluating RAG systems with incomplete or changing ground truth. Key approaches include LLM as a judge with methods like Umbrella and AutoNuggetizer, separating retrieval vs. generation accuracy, and using synthetic data or online sampling for ongoing evaluation.

Summarized by Podsumo

✨ Key Takeaways

💬 Notable Quotes

Get every episode summarized
Delivered to Telegram. Ask questions about any episode.
Start on Telegram