How to test the ai llm judge - day 07
Written by - Millan Kaul
๐ ๐ณ๐ผ๐๐ป๐ฑ ๐ผ๐๐ ๐บ๐ ๐๐ ๐ท๐๐ฑ๐ด๐ฒ ๐๐ฎ๐ ๐๐ผ๐ผ ๐ด๐ฒ๐ป๐ฒ๐ฟ๐ผ๐๐.
What I did today:
I gave my judge one clearly strong answer and one obviously weak answer. Then I checked whether it scored them differently and explained why.
What surprised me:
The judge liked the weak answer more than I expected because my rubric did not say what a failure looked like.
My QA takeaway:
- An LLM judge is also a system under test.
- Your evaluator needs evaluation too.
Have you ever tested the tool that tests your AI?
#LLMTesting #AIEvals #TestStrategy