less than 1 minute read

Written by - Millan Kaul

๐—œ ๐—ณ๐—ผ๐˜‚๐—ป๐—ฑ ๐—ผ๐˜‚๐˜ ๐—บ๐˜† ๐—”๐—œ ๐—ท๐˜‚๐—ฑ๐—ด๐—ฒ ๐˜„๐—ฎ๐˜€ ๐˜๐—ผ๐—ผ ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ผ๐˜‚๐˜€.

Image showing AI judge evaluating AI responses
Your evaluator needs evaluation too.

What I did today:

I gave my judge one clearly strong answer and one obviously weak answer. Then I checked whether it scored them differently and explained why.

What surprised me:

The judge liked the weak answer more than I expected because my rubric did not say what a failure looked like.

My QA takeaway:

  1. An LLM judge is also a system under test.
  2. Your evaluator needs evaluation too.

Have you ever tested the tool that tests your AI?

#LLMTesting #AIEvals #TestStrategy