less than 1 minute read

Written by - Millan Kaul

๐—ง๐—ผ๐—ฑ๐—ฎ๐˜† ๐—œ ๐—ฎ๐˜€๐—ธ๐—ฒ๐—ฑ ๐—ผ๐—ป๐—ฒ ๐—”๐—œ ๐˜๐—ผ ๐—ด๐—ฟ๐—ฎ๐—ฑ๐—ฒ ๐—ฎ๐—ป๐—ผ๐˜๐—ต๐—ฒ๐—ฟ ๐—”๐—œ. ๐—ฆ๐—น๐—ถ๐—ด๐—ต๐˜๐—น๐˜† ๐˜„๐—ฒ๐—ถ๐—ฟ๐—ฑ. ๐—ฉ๐—ฒ๐—ฟ๐˜† ๐˜‚๐˜€๐—ฒ๐—ณ๐˜‚๐—น.

Image showing AI response evaluation
โ€œHelpfulโ€ is not a test case until you define it.

What I did today:

I wrote a simple rubric for helpfulness: answer the question, be accurate, and give a useful next step. Then I gave that rubric to an LLM judge for a handful of responses.

What surprised me:

The hardest part was not running the judge. It was defining what โ€œhelpfulโ€ actually means.

My QA takeaway:

  1. If your quality criteria are vague, your eval results will be vague too.
  2. โ€œHelpfulโ€ is not a test case until you define it.

What word in your quality bar needs a clearer definition?

#LLMEvals #AIQuality #QualityEngineering