I turned a defect into an eval - day 15
I found one AI bug today—and turned it into a permanent test in 10 minutes.
I found one AI bug today—and turned it into a permanent test in 10 minutes.
I tested two models today. One answer was only slightly better and much more expensive.
Putting predictable constraints and standard expectations to the test, evaluating how the LLM judge grades standard formatting and basic rules.
Setting up a basic LLM evaluation only to find the initial results predictable, but discovering why simple tests are the best place to start.
Running the exact same prompt multiple times reveals the inherent variability in LLM responses, a core challenge in AI testing.