Adding a CI pass fail threshold - day 30
I brought my eval journey to the finish line today by integrating our small test dataset directly into our CI/CD pipeline with a pass/fail threshold.
I brought my eval journey to the finish line today by integrating our small test dataset directly into our CI/CD pipeline with a pass/fail threshold.
I found one AI bug today—and turned it into a permanent test in 10 minutes.
I tested two models today. One answer was only slightly better and much more expensive.
Putting predictable constraints and standard expectations to the test, evaluating how the LLM judge grades standard formatting and basic rules.
Setting up a basic LLM evaluation only to find the initial results predictable, but discovering why simple tests are the best place to start.