Recent posts

Five examples are enough to start - day 04

less than 1 minute read

You don’t need thousands of test cases to begin. Five carefully chosen examples are more than enough to boot up your AI evaluation pipeline.

I tested a longer prompt - day 03

less than 1 minute read

Increasing the complexity of prompt inputs to see how the LLM judge responds and analyzing its limits when prompt length grows.

My first AI eval was boring - day 02

less than 1 minute read

Setting up a basic LLM evaluation only to find the initial results predictable, but discovering why simple tests are the best place to start.