Recent posts

I tested a longer prompt - day 03

less than 1 minute read

Increasing the complexity of prompt inputs to see how the LLM judge responds and analyzing its limits when prompt length grows.

My first AI eval was boring - day 02

less than 1 minute read

Setting up a basic LLM evaluation only to find the initial results predictable, but discovering why simple tests are the best place to start.