Recent posts

The ā€˜best’ model lost - day 08

less than 1 minute read

The model everyone calls ā€˜best’ lost my tiny test today. Learn why a cheaper, faster model might outperform state-of-the-art models for specific tasks.

I made AI Judge - Day 06

less than 1 minute read

Evaluating the logic behind having an LLM judge another LLM’s output. Establishing a helpfulness rubric and confronting definitions.

Five examples are enough to start - day 04

less than 1 minute read

You don’t need thousands of test cases to begin. Five carefully chosen examples are more than enough to boot up your AI evaluation pipeline.