I tested a longer prompt - day 03
Increasing the complexity of prompt inputs to see how the LLM judge responds and analyzing its limits when prompt length grows.
Increasing the complexity of prompt inputs to see how the LLM judge responds and analyzing its limits when prompt length grows.
Setting up a basic LLM evaluation only to find the initial results predictable, but discovering why simple tests are the best place to start.
Running the exact same prompt multiple times reveals the inherent variability in LLM responses, a core challenge in AI testing.
Written by - Millan Kaul
How three minor Docker configuration changes shaved 33% off my Jekyll build times, saving hours of developer wait time.