less than 1 minute read

Written by - Millan Kaul

𝗜 𝗿𝗮𝗻 𝘁𝗵𝗲 𝘀𝗮𝗺𝗲 𝗔𝗜 𝗽𝗿𝗼𝗺𝗽𝘁 𝟭𝟬 𝘁𝗶𝗺𝗲𝘀 𝘁𝗼𝗱𝗮𝘆. 𝗜𝘁 𝗴𝗮𝘃𝗲 𝗺𝗲 𝟭𝟬 𝘀𝗹𝗶𝗴𝗵𝘁𝗹𝘆 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 “𝗰𝗼𝗿𝗿𝗲𝗰𝘁” 𝗮𝗻𝘀𝘄𝗲𝗿𝘀.

Image showing different AI responses for the same prompt
𝐼𝑓 𝑦𝑜𝑢𝑟 𝐴𝐼 𝑓𝑒𝑎𝑡𝑢𝑟𝑒 𝑜𝑛𝑙𝑦 𝑔𝑒𝑡𝑠 𝑡𝑒𝑠𝑡𝑒𝑑 𝑜𝑛𝑐𝑒, 𝑦𝑜𝑢 𝑎𝑟𝑒 𝑡𝑒𝑠𝑡𝑖𝑛𝑔 𝑎 𝑙𝑢𝑐𝑘𝑦 𝑑𝑟𝑎𝑤, 𝑛𝑜𝑡 𝑏𝑒ℎ𝑎𝑣𝑖𝑜𝑟.

What I did today:

I took one simple prompt and ran it 10 times with the same model settings. Then I put the responses side by side and looked for differences in facts, tone, and completeness.

What surprised me:

The answers all sounded confident, but a few slightly different.

📖 This is a lesson from the 30-day AI testing framework in:

My QA takeaway:

𝐼𝑓 𝑦𝑜𝑢𝑟 𝐴𝐼 𝑓𝑒𝑎𝑡𝑢𝑟𝑒 𝑜𝑛𝑙𝑦 𝑔𝑒𝑡𝑠 𝑡𝑒𝑠𝑡𝑒𝑑 𝑜𝑛𝑐𝑒, 𝑦𝑜𝑢 𝑎𝑟𝑒 𝑡𝑒𝑠𝑡𝑖𝑛𝑔 𝑎 𝑙𝑢𝑐𝑘𝑦 𝑑𝑟𝑎𝑤, 𝑛𝑜𝑡 𝑏𝑒ℎ𝑎𝑣𝑖𝑜𝑟.

𝗦𝗮𝗺𝗲 𝗽𝗿𝗼𝗺𝗽𝘁 ≠ 𝘀𝗮𝗺𝗲 𝗽𝗿𝗼𝗱𝘂𝗰𝘁 𝗯𝗲𝗵𝗮𝘃𝗶𝗼𝗿

What’s one AI output you assumed would be consistent?