less than 1 minute read

Written by - Millan Kaul

๐—œ ๐—ฟ๐—ฎ๐—ป ๐˜๐—ต๐—ฒ ๐˜€๐—ฎ๐—บ๐—ฒ ๐—”๐—œ ๐—ฝ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜ ๐Ÿญ๐Ÿฌ ๐˜๐—ถ๐—บ๐—ฒ๐˜€ ๐˜๐—ผ๐—ฑ๐—ฎ๐˜†. ๐—œ๐˜ ๐—ด๐—ฎ๐˜ƒ๐—ฒ ๐—บ๐—ฒ ๐Ÿญ๐Ÿฌ ๐˜€๐—น๐—ถ๐—ด๐—ต๐˜๐—น๐˜† ๐—ฑ๐—ถ๐—ณ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐˜ โ€œ๐—ฐ๐—ผ๐—ฟ๐—ฟ๐—ฒ๐—ฐ๐˜โ€ ๐—ฎ๐—ป๐˜€๐˜„๐—ฒ๐—ฟ๐˜€.

Image showing different AI responses for the same prompt
๐ผ๐‘“ ๐‘ฆ๐‘œ๐‘ข๐‘Ÿ ๐ด๐ผ ๐‘“๐‘’๐‘Ž๐‘ก๐‘ข๐‘Ÿ๐‘’ ๐‘œ๐‘›๐‘™๐‘ฆ ๐‘”๐‘’๐‘ก๐‘  ๐‘ก๐‘’๐‘ ๐‘ก๐‘’๐‘‘ ๐‘œ๐‘›๐‘๐‘’, ๐‘ฆ๐‘œ๐‘ข ๐‘Ž๐‘Ÿ๐‘’ ๐‘ก๐‘’๐‘ ๐‘ก๐‘–๐‘›๐‘” ๐‘Ž ๐‘™๐‘ข๐‘๐‘˜๐‘ฆ ๐‘‘๐‘Ÿ๐‘Ž๐‘ค, ๐‘›๐‘œ๐‘ก ๐‘๐‘’โ„Ž๐‘Ž๐‘ฃ๐‘–๐‘œ๐‘Ÿ.

What I did today:

I took one simple prompt and ran it 10 times with the same model settings. Then I put the responses side by side and looked for differences in facts, tone, and completeness.

What surprised me:

The answers all sounded confident, but a few slightly different.

My QA takeaway:

๐ผ๐‘“ ๐‘ฆ๐‘œ๐‘ข๐‘Ÿ ๐ด๐ผ ๐‘“๐‘’๐‘Ž๐‘ก๐‘ข๐‘Ÿ๐‘’ ๐‘œ๐‘›๐‘™๐‘ฆ ๐‘”๐‘’๐‘ก๐‘  ๐‘ก๐‘’๐‘ ๐‘ก๐‘’๐‘‘ ๐‘œ๐‘›๐‘๐‘’, ๐‘ฆ๐‘œ๐‘ข ๐‘Ž๐‘Ÿ๐‘’ ๐‘ก๐‘’๐‘ ๐‘ก๐‘–๐‘›๐‘” ๐‘Ž ๐‘™๐‘ข๐‘๐‘˜๐‘ฆ ๐‘‘๐‘Ÿ๐‘Ž๐‘ค, ๐‘›๐‘œ๐‘ก ๐‘๐‘’โ„Ž๐‘Ž๐‘ฃ๐‘–๐‘œ๐‘Ÿ.

๐—ฆ๐—ฎ๐—บ๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜ โ‰  ๐˜€๐—ฎ๐—บ๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜ ๐—ฏ๐—ฒ๐—ต๐—ฎ๐˜ƒ๐—ถ๐—ผ๐—ฟ.

Whatโ€™s one AI output you assumed would be consistent?