In brief: Prompt sensitivity measures how much the content, sources or recommendations in an AI answer change after a small, meaning-controlled prompt adjustment.
How I use Prompt sensitivity in practice
I do not use Prompt sensitivity as another number in a presentation. First I determine which decision it should clarify, which data or observations support it and who will change something based on the result. I test predefined paraphrases of the same need and separate wording effects from ordinary variability in repeated answers. I also record the baseline, measurement date and limits of interpretation. This makes it possible later to distinguish real progress from a change in the tool, sample or query wording.
What to watch out for
The greatest risk is accuracy that only looks precise. If a paraphrase changes the customer’s intent, this is not sensitivity to wording but a different question and a different market being measured. I therefore compare results over time, on a stable sample and together with business context. If the term does not lead to a concrete next step, no analysis has been created, only a new label.
Questions for a decision
- Do the variants preserve the same intent?
- Which words change the source selection?
- Is the effect repeatable?
- Should sensitive prompts be included in a separate cluster?