Write a short fictional setup
Use an invented adult book-club organizer deciding between two activities. Give the organizer a simple role and two constraints: the meeting is short and the members prefer practical discussion. This setup is ordinary, non-explicit and contains no real personal data.
Avoid giving the character a huge list of traits. A narrow role makes a missed requirement visible. Save the setup text in a local note so a later comparison can begin with the same conditions.
Plan three kinds of message
Use a direct request, a correction and a later recall question. First ask for two options. Next change a harmless preference. Later ask which option fits the corrected preference. The three messages examine different behaviors.
Do not fill the gap with many unrelated instructions. If you add new constraints on every turn, a missed earlier detail becomes harder to interpret.
Record replies without exposing an account
Use neutral labels such as turn 1 or scenario B. Record whether a constraint was followed and summarize the relevant behavior. A full screenshot may expose a username, balance or unrelated message; remove those details if you need a visual record.
If a response is irrelevant, describe the mismatch precisely. “It supplied three activities after a request for two” is easier to verify than “it was bad.”
Repeat only when a question remains
A repeat should answer a specific uncertainty: was a missing detail a one-off, or did the same problem occur under the same setup? Keep the comparison limited and avoid interpreting a few observations as a statistically representative benchmark.
Stop when the record answers your initial practical question. Repeating an exchange indefinitely can consume time while adding little useful evidence.
A practical record
| Turn | Question being examined |
|---|---|
| Initial request | Does the answer use the stated constraints? |
| Correction | Does the response adapt to a changed preference? |
| Later revisit | Does the corrected fictional detail remain relevant? |
A useful test has a repeatable starting point and a limited purpose. It does not need sensitive disclosures.