In our test, AI's quick reply promised things nobody agreed to.
One email, five things to answer
A common shortcut: paste the email and type "Write a reply to this email." Dana's fictional email asks for five answers, and three of them are written as statements, not questions.
What a plain request produced
Real output · GPT-5.5 (OpenAI API) · Oct 2026
With only "Write a reply to this email," GPT-5.5 answered for Sam. He had given it none of these answers.
The model filled gaps with guesses
It knew none of Sam's answers, so it guessed with confidence. A second run made the same three promises, and a wrong "yes, it's included" can cost you money.
Source: Our GPT-5.5 test runs (OpenAI API), Oct 2026
Make it list the questions first
Tested in GPT-5.5 (OpenAI API), Oct 2026
Paste this above the email. The list shows what needs an answer, and [CHECK] marks what the model doesn't know. OpenAI's GPT-5.5 guide also favors placeholders over unsupported specifics.
Try it: Swap in your name, then add "My notes:" and your short answers.
Source: OpenAI GPT-5.5 guide, as of Oct 2026
Add notes, and the gap shows
Real output · GPT-5.5 (OpenAI API) · Oct 2026
With four short notes from Sam added under the prompt, GPT-5.5 used each one. It wrote [CHECK] for the Thursday call, the one answer Sam hadn't given.
It still isn't foolproof
Real output · GPT-5.5 (OpenAI API) · Oct 2026
With no notes, a line claimed the job was done before its placeholder. With notes, it added an opinion Sam never gave: “I think the logo will feel cleaner without it.”
Before you paste, and before you send
Remove names, account numbers and other personal details first, and check the tool’s data policy or your employer’s approved tools before pasting a real email. Then fix every [CHECK] and read each line before sending.
Try it: Add your notes under the prompt and run it on one email.
Sources and assumptions
- OpenAI API docs: GPT-5.5 guide (Creative drafting guardrails): OpenAI's own guidance: 'If there is little or no citable support, write a useful generic draft with placeholders or clearly labeled assumptions rather than unsupported specifics.' This backs the [CHECK] placeholder idea (the guide is written for API builders and is about drafting in general, not email replies specifically). (checked 2026-10-09)
- OpenAI API docs: GPT-5.5 guide (Outcome-first prompts and stopping conditions): 'GPT-5.5 is strongest when the prompt defines the target outcome, success criteria, constraints, and available context, then lets the model choose the path.' This supports giving the model your notes (the context) and a clear rule for what counts as done. The guide also advises against long step-by-step procedures, so the post should present the numbered list as a check the reader can see, not as a required procedure. (checked 2026-10-09)
- Our own test runs, engine/ai-test.js (tests 'before', 'before-rerun', 'after', 'after-with-notes', GPT-5.5 via the OpenAI API): Every quoted output on slides 3, 6 and 7 comes from record.tests. The URL here is only the model's documentation page; the outputs themselves are in tests.json. (checked 2026-10-09)
Assumptions:
- The client email (Dana), the designer (Sam), 22 Harbour Lane, 'since 1987', the 14th launch and the Thursday 10:00 call are all fictional sample material written for this test.
- Only GPT-5.5 via the OpenAI API produced outputs. The two Claude runs in tests.json failed with a CLI error and returned no output, so the post must not name Claude as tested or say the prompt works in Claude.
- Each prompt was run once, except the vague 'before' prompt, which was run twice ('before' and 'before-rerun'). Both runs made the same three promises for Sam (see fact_record 0). Results can vary between runs.
- The before output made three promises for Sam: free second-round revisions, the invoice address already updated, and the Thursday call still at 10. It also suggested allowing a full week for the printer, but hedged that with 'if possible', so it is advice, not a promise, and is not counted. Its view on the tagline is an opinion it gave on Sam's behalf, also not counted as a promise.
- Without notes, the after prompt listed 7 items, including two that aren't really requests (the 'still aiming for the 14th' statement and the weekend well-wish), and answered all of them with [CHECK].
The short version
- A plain "write a reply" can promise for you.
- Have it list the questions first.
- Ask for [CHECK] where it lacks your answer.
- Fix every [CHECK], then read each line before sending.


