GPT-5.5 put my private note into a customer email. Twice.
I said "keep every fact"
I pasted fictional order notes and said "keep every fact exactly as given." We'd shipped her a grey tablecloth by mistake. One note was only for me: "No discount code (she already gets the grey one)."
The private line got into the email
Real output · GPT-5.5 (OpenAI API) · Oct 2026
In 2 of 3 runs, the private note ended up in the draft email. This run also greeted her as "Dana R.", probably because I said keep every fact exactly.
Nothing says who a note is for
Nothing in the notes says which lines are for the customer, so "keep every fact" covers them all. OpenAI's guide says headers can help "mark distinct sections of a prompt."
Source: OpenAI API docs, Prompt engineering guide, Oct 2026
Two headings and one instruction
Tested in GPT-5.5 (OpenAI API), Oct 2026
The prompt now says to use only the facts under the first heading and to never mention the second group in the email.
Try it: Move any note that's for you under INTERNAL ONLY before pasting.
The facts stayed, the leak didn't
Real output · GPT-5.5 (OpenAI API) · Oct 2026
Neither of the 2 runs mentioned a discount code, and both kept every customer fact: the dates, the free grey tablecloth and the $8.50 refund.
Splitting lowers the risk, not to zero
This was a small test: 3 runs before, 2 after. One unsplit run didn't leak, so you can't predict when it will. Read the email once for anything you'd only say to a colleague.
Before you paste real notes
Remove names, account numbers and other personal details first. Then check the tool's data policy or your employer's approved tools before uploading real customer documents.
Sources and assumptions
- OpenAI API docs: GPT-5.5 model page: GPT-5.5 is available in the OpenAI API (snapshot gpt-5.5-2026-04-23) and supported on v1/chat/completions. Reasoning effort supports none, low, medium (default), high and xhigh, so our runs used medium. (checked 2026-10-07)
- OpenAI API docs: Prompt engineering guide: 'Markdown headers and lists can be helpful to mark distinct sections of a prompt, and to communicate hierarchy to the model.' Also: 'you can help the model understand logical boundaries of your prompt and context data using a combination of Markdown formatting and XML tags.' (checked 2026-10-07)
Assumptions:
- All runs used GPT-5.5 through the OpenAI Chat Completions API (engine/ai-test.js) with default settings: no system message, no temperature, and the default reasoning effort (medium, per OpenAI's model page). The ChatGPT app adds its own system prompt, so app results can differ.
- Before prompt: 3 runs (gpt-1, gpt-2, before-3). After prompt: 2 runs (after-1, after-2), all on 2026-10-07 local time (2026-10-08 UTC). This is a small test. Outputs vary run to run, so the post claims what happened in these runs, not a rate.
- Claude was not tested: the Claude runner failed in this session, so the post names only GPT-5.5 and makes no claim about other models.
- Hollow Pine Linen, Maya, Dana R., order #4471, the dates and the $8.50 refund are fictional sample material.
- The headings TELL THE CUSTOMER / INTERNAL ONLY are our own wording. OpenAI's guide supports marking distinct sections with headers in general and does not promise this stops leaks.
The short version
- Notes in one pile leaked in 2 of 3 runs.
- Split notes leaked in 0 of 2 runs.
- Use TELL THE CUSTOMER and INTERNAL ONLY, and say never to mention the second.
- Still read the email before sending.


