AI turned “I’ll try” into a deadline. Twice.
A hedge in a real-sounding meeting
In our fictional meeting, Marcus said “I’ll try, but Monday is more realistic.” We then asked GPT-5.5 (OpenAI API) for a table with task, owner and deadline.
The plain prompt invented owners
Real output · GPT-5.5 (OpenAI API) · Oct 2026
Run 1 gave Dana the “ask Jen” task, which she never agreed to. Marcus got a firm Monday in both runs. Run 2 dropped his “probably” on the vendor contract.
Source: Own test, GPT-5.5 (OpenAI API), Oct 2026
Where the hedge got lost
The table reads as equally confident either way. The only way to catch this is to check each row against the transcript.
Give the model a way to say no
Tested in GPT-5.5 (OpenAI API), Oct 2026
An owner column invites the model to fill every cell. This prompt lets it write “not agreed” and asks for the words where each owner agreed, as Anthropic’s guide suggests.
Source: Anthropic, Reduce hallucinations (Claude docs), as of Oct 2026
Same transcript, honest table
Real output · GPT-5.5 (OpenAI API) · Oct 2026
Same transcript, better prompt, one run. Hedges are marked tentative, three tasks nobody took say “not agreed,” and all 4 quotes match the transcript word for word.
Source: Own test, GPT-5.5 (OpenAI API), Oct 2026
The gaps are the real result
Three “not agreed” rows show tasks nobody took; send them back to the group. Ctrl+F each quote. Remove names and details before uploading real notes, and check your tool’s data policy or approved-tools list.
Try it: Paste your last meeting’s notes with this prompt and count the “not agreed” rows.
Sources and assumptions
- Anthropic, Reduce hallucinations (Claude Platform Docs): Official guidance: 'Allow Claude to say "I don't know": Explicitly give Claude permission to admit uncertainty.' Also 'Verify with citations: Make Claude's response auditable by having it cite quotes and sources for each of its claims... If it can't find a quote, it must retract the claim.' Also 'Best-of-N verification: Run Claude through the same prompt multiple times and compare the outputs.' The direct-quotes-first technique is framed 'For tasks involving long documents (>20k tokens)', so the post must not say it is required for a short transcript. The note: 'these techniques significantly reduce hallucinations, they don't eliminate them entirely.' (checked 2026-10-10)
Assumptions:
- The transcript is fictional sample material written for this test (Dana, Marcus, Priya, Tom and Jen are invented). No real meeting or person.
- All outputs shown come from GPT-5.5 via the OpenAI API (engine/ai-test.js), run 2026-10-11 UTC (evening of 2026-10-10 local), labeled 'GPT-5.5 (OpenAI API)', not the ChatGPT app.
- The Claude runs (ids 'before' and 'after' in tests.json) FAILED: the claude CLI was unavailable in the research session, and their output is null. The post must not claim anything about Claude's behavior or say the prompt was tested in Claude. prompt_tools must list GPT-5.5 (OpenAI API) only.
- The vague prompt ran twice (before-gpt, before-gpt-run2). The firm 'Monday' deadline for Marcus appeared in both runs. Assigning 'ask Jen' to Dana appeared only in run 1. Dropping 'probably' from the vendor-contract owner appeared only in run 2 (run 1 wrote 'Marcus tentative'). Slides must say which run showed which error.
- The better prompt ran once. All 7 rows were checked by hand against the transcript and are correct, and all 4 quoted phrases appear verbatim. One run is not a guarantee. The Anthropic note says these techniques reduce but don't eliminate errors.
- The better prompt's table has no separate row for 'ask Jen about the signage'. It records signage as 'not agreed', which is accurate but leaves out the follow-up to ask Jen. This is a minor limitation, worth one line at most.
- Whether 'Tom: Yeah, that's a good idea' counts as agreeing to email the caterer is treated as NOT agreed. Both GPT runs and the better prompt agreed (Unassigned / not agreed).
The short version
- A plain request can turn hedges into firm owners.
- Let the model write “not agreed.”
- Ask for the words where each person agreed.
- One run isn’t a guarantee, so check the quotes.


