My answer is: restate the job, then check the restatement. Opening a fresh chat is easy; deciding what to carry into it is the difficult part. V recommends first making a portable brief whose goal and constraints you check against the original material. This is an editorial proposal, not a measured remedy for the apps you use today.

The one line a summary can lose

Consider an invented example. You ask for a two-day trip to Busan, then add that you will not drive, then that you must be in Haeundae on Saturday afternoon. Each addition supplies information without replacing an earlier requirement. A self-contained brief must preserve all three. “Suggest a Busan trip” is shorter, but less faithful.

The check is not whether the wording sounds polished. It is whether the duration, transport restriction and required destination survived. A pleasing phrase such as “a relaxed itinerary” cannot substitute for a hard constraint. This is an illustration of brief checking, not an observed model response or a task from the experiment.

Reading the numbers

The 2025-05-09 paper averages accuracy (0–100) across Code/Database/Math/Actions, with ten simulations/instruction. Table 2 compares GPT-4o-mini and GPT-4o.

Concat supplies everything upfront; Sharded distributes requirements; Recap restates them finally. Concat is not a restart after failure.

Condition4o-mini4o
Concat84.490.9
Sharded50.459.1
Recap66.576.6

Recap improved average scores but left roughly half the initial gap. Calculations use rounded averages.

4o-mini: 16.1 recovered (47.4%), 17.9 remaining. 4o: 17.5 recovered (55.0%), 14.3 remaining.
Bar values are percentage points.

Each bar is Concat−Sharded (100%): recovered=Recap−Sharded; remaining=Concat−Recap.

The denominator of this fraction is the initial score gap. It must not be confused with the probability of rescuing an individual failure. Read the recovered portion alongside the remaining distance.

A provenance limit worth keeping

Appendix M describes collecting user utterances.

At pinned commit c865793fe34a929d316119b0451d01bd9183bcfd, recap copies the old trace but builds its added message from the task’s consolidated prompt, rather than collecting utterance strings. The single-turn code sends only system/user messages. The revision and execution path behind the table, and any numerical impact, remain unestablished. The code alone is insufficient grounds to revise the published scores.

Pinned code

The strongest reason to keep the conversation

Moving to a fresh chat can discard approved examples, shared terminology and the reasons for exceptions. During exploration, the conversation itself may be valuable working material. Compressing the reasons you changed your mind into one line can freeze a mistaken interpretation of the goal. V would preserve the exploratory record when the job is not yet clear enough to specify.

A 2026-06-11 preprint trained Qwen2.5-Math-1.5B-Instruct and Qwen3-4B-Thinking with 256-token rolling memory, reporting gains even under full-history inference. It tests training, not manual restarting; cross-family validation remains open.

That countercase argues against treating conversational failure as inevitable. A user’s input-editing habit and a model’s training for conversation are separate interventions to test. Success for one cannot certify the other.

V’s proposal: check the portable brief first

Instead of an always-restart rule, try making a brief you could carry elsewhere. Write the goal, mandatory constraints, locations of supporting material and unresolved questions. Keep approved examples that explain exceptions. If a requirement changed, distinguish its current version from the one it replaced.

Then compare it item by item against the originals. An AI-written draft still needs checking. As a first check, select one indispensable requirement from the source and locate it in the brief. If it is missing, restore it rather than compressing further. Once checked, the brief could be used in the existing conversation or a fresh one. This article has not tested which route is better.

The test that would change my recommendation

Use the same checked brief, model, settings and task instances to compare adding it to existing history with putting it in a clean conversation. Include continuing without a brief. Randomize run order, repeat runs and score every requirement without revealing the condition to evaluators. Count user time and errors caused by lost context too. This is a proposed study, not an experiment we ran.

If the two brief conditions perform similarly, there is less reason to recommend the extra restart step. If a fresh chat loses important context more often, the recommendation should favor keeping the history. For now, the useful exercise is to make the job inspectable by someone else, rather than memorize a restart ritual with an unearned promise.

Read next

Context and memory background

Storing information and recovering a drifting task are different questions. Continue here for the storage concepts.

Sources

2025 paper

Single-turn code

2026 paper