When summaries beat raw history
For long-running chats, raw history grows unbounded and costs grow with it. Periodic summarization replaces older turns with a condensed assistant note, preserving the facts the model needs without the conversational fluff. Done well, this keeps quality high and tokens low.
What to preserve
Good summaries preserve: user-stated facts ('I work in pharma'), explicit preferences ('always reply in Spanish'), unresolved tasks, and conclusions reached. Bad summaries preserve: greetings, banter, and the model's own apologies. Hand-craft your summarization prompt to call out what to keep.
Drift is real
Each summarization pass loses fidelity. Run too many and the model's effective memory degrades into vague impressions. Mitigate by summarizing in tiers — last-N turns verbatim, middle tier summarized lightly, oldest tier heavily summarized — instead of one flat compression.