LLM / Conversation Management
48 valid surveys across Chinese, English, and Spanish, plus 7 semi-structured interviews, reveal how the standard chat UI pattern turns accumulated AI conversations into an archive nobody can retrieve from.

TL;DR
- The only significant predictor of giving up on retrieval was similar-topic scattering (OR = 12.07, p = .003); the wide confidence interval calls for a cautious read.
- Even after controlling for technical skill, organizational initiative still predicted retrieval success (partial r = .345, p = .017).
- Three mental models of conversation management emerged from clustering, and in-conversation bookmarks were the only feature significantly correlated with satisfaction (r = .282).
CONTEXT
From helper to hoard.
LLM chat products accumulate conversations linearly. Titles auto-generate into near-identical labels, and search is keyword-only, useless when all you remember is what the chat was about.
Early observations from my own use and community complaints pointed the same way: people abandon old chats and restart, losing the context that made the conversation valuable in the first place.
The research problem: what actually predicts retrieval failure and dissatisfaction?
“I know I asked this before. I just can't find it, so I start over.”
Interview participant
Role
Lead researcher (instrument design, translation equivalence, analysis)
Type
Independent research study
Duration
~8 weeks
Methods
Survey, interviews, regression, K-means
RESEARCH QUESTIONS
What predicts abandoning retrieval of a past conversation?
Candidate predictors: topic scattering, volume, proficiency, tool features.
What drives satisfaction with conversation management?
Initiative vs. proficiency vs. feature usage.
Are there distinct mental models of "managing" AI conversations?
Cluster analysis on behavior + attitude items.
METHODS
How the study was run.
method one
Survey (n=48 valid)
Instrument in Chinese, English, and Spanish with back-translation checks for equivalence. Likert batteries on behavior, agency, proficiency, features; screening for validity.
method two
Semi-structured interviews (n=7)
Follow-up interviews sampled across response profiles; thematic coding to explain the quantitative signal.
KEY FINDINGS
finding 01
12.07
Similar-topic scattering was the strongest predictor of retrieval abandonment. When chats sprawl across similar topics, people stop even trying to find them.
finding 02
0.345
Even after controlling for self-rated technical skill, organizational initiative still predicted satisfaction (p = .017). The lever is open to users at any skill level.
finding 03
0.282
K-means surfaced three mental models of the history drawer, and in-conversation bookmarks were the only feature significantly correlated with satisfaction.
Three ways people cope.
K-means clustering on behavior and attitude items surfaced three distinct mental models of conversation management, each mapping to a cluster in the scatter.
Mental model: drawer ≈ inbox
Mental model: drawer ≈ toolbox
Mental model: drawer ≈ scroll
The Archivist
Files diligently, folder-first.
Needs hierarchy the UI doesn't offer.
The Restartee
Abandons and restarts.
Needs retrieval that forgives chaos.
The Threader
One eternal conversation.
Needs in-conversation landmarks (bookmarks).
SO WHAT
Design principles for conversation management.
Design for scattering: semantic grouping across similar topics, not just chronological lists.
Finding 1Ship in-conversation bookmarks, the only feature that correlated with satisfaction.
Finding 3Build agency, not features: visible structure beats power-user tooling.
Finding 2Retrieval should tolerate vague memory, like "that chat about the thing last week".
User Interviewsvalid surveys
languages (zh · en · es)
topic scattering odds ratio
initiative partial r (p = .017)
LIMITATIONS & REFLECTION
Where the evidence ends.
The sample is modest for regression, mitigated by reporting effect sizes and confidence intervals, and by keeping claims conservative. Significant findings were treated as signals, not verdicts.
The sample self-selects toward engaged AI users, so the results likely understate how bad retrieval failure is for casual users.
Cross-language equivalence was checked through back-translation, but cultural response styles may still differ across the Chinese, English, and Spanish instruments.
Clustering is descriptive, not prescriptive, the three mental models are a map of how people cope today, not a mandate for how products should segment them.
“People don’t fail at managing conversations, the interface fails at being manageable.”