Academic evaluation of dialogue systems has improved enormously in a decade. We have standard benchmarks, human preference studies, and increasingly sophisticated measures of factuality and safety. What we still lack is a good account of what happens when a conversational system runs continuously with the same user for months. That regime is where the interesting failures live, and almost all of the evidence about it sits inside commercial deployments rather than in published literature.
Table of Contents
ToggleThe benchmark and the deployment diverge
Most evaluation treats dialogue as a short, bounded task. A model is given a prompt or a brief exchange and its output is scored. This is tractable and reproducible, and it has driven real progress.
Deployed conversational systems do not work like that. A user returns daily for six months. The exchange accumulates thousands of turns. The system is expected to remember what was said in week three, to maintain a consistent persona, and to behave sensibly when the user tests its boundaries, which they invariably do. None of that is captured by a benchmark measuring single-turn quality.
The gap produces a familiar pattern. A model scores well in evaluation, ships, and then exhibits problems nobody measured: persona drift across long sessions, contradictions between what it said in different weeks, or degradation once the context window fills and older material is silently dropped.
Memory is an architectural problem, not a feature
The practical solutions in production are less elegant than the literature might suggest. Few systems keep full history in context, because cost and latency make that impossible at scale. Instead they compress: periodic summarisation of past exchanges, retrieval of relevant fragments, and structured extraction of durable facts into a separate store.
Each approach introduces its own failure mode. Summarisation loses specifics and tends to flatten emotional nuance. Retrieval surfaces the wrong fragment when the query is ambiguous. Structured extraction captures facts but not tone, which produces systems that remember a user’s stated preferences while forgetting entirely how the relationship has felt.
These trade-offs are being explored empirically by product teams under commercial pressure, mostly without publication. Consumer companion platforms are the clearest case, since long-running memory is the core of the value proposition. Services such as justsext.com maintain continuity across text and voice over extended periods, which makes them an unusually rich source of evidence about what long-horizon dialogue actually requires.
Voice changes the requirements
Adding speech is often treated as a presentation layer, as though the dialogue problem is solved and only the output format changes. In practice voice alters user expectations substantially.
Latency tolerance collapses. A two-second pause is unremarkable in text and conspicuous in speech. Turn-taking becomes an explicit engineering problem: when a user interrupts, the system must decide whether to yield, and the correct behaviour depends on prosody as much as content. Emotional expectation rises as well, because a flat delivery of an otherwise appropriate response reads as indifference in a way that plain text does not.
These are researchable questions, and largely unresearched ones outside industrial labs.
Evaluation methods worth borrowing
Several practices from deployment deserve academic attention.
Longitudinal consistency testing. Replay a long conversation and check whether the system contradicts its own earlier statements. Simple to implement, rarely done in evaluation, and it catches a class of failure that single-turn scoring cannot see.
Retention as a quality signal. Whether a user returns is a coarse measure, but it aggregates factors that no rubric captures well. It is also noisy and confounded by design choices unrelated to dialogue quality, so it should complement rather than replace controlled evaluation.
Adversarial long-horizon probing. Users spend months finding the edges of a system. Systematic, extended probing during evaluation would surface those edges before release rather than after.
An argument for engagement
There is understandable reluctance in the research community to engage with consumer applications in this space, particularly those serving adult markets. The reluctance is not costless. These are the systems with the longest running conversations, the most demanding memory requirements and the largest volume of genuine long-horizon interaction data in existence.
Ignoring them means the field continues to theorise about long-term human and machine dialogue while the actual evidence accumulates entirely in private hands, unpublished and unexamined. That is a poor outcome for a research community that cares about understanding these systems rather than only about building them.



