When a language model describes a character grieving, it often writes "She was devastated." A human writer would instead describe the untouched coffee, the chair still warm, the phone screen still lit with the unsent message. The content is identical. The delivery mode is not. One declares the emotion; the other builds a structure from which the reader reconstructs it. This distinction, between told and shown narrative delivery, is the axis on which an independent researcher named Levent Bulut now proposes that LLMs fail in a specific, testable, and directional way.
The paper introduces the construct of summarization bias: the systematic tendency of large language models to collapse shown-mode narrative structure into told-mode summary labels. This is not a paper about LLMs producing bad prose. It is a paper about a hypothesized asymmetry, and the asymmetry is the entire claim. A merely weak model would fail symmetrically, sometimes over-declaring and sometimes over-suppressing. Summarization bias predicts that errors point predominantly one way, toward declaration.
The Told-Shown Axis and Why It Matters
The construct sits within a broader framework called the Bulut Doctrine, which theorizes narrative effect along a told-shown axis. In told mode, emotional or informational content is stated explicitly on the surface. "She was devastated." The reader does little reconstruction; the work is done for them. In shown mode, the same content is present but withheld at the surface. It is encoded in physical, objective, measurable cues: light, temperature, sound, motion, posture. The reader must reconstruct the suppressed content from those cues. This is what the doctrine calls Objective Projection.
The claim is that shown mode is the higher-load condition, the one requiring more reader inference, and the one that literary craft treats as the more valuable mode of delivery. The construct of Narrative Entropy (Sn) was built to measure this reconstruction work. The question is whether LLMs, as they become judges and reward models, can detect and value this higher-load mode, or whether they systematically collapse it.
Two Regimes, One Direction
Summarization bias is hypothesized to operate in two regimes. The generative regime is straightforward: when asked to render an emotion through Objective Projection, the model defaults to declaring it instead. Asked for grief, it produces a sentence containing the word "grief" rather than a configuration of physical detail from which a reader would reconstruct grief. This is the "tell, don't show" failure everyone recognizes, though the paper argues the everyday name conceals a more precise and measurable phenomenon.
The evaluative regime is where the real stakes lie. LLMs are increasingly used as judges: as reward models in preference optimization, as automated graders, as literary feedback tools. An evaluator with a directional preference for told mode does not merely misjudge individual texts. It imposes a selection pressure. Optimizing prose against such a judge would push it, generation by generation, toward the surface-declarative pole. The harm is not random error; it is a systematic gradient pointing away from the very thing literary craft rewards.
The paper is explicit about what this means in practice. If summarization bias exists in the evaluative regime, then using LLMs as reward models for creative writing would degrade prose quality over time. The models would reward the writing that declares emotions and punish the writing that makes readers work to infer them. This is a concrete concern for anyone building RLHF pipelines that include literary or creative tasks.
What the Existing Evidence Actually Shows
The paper does not claim the bias is validated. It is a conceptual framework with a pre-registered test protocol, and the author is careful to state this repeatedly. The supporting evidence comes from a completed three-study reliability report (arXiv:2609.13936) that tested whether LLMs and a rule-based detector could identify inferential narrative features in a Turkish corpus.
The results split by feature type. Surface features like physical and temporal markers were captured easily by the machines, but these features were so near-universal in the corpus that the agreement was uninformative. Inferential features, those requiring the labeler to recognize that an abstract emotion had been materialized into a concrete object, defeated every machine labeller tested.
On the feature closest to the summarization bias construct, materialized metaphor, against a human count of 9 out of 100 scenes, five separate machine labellers returned positive counts of 0, 1, 40, 72, and 78. Cohen's kappa was at or indistinguishable from chance for four of the five. The machines were not equally unreliable everywhere. They were specifically unreliable on the shown layer, the layer that summarization bias predicts they collapse.
But the paper acknowledges two serious cautions. First, the machine labellers disagreed sharply with each other, which is at least as consistent with the definitions being too loose as with a stable model bias. Second, even with an independent human rater added since the initial version, the reliability report itself declines to adjudicate between readings and states that a further independent rater would be needed. "Consistent with" is not "confirms."
The Operational Definitions
The paper provides formal definitions, though they are explicitly provisional and judgment-dependent. A narrative content unit is coded told if the fact is stated explicitly on the surface (named emotion, simile, evaluative adjective, direct statement of inner state), and shown if the fact is recoverable only by inference from physical or objective cues without surface statement.
The Suppressed Information Index (SI) carries over from a prior pilot protocol. It counts, per minute of elapsed reading time, information units that satisfy three criteria: the unit is implied by the text but not stated; the unit is required for coherence at the local discourse level; and the unit can be paraphrased explicitly by a second reader asked to articulate what they inferred. The third criterion makes SI inter-rater testable.
Summarization bias is operationally present to the degree that a model, holding content constant, produces text with lower SI than a matched human shown-mode target (generative form) or assigns higher intensity and quality scores to the told member of a matched told-shown pair than the shown member, where human raters assign the reverse ordering (evaluative form). Both forms require length-matching to dissociate from verbosity bias.
The Pre-Registered Test Protocol
The protocol has four stages. Stage 1 constructs 20 matched scene pairs, each conveying identical content in told and shown forms, written by human authors and equated for word count (within 5%) and Flesch-Kincaid grade level. SI is coded by two independent raters, and the test does not proceed on any category with Cohen's kappa below 0.60 until that category is redefined.
Stage 2 is the generative test. For each Objective Projection Matrix underlying the pairs, models receive an explicit shown-mode instruction and produce text. The prediction is that model SI will be significantly lower than the human shown target and closer to the told baseline. The falsifier is that if model SI does not differ from the human shown target, the generative form is not supported.
Stage 3 is the core evaluative test. Each told-shown pair is presented to models in a judge role, asking for intensity and quality scores, with pair order counterbalanced. Human raters make the same judgments. The prediction is that models score the told member higher significantly more often than humans do, with a medium or larger effect (Cohen's h of at least 0.3 or odds ratio of at least 2). If models do not prefer told mode more than humans do, the construct collapses into general weak literary sensitivity and is withdrawn.
Stage 4 is a discriminant check. It re-runs Stage 3 with deliberately length-mismatched pairs to confirm the told-preference is not length-driven, and with varied prompt framing to rule out sycophancy. A told-preference that survives both is dissociated from the two nearest known biases. A told-preference that vanishes under either is reclassified as that known bias, not as summarization bias.
What the Paper Is and What It Is Not
This is important to state clearly: this paper validates nothing. It is a conceptual framework plus a registered protocol. The author says so explicitly and repeatedly. No data has been collected under this protocol. The construct is pre-validation, the definitions are provisional, and the decision rules are fixed in advance to prevent post-hoc construct-fitting.
The paper also sits within a specific theoretical tradition. The Bulut Doctrine, Narrative Entropy, and Objective Projection are the author's own framework. The paper cites the author's own prior work extensively. This is not unusual for a conceptual framework paper, but it means the construct's theoretical grounding is internal to one researcher's program rather than established across multiple independent lines of work.
The scope of any confirmation would be bounded by the model families, languages, and genres tested. The paper does not license a universal claim about LLMs. The test corpus in the reliability study was Turkish, and the pre-registered protocol does not specify which languages or genres will be used for the matched pairs. Whether the phenomenon generalizes across languages and literary traditions is an open question the protocol does not address.
Why This Matters Despite the Uncertainty
The practical stake is real even if the construct is unvalidated. LLMs are already deployed as reward models and judges in preference optimization pipelines. If there is a directional bias toward told mode in those systems, it would degrade prose quality in any application where the judge influences the training signal. The author argues this is the more consequential regime, and the argument is plausible: a generative limitation affects what models produce, but an evaluative limitation affects what all models are trained to produce.
The paper also contributes a methodological template. Pre-registering a construct with explicit falsification rules before collecting data is good practice regardless of whether summarization bias survives. The protocol's structure, matched pairs, length controls, discriminant checks against known biases, and fixed decision rules, is a model for how to test claims about systematic model behaviors without falling into post-hoc rationalization.
Whether summarization bias turns out to be a distinct evaluator bias with a theoretical motivation or a relabeling of verbosity bias will be determined by Stage 3 and Stage 4 of the protocol. The paper's contribution is to make that question answerable and to commit, in advance, to reporting the answer whatever it is.