2 From “Stochastic Parrots” to “An Interpretation Has Been Established”
2.1 A Question That Has Already Been Advanced
Academic controversy over generative AI often begins with a question that is at once forceful and easily closed: does a language model actually understand the text it generates? The value of the “stochastic parrots” critique lies in its refusal to infer meaning, communicative intent, and the capacity for responsibility directly from formal linguistic success. A model can generate coherent, grammatical, prompt-responsive text on the basis of statistical relations in its training corpus, but this formal success by itself cannot show that the text bears the meaning carried by human utterance, nor can it show who should be responsible for the judgments it contains. Bender and her co-authors (2021) at the same time direct attention to training materials, deployment environments, and relations of power: how meaning, knowledge, and responsibility come to be ascribed to machine outputs depends on a set of sociotechnical conditions, not merely on the scale of a model’s parameters.
This critique remains necessary, but it is insufficient for dealing with today’s long-form research outputs. If the question remains “does the model understand,” all research risks are easily compressed into whether some capability exists inside the machine; if the question is changed to “does the text contain errors,” the selection of materials, the relations among sources, and the inferential processes that occur before an interpretive judgment is decomposed are hidden from view. This paper starts from a fact that technical research has already established: the statements, sources, and long-form organization of generative systems can each be evaluated separately. The question thereby shifts from whether the model possesses meaning to what entitles a finished text to be recognized.
2.2 From Hallucination to Sources and Support Relations
The first object to come into view was hallucination. Hallucination research singles out the relation between a statement and external facts, requiring that one judge whether a sentence can be supported by available evidence; source-attribution research goes further, asking whether the listed literature actually exists, whether it in fact supports the statement, and whether the citation is placed in a position commensurate with its role. That a page really exists does not mean that the page supports every judgment attached to it; that a statement is broadly correct does not mean that its chain of sources is complete. Retrieval-augmented generation and the corresponding RAG evaluations place retrieved materials, candidate evidence, and generated statements within a single framework of checking, so that omissions, mismatches, and insufficient support can be identified item by item.
This work has changed the meaning of “factuality.” A fact is no longer a reader’s overall impression of an article; one must first decompose the statements, confirm the role of each piece of evidence, and then judge whether, within the given range of materials, the support is sufficient. The relational checks established by ALCE (Gao et al. 2023), RAGChecker (Ru et al. 2024), and research on source fidelity (Rashkin et al. 2023) can show that an answer’s citations fail to support the corresponding claims, or that the retrieved context omits decisive materials. They therefore constitute an indispensable foundation for research with generative systems. Yet the verifiability of source relations still only answers the question “what is the relation between this sentence and these materials”; it has not answered “why was the question formed in this way,” “which conceptual distinctions changed the judgment,” or “did competing interpretations have a chance to defeat the original conclusion.”
2.3 Long-Form Text Makes Local Strengths and Weaknesses Easier to See
As systems began to generate long-form reports, a single rate of factual correctness could no longer bear the entire burden of evaluation. LongFact decomposes long answers into independently verifiable factual claims, SAFE checks these claims through search and item-by-item inspection, and VeriScore takes verifiable claims as the unit of analysis (Wei et al. 2024; Song et al. 2024). These methods do not claim that a report consists only of facts; their role is to separate out, in advance, the checkable relations otherwise blended into fluent narration, so that researchers know which sentences are stating facts and which are advancing inferences that go beyond the facts.
Coverage and report structure thereby became another set of objects of evaluation. ICAT attends to whether an answer addresses the aspects prescribed in advance by the task (Samarinas et al. 2025); DeepResearch Bench goes further, incorporating source quality, synthesis of materials, causal connection, multiple perspectives, comparative explanation, and insight into expert criteria (Du et al. 2025). ResearchRubrics acknowledges that deep-research answers may have multiple valid paths, does not require a system to reproduce a single reference text, and instead uses fine-grained criteria to judge whether factual grounding, reasoning, and expression meet the requirements (Sharma et al. 2025). ReportLogic takes as its object of judgment whether readers can follow a report’s analytic arc, its context, its claim–support relations, and its bridging reasons, and it uses adversarial examples to show that adding headings, numbering, summaries, or “evidence” labels may improve surface organization without necessarily adding underlying support (Zhao et al. 2026).
It is important to acknowledge that this work carries real information. It has neither crudely reduced complex research to error-checking nor abandoned evaluation because open-ended tasks are hard to exhaust; it can show on which dimensions of factuality, sourcing, coverage, reasoning, or expression a system has improved, and it can help researchers compare the local differences among systems on the same task. The evaluation contract discussed later presupposes precisely that such local evaluations can hold. The question that remains to be added is this: once a local evaluation has been confirmed, under what conditions is the result understood as an interpretation that has been formed?
2.4 Open-Ended Tasks Also Need Their Scope Stated First
DeepResearch Bench II displays this duality in open-ended tasks most clearly (Li et al. 2026). It extracts facts and inferences from survey articles written by experts and, through automatic extraction, manual cleaning, and revision by domain experts, forms fine-grained evaluation criteria; some of the expert articles are designated as shielded sources, and the system must reconstruct their content from other materials. Such a design allows answers to remain open while still specifying which facts must appear, which inferences count as non-trivial analysis, and which sources may not be reproduced directly. What the system is asked first is whether, given the question, the materials, and the criteria, it can find the necessary evidence and establish the corresponding relations; it has not yet been asked whether it can re-specify the question, change the identity of the evidence, or propose a different scope and basis of evaluation.
This is not an accusation against the benchmark but an accurate description of what it measures. The expert articles and fine-grained criteria make comparison possible and provide localizable failure information for system improvement; at the same time, they provisionally convert openness into a set of checkable task conditions. ResearchRubrics permits multiple valid paths, ReportLogic attends to the internal relations of a report, and DeepResearch Bench II prevents simple reproduction by shielding sources (Sharma et al. 2025; Zhao et al. 2026; Li et al. 2026); the differences among the three show precisely that “openness” does not mean the absence of boundaries. An answer can display very high quality within a task and still have obtained only a local pass within that task.
2.5 Local Evaluation Cannot by Itself Show That an Interpretation Is Established
From the discussions of construct validity and benchmark design, any score must answer two questions: what exactly it measures, and which interpretation and which decision it is used to support. The work of Raji and her co-authors on benchmark audits (Raji et al. 2021) reminds us that an evaluative construct cannot be detached from the goals of the task and the actual scenarios of use; Schlangen’s critique of language evaluation (Schlangen 2023) likewise points out that the comparability of scores is not equivalent to the construct having been adequately confirmed. This paper accepts this methodological vigilance: it neither writes local evaluation off as a meaningless proxy, nor writes open interpretation off as arbitrary activity without standards.
Likewise, the neighboring research on Proxy-Based Evaluation has made it clear that performance metrics, benchmarks, transparency documentation, and auditing procedures are all proxy evidence for governance goals; even when the proxies are satisfied, distributional shifts in deployment, hidden harms, and runtime failures may remain undiscovered. This paper follows the demarcation that “a proxy is not the target” and goes on to trace a more specific cross-vehicle question: after a proxy result leaves its task, is it rewritten by papers, platforms, registries, or dissemination materials into “the interpretation has been established” or “the system already possesses general research capability”? Within the original evaluation scope, a local pass may be entirely genuine; the problem occurs when the result is endowed with a larger identity without the addition of commensurate materials and bridging reasons.
2.6 The Difference Between Local Testability and Overall Interpretation
The current literature has therefore established a floor that cannot be bypassed: without minimal testability of facts, sources, coverage, and report logic, no research output should ask a community to treat it as an interpretation. But this floor is not a sufficient condition. A report can pass every prescribed item of statement, citation, coverage, and structure and still have completed only a checkable answer within a bounded task; it may not have explained why the materials belong to the relevant scope, may not have engaged evidence that runs against its main line, and may not have displayed the bridge from local relations to an overall interpretation. If these limits are preserved, a local pass can function accurately; if they are omitted in subsequent circulation, the finished appearance of the output may acquire standing ahead of the process of its formation.
The research object of this paper is thereby confined to a stretch of relations that technical evaluation has opened but not yet closed: rather than inventing a replacement concept for meaning, factuality, sourcing, or report logic, it examines how these local relations are organized, named, and carried into a community’s interpretive judgment. The next section will define the divergence between the finished appearance of the output and the process of its formation as “interpretive appearance,” and will explain how it differs from hallucination, plausibility, trustworthy interfaces, and illusions of understanding. Only after this formation layer has been distinguished can one further ask how the evaluation contract translates a local pass into interpretive standing.