OER·harvester

← Back to the library
arXiv HTML resource

When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI

Generative AI research has increasingly evaluated factuality, citation, coverage, and report structure. Yet passing such local checks does not by itself show that a humanistic interpretation has been established. This paper asks how an interpretation comes to be recognized within sociotechnical processes. It introduces three connected concepts. Interpretive appearance names the gap between the finished form of an ou…

Licence
OPEN CC-BY-4.0
Authors
Deyu Jing
Published
2026-09-04 · arXiv
Language
en
Length
19362 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

1 Introduction — When Does Interpretive Appearance Cease to Suffice as Evidence That a Judgment Has Been Formed

In competitive sports, spectators ordinarily do not conclude from a single brilliant move that an athlete already possesses a stable ability. A performance can count as evidence of ability because it is usually connected, in a traceable way, to training, failure, rules, and testing by a community; this connection does not guarantee that every performance is correct, but it allows the performance to be placed within a longer relation of judgment. Generative artificial intelligence is changing the visibility of this connection: finished text can circulate before the process of its formation has been accounted for and before the boundaries of its evaluation have been preserved.

Current academic controversies over generative AI have formed two important but individually limited paths. One asks whether language models understand, whether they have intentions, or whether they can bear responsibility. The other decomposes outputs into facts, sources, verifiability, coverage, reasoning, and report structure, and confirms local quality through finer-grained evaluation. The latter path has made it possible to judge long-form research outputs by more than overall impression, and has made it possible to locate errors, mismatched citations, and insufficient support. Yet even when an output passes established checks of factuality, citation, coverage, and logic, one question does not thereby disappear: when are these local passes sufficient for a community to recognize that an interpretation has been formed?

The question can be presented through a minimal counterexample that does not depend on the inner states of any model. Suppose that every citation in a historical report is genuine, that every page can be linked back to its source, that factual verification and source attribution have both passed, and that the report’s structure satisfies the required claim–support relations; but the report’s key causal judgment merely repeats a piece of secondary scholarship, without reworking the counterevidence in the primary materials and without explaining why the secondary author’s generalization applies to the present problem. If this report is subsequently named by papers, platforms, or institutional registries as “a historical interpretation that has been formed,” its defect cannot be exhausted by the categories of “hallucination” or “citation error.” The local evaluations may be genuine; the problem occurs where a local pass is endowed with a stronger interpretive standing.

This paper accordingly poses a narrower research question: how does an interpretation, presented in the appearance of a finished product, acquire the standing of being “already established” in sociotechnical circulation, and how can this standing come apart from conditions of formation, evaluation boundaries, and responsibility for judgment? Interpretive standing here is not a synonym for interpretive truth, nor is it an inference about whether an author inwardly read, reflected, or revised. What this paper examines are public conditions: whether readers can return to the materials and versions; whether they can distinguish facts from inferences; whether they can see how counterexamples changed the judgment; whether they can know under what circumstances a conclusion should be downgraded or withdrawn; and whether there exists an agent or structure that can give reasons and bear consequences.

The argument of this paper proceeds on three interconnected but mutually irreplaceable levels. First, building on Adhya’s pedagogical observation of “interpretive appearance” in the classroom, we translate it into a question of formation at the level of the research community, and we define interpretive appearance as the divergence between the finished appearance of the output and publicly traceable conditions of formation. This concept does not simply relabel fluent text as mere appearance, nor does it claim that all research using generative systems lacks a process of formation; it examines only whether, when a text asks its readers to believe that an interpretation has been settled, the public record provides conditions of formation commensurate with that demand.

Second, this paper proposes the working term “evaluation contract,” which refers to the scope and basis of an evaluation. It explains why local evaluation must be provisionally closed within bounded materials, tasks, scoring, permitted inferences, and failure conditions, and it proposes “standing substitution” to describe a specific conversion from one judgment to another: a genuine local pass is treated by later papers, platforms, or dissemination materials as sufficient warrant for a stronger claim of interpretation or achievement, without commensurate new evidence or bridging arguments. Treating a local pass as if the whole were established is not a rejection of evaluation, nor another name for inadequate measurement in general; it requires that the original evaluation be genuine and valid, that the subsequent claim be stronger, that the original evaluation be used as the key warrant, that no commensurate new support has been added, and that the original boundaries and conditions for downgrading no longer constrain the subsequent, stronger claim.

Third, this paper turns the question of responsibility away from the ontological controversy over “whether machines can be responsible” and back toward the public practical relations of judgment. A generated text may acquire recognition while the judgment as a whole still has no identifiable attribution, no statable reasons, no procedure by which counterexamples can change the judgment, no executable path of downgrading or withdrawal, and no person or explicit structure that bears the consequences. The “delayed closure” proposed on this basis is not a matter of the slower the better; it means postponing the moment at which an interpretation is made irrevocable, and keeping the record of formation, the scope and basis of evaluation, and the relations of responsibility open to re-examination, so that revision and withdrawal remain possible. Section 6 will translate this requirement into five public conditions, while acknowledging that they only increase the likelihood of discovery, contestation, correction, and withdrawal, and do not guarantee that an interpretation is true.

The paper unfolds as follows. Section 1 reviews the progression from the stochastic-parrots critique, through hallucination and source attribution, to the evaluation of long-form research, and argues that local verifiability is a necessary floor for the recognition of an interpretation, not a sufficient condition. Section 2 poses the problem of formation and delimits it against hallucination, plausibility, trustworthy interfaces, capability inference, and illusions of understanding. Section 3 defines the scope and basis of evaluation and discusses warranted elevation of standing together with two strict counterexamples, in order to show the boundaries of “treating a local pass as the whole being established.” Section 4 turns to responsibility, reconstructing responsibility for judgment through footnotes, context, formation, and practical standards, while keeping its distance from neighboring problems such as the responsibility gap. Section 5 takes humanistic scholarship as the testing ground, shows how materials and counterexamples can rearrange questions and criteria, and defines delayed closure. Section 6 then converts the analysis of these three levels into several concrete requirements: that one can return to the original materials and the original judgment, that different evidential levels can be distinguished, that failure can be acknowledged, that revision is possible, and that a bearer of responsibility can be found.

To avoid conceptual conflation, this paper draws the following distinctions. “Cross-level translation” is the neutral process by which an evaluation result enters a new vehicle; “warranted elevation of standing” means that a subsequent, stronger claim has acquired commensurate new materials, independent review, and bridging arguments; “standing substitution” is the improper form of this process. What this paper means by an interpretation being recognized is, at a minimum, that some vehicle permits an output to serve as an acceptable premise for subsequent judgment, evaluation, citation, teaching, resource allocation, or public knowledge activity. This does not mean that the interpretation is true, nor does it require that the community has reached a final consensus.