OER·harvester

← Back to the library
arXiv HTML resource

When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI

Generative AI research has increasingly evaluated factuality, citation, coverage, and report structure. Yet passing such local checks does not by itself show that a humanistic interpretation has been established. This paper asks how an interpretation comes to be recognized within sociotechnical processes. It introduces three connected concepts. Interpretive appearance names the gap between the finished form of an ou…

Licence
OPEN CC-BY-4.0
Authors
Deyu Jing
Published
2026-09-04 · arXiv
Language
en
Length
19362 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

5 Who Explains and Answers for the Judgment — Interpretive Credit and Responsibility for Judgment

5.1 What Remains After an Interpretation Has Been Recognized

The standing substitution discussed in Section 3 concerns how a local evaluation, in cross-vehicle circulation, comes to be endowed with a stronger interpretive status; it does not automatically follow that a structure of responsibility is missing. A researcher may, after a bounded evaluation, add materials, engage counterexamples, and make the bridging reasons public, so that the original local pass becomes one link in a stronger judgment. Conversely, even if an interpretation has been provisionally recognized by journals, platforms, or peers as citable, teachable, or usable as a basis for further work, the interpretive credit it carries may no longer be connected to a person or structure capable of answering for the overall judgment. What this paper means by “interpretive credit” is only the epistemic and institutional credit that a text or a signed achievement acquires because it is regarded as “an interpretation that has been formed”; it is not a synonym for interpretive truth, nor is it identical with the long-term prestige acquired by an author.

The question here is therefore neither whether a generative system can become a moral agent, nor a post hoc search for an individual who can bear blame. A judgment in humanistic scholarship can acquire provisional authority usually also because it is placed within a practical relation to which one can return: someone can indicate the scope and versions of the materials, explain what role a given piece of material actually plays in the argument, and account for the intermediate reasons leading from materials to concepts and from concepts to historical generalization; others can accordingly advance contrary materials, different contexts, or competing interpretations. If this relation is deleted, the finished product may remain coherent and the genuine citations may still exist, but what the community receives is a judgment already completed, not a judgment whose formation, reasons, and limits can still be queried.

A distinction should be drawn here between “there were in fact people behind the conclusion” and “the relations of responsibility have been made public.” A paper produced by multiple collaborators, through model generation and editorial processing, may of course have many actual participants; the question is whether readers can tell which agents or institutional links are responsible for which judgments, whether they can demand reasons from them, and whether they can know under what counterexamples, version problems, or insufficiencies of evidence the conclusion will be downgraded, rewritten, or withdrawn. What is to be identified here is not some hidden psychological fact, but whether this relation has, in the public record, an identifiable form sufficient to support interpretive credit.

5.2 Footnotes, Context, and the Constraint of Materials

Grafton’s account of the history of the footnote provides a relatively concrete point of entry. The footnote has never been merely a typographic device that decorates existing knowledge into professional knowledge: it connects the author’s assertions to materials that can be sought out, checked, and differently interpreted, and it places the author under the continuing scrutiny of peers and readers. The early historiographical controversies Grafton cites make this especially clear: listing “proofs” does not oblige readers to agree with the author; precisely because the materials are indicated, readers may treat the same materials in different ways and demand that the author revise. The value of the footnote therefore lies not in guaranteeing truth, proving how much labor the author actually expended, or replacing complete argument, but in making the assertion a challengeable commitment: it must accept questions about whether the sources genuinely support it, whether omissions change the conclusion, and whether citations have been selectively used.

This also shows that the mere existence of footnotes cannot satisfy the conditions of responsibility. Footnotes can be piled up, can display only the range of one’s reading, and can miswrite secondary generalizations as primary evidence; a complete citation interface may even strengthen the appearance that an interpretation has been completed. Only when readers can follow the footnotes back to the materials, and when this return can change the acceptability of the argument, do footnotes participate in responsibility for judgment. This paper does not understand such a return as the recovery of an absolutely infallible evidential foundation; it requires only that the judge cannot, after the citations have been challenged, continue to maintain the conclusion as an established fact unaffected by the challenge.

Skinner’s critique of method in intellectual history gives this point another qualification. He opposes treating the text as a self-sufficient object, and he also opposes treating the macro-social context as a background that can directly determine the text’s meaning; to understand an utterance, one must examine the speech acts conventionally performable on that occasion, and make the reconstruction of vocabulary, questions, audiences, and modes of expression a constraint among competing interpretations. This paper draws no ontological conclusions about artificial intelligence from this; it borrows the account only to show that humanistic interpretation does not confront a textual surface that can be filled in at will. The dating of materials, the availability of words at the time, the historical position of the topic, and neglected alternative formulations can all deprive an apparently smooth generalization of its conditions of holding. Context here is not supplementary background material; it is the constraint that materials impose on judgment, and when new contextual evidence appears, the original judgment should be open to change.

5.3 Judgment Is Not Shown by Unverifiable Inner Processes

If responsibility is understood merely as whether the signatory can answer questions after the fact, the temporality of the formation of judgment is still missed. At this point the paper makes limited use of Gadamer: understanding is not a one-directional act in which an already completed subject imposes meaning on an object; the process of understanding also changes the questioner’s own ways of asking and anticipating. This dimension of formation does not require researchers to disclose all their private reflection, still less does it equate visible process directly with genuine understanding; it only suggests that the person or structure capable of answering for a judgment should allow materials, discussion, and counterexamples to change the conclusion it sets out to defend.

The internal standards of practice emphasized by MacIntyre supply the public face of this point. Competence is not an endowment possessed once and for all, detached from practice; it is formed gradually within the standards of a shared activity, the experience of failure, and the testing of a community. This paper does not turn this view of practice into a threshold of “only those unassisted by AI can form judgment”; human–machine collaboration can likewise form a structure of responsibility within a clear division of labor, community testing, and withdrawable conclusions. What it rejects is only another over-hasty inference: that so long as a text is formally proficient, or so long as someone signs the cover, one can dispense with an account of how the judgment acquired its standing through materials, failures, and revisions.

Gadamer and MacIntyre perform only one delimiting task in this paper: the former shows that understanding in formation is reshaped by its object and by the process, and the latter shows that the capacity for judgment must undergo failure within testable practical standards. Neither can adjudicate whether machines understand or whether they can be responsible, nor can either prove that any particular user has not understood. What this paper needs is not a chain of reasoning from philosophical anthropology to a technological prohibition, but a narrower institutional judgment: if interpretive credit is to continue to be recognized, the fact that a judgment, in the process of its formation, remains open to being changed by materials, discussion, and counterexamples cannot wholly disappear from public relations of responsibility.

5.4 Why Errors, Hesitations, and Drafts Matter

Adhya’s (2026) observation of the classroom is directly, though limitedly, instructive here. She notes that when learning is compressed into a submittable product, fluency, orderliness, and finish may conceal the fact that reading, discussion, and revision have not occurred; and the errors, hesitations, drafts, annotations, and discussions of in situ writing are worth including in evaluation not because imperfection is inherently more authentic, but because they sometimes preserve traces of a judgment still in formation, still open to being changed by others. This observation cannot be directly transferred as an empirical conclusion about the research community: drafts can be forged, hesitation can be mere irrelevant delay, and a refined final draft may equally well rest upon rigorous revision.

This paper therefore does not take “leaving traces of process” as a sufficient condition for an interpretation to be established. Process materials are relevant to responsibility only when they can show how materials entered the judgment, how counterexamples were handled, which intermediate judgments were abandoned, and what conditions would trigger downgrading or rewriting. What Adhya (2026) resists is the complete replacement of the learning process by the product; what this paper further asks is whether, when the product leaves the classroom and enters papers, databases, platforms, or public circulation, there is still some public structure by which interpretive credit can be reconnected to the process in which the judgment was formed. The two are connected, but the pedagogical observation of process is not thereby expanded into suspicion of all academic texts.

5.5 Who Explains the Judgment, and Who Bears the Consequences

This paper uses “responsibility for judgment” as the general term, but does not conflate it with answerability or accountability. The former means that the bearer can be required to state the reasons for a judgment and to respond to materials, inferences, and counterexamples; the latter means that such accounting, correction, downgrading, withdrawal, and the bearing of consequences have been embedded in some executable institutional relation. Responsibility for judgment is the wider practical relation: who advances and maintains a judgment, on what reasons it is maintained, which counterexamples can change it, and who undertakes the corresponding actions when the judgment needs revision.[^4]

On this basis, the third level can be defined as follows: an output, even after it has acquired recognition, may have no public, identifiable structure of responsibility through which its overall judgment can be explained, questioned, revised, or withdrawn. In other words, An output may acquire interpretive standing even when no publicly identifiable structure of responsibility exists through which the judgment as a whole can be justified, challenged, revised, or withdrawn. The divergence here occurs between interpretive credit and responsibility for judgment, not between two ontologies called “human” and “machine.”

The five conditions of the responsibility layer jointly specify a re-examinable relation of responsibility. Different disciplines may adopt different procedures of material preservation, authorial division of labor, anonymous review, and errata, and an exploratory judgment may be provisionally advanced before the evidence has reached its strongest state.[^5] The problem is this: a conclusion has acquired stable credit through publication, platform, or institutional endorsement, yet readers cannot discern how responsibility is allocated, cannot demand reasons from the bearer, and cannot make new materials actually change that credit. The responsibility conditions do not freeze an interpretation into a qualified product; they preserve the relations in which an interpretation, after acquiring standing, can still be re-examined.

The core of the responsibility relation consists, first, of three conditions: the judgment must be attributable, the reasons must be statable, and counterexamples must be able to change the judgment. First, the judgment must be attributable. A signature, a project name, or a model name is by itself insufficient to complete attribution; readers should be able to tell whether the key judgments were advanced and maintained by an individual, a team, journal editors, a research institution, or a human–machine collaborative structure with an explicit division of labor. Second, the reasons must be statable. The bearer need not hand over every private note, but should be able to state the scope of materials, versions, the roles of evidence, the bridging inferences, and the key alternative paths not adopted. Third, counterexamples must be able to change the judgment. If feedback can only be collected without being able to affect the argument, the evaluation, or the conclusion, so-called responsiveness is merely a form.

Beyond these three core conditions, two executive conditions are needed for responsibility to take effect institutionally. Fourth, the conclusion must be capable of being downgraded or withdrawn. This does not require researchers to suspend every provisional judgment indefinitely; it requires that versions, qualifications, and withdrawal conditions remain visible while the conclusion still has practical effect, so that “candidate leads,” “bounded generalizations,” and “interpretations already defeated by counterevidence” do not continue to circulate with the same standing. Fifth, there must be a person or an explicit structure that bears the consequences. The consequences here include scholarly errata, reputation, and re-evaluation, and also the answerable effects produced when a judgment is used in curricula, research decisions, or public discourse; this does not presuppose that all consequences should be borne by a single individual, but it excludes the situation in which responsibility is diluted without limit among system, platform, authors, editors, and institutions.

What the five conditions jointly specify is the practical form of responsibility, not an individualist model of ownership. An individual author can satisfy it; a research team can satisfy it through a clear division of labor; journals, laboratories, or archival institutions can bear some of its conditions; in human–machine collaboration, the system can participate in retrieval, generation, comparison, and revision, while it must still be made clear who is responsible for explaining the overall reasons, who has the authority to revise or withdraw, and what procedure enables counterexamples to change the judgment. The bearers of responsibility may be plural, but the structure of responsibility must not thereby become unidentifiable. Conversely, if a team clearly marks, in the public record, the role of the model, the boundaries of the materials, the points of human judgment, and the withdrawal procedure, the use of a generative system by itself does not constitute a deficiency of responsibility.

The re-examination of code usually occurs not at the first run, but in maintenance, requirements changes, dependency upgrades, vulnerability disclosures, production incidents, and system migrations. What then needs to be re-explained is not only whether a function returns the expected result, but also why this architecture was chosen in the first place, which evidence would force it to change, and who has the authority to modify, suspend, roll back, and withdraw. A legacy system that runs but that no one can explain and no one can safely modify displays precisely the code variant of the responsibility layer: the artifact has acquired institutional standing, while the overall judgment lacks a public structure of responsibility through which it can be explained, questioned, revised, and withdrawn. A successful run has not eliminated the problem of responsibility; it has merely postponed the moment of re-examination to the stages of maintenance and incident response.

5.6 The Distance from Neighboring Responsibility Problems

This definition differs, first, from the question posed by the responsibility-gap literature. The classic formulation of Matthias (2004) and the extended analysis of Santoni de Sio and Mecacci (2021) mainly address how fault, accountability, and active responsibility should be ascribed after the behavior of learning automata and its adverse consequences have occurred. This paper of course borrows their wariness toward a single answer to responsibility ascription, but it does not discuss the moral and legal imputation of accidents, harms, or automated behavior. What this paper asks about is another moment in time: after an interpretation has acquired standing in academic or public circulation, who is responsible for its structure of reasons as a whole. The former can occur in automation scenarios in which no interpretation is recognized at all; the latter can occur in academic writing in which there is no imputable harmful event. The two are structurally similar but cannot replace each other.

A further self-limitation concerns agency and authorship. This paper does not measure genuine human agency, construct an agency score, or reduce distributed scholarly work to a single authorial signature. Its object remains narrower: whether an interpretation in circulation has a public structure of reasons, revision, and withdrawal. Authorship and signature therefore show neither that responsibility is identical with single authorship nor that a signature alone completes the responsibility of judgment.

These two sets of distinctions at the same time delimit this paper’s diagnosis. The responsibility layer does not consist in discovering that no human being participated in the interpretation, nor does it infer from a finished text that some author in fact did not read, reflect, or hesitate. It requires only that, on occasions where the community grants credit to an interpretation, the public record can answer five practical questions: who, or what structure, advanced the judgment; on what reasons it is maintained; which counterexamples can change it; when it can be downgraded or withdrawn; and who bears the ensuing consequences. If these questions cannot be answered at the public level, there is reason to say that interpretive credit and responsibility for judgment have come apart; if they can be answered, the interpretation may still be wrong, but it should not be diagnosed as a deficiency of responsibility merely because generative tools were used.

5.7 Conclusion: Responsibility Is Not a Label Attached to a Finished Product

The reason interpretive standing cannot be automatically granted by textual finish, genuine citations, or local evaluation scores is not only that these indicators may still miss errors. The deeper difficulty is that they can allow a judgment to circulate before any public structure of bearing has been established: what readers see is an interpretation already completed and thus seemingly already answered for by someone, while what is actually visible may be only the product, the signature, or the evaluation label. The footnote’s commitment as revealed by Grafton (1997) and the contextual constraint demanded by Skinner (1969) remind us that an interpretation must face the counter-pressure of materials and readers; Gadamer (2004), MacIntyre (2007), and Adhya (2026) respectively show, from the formation of understanding, practical standards, and traces of process, that a judgment is not a static product that can be wholly separated from the possibility of being changed.

A structure of responsibility can exist while the judgment remains wrong, and footnotes, version records, and community procedures may be executed only formally. What is required here is that interpretive credit be connected to the structure of bearing described in this section; where that structure is missing, it should not be treated as an already stabilized interpretive standing. The next section will explain why humanistic scholarship makes this requirement especially visible: what it needs is not the indefinite postponement of conclusions, but the capacity to keep the closure of an interpretation open to re-examination by materials, concepts, and counterexamples.