6 How Humanistic Scholarship Tests the Conditions Under Which an Interpretation Is Established — The Formation of Judgment and the Revisability of Conclusions
6.1 Humanistic Scholarship Does Not Refuse Evaluation
The distinctions of the preceding three sections do not require imagining humanistic scholarship as an activity naturally superior to technical evaluation. Historical interpretation, textual explication, and conceptual-history research likewise require a scope of materials, a basis in versions, norms of citation, peer review, and responses to counterexamples; a beautifully written article that cannot explain where its materials come from or what role its citations play cannot acquire recognition on the strength of its style alone. To treat humanistic scholarship as a testing ground does not mean that it rejects standards; it means that it more often exposes a relation that standardized evaluation cannot settle in one pass: materials may change the question, new contexts may change the boundaries of concepts, and competing interpretations may force researchers to decide anew what counts as relevant evidence.
In laboratory-style closed tasks, the question, the materials, and the scoring criteria can usually be specified first, and only then is the result judged for satisfaction. Humanistic research of course makes similar provisional specifications, but these specifications themselves often become objects of controversy. A difference between versions may dissolve the original “same text”; a neglected usage of a word may invalidate an apparently clear conceptual distinction; a new archival document may not add one more fact to an old interpretation, but change what the researcher takes to be the phenomenon in need of explanation. The openness here is not the absence of boundaries; it is that the boundaries can still be rearranged in contact with the object.
6.2 Why an Interpretation Should Not Be Settled Prematurely
This paper calls this state of remaining rearrangeable “delayed closure,” that is, not settling an interpretation prematurely: what it delays is not writing, but the moment at which an interpretation is made irrevocable.[^6] A conclusion may hold provisionally on the current evidence, may be cited, or may be used to pose the next question, but its material boundaries, competing interpretations, unhandled counterexamples, and withdrawal conditions must remain visible, and new evidence must genuinely be able to change it.
Delayed closure is therefore a combination of practices, not an index of time. Version comparison lets researchers know whether the text before them is the same; the return through footnotes re-exposes citations to sources and to readers’ different handling; conceptual discrimination keeps a familiar word from sliding silently across epochs; contextual rechecking keeps the present question from arbitrarily overrunning the historical object’s own questions; and the handling of counterexamples lets negative materials actually change the scope of the claim rather than merely appearing in a “limitations” section. The capacity to downgrade, suspend, and withdraw is the condition under which these practices have institutional consequences. A “reservation of judgment” without consequences may still be only part of the finished appearance of the output.
6.3 Leaving Room for Revision So That Judgment Can Form
The formation of judgment is not a matter of there first being a complete subject who then projects meaning onto passive materials. In working through materials, a researcher may find that the original question does not hold, or that a different set of concepts must be adopted; the community’s objections may deprive evidence that once seemed sufficient of its force. It is precisely in this process — affected by the object and by others, and continually changed — that a judgment gradually takes on a form that can be explained and revised. The “formation” spoken of here does not appeal to unverifiable inner processes; it refers to a relation that can be partially tracked in the public record: which materials entered the judgment, which materials were excluded, which counterexample forced the conclusion to be downgraded, and which bridging reasons remain provisional.
This also explains why technical processing can precede judgment but cannot replace it. OCR, chunking, retrieval, and ranking can make some materials appear earlier and make others harder to bring into view; a generative system can also rapidly organize local associations into a complete narrative. But research judgment is not merely choosing the smoothest line among candidate materials; it must also withstand the moments when materials fail to meet expectations, and allow the question itself to be rewritten. Delayed closure does not refuse technical processing; it requires that technical processing leave entry points for re-ranking, adding materials, and changing the claim. If a system preserves only the final coherent narrative and cannot say what materials would stop it, turn it, or withdraw it, then the formation of judgment has been concealed by the product.
6.4 Why Generative Systems Make the Problem More Visible
Generative systems can rapidly produce texts that pass within existing evaluation systems: they can organize facts, citations, paragraph structure, and a cautious tone into a single report, so that a stage judgment quickly acquires the appearance of completion. The normative risk here is not the empirical assertion that “machines necessarily produce middling texts,” but that institutions may prefer a form of result that is easy to compare, easy to disseminate, and easy to archive. When a qualified finished appearance is repeatedly treated as a stable proxy for research capability, researchers may reduce those practices that temporarily lower clarity but allow the question to be rewritten. What is ultimately crowded out is not just some erroneous step, but the space for posing a different question, keeping undecided materials in view, and admitting the failure of an interpretation.
The value of humanistic scholarship should therefore not be reduced to preserving handwritten drafts or maintaining a purely human style. A machine-assisted research process can entirely form responsible judgment, so long as it lets materials interrupt an existing narrative, lets counterexamples change evaluations, and keeps the judgment nodes and withdrawal authorities of different participants visible. Conversely, a hand-written text full of hesitations may hide all its qualifications. What leaving room for revision protects is the possibility that the exploratory process can be readjusted, not the purity of some medium or authorial identity.
The significance of this test lies precisely here: it brings back into view the temporality of formation that can be provisionally set aside in most evaluation environments. Scoring often requires that a judgment present itself at some deadline as a comparable result, and publication and dissemination require that the result bear a name that can be briefly restated; humanistic research, by contrast, often must maintain a working tension between completion and incompletion. When materials are insufficient, the most responsible result may be to shrink the claim; when a concept remains ambiguous, the most valuable advance may be to change the question rather than to complete the conclusion; when counterexamples have not been handled, suspending judgment is not an evaluative failure but a way of avoiding writing a candidate interpretation prematurely into public fact. If institutions reward only answers that have been closed, these actions will be recorded as insufficient efficiency, and interpretive standing will favor the texts that most easily present themselves as finished.
This does not mean that all uncertainty is worth preserving. Hesitation may indicate insufficient evidence, but it may also be an excuse for not having completed the necessary work; openness can likewise be used to evade responsibility for materials and reasons. Delayed closure must therefore attach to concrete, checkable actions: indicate which kind of material is still missing, state which counterexample remains unresolved, mark what evidence would downgrade the conclusion, and actually revise the text when new evidence arrives. Only then is delay not the indefinite postponement of standing, but a way of keeping the acquisition, maintenance, and withdrawal of standing explicable.
Code — and in particular AI-assisted coding practices that take run results as their main feedback, colloquially “vibecoding” — provides a reverse test. Compilation, execution, and testing can quickly close local functional questions, but they cannot in the same way close questions of requirements, architecture, and responsibility; from passing tests to a system ready for production, one must still pass through specification interpretation, untested conditions, security review, version records, and maintenance authority. The point of conversion here is therefore very clear: the problem of appearance has not disappeared; it has migrated from whether the code works to what overall standing can be inferred from this working result.
The difference between humanistic scholarship and code is therefore not a difference between substance and appearance. Humanistic interpretation lacks a single public runtime that can quickly and repeatedly exhaust its meaning-claims, and the cycles in which materials, versions, and counterexamples change judgments are usually longer; software engineering, through version control, code review, versioned testing, rollback, and maintainer institutions, embeds part of the requirement of re-examination in engineering practice. The latter set of institutions is not a sufficient guarantee either: only when they can actually change evaluations, block deployment, initiate revision, or locate consequences do they perform the function of delayed closure. This comparison is not a scale for ranking disciplines; it shows that the same problem of standing appears at different positions under different verification conditions.
The athlete example from the introduction can be seen more clearly within this comparison. A goal or a beautiful move can be confirmed on the spot, but cannot by itself prove a stable ability; as far as on-the-spot confirmation is concerned, the run result of code is similar. But there is a deeper difference between the two: the athlete’s conditions of formation are enforced by the body — training, failure, and correction cannot be generated or forged in the same symbolic medium as the performance — and the finished appearance and the process of formation are in principle inseparable. The artifacts, tests, comments, and commit records of code, by contrast, all reside in a medium that language models can natively generate, and the record of formation itself can become part of the finished appearance; the runtime environment only partially restores the enforcing function of the body, and it covers only the functional layer. Humanistic interpretation lacks even this layer of substitution; its resistance to divergence can come only from materials, counterexamples, community controversy, and the concrete practices discussed in this paper. Delayed closure thereby acquires its exact position: in the absence of body and runtime, humanistic scholarship, if it is to keep interpretive standing re-examinable, can only maintain this substitute structure on its own.
6.5 From Three Questions to a Repeatable Process of Testing
Placing delayed closure back into this paper’s three-layer framework makes its institutional function clearer. The formation layer requires that the record allow return; the evaluation layer requires that a local pass remain constrained by the contract, the use, and the permitted inferences; the responsibility layer requires that a judgment that has acquired standing remain connected to an identifiable structure of bearing. Delayed closure is not a fourth independent threshold, but the common condition that keeps the relations among these three layers reversible: records of formation can prompt re-evaluation, evaluative controversies can push a claim to be rewritten, and responsibility procedures can prompt a conclusion to be downgraded.
This cycle also shows why transparency by itself is not enough. Publishing more logs, if the logs can only prove that the system ran but cannot show how materials changed the judgment, cannot repair the formation layer; publishing scores, if the scores are still treated as general capability after leaving the original task, cannot repair the evaluation layer; listing a responsible person, if that person has no authority to revise or withdraw the judgment, cannot repair the responsibility layer. What delayed closure requires is relational revisability, not an increase in the quantity of information. It can be realized through version records, registries of counterexamples, revisions of the scope and basis of evaluation, and explicit withdrawal procedures, but these forms have interpretive significance only when they can genuinely change subsequent standing.
6.6 A Further Distinction Concerning Distributed Agency
A distinction must also be kept between this paper’s institutional question and broader theories of distributed agency in AI-mediated writing. This paper does not attempt to measure “genuine human agency,” nor does it construct an agency score capable of comparing different writing subjects. What this paper means by delayed closure is a concrete set of practices: whether the community can trace the conditions of formation, whether the evaluation contract can be re-examined in light of new materials and counterexamples, and whether the structure of responsibility can demand explanation, correction, and withdrawal.
This distance is also a self-limitation. If future research shows that a distributed system can preserve adequate records of formation, allow evaluation criteria to change, and bear the consequences of judgment, this paper has no reason to deny its interpretive standing merely because machine participation is involved; if a human author’s text has no structure of reasons open to re-examination, the signature alone cannot automatically satisfy the three layers of conditions. Delayed closure provides only a way of testing public arrangements; it does not provide a final answer about consciousness, personhood, or human uniqueness.
6.7 Conclusion: Keeping Conclusions Open to Revision
Humanistic scholarship is an especially telling test of this problem not because it enjoys a privilege of exemption from evaluation, but because it pushes to the foreground those relations in the formation of judgment that are easily concealed by the finished product: materials resist existing questions, contexts restrict the migration of concepts, counterexamples change conclusions, and researchers change their own understanding in the process. Generative systems make it easier for the finished appearance of the output to circulate first, and therefore make it all the more necessary to preserve the possibility of revision and withdrawal, lest the product be prematurely treated as an unchangeable conclusion.
What this paper means by delayed closure ultimately points to a limited normative requirement: stage conclusions may be advanced, but they cannot continue to circulate at the same strength after losing their material boundaries, evaluative uses, and structures of responsibility. It does not demand infinite slowness, nor does it treat errors, hesitations, drafts, or human signatures as guarantees of truth; what it requires is that when new materials, counterexamples, or community controversies appear, an interpretation still have a real path of re-examination. Section 6 converts this requirement into five more concrete public conditions, while preserving their character as non-sufficient guarantees.