7 Five Public Requirements — From Principles to Practices
7.1 Making “Revisability” a Public Practice
The following five requirements are not a uniform technical checklist that all disciplines must adopt; they are minimum requirements commensurate with the risk of the judgment. They can be realized through different forms — footnotes, version notes, research logs, editorial procedures, or platform markers; the forms may differ, but they must bring key materials, bridging reasons, limiting conditions, and revision authority back within the community’s range of inspection.
If the delayed closure discussed in Section 5 remains at the level of an attitudinal expression that “research should stay open,” it is still insufficient to change how an interpretation circulates in the community. A system may declare in its documentation that its results are provisional, and a paper may list limitations at its end, but if readers cannot know where the materials came from, what translations the judgment passed through, or when the conclusion should be downgraded, the professed openness will impose no actual constraint on the interpretive credit already acquired. Re-examination is therefore not an ethical reminder appended to the finished product; it is the preservation, within the result and its subsequent papers, platforms, or dissemination materials, of the re-examinable relations among formation, evaluation, and responsibility.
The minimum objects of the public conditions are the records directly relevant to interpretive standing that the community can return to: material boundaries and versions, the roles of evidence in the argument, the translation between model-generated content and the researcher’s judgment, the applicable scope of the evaluation contract, counterexamples and failure states, and the structure of responsibility with authority to revise or withdraw the conclusion.[^7] So long as these relations remain traceable and can actually take effect when necessary, the records may take different technical and disciplinary forms.
7.2 Requirement One: State Which Materials and Which Versions Were Used, and Whether Others Can Access Them
An interpretation cannot carry only a string of source names; it must also state which materials actually entered the judgment, in which versions the materials stand, and whether readers, at the time the judgment was made, could return to them under the same or comparable conditions of access. “Scope” includes not only the literature, archives, corpora, and data that were included, but also the categories of excluded materials and the reasons for exclusion; “version” includes not only file dates, but also revisions, translations, abridgments, scan quality, database updates, and retrieval times — conditions that change the identity of the evidence; “conditions of access” concern permissions, link rot, regional restrictions, changes in retrieval interfaces, and whether the model could actually read the materials. Without these qualifications, the same statement “cited a certain work” may represent different evidential relations at different times, in different versions, and under different conditions of availability.
This requirement does not mistake repeatable access for a sufficient foundation of interpretation. A fully open dataset can still be misunderstood, and an archive with complete versions still requires contextual and conceptual judgment. What it prevents is another, more basic divergence: the product retains the names of sources while readers cannot know how the material boundaries affected the conclusion, or cannot tell whether a later version has replaced the original basis. For generative systems, the minimal conditions of retrieval results, context truncation, and model version should also be recorded; if these cannot all be disclosed for reasons of privacy, copyright, or security, one should at least state which inferences lose traceability because of the missing parts, and correspondingly lower the strength of the claim.
7.3 Requirement Two: Record Facts, System Outputs, and the Researcher’s Interpretation Separately
Re-examination also requires that different evidential levels be kept separate in the record. A textual fact is a statement locatable in the materials; a model result is a candidate output produced by the system under specific prompts, contexts, and parameters; a textual interpretation is an argument about words, structures, and relations of meaning; a historical generalization is an induction about relations across texts, periods, or practices; and a causal mechanism further claims how a relation arises, is maintained, or changes. These may be interrelated within the same study, but they do not acquire the same interpretive standing merely because they are smoothly written together in the final paragraphs.
The function of this distinction is not to prescribe a rigid pipeline, but to make cross-level bridging a checkable object. A model’s indication that two words co-occur in a corpus can serve first only as a model result or a candidate association; if the researcher wants to write it as a textual interpretation, they must explain how word senses, genre, and context support this relation; if it is further generalized into a historical mechanism, they must also account for chronology, alternative materials, and the inference from association to mechanism. Each translation may change the scope of the question and the burden of evidence, and should therefore leave, in the version record or an argument map, nodes from which one can return to the original materials and the original judgment. Conversely, labeling a model-generated historical generalization as a “fact,” or placing a model inference that the researcher has not rechecked at the same level as a quotation from the source text, lets the finished appearance of the output conceal the conditions of formation.
Layering does not require pre-separating human judgment and machine output into two pure blocks. A researcher may re-retrieve materials on the basis of candidate leads proposed by the model, and a model may complete summaries and comparisons within a conceptual framework specified by humans; what genuinely needs to be marked is which judgment was advanced by which materials and which participant, and which step remains only a bridge awaiting testing. In this way, machine participation by itself is not treated as disqualifying, nor is the coherence of machine prose treated as a completed interpretation.
7.4 Requirement Three: Failure, Downgrading, Suspension, and Withdrawal Must Actually Work
If an evaluation regime can produce only “pass” or “generation complete,” any promise of openness is easily flattened in circulation. Re-examination requires that states such as miss, insufficient evidence, unresolved counterexample, inaccessible materials, task mismatch, and inability to judge be written as formal results, not as noise awaiting cleanup. A system’s failure to find materials supporting a claim may mean that retrieval failed, or it may mean that the claim itself needs to be abandoned; neither can be automatically rewritten as “pending polish.” Likewise, if a counterexample means that the original conclusion can be retained only within a narrower range, downgrading is the judgment successfully responding to the materials, not a shameful record of declining system performance.
This requirement also concerns the institutional status of suspension and withdrawal. Suspension is not hiding the result; it is explicitly stating that the current evidence is insufficient to maintain the original standing, while preserving what materials could restart the judgment. Withdrawal is not erasing a text once published; it is stopping the old version’s interpretive credit from continuing to circulate at its original strength, and letting readers see the reasons for withdrawal, the scope of impact, and the replacement conclusion. Only when these states can trigger version updates, citation notices, database flags, evaluation reruns, or editorial procedures do they possess actual revisability. If “limitations,” “pending verification,” and “provisional” never change the title, abstract, signature, or mode of dissemination, they are merely rhetorical buffers within a finished text.
One must still avoid romanticizing failure here. A miss does not automatically signify depth, and hesitation does not automatically signify responsibility; a team can use openness to evade verification it ought to have done, and can also design withdrawal procedures so cumbersome that they exist in name only. The criterion is whether failure states are connected to concrete evidential gaps, counterexamples, and acts of revision, and whether they can genuinely change interpretive standing — not whether they look humble.
7.5 Requirement Four: The Evaluation Contract Must Be Versioned and Changeable in Light of New Materials
The scope and basis of evaluation were defined in Section 3 as the operative constraints of materials, tasks, scoring, permitted inferences, and failure conditions. If a contract has only an initial version, and the applicable scope of a result is no longer recorded once it passes, the local evaluation acquires an unrestricted extension in circulation. Re-examination requires preserving at least the contract’s version number, date of formulation, snapshot of materials, scoring dimensions, permitted inferences, and failure handling; when new materials, counterexamples, task purposes, or community criticism change these elements, a revision should be issued, stating within what range the results under the old version remain valid.
Versioning does not replace theoretical judgment with administrative procedure. New materials need not overturn an old contract, and a contract revision need not invalidate all previous results; what needs to be stated is which inference’s supporting conditions have changed, which conclusions should be re-evaluated, and which can still be retained as qualified historical records. For open-ended humanistic tasks, a contract revision may especially involve the question itself: a new version or new contextual materials may force researchers to redraw the object, rather than merely adding an item to the existing scoring table. If an evaluation system does not permit modification of the question and the criteria, it cannot record this kind of event that genuinely changes the judgment.
At the same time, the modifiability of the contract cannot become a license to tailor standards to results after the fact. A revision should leave a record of the time, the reasons, and the participants, and state which results before and after the revision are comparable and which are not; when necessary, the original results under the old contract should be preserved, so that the community can ask whether the change of standards was made only to protect existing standing. In this way, the evaluation contract can provide a stable boundary within a single judgment while acknowledging, amid controversy and new evidence, that the boundary itself may need to be redrawn.
7.6 Requirement Five: State Who Makes the Judgment, Who Explains the Reasons, Who Can Withdraw, and the Status of System-Generated Text
Finally, the result and its signature must make responsibility for judgment identifiable. At a minimum, it should be clear: who advanced the key judgments; who bears the explanation of reasons and materials; who has the authority to revise, downgrade, or withdraw when counterexamples or version problems appear; and what standing the system-generated text, candidate relations, summaries, or code currently has within the result. The “who” here may be an individual, co-authors, an editorial board, a research institution, or a human–machine collaborative structure with an explicit division of labor, but attribution cannot end with only a model name, a project name, or a generic “team.” Nor does the structure of responsibility mean pressing all consequences onto the last person to sign: if a platform, institution, or publisher holds substantive authority over review, dissemination, and withdrawal, it too should be recorded within the corresponding scope.
The standing of system-generated content must at least distinguish among candidate leads, facts pending verification, verified citations, arguments rewritten by the researcher, and interpretations accepted by the community. Marking something “generated by the model” by itself does not explain what role the content plays in the judgment; likewise, the absence of the model from the signature does not mean that the researcher has assumed the reasons for the sentences the model produced. Only when the researcher can explain how they checked, modified, or rejected the system’s content, and has the authority to change the relevant judgments when new evidence appears, is system participation placed within an answerable structure of responsibility. If content is directly incorporated into the result while no agent can explain its bridging inferences or initiate a withdrawal procedure, it may simultaneously reproduce the interpretive appearance of the formation layer and the responsibility gap of the responsibility layer.
A contribution table has effect only when it is connected to the statement of reasons, the version record, and withdrawal authority.[^8] An anonymous review system can protect critics, and a collective signature can express shared bearing, but both need corresponding editorial, institutional, or team procedures to take up the consequences; otherwise anonymity and collectivity become merely another surface form under which responsibility cannot be located.
7.7 How the Five Requirements Work Together, and Their Limits
The requirements above are not a compliance checklist that can be ticked off once before publication. The first provides the path of return to materials; the second makes cross-level translation visible; the third gives negative results actual effect; the fourth lets the evaluation boundary be revised with new evidence; and the fifth connects the judgment and its consequences to an identifiable structure of bearing. The absence of any one may weaken the effect of the others: with complete materials but no layered record, readers still do not know which sentences are model inferences; with a versioned contract but no withdrawal authority, a change of standards still cannot change existing credit; with a responsible person but no accessible materials, responsibility cannot state what its judgment rests on.
The five requirements likewise do not constitute a sufficient condition of interpretive truth. Materials can be misread, bridging arguments can be formally complete yet conceptually fail, and counterexamples may not yet have appeared; a structure of responsibility can honestly acknowledge error, but it cannot thereby turn error into correctness. What this paper can offer is only a narrower institutional proposition: when conditions of formation, evaluation boundaries, and relations of responsibility all have public forms that are traceable, layered, capable of failure, revisable, and attributable, the process by which an interpretation acquires standing is more likely to be discovered, questioned, and changed by the community; when these forms are compressed or disabled in the dissemination of results, the appearance of completion is more easily taken for an established interpretation.
Boundaries of application must also be preserved. Different disciplines’ material permissions, anonymity rules, data protection, and scales of collaboration cannot be unified into a single archival format; some historical materials cannot be made public, some model services cannot preserve full run states, and some exploratory projects cannot foresee all counterexamples at the outset. What should be done then is not to pretend the conditions have been satisfied, but to state the gaps and their consequences, and to lower the strength of the claims that can be advanced. If a gap makes a key judgment impossible to trace, question, or withdraw, the most appropriate result may be a candidate lead, internal working material, or suspended dissemination, rather than hiding institutional limits behind a finished report.
7.8 Conclusion: Keeping Interpretation Open to Re-examination
Re-examining the conditions of interpretation does not mean restoring, outside generative systems, a pure research process unchanged by technical mediation. Materials will still pass through OCR, retrieval, ranking, and model organization, and judgments can be formed by individuals, teams, and human–machine collaboration. The requirement is that these processes must not extract the standing of a subsequent interpretation from its conditions of formation, evaluation, and responsibility. Material scope and conditions of access must be returnable; facts, model results, and interpretations must be distinguishable; misses and withdrawals must genuinely change standing; the evaluation contract must be modifiable by new materials; and the standing of the judgment and of system-generated content must be clearly attributable.
These arrangements increase the likelihood of discovery, contestation, correction, and withdrawal; they do not guarantee that any particular interpretation is true. Nor do they elevate human signatures, handwritten drafts, complete logs, or high evaluation scores into unquestionable credentials. The significance of leaving room for revision lies precisely here: an interpretation may hold provisionally at a given moment, but it must not, merely because its finished appearance has already circulated, lose the real opportunity to be re-examined by materials and by the community. At this point, this paper’s three-layer framework forms a bounded cycle: conditions of formation constrain interpretive appearance, the evaluation contract restricts the translation of standing, the structure of responsibility preserves the path of withdrawal, and re-examination keeps these three constraints effective in cross-vehicle circulation.