4 When Does a Local Evaluation Get Taken as the Whole Being Established
4.1 Evaluation Is Not the Enemy of Interpretation
The long-form research outputs of generative AI have deprived an old dichotomy of its explanatory power: at one end, interpretation is reduced to whether the model truly understands; at the other, the evaluation of interpretation is reduced to whether the text contains factual errors. Recent research on source attribution, long-form factuality, citation verifiability, and report logic has shown that the relations between statements and facts, between statements and sources, and between report structure and support relations can be separated into distinct objects of evaluation. Evaluation is therefore not a crude substitute external to humanistic scholarship; it can reveal what a report has omitted, whether a source really supports the attached claim, and whether a causal connection lacks a bridging reason, and it can also display the local differences among systems on the same task.
The problem appears at the next step. An evaluation result usually has to leave a single run, a dataset, or a set of scoring criteria and enter papers, platforms, institutional registries, project descriptions, and dissemination materials. In the course of this movement, whether the material boundaries, task uses, permitted inferences, and failure conditions originally attached to the result remain visible is not automatically decided by the score. That a report is genuinely qualified on the present task cannot ground the conclusion that a stronger judgment has acquired the standing of open interpretation. To analyze this conversion, one must first explain how local evaluation can hold at all, and why it must be provisionally closed.
4.2 What the Scope and Basis of Evaluation Is
This paper calls the materials, tasks, scoring, permitted inferences, and failure conditions that enable local evaluation to operate an “evaluation contract”: the evaluator must know which batch of materials to check, which question to address, which standards to judge by, what may be inferred from the result, and which circumstances should be recorded as partial satisfaction, insufficient evidence, or inability to answer.[^2]
Without such constraints, open-ended outputs cannot be stably compared. ReportLogic places a given query, a retrieval context, and a candidate report within a single evaluation unit, checking whether the report forms a unified analytic arc, provides the necessary context, and establishes explicit relations between claims and support; it also uses adversarial examples to show that adding headings, numbering, lengthy explanations, or “evidence” labels may improve surface presentation without repairing the underlying support relations. Being qualified here is not a natural property of the report, but the result of the report within a specific query, context, set of dimensions, and judgment procedure. This boundary does not diminish the value of ReportLogic; on the contrary, it enables ReportLogic to state clearly what it measures.
ResearchRubrics addresses another kind of openness. It acknowledges that deep-research answers are long, formally diverse, and may have multiple valid paths, and therefore does not require the system to reproduce a single complete reference text; instead it uses fine-grained criteria written by experts to check factual grounding, whether the reasoning holds, and clarity of expression. Its abstract reports that even the leading system’s rate of criterion compliance remains below 68%, with the main defects being the omission of implicit context and insufficient reasoning over retrieved information. This result shows that fine-grained criteria can yield genuine diagnostic information, and also that “openness” does not mean the absence of an evaluation contract: the answer need not be unique, but the scope of materials, the task goals, and the acceptable inferences are still specified in some way.
DeepResearch Bench II displays the pre-set conditions within open-ended tasks even more directly. The benchmark extracts facts and inferences from survey articles written by experts and, through automatic extraction, manual cleaning, and revision by domain experts, forms fine-grained evaluation criteria; some of the expert articles are listed as shielded sources, and the system must reconstruct their content from other materials. This does not treat the expert articles as eternally correct answers; it confirms, within a specified task, whether the system found the necessary evidence and produced the corresponding analysis. An open-ended task therefore still needs to specify provisionally what counts as a relevant fact, what counts as non-trivial analysis, and which sources may not be reproduced directly. The existence of a scope and basis of evaluation is a condition for comparison to begin, not a proof that interpretation has been completed.
Software testing makes the point about the “evaluation contract” especially clear. A test suite specifies bounded inputs, task goals, pass criteria, and failure conditions, so that “the tests pass” is, first of all, only a local pass of the artifact within that evaluation scope. It cannot by itself yield the conclusions that the system is globally correct, that the requirements have been correctly understood, that the code is safe enough to deploy, or that the coder should thereby be recognized as possessing a stable software-engineering capability.
The conversion that occurs here is structurally analogous to what this paper calls standing substitution: the original evaluation result is genuine, but the subsequent, stronger claim adds new burdens of specification interpretation, architectural judgment, untested conditions, maintenance, and responsibility. If these burdens are not added, testing turns from a useful local judgment into an unduly enlarged credential of standing.
4.3 Provisionally Closed, but Open to Re-examination
The “provisional closure” of the scope and basis of evaluation is necessary: if materials, questions, scoring, and failure records changed without limit in every judgment, local results would not be comparable. But the demand of open interpretation is not to remain forever inside the evaluation scope. New materials may make the original question no longer the most important one; counterexamples may invalidate the old evaluation criteria; differences among versions may change the identity of the evidence; competing interpretations may reject the original decomposition of the question. The scope and basis of evaluation must therefore be closed within a single act of judgment and yet remain open to re-examination within the community’s controversies.
Validity theory can state this distinction more clearly. The Messick tradition of construct validity (Messick 1989) requires examining the interpretation that a measurement supports together with its social consequences; Kane’s (2013) argument-based validity requires explicitly listing the inferences from observations to interpretation and use, together with their supporting assumptions.[^3] Kane emphasizes in particular that the object of validity is the proposed interpretation and use, not a test or a score detached from its use; more ambitious claims require more evidence, and interpretations and uses change with new needs and understandings. This paper accepts this work and does not write the “evaluation contract” as a replacement for validity theory.
What this paper adds lies at the two ends of the evaluation chain. At one end is how the contract, before evaluation, selects materials, organizes the question, and specifies the permitted inferences; at the other end is whether, after evaluation, when the result enters papers, platforms, registries, and dissemination, the original boundaries, uses, and conditions for downgrading continue to constrain subsequent claims. A validity argument can show how a score or a report is supported under a specific interpretation and use, but it does not automatically show whether, once the result enters a new vehicle, it has been exchanged for a larger identity. The latter requires tracking how wording, evidence, and responsibility move among vehicles.
4.4 Treating a Local Pass as the Whole Being Established: Five Conditions
This paper calls this specific conversion “standing substitution,” that is, treating a local pass as if the whole were established. More precisely, standing substitution is the unwarranted conversion of local evaluation into interpretive standing: a result genuinely passes a local check within a specific evaluation scope, but is treated by subsequent papers, platforms, registries, or dissemination vehicles as sufficient warrant for a stronger claim of interpretation or scholarly achievement, without the addition of new materials or bridging arguments commensurate with the stronger claim, and when the original material boundaries, uses, and conditions for downgrading or withdrawal no longer actually constrain the subsequent, stronger claim.
“Substitution” here does not mean that the original evaluation score carried no information, nor that the subsequent, stronger claim is necessarily wrong. What it describes is a change in the relation of reasons: a result that could only support “a local pass under these materials, this question, and this set of standards” is rewritten as “therefore the interpretation has been formed,” “the system possesses a certain research capability,” or “the result has been established.” To avoid conflating ordinary extrapolation, inadequate measurement, and exaggeration in dissemination, five conditions are jointly necessary.
First, the local pass in the original evaluation must be genuine. If the citations are fabricated, if factual verification failed, or if the report never satisfied the original criteria, the first diagnosis should be error, hallucination, or evaluation failure; not every failure can be called standing substitution. The distinctive pressure of standing substitution lies precisely in the fact that the original evaluation result can be genuine.
Second, the subsequent level must advance a stronger claim. Rewriting “passed within this task” as “this interpretation is provisionally acceptable” may not yet cross a level; rewriting it as “an open interpretation has been formed,” “it represents a historical object,” or “it demonstrates that the system possesses general research capability” significantly raises the strength of the subsequent claim. Third, the subsequent, stronger claim must also take the original evaluation pass as its key warrant. If the subsequent claim relies entirely on new historical materials, independent review, and the researcher’s own bridging argument, so that the original evaluation score is merely background information, there is no chain of reasons that treats a local pass as the whole being established.
Fourth, the subsequent level has not added commensurate new materials or bridging arguments. Stronger claims require stronger support — a principle already clearly expressed in Kane’s validity argument. New primary materials, version comparison, contextual analysis, competing interpretations, the handling of counterexamples, and causal bridging can all bear this added burden; without them, the original evaluation pass cannot be upgraded into the recognition of an open interpretation. Fifth, the original boundaries, uses, and conditions for downgrading or withdrawal must no longer actually constrain the claim at the subsequent level. If the dissemination materials still explicitly state that the evaluation addressed only a certain task, a certain range of knowledge, and a certain use, and allow new evidence to downgrade the conclusion, standing substitution cannot be diagnosed merely from an enlarged name. “Actual constraint” here is not a matter of whether the qualifications still linger in a footnote or appendix; it depends on whether they genuinely change the title, abstract, capability naming, scope of application, dissemination wording, or use decisions at the subsequent level. A counterfactual question can serve as a test: if the material boundaries, uses, and downgrading conditions of the original evaluation were carried into the subsequent level unchanged, would the subsequent, stronger claim have to shrink? If the answer is yes, while the subsequent level still maintains the broader naming of standing, then the original qualifications, though perhaps still verbally present, have lost their actual constraining force.
These five conditions distinguish standing substitution from the general point that “a proxy is not the target.” Sharma’s summary of Proxy-Based Evaluation has already made clear that metrics, benchmarks, transparency documentation, and auditing procedures are proxy mechanisms of oversight; that the proxies are satisfied still does not prove that the deployment goals have been achieved. This paper does not rediscover this point, but traces how proxy results acquire recognition among academic vehicles. Nor does standing substitution require that any agent deliberately game the metrics. The tradition of metric reactivity represented by Goodhart (1975) and Campbell (1979) mainly warns that the objects of evaluation respond to metrics and institutions; standing substitution, by contrast, can be completed step by step across multiple vehicles without deliberate catering, without score fabrication, and even without any single actor making the complete erroneous inference.
A further boundary must be drawn between the validation of an output and the recognition of a producer or achievement. This paper does not treat those as the same process. An output may pass a specified validation loop without its producer acquiring durable credibility, while an interpretation may receive provisional standing without becoming part of a producer’s long-term reputation. The two chains should therefore not be conflated.
4.5 Two Counterexamples and One Case in Which Standing Substitution Does Not Occur
The first counterexample is a long-form historical report: all its citations are genuine, the pages do exist, and factual verification, source attribution, coverage, and the report’s internal logic all pass the established checks. But the report’s key causal judgment merely repeats a piece of secondary scholarship, without reworking the counterevidence in the primary materials and without explaining why the secondary author’s generalization applies to the present problem. The local pass in the original evaluation is genuine; if the subsequent discourse names it “a historical interpretation that has been formed,” it simultaneously lacks new materials and bridging arguments commensurate with the stronger claim, and if the report no longer preserves the boundaries of “secondary generalization, limited corpus, pending review,” it constitutes standing substitution. This example shows that standing substitution cannot be exhausted by hallucination or citation error: the citations may all be genuine while the relation of reasons remains insufficient.
The second counterexample is a digital-humanities analysis: the statistics, the co-occurrence network, and the model’s numerical values can all be rechecked against public code and data, and the numerical changes remain stable across different runs. A subsequent paper, however, directly names a certain numerical relation a “conceptual-history event,” and from it infers a social mechanism of an intellectual tradition. Without sense analysis, controls for chronology and genre, alternative orderings, contrary materials, and a bridging argument from numerical relations to historical practice, the recheckability of the numbers supports only a computational result or a candidate association; it cannot support the stronger conceptual-historical and causal claims. What is genuinely missing here is not more decimal places, but the materials and arguments needed to change the level of interpretation.
It is also possible that no treating of a local pass as the whole being established occurs. Suppose a research system always marks its generated results as “candidate leads,” with each result accompanied by the scope of materials, versions, the evaluative use, and an account of expert involvement; and the dissemination documents explicitly state that the report passed only within a specific task and that competing interpretations and new materials can downgrade or withdraw the conclusion. The researchers then add independent primary materials, version comparison, contextual analysis, and the handling of counterexamples, and in the paper they set out, step by step, the bridging process from candidate relations to a historical interpretation. In this case, standing may indeed be legitimately elevated from a local result to a stronger interpretation, but the elevation comes from the added evidence and bridging arguments, not from the original evaluation pass being smuggled in as sufficient warrant; it therefore does not constitute treating a local pass as the whole being established.
4.6 What Happens After Results Enter Other Settings
The framework of judgment here is a diagnostic framework, not an empirical conclusion already completed for all projects. It can generate three propositions for subsequent research. First, when an evaluation result moves from a system report into papers, platforms, and dissemination materials, the material boundaries, uses, and failure conditions may gradually diminish; this requires paragraph-by-paragraph comparison of versioned documents and cannot be inferred from names alone. Second, the dissemination of results may prefer capability names with broader extensions that are more easily recognized by the public, making local task capabilities more readily understood as general research capability; this is a testable hypothesis about choices in dissemination, not a fact about all projects. Third, the more a project lacks records of boundary migration, the more likely its local evaluations are to be used by later papers, platforms, or dissemination materials for stronger claims of standing; this requires cross-project comparison of evaluation records, dissemination texts, and audience use, and cannot be proven by this paper’s conceptual analysis alone.
These propositions do not prearrange open-ended and closed tasks into risk grades. Closed questions can measure bounded capabilities accurately, and open-ended tasks can use expert criteria to yield meaningful local information; what really needs to be tracked is whether a pass within the evaluation scope is endowed with a new claim when it leaves the original task, and whether the original constraints move along with it. Espeland and Stevens’s (2008) discussions of quantification and commensuration, Power’s (1997) warning about the audit society, and Muller’s (2018) analysis of metric fixation can provide a sociological vocabulary for this institutional background, but none of them can replace the examination of concrete vehicles, versions, and chains of reasons.
4.7 Conclusion: A Local Pass Is Not an Established Interpretation
The scope and basis of evaluation make comparison possible, and they also provisionally fix materials, tasks, scoring, and failure conditions as objects the community can check. Its value lies not in compressing open interpretation into a score table that never changes, but in making the scope and reasons of a local judgment statable. Kane (2013) and Messick (1989) have shown that claims about interpretation and use must accept evidential scrutiny commensurate with their strength; research on complex evaluation has further separated facts, sources, coverage, analysis, and report logic into diagnosable relations. Sharma’s (2026) discussion of proxy-based evaluation reminds us that satisfying a proxy does not equal achieving the target; the same distinction matters here because an output passing a bounded evaluation does not by itself confer higher status on the producer or the result.
The increment of this paper is therefore limited and concrete: it traces how a local evaluation is converted into interpretive standing among papers, platforms, registries, and dissemination. Only when the original evaluation is genuine and valid, the subsequent claim is stronger, the original evaluation is used as the key sufficient warrant, no commensurate new evidence or bridging argument has been added, and the original boundaries and downgrading conditions no longer constrain the subsequent level is there reason to diagnose standing substitution. If the scope and basis of evaluation can be re-examined by the community, and if new materials, counterexamples, and bridging arguments do change the claim, then an elevation of standing can hold; if these conditions appear only in words without being able to change the conclusion, one must still return to the responsibility layer and ask whether the judgment is genuinely questionable, revisable, and withdrawable.