A common failure in empirical AI writing occurs when authority moves quietly from a measured protocol to a claim with wider scope: a table reports a number, the prose assigns that number a name, and the reader is asked to treat the name and the measurement as if they referred to the same thing. In many AI papers, that transfer is where much of the evidential work occurs: a benchmark score becomes instruction-following, a win rate becomes alignment, a judge preference becomes reasoning quality, and a set of high-scoring edges becomes a circuit. Numbers can make an experiment precise, but the sentence surrounding a number can still exceed what the experiment warrants.
From score to claim
Quantitative evaluation gives researchers a shared object to inspect because two models can be compared on the same prompt set, two prompting strategies can be evaluated under the same judge, and two localization methods can be ranked against the same intervention target. A ten-point gap on a stable evaluation, or a better ranking of exact-trace edges by one method than by another, gives real information about behavior under the conditions of the experiment. That information becomes evidence for a larger claim only when the paper explains why the protocol measures the property named in the conclusion.
The meaning of a score comes from the full pipeline that produced it, including the prompt distribution, the judge or target, the perturbation or intervention, the aggregation rule, and the treatment of borderline cases. A refusal rate means little unless the report states what counted as refusal, what counted as unsafe behavior, and how mixed answers were handled. A judge preference means little unless the reader knows the rubric, the judge, the answer format, and the cases where style can be rewarded over correctness. An edge score means little unless the scorer, the intervention, and the downstream selection rule are part of the claim.
Protocol boundaries
The instruction-following case shows how quickly a useful number can become an overextended claim when a score is separated from the judge, rubric, and aggregation rule that produced it. Suppose a model obtains 87 on a benchmark judged by an LLM. That value supports a claim about performance on that prompt set under that judge, rubric, and aggregation rule. A reliability claim across the cases where instruction-following is most operationally important would require evidence about the relevant boundary cases. The judge may reward polished formatting, overvalue verbosity, or miss brittle behavior around clarifications and refusals. The model may perform well on ordinary requests while failing on cases that the average mostly hides. The score is still informative, but the claim has to stay close to what was measured.
Mechanistic interpretability has the same problem in a less familiar form when a localization method assigns high scores to edges, heads, or representation components on an indirect object identification task, and those scores become evidence about local measured influence under the specified intervention. The selected circuit is a different object. A positive-only selector, a magnitude selector, a top-k budget, a protected shortlist, or a path-supported subset can all turn the same score field into different graphs. Once the prose says the selected graph is the mechanism, the evidential target has moved from local scoring to circuit-level behavior.
In mechanistic interpretability, the movement from score field to selected graph is easy to miss because the words are close together. High edge scores can guide circuit identification, and a selected circuit may be tested after selection, but the score field does not validate the induced graph on its own. Edge-wise faithfulness and circuit-wise faithfulness are related objectives with different evidential requirements. A paper that optimizes agreement with a local exact-intervention target and then claims recovery of the best circuit under a circuit-level objective needs evidence at the circuit level; without that evidence, the method has shown that it ranks local influence well under one scorer, which is a narrower but still meaningful result.
Objective mismatch
Objective mismatch generalizes the same error because a metric can measure one target while the text uses it to support another. An LLM judge may prefer answers that are verbose, deferential, or neatly formatted. A reward model may prefer the answer that sounds better to the evaluator over the one that is more accurate. A compliance score can reward visible obedience to surface instructions while missing whether the substantive request was satisfied. These metrics may be appropriate for the questions they were built to answer, but they do not become evidence for every property that sounds adjacent to the metric label.
The writing problem is tightly connected to the measurement problem because the name of a metric can become the name of the conclusion. When that happens, the reader loses the difference between a protocol result and an empirical property. “The judge preferred this answer” is narrower than “the model reasoned better.” “The selected components received high scores” is narrower than “this is the mechanism.” “The model scored higher under this rubric” is narrower than “the model is more aligned.” The narrower sentences may be less dramatic, but they are easier to evaluate and harder to misuse.
Reporting standard
Good reporting keeps the object of measurement visible from the first claim to the last sentence by naming the property being claimed, stating what was scored, describing the distribution and intervention or judging setup, and reporting the aggregation rule. Filtering, thresholding, weighting, reranking, top-k selection, and judge normalization belong in the evidential account when they affect the reported value. If the claim concerns tail behavior, boundary cases, refusal thresholds, safety, or circuit structure, the report should show evidence at that level and avoid letting the average score carry the claim indirectly.
The strongest empirical writing treats scores as protocol-bound evidence, so a score can support a comparison, motivate an ablation, guide a selection, or expose a failure mode only through the match between the protocol and the claim. When that match is explicit, the number becomes more useful, because the reader can see what was measured, what was inferred, and where the experiment stops.