Abstract
The paper audits IFEval’s strict and loose scoring modes with human adjudication of every case where the two scoring modes disagree. The study measures how often loose scoring turns strict failures into valid corrections, false positives, or ambiguous cases.
Across 169 strict-fail/loose-pass flips, only 23 are valid corrections, while 102 are loose false positives and 44 are ambiguous cases. The study turns evaluation disagreement into a human-adjudicated evidence surface, clarifying what loose scoring adds and where it weakens the reliability of benchmark claims.