Skip to content
José Luis Delgado
Paper

Auditing Loose Scoring in IFEval: A Human-Adjudicated Study of Verifiable Instruction Following

A human-adjudicated audit of IFEval strict-fail/loose-pass disagreements, separating valid corrections, loose false positives, and ambiguous cases.

LLM Evaluation Instruction Following Human Adjudication Benchmarks AI

Status

Under review

Authors

José Luis Delgado

Year

2026

Abstract

The paper audits IFEval’s strict and loose scoring modes with human adjudication of every case where the two scoring modes disagree. The study measures how often loose scoring turns strict failures into valid corrections, false positives, or ambiguous cases.

Across 169 strict-fail/loose-pass flips, only 23 are valid corrections, while 102 are loose false positives and 44 are ambiguous cases. The study turns evaluation disagreement into a human-adjudicated evidence surface, clarifying what loose scoring adds and where it weakens the reliability of benchmark claims.