IFEval evidence and audit

The complete OLMoE 0924 IFEval population contains 541 prompts and 834 evaluated instructions. Termination: 467 natural EOS and 74 length-limit endings.

MetricPassed / evaluatedPercent
Prompt strict217/54140.11%
Prompt loose247/54145.66%
Instruction strict433/83451.92%
Instruction loose474/83456.83%

E1 standalone_benchmark, with separate public admission. This experiment is separate from the earlier 120-pair BF16/Horizon study. The CPU verifier recomputes the frozen scores from retained responses; it does not rerun runtime generation. Factual correctness was not graded.

Explorers and raw evidence