IFEval evidence and audit
The complete OLMoE 0924 IFEval population contains 541 prompts and 834 evaluated instructions. Termination: 467 natural EOS and 74 length-limit endings.
| Metric | Passed / evaluated | Percent |
|---|---|---|
| Prompt strict | 217/541 | 40.11% |
| Prompt loose | 247/541 | 45.66% |
| Instruction strict | 433/834 | 51.92% |
| Instruction loose | 474/834 | 56.83% |
E1 standalone_benchmark, with separate public admission. This experiment is separate from the earlier 120-pair BF16/Horizon study. The CPU verifier recomputes the frozen scores from retained responses; it does not rerun runtime generation. Factual correctness was not graded.
Explorers and raw evidence
- All 541 original responses (English explorer)
- Todas as 541 respostas (explorador em português)
- Four metrics (JSON)
- Retained responses (JSONL)
- Frozen execution contract (JSON)
- Public admission (JSON)
- Audit method and report (Markdown)
- Manifest and SHA-256 hashes
- CPU verifier (Python)
- Complete evidence bundle (ZIP)