Research documentation

Quality Evaluation

These two experiments ask different questions and use separate populations and scorers. Instruction following and structural conformity do not establish general answer correctness or equivalence to original BF16 output.

IFEval standalone benchmark

541 prompts and 834 instructions. Evidence level E1; experiment type standalone_benchmark; validation valid; publication public; independent reproduction not_attempted.

MetricPassed / total
Prompt strict217 / 541
Prompt loose247 / 541
Instruction strict433 / 834
Instruction loose474 / 834

CPU recomputation using frozen dataset and scorer over retained responses; no generation rerun. The published benchmark is Horizon-only; contextual model references are not a paired BF16 control.

Explore all retained responses · Summary JSON · Experiment contract · Method · Verifier · Manifest · Evidence ZIP

Earlier paired BF16 / Horizon structural comparison

120 paired prompts. Evidence level E3; validation valid; publication public; independent reproduction not_attempted.

Frozen structural checks were satisfied by 29 BF16 responses and 28 Horizon responses. Horizon had 8 losses and 7 gains under those checks. These are structural categories, not answer-correctness grades.

CPU recomputation of frozen classifications from retained response pairs; no answer-correctness grade or generation rerun.

Paired results JSON · Original report · Method · Contract · Verifier · Manifest · Evidence ZIP