Research documentation

Horizon RunMap - experimental evidence

Reviewed 2026-09-16. Values below are generated from the existing public editions, without new model execution. Machine-readable claims.

Separate experiments

ExperimentPopulationEvidence / validation / publication / reproduction
Residency and executors13 configurations; 39 workers; 468 measured responsesE3 / valid / public / not_attempted
Historical BF16/Horizon comparison120 paired promptsE3 / valid / public / not_attempted
IFEval541 prompts; 834 instructionsE1 standalone_benchmark / valid / public / not_attempted

These are separate populations with separate scorers and questions. Public admission, evidence maturity, verification and independent reproduction are distinct.

Performance contract and results

PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.

Pooled decode = sum(N - 1) / sum(internal decode seconds); it excludes first-token wait. Overall output rate uses sum(N) / sum(internal response runtime). TTFT and response time are arithmetic request means. VRAM is a device-wide whole-worker driver peak including preparation, warmup and shared desktop use; RSS is a separate whole-worker process peak. H2D is a sum over measured requests, not a rate. Isolated prefill is UNKNOWN; OLMoE HTTP observation is UNSUPPORTED. Block spread is a descriptive range, not a confidence interval.

Configuration / sourceExecutorDecode token/sWhole-worker VRAM GiB
OLMoE 0924 c30/64PRODUCT-core scalar23.004.11
OLMoE 0924 c40/64PRODUCT-core scalar27.694.62
OLMoE 0924 c50/64PRODUCT-core scalar31.505.15
OLMoE 0924 c60/64PRODUCT-core scalar35.015.67
OLMoE 0924 c64/64Grouped T361.265.99
OLMoE 0125 c30/64PRODUCT-core scalar23.044.15
OLMoE 0125 c40/64PRODUCT-core scalar27.124.62
OLMoE 0125 c50/64PRODUCT-core scalar31.225.15
OLMoE 0125 c60/64PRODUCT-core scalar35.445.67
OLMoE 0125 c64/64Grouped T362.076.01
Qwen c20/60top4-fused-v115.227.11
Qwen c30/60top4-fused-v117.178.08
Qwen c39/60top4-fused-v119.958.99

Capacities are per routed layer. OLMoE c30-c60 use PRODUCT-core scalar; c64 uses Grouped T3, so c30-to-c64 or c60-to-c64 is not a residency-only ablation. Qwen c20/c30/c39 use top4-fused-v1. Keep model revisions and executors attached to every value; no cross-runtime ranking is measured.

Performance manifest · Contract/report · Verification method · Verifier

Historical behavioral comparison

Under 120 paired structured prompts, BF16 produced 29/120 structurally valid responses and Horizon produced 28/120. There were 8 paired structural regressions, 7 improvements, 19 observable-behavior parity cases and 86 directionless divergences. Structural conformity is not factual-answer correctness. The original campaign's publication sanitizer failure is retained in provenance; the CPU-only successor recovers the original responses without changing the frozen scorer.

Paired manifest · Contract · Summary · Audit method · ZIP

IFEval 541

Model: allenai/OLMoE-1B-7B-0924-Instruct, revision 7f1c97f440f06ce36705e4f2b843edb5925f4498. The complete declared population was scored using frozen evaluator inputs. All 467 EOS and 74 length-limit endings remain included.

MetricNumerator / denominatorPercent
Prompt strict217 / 54140.11%
Prompt loose247 / 54145.66%
Instruction strict433 / 83451.92%
Instruction loose474 / 83456.83%

The original frozen contract keeps its historical internal_only field; separate publication metadata records later public admission as E1 standalone_benchmark. These instruction-following scores do not establish matched BF16 preservation. External model-card numbers are contextual, non-controlled references; one reference does not specify the IFEval variant.

IFEval manifest · Frozen contract · Public admission · All responses · Audit method · ZIP

Verification and reproduction

The performance verifier checks hashes and aggregate-copy consistency; that release does not include the complete raw-run bundle. Behavioral and IFEval verifiers recompute their declared scores from retained responses on CPU. None reruns the private inference runtime. Independent reproduction is not_attempted for all three editions. UNKNOWN denotes an absent observation; UNSUPPORTED denotes unavailable observation capability.