Research documentation

Review edition: horizon-ai-review-v1 · Reviewed 2026-09-17

Horizon RunMap - experimental evidence

Reviewed 2026-09-17. Values below are generated from the existing public editions, without new model execution. Machine-readable claims.

Separate experiments

ExperimentPopulationEvidence / validation / publication / reproduction
technical-preview-residency39-v313 configurations; 39 workers; 468 measured responsesE3 / valid / public / not_attempted
olmoe-0924-fidelity-v2120 paired promptsE3 / valid / public / not_attempted
olmoe-0924-ifeval541-v1541 prompts; 834 instructionsE1 / valid / public / not_attempted

These are separate populations with separate scorers and questions. Public admission, evidence maturity, verification and independent reproduction are distinct.

Performance contract and results

PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.

Pooled decode = sum(N - 1) / sum(internal decode seconds); it excludes first-token wait. Overall output rate uses sum(N) / sum(internal response runtime). TTFT and response time are arithmetic request means. VRAM is a device-wide whole-worker driver peak including preparation, warmup and shared desktop use; RSS is a separate whole-worker process peak. H2D is a sum over measured requests, not a rate. Isolated prefill is UNKNOWN; OLMoE HTTP observation is UNSUPPORTED. Block spread is a descriptive range, not a confidence interval.

Configuration / sourceExecutorDecode token/sWhole-worker VRAM GiB
allenai/OLMoE-1B-7B-0924-Instruct c30/64PRODUCT-core scalar23.004.11
allenai/OLMoE-1B-7B-0924-Instruct c40/64PRODUCT-core scalar27.694.62
allenai/OLMoE-1B-7B-0924-Instruct c50/64PRODUCT-core scalar31.505.15
allenai/OLMoE-1B-7B-0924-Instruct c60/64PRODUCT-core scalar35.015.67
allenai/OLMoE-1B-7B-0924-Instruct c64/64Grouped T361.265.99
allenai/OLMoE-1B-7B-0125-Instruct c30/64PRODUCT-core scalar23.044.15
allenai/OLMoE-1B-7B-0125-Instruct c40/64PRODUCT-core scalar27.124.62
allenai/OLMoE-1B-7B-0125-Instruct c50/64PRODUCT-core scalar31.225.15
allenai/OLMoE-1B-7B-0125-Instruct c60/64PRODUCT-core scalar35.445.67
allenai/OLMoE-1B-7B-0125-Instruct c64/64Grouped T362.076.01
Qwen/Qwen1.5-MoE-A2.7B-Chat c20/60top4-fused-v115.227.11
Qwen/Qwen1.5-MoE-A2.7B-Chat c30/60top4-fused-v117.178.08
Qwen/Qwen1.5-MoE-A2.7B-Chat c39/60top4-fused-v119.958.99

OLMoE c30-c60 use PRODUCT-core scalar; c64 uses Grouped T3. This is not a residency-only comparison. Qwen c20/30/39 use top4-fused-v1. Capacities are per routed layer; no cross-runtime ranking is measured.

Performance manifest · Contract/report · Verification method · Verifier

Historical behavioral comparison

Under 120 paired prompts, BF16 produced 29/120 structurally conforming responses and Horizon produced 28/120. There were 8 Horizon conformity losses, 7 gains, 16 identical output-token sequences and 3 equivalent JSON pairs under the frozen tolerance.

Structural conformity is not factual-answer correctness. The original campaign's publication sanitizer failure is retained in provenance; the CPU-only successor recovers the original responses without changing the frozen scorer.

Paired manifest · Contract · Summary · Audit method · ZIP

IFEval

Model: allenai/OLMoE-1B-7B-0924-Instruct, revision 7f1c97f440f06ce36705e4f2b843edb5925f4498. The complete declared population was scored using frozen evaluator inputs. All 467 EOS and 74 length-limit endings remain included.

MetricNumerator / denominatorPercent
Prompt strict217 / 54140.11%
Prompt loose247 / 54145.66%
Instruction strict433 / 83451.92%
Instruction loose474 / 83456.83%

The original frozen contract keeps its historical internal_only field. Separate publication metadata records public admission as E1 standalone_benchmark. These instruction-following scores do not establish matched BF16 preservation. External model-card numbers are contextual, non-controlled references; one reference does not specify the IFEval variant.

IFEval manifest · Frozen contract · Public admission · All responses · Audit method · ZIP

Verification and reproduction

The performance verifier checks hashes and aggregate-copy consistency; that release does not include the complete raw-run bundle. Behavioral and IFEval verifiers recompute their declared scores from retained responses on CPU. None reruns the private inference runtime. Independent reproduction status is technical-preview-residency39-v3: not_attempted; olmoe-0924-fidelity-v2: not_attempted; olmoe-0924-ifeval541-v1: not_attempted. UNKNOWN denotes an absent observation; UNSUPPORTED denotes unavailable observation capability.