Horizon RunMap - Architecture, Results and Public Evaluation
Current edition · 2026-10-01 · Lucas Ribeiro
What Horizon is
Horizon RunMap is an MoE inference architecture conceived and implemented independently by Lucas Ribeiro. It combines selective compression of routed experts, custom compact-weight GPU execution and controlled expert residency. Full residency is a complete operating path; when the GPU memory budget calls for partial residency, the documented cache contract coordinates movement from immutable host backing. A public desktop app provides an external evaluation path.
The published RTX 5070 campaign records execution, memory and speed. In its OLMoE 0924 full-residency configuration, Grouped T3 recorded 61.26 post-first-token decode tokens/s at 5.99 GiB whole-worker peak VRAM (source result). A separately published IFEval edition contains all 541 responses and its scored results. The public app also completed the full IFEval process on another computer with an RTX 4060. Explore the results · Get the app.
Engineering question
How can storage precision, execution precision, expert residency and data movement be coordinated under a constrained GPU-memory budget, while making memory, latency, output rate and output behavior separately inspectable?
Architecture invariants
The mixed-precision path keeps routed expert weights compact as INT4 qpacked data with FP16 scales in the selected campaign. Routing selects the needed experts, and custom GPU kernels execute them with FP16 activations. Non-routed checkpoint BF16 values are materialized as FP16 in the measured paths. When all relevant compact experts fit on the GPU, execution is fully resident. The documented resident-qpacked-cache-v1 contract adds bounded movement for partial residency: immutable host backing holds every routed expert, and disposable GPU copies are published only after complete transfer. See Architecture for the profile-specific lifecycle.
What is implemented
The published paths execute selectively compressed INT4 routed experts using FP16 activations (W4A16); weights remain compact in GPU storage rather than persisting as fully expanded expert matrices. Original checkpoint dtype, quantized source, scales, activation dtype and accumulation dtype remain distinct fields.
The selected campaign records partial residency with PRODUCT-core scalar OLMoE execution, full OLMoE residency with Grouped T3, and partial Qwen residency with top4-fused-v1. The earlier offload-qpacked expanded-slot mechanism retains its different meaning. The desktop app gives an external user a guided way to prepare the public model and run IFEval.
What evidence demonstrates
The RTX 5070 performance edition reports execution, internal timing, resource and transfer observations for 13 configurations. The published IFEval edition reports a complete 541-prompt instruction-following evaluation, including 45.66% prompt-level loose accuracy and every retained response. The earlier behavioral study compares 120 paired BF16 and Horizon responses under predefined structural checks. The app's complete IFEval execution on another RTX 4060 computer adds an external-machine milestone. Exact metrics, contracts and source files are in Evidence and Claims.
Scope of the published results
Each numerical result belongs to its recorded model, executor, hardware and workload. The performance campaign, published IFEval score, paired study and RTX 4060 app execution are distinct records. The public artifacts support inspection and their stated verification methods. The runtime source remains private by author decision; the public app enables execution on supported setups.
Closest prior art
Expert offloading, mixed quantization, GPU caching and heterogeneous inference are established techniques. Eliseev and Mazur's Mixtral offloading combines expert caching and speculative loading. MoE-Infinity documents activation-aware caching and prefetching. KTransformers documents CPU/GPU expert computation. llama.cpp exposes configurable tensor/layer placement. The architecture comparison matrix records source-scoped properties, including overlap and unresolved details, without a throughput ranking.
Key architectural differences
Horizon brings compact expert representation, complete immutable RAM authority, H2D-only miss transport, discard eviction and complete-generation publication into one explicit cache contract. The source-scoped matrix compares these design choices with related implementations and marks unresolved properties as UNKNOWN.
Experimental contracts
- Performance: fixed prompt order, repeated blocks, greedy selection, EOS suppressed, excluded warmups and internal timing boundaries. Each configuration retains its executor and per-layer residency.
- Historical behavioral comparison: matched retained BF16/Horizon responses with a frozen structural scorer; answer correctness was not graded.
- IFEval: a complete standalone benchmark under its frozen dataset, scorer and generation policy; external model-card values are contextual references rather than matched control arms.
Measurement notes
The RTX 5070 campaign's whole-worker VRAM includes preparation, warmup and shared desktop use; process RSS has a separate scope. Its pooled decode rate excludes first-token wait. IFEval scores follow their frozen instruction checks, while the earlier 120-pair study uses structural checks. The release pages retain the full definitions and verification methods.
How to compare the recorded results
Compare OLMoE c30–c60 within the PRODUCT-core scalar family; c64 uses Grouped T3 and is a separate executor result. Read IFEval and the earlier paired study as separate populations with different scoring questions. Keep FP16 activations and execution values distinct from original checkpoint values. Read decode rate alongside first-token wait and total response time, and use measured transfers when analyzing movement.
Reviewer notes
Reconstruct the implemented system before judging its contribution. Assess the integrated architecture, app, recorded results and relevant prior work at their own levels. Keep the executor identity attached to each performance observation and the published IFEval score attached to its own edition. Reviewer brief · Public app.
Current report (HTML) · Markdown · PDF · Human presentation · Historical report v1
Horizon RunMap - architecture
Current overview · 2026-10-01. Horizon RunMap coordinates model representation, GPU execution, expert residency and evaluation as one implemented MoE inference architecture. The public technical source map preserves the contract details of the original performance edition.
The complete execution path
- Keep non-routed components on the measured FP16 execution path while selectively representing routed expert weights as compact INT4 qpacked data with FP16 scales.
- Route each token to its selected experts. Custom W4A16 GPU kernels use compact expert weights with FP16 activations; the recorded executor identifies the arithmetic path.
- Keep all relevant compact experts resident when the GPU budget permits. This is a complete operating regime of the architecture.
- Use partial residency when the budget is smaller. The documented immutable-cache contract keeps every routed expert available in RAM and moves needed compact copies to the GPU through bounded staging.
- Evaluate the resulting behavior through recorded campaigns, retained outputs and the public desktop app. The app completed the full IFEval process on another computer with an RTX 4060.
The RTX 5070 campaign records memory and speed for selected configurations; the published IFEval edition records scored responses. These observations retain their own execution contracts.
Representation and computation
| Dimension | Declared meaning |
|---|---|
| Routed expert source | INT4 qpacked; 256-value blocks and FP16 scales in the selected performance campaign |
| GPU expert representation | Compact qpacked weights on the W4A16 paths |
| Activation precision | FP16 activations; original checkpoint dtype remains a separate fact |
| Non-routed path | Checkpoint BF16 values materialized as FP16 in measured paths |
| Persistent full expert expansion | No persistent expanded FP16/BF16 expert cache in this contract; temporary kernel values and workspaces are separate |
| Accumulation | Executor-specific; UNKNOWN where not declared by a public experiment |
| CPU role | Host authority, staging and coordination in the immutable-cache contract; CPU utilization is a separate measurement |
Selective compression applies to routed experts. It is not whole-model INT4 quantization. Dequantization inside computation does not recover information lost during quantization.
Partial-residency contract: immutable authority and miss lifecycle
The resident-qpacked-cache-v1 contract defines a complete process-owned, page-touched host source before request admission. Identity, inventory, offsets and layout become immutable; the preparation artifact reader is closed. The complete RAM source remains available even when all expert copies fit on the GPU.
- A request captures a complete cache generation and acquires its routed expert group.
- An all-hit group reads existing compact GPU copies.
- A miss reserves bounded pinned staging and unpublished GPU spare rows.
- Transport runs from immutable pageable RAM to pinned staging to GPU. Staging is reused after its transfer completion event retires.
- Transfer completion precedes atomic publication of the complete routed group as a new generation.
- Prior-generation consumers retain their leases. Victim storage becomes reusable only after those consumers retire.
- Eviction discards GPU copies. It neither returns expert weights to RAM nor rewrites the source.
The first contract serializes miss transactions. Capacity is fixed before admission, and insufficient staging/spares fail startup. Before publication, a failed transaction leaves the old generation authoritative; ambiguous failure after publication faults the cache. These publication/lifetime guarantees are contract properties, not measurements of universal overlap or concurrency performance.
Regimes and executor boundaries
Partial and full residency are regimes within the compact-expert design. They do not imply identical kernels in each recorded experiment.
| Recorded path | Resident experts per routed layer | Executor |
|---|---|---|
| OLMoE 0924 and 0125 partial | 30, 40, 50, 60 of 64 | PRODUCT-core scalar |
| OLMoE 0924 and 0125 full | 64 of 64 | Grouped T3 |
| Qwen selected partial configurations | 20, 30, 39 of 60 | top4-fused-v1 |
An expert being resident does not mean it is routed for the current token. Physical staging/spare allocations do not count as published residency. OLMoE c60-to-c64 changes both residency and executor; no isolated causal residency effect follows from that transition.
Historical contracts remain distinct
offload-qpacked dequantizes a low-bit source into bounded FP16/BF16 execution slots. The older direct-INT4 mutable-tier mechanism used a different ownership and writeback lifecycle. Neither is renamed to the immutable cache. ADR 0009's original authorization was a default-off Qwen P0 slice; later OLMoE measurements have their own source and experimental authority. A diagram of the cache contract does not prove that every historical executor implements it.
Observation boundaries
The cache contract requires directly observed zero expert D2H, writeback, host rewrite, request SHA work and request source reads. These are required invariants, not zero values invented for an uninstrumented run. Missing observations remain UNKNOWN; unavailable platform or dependency capabilities remain UNSUPPORTED. See Evidence for what each public release actually verifies and machine-readable architecture for a compact representation.
Horizon RunMap - results and evidence
Current overview · 2026-10-01. The RTX 5070 campaign records execution, memory and speed across 13 configurations. The published IFEval edition evaluates all 541 prompts and retains every response; prompt-level loose accuracy is 45.66%. The public app completed the full IFEval process on another computer with an RTX 4060. Public app · Machine-readable claims.
Public app and external execution
The downloadable desktop app gives an external user a guided path to prepare the pinned public model and run IFEval. Its complete execution on the other RTX 4060 computer is a project milestone. The campaign's memory and speed numbers below belong to the RTX 5070; the published IFEval score belongs to its own retained-response edition. Project presentation · Download page.
Separate experiments
| Experiment | Population | Evidence / validation / publication / reproduction |
|---|---|---|
| Residency and executors | 13 configurations; 39 workers; 468 measured responses | E3 / valid / public / not_attempted |
| Historical BF16/Horizon comparison | 120 paired prompts | E3 / valid / public / not_attempted |
| IFEval | 541 prompts; 834 instructions | E1 standalone_benchmark / valid / public / not_attempted |
The three editions have separate populations, scorers and questions. Their original manifests and responses remain available for inspection.
Performance contract and results
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled decode = sum(N - 1) / sum(internal decode seconds); it excludes first-token wait. Overall output rate uses sum(N) / sum(internal response runtime). TTFT and response time are arithmetic request means. VRAM is a device-wide whole-worker driver peak including preparation, warmup and shared desktop use; RSS is a separate whole-worker process peak. H2D is a sum over measured requests, not a rate. Isolated prefill is UNKNOWN; OLMoE HTTP observation is UNSUPPORTED. Block spread is a descriptive range, not a confidence interval.
| Configuration / source | Executor | Decode token/s | Whole-worker VRAM GiB |
|---|---|---|---|
| OLMoE 0924 c30/64 | PRODUCT-core scalar | 23.00 | 4.11 |
| OLMoE 0924 c40/64 | PRODUCT-core scalar | 27.69 | 4.62 |
| OLMoE 0924 c50/64 | PRODUCT-core scalar | 31.50 | 5.15 |
| OLMoE 0924 c60/64 | PRODUCT-core scalar | 35.01 | 5.67 |
| OLMoE 0924 c64/64 | Grouped T3 | 61.26 | 5.99 |
| OLMoE 0125 c30/64 | PRODUCT-core scalar | 23.04 | 4.15 |
| OLMoE 0125 c40/64 | PRODUCT-core scalar | 27.12 | 4.62 |
| OLMoE 0125 c50/64 | PRODUCT-core scalar | 31.22 | 5.15 |
| OLMoE 0125 c60/64 | PRODUCT-core scalar | 35.44 | 5.67 |
| OLMoE 0125 c64/64 | Grouped T3 | 62.07 | 6.01 |
| Qwen c20/60 | top4-fused-v1 | 15.22 | 7.11 |
| Qwen c30/60 | top4-fused-v1 | 17.17 | 8.08 |
| Qwen c39/60 | top4-fused-v1 | 19.95 | 8.99 |
Capacities are per routed layer. Compare OLMoE c30-c60 within the PRODUCT-core scalar family; c64 is a separate Grouped T3 executor result. Qwen c20/c30/c39 use top4-fused-v1. Keep model revisions and executors attached to every value; a runtime-to-runtime ranking calls for a matched comparison.
Performance manifest · Contract/report · Verification method · Verifier
Historical behavioral comparison
Under 120 paired structured prompts, BF16 produced 29/120 structurally valid responses and Horizon produced 28/120. There were 8 paired structural regressions, 7 improvements, 19 observable-behavior parity cases and 86 directionless divergences. Structural conformity is not factual-answer correctness. The original campaign's publication sanitizer failure is retained in provenance; the CPU-only successor recovers the original responses without changing the frozen scorer.
Paired manifest · Contract · Summary · Audit method · ZIP
IFEval 541
Model: allenai/OLMoE-1B-7B-0924-Instruct, revision 7f1c97f440f06ce36705e4f2b843edb5925f4498. The complete declared population was scored using frozen evaluator inputs. All 467 EOS and 74 length-limit endings remain included.
| Metric | Numerator / denominator | Percent |
|---|---|---|
| Prompt strict | 217 / 541 | 40.11% |
| Prompt loose | 247 / 541 | 45.66% |
| Instruction strict | 433 / 834 | 51.92% |
| Instruction loose | 474 / 834 | 56.83% |
The complete IFEval population and all scored responses are public. Its frozen scorer defines these instruction-following metrics. External model-card numbers provide context for readers; the original contract and publication metadata retain the formal evidence classification.
IFEval manifest · Frozen contract · Public admission · All responses · Audit method · ZIP
Verification and reproduction
The performance verifier checks hashes and aggregate-copy consistency; that release does not include the complete raw-run bundle. Behavioral and IFEval verifiers recompute their declared scores from retained responses on CPU. None reruns the private inference runtime. Independent reproduction is not_attempted for all three editions. UNKNOWN denotes an absent observation; UNSUPPORTED denotes unavailable observation capability.
Architecture comparison matrix
Checked 2026-09-16. This is a source-scoped architecture comparison, not a benchmark. Cells summarize the specific documentation linked below; they are not exhaustive statements about every version, backend or configuration. UNKNOWN means the consulted source does not settle the property, not that the project lacks it. Mutable upstream documentation may change.
| Dimension | Horizon immutable cache [H] | Eliseev / Mazur [E] | MoE-Infinity [M] | KTransformers [K] | llama.cpp placement [L] |
|---|---|---|---|---|---|
| Expert source representation | INT4 qpacked, FP16 scales | HQQ mixed low-bit experts | Format-dependent; INT4 and FP4 paths documented | CPU INT4/INT8; other formats supported | GGUF; format depends on model |
| GPU cached representation | Qpacked copies | Quantized expert storage | Path-dependent; compact expert paths documented | GPU GPTQ supported; placement-dependent | Backend/format-dependent; dynamic cache UNKNOWN |
| Expert execution | GPU W4A16, recorded executor | GPU, HQQ-based path | GPU kernels; Marlin INT4 and FP4 paths listed | CPU/GPU hybrid expert computation | Configured CPU/GPU backend |
| Persistent full expansion | No expanded expert cache in this contract | UNKNOWN | UNKNOWN across all paths | UNKNOWN across all paths | UNKNOWN across all backends |
| Source authority | Complete immutable process RAM | Host expert backing; immutable-copy contract UNKNOWN | Host/SSD backing; immutable-copy contract UNKNOWN | Heterogeneous placement; immutable-copy contract UNKNOWN | Model loading/mapping; immutable-copy contract UNKNOWN |
| Eviction writeback | Discard copy; no expert D2H writeback | UNKNOWN | UNKNOWN | UNKNOWN | Dynamic expert eviction contract UNKNOWN |
| Partial/full relationship | Shared design; measured executors differ | Cache budget is configurable | Activation-aware GPU cache | Hot/cold CPU-GPU placement | GPU layer count and tensor placement controls |
| Publication/staging semantics | Bounded staging; completion before atomic generation publication | Per-layer LRU and speculative next-layer loading; atomic generation UNKNOWN | Tracing and prefetching; atomic generation UNKNOWN | Scheduling depends on selected path; atomic generation UNKNOWN | Dynamic publication contract UNKNOWN |
Sources and scope
- H: Horizon technical source map: immutable-cache contract, its original P0 scope and campaign executor boundaries. Current architecture explanation.
- E: Eliseev and Mazur, paper v1, sections 3-4: Mixtral expert offloading, LRU, speculative loading and mixed quantization. Author implementation. This paper is an especially close architectural precedent, not a matched Horizon baseline.
- M: MoE-Infinity README: host/SSD offload, activation-aware caching, prefetching and compact execution paths. The README explicitly distinguishes the released HuggingFace-oriented implementation from the paper version. This table describes the consulted implementation documentation.
- K: KTransformers README, inference capabilities: heterogeneous expert placement, CPU AMX/AVX computation, CPU quantization and GPU GPTQ. CPU computation is part of this documented path; Horizon's documented host role is authority and transfer coordination.
- L: llama.cpp server options:
--cpu-moe,--n-cpu-moe,--gpu-layersand tensor overrides. These placement controls alone do not specify a dynamic expert cache. No unnamed fork or proposed cache is attributed to upstream llama.cpp here.
Reading the comparison
Caching, quantization and offloading overlap across these projects. Horizon's explicit ownership, representation and publication rules are useful comparison coordinates; an unresolved cell cannot establish novelty. Matching those contracts requires inspecting a specific implementation path. Comparing performance additionally requires the same model revision, hardware, workload, precision, output policy and timing boundary. No such cross-runtime experiment is included in the published Horizon population.