Horizon RunMap - Architecture, Results and Public Evaluation

Current edition · 2026-10-01 · Lucas Ribeiro

What Horizon is

Horizon RunMap is an MoE inference architecture conceived and implemented independently by Lucas Ribeiro. It combines selective compression of routed experts, custom compact-weight GPU execution and controlled expert residency. Full residency is a complete operating path; when the GPU memory budget calls for partial residency, the documented cache contract coordinates movement from immutable host backing. A public desktop app provides an external evaluation path.

The published RTX 5070 campaign records execution, memory and speed. In its OLMoE 0924 full-residency configuration, Grouped T3 recorded 61.26 post-first-token decode tokens/s at 5.99 GiB whole-worker peak VRAM (source result). A separately published IFEval edition contains all 541 responses and its scored results. The public app also completed the full IFEval process on another computer with an RTX 4060. Explore the results · Get the app.

Engineering question

How can storage precision, execution precision, expert residency and data movement be coordinated under a constrained GPU-memory budget, while making memory, latency, output rate and output behavior separately inspectable?

Architecture invariants

The mixed-precision path keeps routed expert weights compact as INT4 qpacked data with FP16 scales in the selected campaign. Routing selects the needed experts, and custom GPU kernels execute them with FP16 activations. Non-routed checkpoint BF16 values are materialized as FP16 in the measured paths. When all relevant compact experts fit on the GPU, execution is fully resident. The documented resident-qpacked-cache-v1 contract adds bounded movement for partial residency: immutable host backing holds every routed expert, and disposable GPU copies are published only after complete transfer. See Architecture for the profile-specific lifecycle.

What is implemented

The published paths execute selectively compressed INT4 routed experts using FP16 activations (W4A16); weights remain compact in GPU storage rather than persisting as fully expanded expert matrices. Original checkpoint dtype, quantized source, scales, activation dtype and accumulation dtype remain distinct fields.

The selected campaign records partial residency with PRODUCT-core scalar OLMoE execution, full OLMoE residency with Grouped T3, and partial Qwen residency with top4-fused-v1. The earlier offload-qpacked expanded-slot mechanism retains its different meaning. The desktop app gives an external user a guided way to prepare the public model and run IFEval.

What evidence demonstrates

The RTX 5070 performance edition reports execution, internal timing, resource and transfer observations for 13 configurations. The published IFEval edition reports a complete 541-prompt instruction-following evaluation, including 45.66% prompt-level loose accuracy and every retained response. The earlier behavioral study compares 120 paired BF16 and Horizon responses under predefined structural checks. The app's complete IFEval execution on another RTX 4060 computer adds an external-machine milestone. Exact metrics, contracts and source files are in Evidence and Claims.

Scope of the published results

Each numerical result belongs to its recorded model, executor, hardware and workload. The performance campaign, published IFEval score, paired study and RTX 4060 app execution are distinct records. The public artifacts support inspection and their stated verification methods. The runtime source remains private by author decision; the public app enables execution on supported setups.

Closest prior art

Expert offloading, mixed quantization, GPU caching and heterogeneous inference are established techniques. Eliseev and Mazur's Mixtral offloading combines expert caching and speculative loading. MoE-Infinity documents activation-aware caching and prefetching. KTransformers documents CPU/GPU expert computation. llama.cpp exposes configurable tensor/layer placement. The architecture comparison matrix records source-scoped properties, including overlap and unresolved details, without a throughput ranking.

Key architectural differences

Horizon brings compact expert representation, complete immutable RAM authority, H2D-only miss transport, discard eviction and complete-generation publication into one explicit cache contract. The source-scoped matrix compares these design choices with related implementations and marks unresolved properties as UNKNOWN.

Experimental contracts

Measurement notes

The RTX 5070 campaign's whole-worker VRAM includes preparation, warmup and shared desktop use; process RSS has a separate scope. Its pooled decode rate excludes first-token wait. IFEval scores follow their frozen instruction checks, while the earlier 120-pair study uses structural checks. The release pages retain the full definitions and verification methods.

How to compare the recorded results

Compare OLMoE c30–c60 within the PRODUCT-core scalar family; c64 uses Grouped T3 and is a separate executor result. Read IFEval and the earlier paired study as separate populations with different scoring questions. Keep FP16 activations and execution values distinct from original checkpoint values. Read decode rate alongside first-token wait and total response time, and use measured transfers when analyzing movement.

Reviewer notes

Reconstruct the implemented system before judging its contribution. Assess the integrated architecture, app, recorded results and relevant prior work at their own levels. Keep the executor identity attached to each performance observation and the published IFEval score attached to its own edition. Reviewer brief · Public app.

Current report (HTML) · Markdown · PDF · Human presentation · Historical report v1

Horizon RunMap - architecture

Current overview · 2026-10-01. Horizon RunMap coordinates model representation, GPU execution, expert residency and evaluation as one implemented MoE inference architecture. The public technical source map preserves the contract details of the original performance edition.

The complete execution path

  1. Keep non-routed components on the measured FP16 execution path while selectively representing routed expert weights as compact INT4 qpacked data with FP16 scales.
  2. Route each token to its selected experts. Custom W4A16 GPU kernels use compact expert weights with FP16 activations; the recorded executor identifies the arithmetic path.
  3. Keep all relevant compact experts resident when the GPU budget permits. This is a complete operating regime of the architecture.
  4. Use partial residency when the budget is smaller. The documented immutable-cache contract keeps every routed expert available in RAM and moves needed compact copies to the GPU through bounded staging.
  5. Evaluate the resulting behavior through recorded campaigns, retained outputs and the public desktop app. The app completed the full IFEval process on another computer with an RTX 4060.

The RTX 5070 campaign records memory and speed for selected configurations; the published IFEval edition records scored responses. These observations retain their own execution contracts.

Representation and computation

DimensionDeclared meaning
Routed expert sourceINT4 qpacked; 256-value blocks and FP16 scales in the selected performance campaign
GPU expert representationCompact qpacked weights on the W4A16 paths
Activation precisionFP16 activations; original checkpoint dtype remains a separate fact
Non-routed pathCheckpoint BF16 values materialized as FP16 in measured paths
Persistent full expert expansionNo persistent expanded FP16/BF16 expert cache in this contract; temporary kernel values and workspaces are separate
AccumulationExecutor-specific; UNKNOWN where not declared by a public experiment
CPU roleHost authority, staging and coordination in the immutable-cache contract; CPU utilization is a separate measurement

Selective compression applies to routed experts. It is not whole-model INT4 quantization. Dequantization inside computation does not recover information lost during quantization.

Partial-residency contract: immutable authority and miss lifecycle

The resident-qpacked-cache-v1 contract defines a complete process-owned, page-touched host source before request admission. Identity, inventory, offsets and layout become immutable; the preparation artifact reader is closed. The complete RAM source remains available even when all expert copies fit on the GPU.

  1. A request captures a complete cache generation and acquires its routed expert group.
  2. An all-hit group reads existing compact GPU copies.
  3. A miss reserves bounded pinned staging and unpublished GPU spare rows.
  4. Transport runs from immutable pageable RAM to pinned staging to GPU. Staging is reused after its transfer completion event retires.
  5. Transfer completion precedes atomic publication of the complete routed group as a new generation.
  6. Prior-generation consumers retain their leases. Victim storage becomes reusable only after those consumers retire.
  7. Eviction discards GPU copies. It neither returns expert weights to RAM nor rewrites the source.

The first contract serializes miss transactions. Capacity is fixed before admission, and insufficient staging/spares fail startup. Before publication, a failed transaction leaves the old generation authoritative; ambiguous failure after publication faults the cache. These publication/lifetime guarantees are contract properties, not measurements of universal overlap or concurrency performance.

Regimes and executor boundaries

Partial and full residency are regimes within the compact-expert design. They do not imply identical kernels in each recorded experiment.

Recorded pathResident experts per routed layerExecutor
OLMoE 0924 and 0125 partial30, 40, 50, 60 of 64PRODUCT-core scalar
OLMoE 0924 and 0125 full64 of 64Grouped T3
Qwen selected partial configurations20, 30, 39 of 60top4-fused-v1

An expert being resident does not mean it is routed for the current token. Physical staging/spare allocations do not count as published residency. OLMoE c60-to-c64 changes both residency and executor; no isolated causal residency effect follows from that transition.

Historical contracts remain distinct

offload-qpacked dequantizes a low-bit source into bounded FP16/BF16 execution slots. The older direct-INT4 mutable-tier mechanism used a different ownership and writeback lifecycle. Neither is renamed to the immutable cache. ADR 0009's original authorization was a default-off Qwen P0 slice; later OLMoE measurements have their own source and experimental authority. A diagram of the cache contract does not prove that every historical executor implements it.

Observation boundaries

The cache contract requires directly observed zero expert D2H, writeback, host rewrite, request SHA work and request source reads. These are required invariants, not zero values invented for an uninstrumented run. Missing observations remain UNKNOWN; unavailable platform or dependency capabilities remain UNSUPPORTED. See Evidence for what each public release actually verifies and machine-readable architecture for a compact representation.

Horizon RunMap - results and evidence

Current overview · 2026-10-01. The RTX 5070 campaign records execution, memory and speed across 13 configurations. The published IFEval edition evaluates all 541 prompts and retains every response; prompt-level loose accuracy is 45.66%. The public app completed the full IFEval process on another computer with an RTX 4060. Public app · Machine-readable claims.

Public app and external execution

The downloadable desktop app gives an external user a guided path to prepare the pinned public model and run IFEval. Its complete execution on the other RTX 4060 computer is a project milestone. The campaign's memory and speed numbers below belong to the RTX 5070; the published IFEval score belongs to its own retained-response edition. Project presentation · Download page.

Separate experiments

ExperimentPopulationEvidence / validation / publication / reproduction
Residency and executors13 configurations; 39 workers; 468 measured responsesE3 / valid / public / not_attempted
Historical BF16/Horizon comparison120 paired promptsE3 / valid / public / not_attempted
IFEval541 prompts; 834 instructionsE1 standalone_benchmark / valid / public / not_attempted

The three editions have separate populations, scorers and questions. Their original manifests and responses remain available for inspection.

Performance contract and results

PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.

Pooled decode = sum(N - 1) / sum(internal decode seconds); it excludes first-token wait. Overall output rate uses sum(N) / sum(internal response runtime). TTFT and response time are arithmetic request means. VRAM is a device-wide whole-worker driver peak including preparation, warmup and shared desktop use; RSS is a separate whole-worker process peak. H2D is a sum over measured requests, not a rate. Isolated prefill is UNKNOWN; OLMoE HTTP observation is UNSUPPORTED. Block spread is a descriptive range, not a confidence interval.

Configuration / sourceExecutorDecode token/sWhole-worker VRAM GiB
OLMoE 0924 c30/64PRODUCT-core scalar23.004.11
OLMoE 0924 c40/64PRODUCT-core scalar27.694.62
OLMoE 0924 c50/64PRODUCT-core scalar31.505.15
OLMoE 0924 c60/64PRODUCT-core scalar35.015.67
OLMoE 0924 c64/64Grouped T361.265.99
OLMoE 0125 c30/64PRODUCT-core scalar23.044.15
OLMoE 0125 c40/64PRODUCT-core scalar27.124.62
OLMoE 0125 c50/64PRODUCT-core scalar31.225.15
OLMoE 0125 c60/64PRODUCT-core scalar35.445.67
OLMoE 0125 c64/64Grouped T362.076.01
Qwen c20/60top4-fused-v115.227.11
Qwen c30/60top4-fused-v117.178.08
Qwen c39/60top4-fused-v119.958.99

Capacities are per routed layer. Compare OLMoE c30-c60 within the PRODUCT-core scalar family; c64 is a separate Grouped T3 executor result. Qwen c20/c30/c39 use top4-fused-v1. Keep model revisions and executors attached to every value; a runtime-to-runtime ranking calls for a matched comparison.

Performance manifest · Contract/report · Verification method · Verifier

Historical behavioral comparison

Under 120 paired structured prompts, BF16 produced 29/120 structurally valid responses and Horizon produced 28/120. There were 8 paired structural regressions, 7 improvements, 19 observable-behavior parity cases and 86 directionless divergences. Structural conformity is not factual-answer correctness. The original campaign's publication sanitizer failure is retained in provenance; the CPU-only successor recovers the original responses without changing the frozen scorer.

Paired manifest · Contract · Summary · Audit method · ZIP

IFEval 541

Model: allenai/OLMoE-1B-7B-0924-Instruct, revision 7f1c97f440f06ce36705e4f2b843edb5925f4498. The complete declared population was scored using frozen evaluator inputs. All 467 EOS and 74 length-limit endings remain included.

MetricNumerator / denominatorPercent
Prompt strict217 / 54140.11%
Prompt loose247 / 54145.66%
Instruction strict433 / 83451.92%
Instruction loose474 / 83456.83%

The complete IFEval population and all scored responses are public. Its frozen scorer defines these instruction-following metrics. External model-card numbers provide context for readers; the original contract and publication metadata retain the formal evidence classification.

IFEval manifest · Frozen contract · Public admission · All responses · Audit method · ZIP

Verification and reproduction

The performance verifier checks hashes and aggregate-copy consistency; that release does not include the complete raw-run bundle. Behavioral and IFEval verifiers recompute their declared scores from retained responses on CPU. None reruns the private inference runtime. Independent reproduction is not_attempted for all three editions. UNKNOWN denotes an absent observation; UNSUPPORTED denotes unavailable observation capability.

Architecture comparison matrix

Checked 2026-09-16. This is a source-scoped architecture comparison, not a benchmark. Cells summarize the specific documentation linked below; they are not exhaustive statements about every version, backend or configuration. UNKNOWN means the consulted source does not settle the property, not that the project lacks it. Mutable upstream documentation may change.

DimensionHorizon immutable cache [H]Eliseev / Mazur [E]MoE-Infinity [M]KTransformers [K]llama.cpp placement [L]
Expert source representationINT4 qpacked, FP16 scalesHQQ mixed low-bit expertsFormat-dependent; INT4 and FP4 paths documentedCPU INT4/INT8; other formats supportedGGUF; format depends on model
GPU cached representationQpacked copiesQuantized expert storagePath-dependent; compact expert paths documentedGPU GPTQ supported; placement-dependentBackend/format-dependent; dynamic cache UNKNOWN
Expert executionGPU W4A16, recorded executorGPU, HQQ-based pathGPU kernels; Marlin INT4 and FP4 paths listedCPU/GPU hybrid expert computationConfigured CPU/GPU backend
Persistent full expansionNo expanded expert cache in this contractUNKNOWNUNKNOWN across all pathsUNKNOWN across all pathsUNKNOWN across all backends
Source authorityComplete immutable process RAMHost expert backing; immutable-copy contract UNKNOWNHost/SSD backing; immutable-copy contract UNKNOWNHeterogeneous placement; immutable-copy contract UNKNOWNModel loading/mapping; immutable-copy contract UNKNOWN
Eviction writebackDiscard copy; no expert D2H writebackUNKNOWNUNKNOWNUNKNOWNDynamic expert eviction contract UNKNOWN
Partial/full relationshipShared design; measured executors differCache budget is configurableActivation-aware GPU cacheHot/cold CPU-GPU placementGPU layer count and tensor placement controls
Publication/staging semanticsBounded staging; completion before atomic generation publicationPer-layer LRU and speculative next-layer loading; atomic generation UNKNOWNTracing and prefetching; atomic generation UNKNOWNScheduling depends on selected path; atomic generation UNKNOWNDynamic publication contract UNKNOWN

Sources and scope

Reading the comparison

Caching, quantization and offloading overlap across these projects. Horizon's explicit ownership, representation and publication rules are useful comparison coordinates; an unresolved cell cannot establish novelty. Matching those contracts requires inspecting a specific implementation path. Comparing performance additionally requires the same model revision, hardware, workload, precision, output policy and timing boundary. No such cross-runtime experiment is included in the published Horizon population.

Technical brief · Recorded experiments