An architecture for AI with less graphics-card memory.
Some models have many specialized parts, but activate only a few at each step. Keeping them all on a graphics card takes a great deal of memory.
Horizon RunMap stores the parts compactly, computes without keeping expanded copies and controls what stays on the card; an intact source remains in computer memory.
Built independently by Lucas Ribeiro, without outside investment and on limited hardware.
The illustrated workshop accompanies the research; the memory arithmetic comes next.
01 / The idea
Why AI needs so much memory.
A GPU does the AI calculations. Its memory, called VRAM, holds the model weights — learned numbers that guide the response — and temporary space for the computation. Some models have many specialized parts, called experts, although only a few are activated at each step.
If the weights do not fit on one card, keeping them entirely on GPUs requires more capacity. Another option is to bring them from computer RAM when needed, but that movement can also affect response time. The example below shows the capacity arithmetic before discussing the solution.
Why does memory matter? Explore the large-model context
In MoE models, a selection of experts works on each token while the complete parameter set needs storage. Compare BF16 scale with each project’s published representation.
The scale of memory
The same model. Two memory scenarios.
GLM-5.3744 billion total parameters · 40 billion active per token
BF16 comparison
16 bits · 2 bytes per parameter
Estimated weight memory
≈ 1.49 TB
Mathematical minimum · weights only
9 B200
Capacity for execution planning
16 B200
2 servers with 8 GPUs
Calculated scenario reserving 20%–30% of total memory and using complete servers. Context and workload guide final sizing.
Published format
FP8 · 8 bits
Idealized weight memory
≈ 744 GB
Mathematical minimum · weights only
5 B200
Execution reference
8 B200 · documented configuration
The vLLM recipe describes FP8 weights and FP8 KV cache for up to 1 million context tokens. Concurrency is tuned to the memory budget.
FP8 is the default checkpoint; Z.ai also publishes a separate BF16 version.
The comparison assumes all weights reside in GPU memory, ideally distributed across devices. Weights ≈ total parameters × bits ÷ 8. Each B200 provides a nominal 180 GB; each server groups 8 GPUs, totaling 1.44 TB. GB and TB use decimal units.
Minimum GPUs = weights ÷ 180 GB, rounded up. Servers with headroom = weights ÷ (1.44 TB × 0.80) through weights ÷ (1.44 TB × 0.70), always rounded up. The reserve is a planning assumption selected for this comparison.
Calculations use rounded published parameter counts rather than actual file sizes. Scales, mixed-precision tensors, buffers, KV cache (context memory) and parallel layout contribute to the execution budget. GPU count and response speed are separate dimensions.
For GLM-5.3, we use Z.ai’s stated 744 billion total and 40 billion active parameters. The vLLM recipe describes approximately 743/39 billion; this difference leaves the rounded GPU counts unchanged in this comparison.
Frontier models show the scale of the memory and infrastructure challenge motivating Horizon RunMap. OLMoE and Qwen are the models used to investigate storage, transfer and execution on available hardware. Applying the architecture at larger scales requires separate validation; the rates below belong to the identified campaign.
Cost references per platform
Offers recorded on 2026-09-10, per platform or instance tier. The purchase and rental figures below have their own scope; a multi-server deployment also depends on networking and contracted services.
NVIDIA B200 (180 GB product) · recorded sources and estimates
GPU capacity, platform acquisition, rental, networking, electricity, cooling and operations are separate cost planes. This is not a complete quote, total cost of ownership or purchase recommendation.
02 / The solution
How Horizon makes room.
Horizon compacts routed experts while keeping the model's other parts outside that compression. The router selects which experts to use; custom kernels compute with their compact weights and scales.
Execution works with experts resident on the GPU. When memory is limited, the runtime can also keep the quantized source in RAM and bring over needed copies. The comparison below shows the capacity difference for the same model.
01
Separate the parts
Non-routed parts stay outside the compression; selectively used experts receive a smaller representation.
02
Compact the experts
Expert weights and scales use distinct formats: compact weights and FP16 scales.
03
Select and compute
The router chooses experts; the custom kernel combines their weights and scales with FP16 activations.
04
Adapt to memory
With room available, experts stay on the GPU. When needed, the runtime keeps the quantized source in RAM and brings over copies on demand.
The model combines non-routed parts executed in FP16 with experts whose weights are compact and whose scales are FP16. The router selects experts for each step; the custom kernel uses their activations and returns the result to the rest of the model.
The main execution path
One model, two ways to store its weights
In the measured configurations, compression is selective. It reduces routed experts while weights in the other parts follow the FP16 path.
Non-routed partsFP16
Attention, embeddings and other components do not enter expert compression.
ExpertsINT4 + scales FP16
Weights stay compact on the GPU. Scales determine how those values are interpreted during computation.
01
The router selects which experts work on this piece of the response.
02
The custom kernel reads compact weights and scales, uses FP16 activations and returns the result to the rest of the model. It does not keep an entire expanded expert on the GPU.
When the weights and execution space fit on the GPU, all compact experts can stay resident. The kernel works without fetching experts during the response.
Explore the technical details
In the partial-residency flow shown, the source of expert weights in RAM stays complete and immutable. A copy passes through bounded pinned host staging before transfer to the GPU. Eviction discards the GPU copy, with no transfer back or source writeback.
Residency is not activation: the cache may keep experts that this step does not use. This diagram explains organization; the campaign below measures partial and complete residency while preserving executor differences.
A copy enters the cache only when ready
Optional cache contract · resident-qpacked-cache-v1 · ADR 0009. Original scope: experimental, default-off P0 Qwen. The executors measured below are described by their specific contracts.
Immutable source
The complete qpacked set stays in process RAM, with validated identity and layout.
Bounded staging
A temporary pinned arena prepares transfer. It becomes reusable only after completion.
Unpublished spare
The qpacked row reaches temporary GPU storage; readers still see the previous generation.
Atomic publication
After routed-group completion, one transition publishes the new complete cache generation.
Use and discard
Readers use a valid generation. An evicted copy is reused only after its consumers retire; the source is not rewritten.
Resource admission
Capacity is fixed before requests. Insufficient staging or spare capacity rejects startup; failures before publication preserve the previous generation.
Earlier profiles have a different contract: offload-qpacked materializes FP16/BF16 execution slots from a quantized source. In the qpacked cache described here, experts remain compact. The kernel converts the values needed during computation without retaining a complete expanded expert slot. That conversion does not recover information lost during quantization.
The execution, memory and speed figures below come from a published RTX 5070 campaign covering OLMoE and Qwen configurations. Each result identifies its model, executor and measurement conditions.
Architecture in operation
The same principle was measured with all specialized copies on the card and with only some of them on the card.
The cards show different models and computation paths. Read each rate with its own model, memory and setup; they are not a direct model-to-model comparison.
Full residency
OLMoE 0924
The compact parts stayed on the graphics card during generation.
Computation path: compact experts computed in groups
Highest measured residency for each model · RTX 5070 · 36 responses of 128 tokens per configuration. OLMoE and Qwen use different executors; each card describes its own run.
Measured campaign
468
measured responses · campaign volume
39/39valid worker runs
13configurations
1h 50min 21scampaign duration
Public measurement under contract · Independent reproduction not attempted
36 responses of 128 tokens per configuration · 12 prompts × 3 blocks · 39 warmups excluded from rates.
Decode divides the sum of tokens after the first by the sum of internal generation times. It is a pooled, time-weighted rate. First-token wait and response duration are request means, with internal boundaries specific to each executor.
Decode
37min 06s
Until first token
35min 10s
Other stages (by subtraction)
38min 05s
Observed relative block spread: 0.27%–4.78%. This is observed variation, not a confidence interval. Recorded maximum temperature: 58 °C.
The consolidation recorded matching token sequences between repetitions of each configuration. This observation describes within-configuration consistency; quality evaluation and output comparisons across configurations have their own scope. Other-stage time is calculated as the remainder of controller duration.
Tokens are small pieces of text, not always whole words.
05 / Memory and movement
Keeping more parts ready: the trade-off.
Keeping more parts on the graphics card uses memory and can reduce transfers. The recorded configurations show how memory, movement and response times changed together. Each result identifies its computation path.
Contract validation: validE3 · controlled researchPublic measurement under contractIndependent reproduction not attemptedSep 11, 2026
36 responses per configuration · 128 tokens per response · 3 blocks
Hardware · NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Post-first-token decode · tokens/s Pooled rate: Σ(tokens − 1) / Σ(internal decode time). Excludes first-token wait; not a mean of individual rates.
Measurements after warmup, with fixed prompts. Research scope
Resident experts per layer · discrete measurements, without interpolation or cross-model comparison.
PRODUCT-core scalar
30/6423.00 tokens/s
40/6427.69 tokens/s
50/6431.50 tokens/s
60/6435.01 tokens/s
Grouped T3 · different executor
Between c30–c60 (scalar) and c64 (Grouped T3), residency and executor change together. The speed difference isolates neither offloading cost nor kernel speedup.
64/6461.26 tokens/s
View all 13 configurations
The table includes every configuration. On narrow screens, scroll the table sideways.
September 11, 2026 campaign · 36 responses of 128 tokens per configuration · E3 · controlled research Public measurement under contract · Independent reproduction not attempted NVIDIA GeForce RTX 5070 (frozen campaign declaration); shared Windows desktop. Decode = Σ(tokens − 1) / Σ(internal decode time).
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0924-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.95 GiB
RAM → GPU transfer (sum of measured responses)
1,108.10 GiB
Three block rates
23.12 / 23.14 / 22.74 tokens/s
Observed relative spread (not a confidence interval)
1.74%
Output rate including first-token wait
10.95 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0924-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.95 GiB
RAM → GPU transfer (sum of measured responses)
576.17 GiB
Three block rates
27.76 / 27.63 / 27.69 tokens/s
Observed relative spread (not a confidence interval)
0.45%
Output rate including first-token wait
14.11 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0924-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.95 GiB
RAM → GPU transfer (sum of measured responses)
230.88 GiB
Three block rates
31.75 / 31.57 / 31.19 tokens/s
Observed relative spread (not a confidence interval)
1.75%
Output rate including first-token wait
16.73 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0924-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.96 GiB
RAM → GPU transfer (sum of measured responses)
37.62 GiB
Three block rates
35.47 / 34.72 / 34.85 tokens/s
Observed relative spread (not a confidence interval)
2.14%
Output rate including first-token wait
19.31 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
OLMoE 0924 · 64/64 · Grouped T3 — contract and additional measurements
Metric
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0924-Instruct
Executor
Grouped T3
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.98 GiB
RAM → GPU transfer (sum of measured responses)
0.00 GiB
Three block rates
59.72 / 62.65 / 61.47 tokens/s
Observed relative spread (not a confidence interval)
4.78%
Output rate including first-token wait
45.41 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0125-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.97 GiB
RAM → GPU transfer (sum of measured responses)
1,115.87 GiB
Three block rates
22.65 / 22.98 / 23.51 tokens/s
Observed relative spread (not a confidence interval)
3.74%
Output rate including first-token wait
11.05 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0125-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.95 GiB
RAM → GPU transfer (sum of measured responses)
580.72 GiB
Three block rates
27.21 / 26.75 / 27.42 tokens/s
Observed relative spread (not a confidence interval)
2.47%
Output rate including first-token wait
13.90 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0125-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.95 GiB
RAM → GPU transfer (sum of measured responses)
235.35 GiB
Three block rates
31.10 / 31.31 / 31.25 tokens/s
Observed relative spread (not a confidence interval)
0.66%
Output rate including first-token wait
16.71 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0125-Instruct
Executor
PRODUCT-core scalar
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.96 GiB
RAM → GPU transfer (sum of measured responses)
40.81 GiB
Three block rates
35.45 / 35.49 / 35.39 tokens/s
Observed relative spread (not a confidence interval)
0.27%
Output rate including first-token wait
19.54 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
OLMoE 0125 · 64/64 · Grouped T3 — contract and additional measurements
Metric
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
allenai/OLMoE-1B-7B-0125-Instruct
Executor
Grouped T3
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
4.96 GiB
RAM → GPU transfer (sum of measured responses)
0.00 GiB
Three block rates
62.69 / 61.39 / 62.15 tokens/s
Observed relative spread (not a confidence interval)
2.10%
Output rate including first-token wait
46.67 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
Qwen · 20/60 · top4-fused-v1 — contract and additional measurements
Metric
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
Qwen/Qwen1.5-MoE-A2.7B-Chat
Executor
top4-fused-v1
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
7.97 GiB
RAM → GPU transfer (sum of measured responses)
2,702.75 GiB
Three block rates
15.08 / 15.03 / 15.56 tokens/s
Observed relative spread (not a confidence interval)
3.49%
Output rate including first-token wait
7.21 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
Qwen · 30/60 · top4-fused-v1 — contract and additional measurements
Metric
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
Qwen/Qwen1.5-MoE-A2.7B-Chat
Executor
top4-fused-v1
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
7.99 GiB
RAM → GPU transfer (sum of measured responses)
1,852.07 GiB
Three block rates
17.45 / 17.02 / 17.04 tokens/s
Observed relative spread (not a confidence interval)
2.51%
Output rate including first-token wait
8.59 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
Qwen · 39/60 · top4-fused-v1 — contract and additional measurements
Metric
Pooled rate: sum of tokens after the first / sum of internal decode times. Excludes first-token wait, preparation and warm-up; not a mean of individual rates.
Model / revision
Qwen/Qwen1.5-MoE-A2.7B-Chat
Executor
top4-fused-v1
Storage / execution
INT4 / block 256 / FP16 scales. FP16 execution values from low-bit source; original checkpoint dtype is a separate plane.
Peak process RAM (whole worker)
7.97 GiB
RAM → GPU transfer (sum of measured responses)
1,172.50 GiB
Three block rates
20.31 / 19.78 / 19.76 tokens/s
Observed relative spread (not a confidence interval)
2.74%
Output rate including first-token wait
10.52 tokens/s
Hardware
NVIDIA GeForce RTX 5070, as recorded in the campaign; shared Windows desktop.
Run contract
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
Campaign extract with original identifiers and source text in English. The full bundle and raw observations have separate availability.
06 / Follow a response
Follow one piece of a response.
A response is built one small piece of text at a time. The model selects parts for each step, and Horizon coordinates where those parts are kept and how they are used.
From input to response: follow how the work is organized. The next diagram shows what happens to expert copies. See hits, misses and eviction
Read the input
Text is split into tokens and processed by the model.
Select experts
Routing chooses experts at each routed layer.
Check the cache
A GPU copy can be used when present. On a miss, it arrives from RAM through staging.
Produce the next token
Layer calculations contribute to the next output. The process repeats.
07 / Back to the workshop
A reliable source in the workshop.
Now the workshop helps visualize the mechanism: the GPU is the workbench, RAM is the warehouse and experts are tools. A copy may be ready, need to be brought over or make room for another; the source remains intact.
The workbench
The GPU does the calculations. VRAM is the memory workspace it uses.
The warehouse
RAM keeps the complete, immutable source of the expert weights.
The tools
Experts are parts of the model. Each step uses a selection of them.
Imagine a workshop: the workbench has limited room and tools stay in storage. Here, each tool represents a set of model calculations.
One complete source. Copies to work with.
In every situation, the compact experts remain stored in RAM. This is the immutable source.
Hit
Use the ready copy
RAMSource intact
Already on GPU
GPUCopy available
The requested expert already has a ready copy on the workbench. Computation uses that copy.
The source stays intact; this lookup needs no new transfer of the expert.
Miss
Bring a copy
RAMSource intact
Copy to GPU
GPUCopy received
The requested copy comes from RAM through transfer staging. It becomes usable when transfer finishes and the new complete cache state is published.
The source stays ready in RAM, including after the transfer.
Eviction
Free space
RAMSource intact
Release on GPU
GPUCopy released
When it can safely be released, the copy leaves the workbench. Its location is reused only after the computations using that copy finish.
The source stays ready for another copy; weights do not need to return to the CPU.
Three possible situations, not a required sequence. The diagram explains the mechanism; quantities and timings are reported in each configuration’s results.Explore cache transfer and publication →
08 / Model behavior
How did the model respond in the tests?
To use less graphics-card memory, Horizon keeps expert weights compact and uses custom routines to compute with them. Other parts of the model follow a higher-precision path. This change can affect responses, so we tested whether the model still follows instructions.
How weights are stored and used in computation
In the measured qpacked architecture, routed expert weights remain compact INT4 in RAM and GPU memory. Kernels use FP16 activations. Other model components stay outside this INT4 quantization.
Stored format, computation format and accumulation precision are separate decisions. Compact weights are converted inside the kernel when needed; the complete expert does not occupy a persistent expanded slot.
01 / Store
Compact experts
INT4
Weights of routed experts stay compact in RAM and on the GPU. Attention, embeddings and other components remain outside this quantization.
Kernels are GPU computation routines. Horizon’s kernels use compact weights and convert the values needed during each operation, without retaining a complete expanded expert.
W4: routed expert weights in INT4, with FP16 scales. The GPU cache retains that compact format.
Compute
A16: FP16 activations. Kernels convert weights during computation, with FP32 accumulation stages and explicit rounding. W4A16 alone does not describe all arithmetic.
Other components
Attention, embeddings, routing and other components stay outside INT4. In this campaign, non-routed weights are materialized in FP16 from the BF16 checkpoint.
IFEVAL / OLMoE 0924
Did the model follow the instructions?
IFEval gives the AI written requests, such as including a required word or using a set format, then checks each response. This published run completed all 541 requests, so we can inspect how often the model followed every instruction in a request.
Requests meeting every instruction · allowed formatting adjustments45.66%
247 of 541 responses satisfied every instruction under loose evaluation.
Score from the published OLMoE run, separate from the other RTX 4060 computer.
What this result means
All 541 requests were completed, and every response can be inspected. The score shows how often this configuration met the written requirements in this test.
IFEval checks requirements such as required words, section counts and formatting. A prompt passes when all its instructions pass. Instruction-level scores count each requirement separately. Strict checks the response directly; loose applies predefined presentation adjustments.
The published model results provide context.
Horizon’s 45.66% sits numerically between two published IFEval figures for OLMoE 0924: 45.29 and 48.1. This is an encouraging reference point for the compact-expert configuration.
External references, evaluated under separate contracts. The later table does not specify the IFEval variant; these figures contextualize the result rather than establish a matched BF16 comparison.
How this benchmark was run
OLMoE-1B-7B-0924-Instruct · 64/64 experts per layer · Grouped T3 · INT4 expert sources / FP16 activations. Greedy decoding, no input truncation, up to 1,280 new tokens per prompt. All 541 prompts are included in the scores.
467/541 responses ended naturally at EOS; 74/541 reached the token limit. All 11 execution shards completed, with zero recorded per-request fallback, cache misses/fills or expert H2D transfers.
The benchmark measures instruction following. Factual correctness and quality equivalence to a matched BF16 run are separate evaluation questions.
E1 · standalone benchmark · verified and published under its contract. Independent reproduction of generation: not attempted.
The bundle contains the original responses, prompts, per-instruction scores, frozen evaluator sources, licenses and a CPU verifier. With the pinned Python dependencies installed, the verifier recalculates all four scores offline, without a model or GPU.
An earlier study gave the same 120 requests to Horizon and the original BF16 model. It checked whether each response followed the required structure and ended normally.
Conformity here means a completed, non-empty response containing one strict JSON object and ending naturally at EOS. It measures the requested format and termination; answer correctness was not graded.
This study is separate from IFEval and the performance campaign.
Explore the 120 pairs, conditions and audit
Where the outcomes changed
Both conform
21/120
Only BF16 conforms
8/120
Only Horizon conforms
7/120
Neither conforms
84/120
The 8 losses are 8/29 of the cases where BF16 conformed (8/120 overall). All eight Horizon outputs fail strict JSON parsing; three also reached the token cap. These are format or termination regressions, with no verdict on answer correctness.
A shared failure is not a regression attributed to Horizon. Of the 86 divergences without a quality direction, 82 fail conformity in both arms and 4 conform in both with different JSON values. Different wording or content alone does not establish worse quality.
How the 120 pairs are classified
How the 120 pairs are classified
Classification
Pairs / 120
Identical output tokens
16
Horizon conformity gain
7
Horizon conformity loss
8
Different outputs; neither conforms
82
Different JSON values; both conform
4
Equivalent JSON values
3
19/120 pairs have observable parity: 16 identical token sequences and 3 equivalent JSON values under the frozen numeric tolerance. Parity alone does not establish correctness.
Completion and termination
120/120 requests completed in each arm. All 120 Horizon records show zero fallback, cache misses/fills and expert transfers to the GPU.
Completion and termination
Observation
BF16
Horizon
Natural EOS
59
54
Reached 128-token cap
61
66
Empty responses
0
0
Conditions of this comparison
OLMoE-1B-7B-0924-Instruct, identical prompt order and tokenizer input hashes, greedy decoding, up to 128 new tokens. Horizon: c64, all experts resident, Grouped T3, INT4 expert sources and FP16 activations. Reference: original BF16 checkpoint values through Transformers/Accelerate with the recorded CPU/GPU placement.
This E3 comparison describes this model, prompt set and configuration. The performance campaign above is a separate experiment. Answer correctness, longer outputs and general quality preservation require their own evaluation.
Extract the ZIP, inspect the Python files, then run: python -B verify-bundle.py . — Python 3.10+, no packages, GPU or model required. It checks the original response hashes and recalculates all 120 paired classifications and the summary.
Publication record
The original publication failed because its sanitizer misread JSON escape characters. This successor corrects that post-processing in CPU, with unchanged scoring and all 240 original generations preserved. Generation validation and this bundle’s verification passed. No new generations were run; independent reproduction has not been attempted.
For organizations and individuals, the potential is to run AI models locally on a computer or infrastructure they control. Requests, documents and responses can stay in that environment instead of passing through an external inference service.
By reducing the memory required by experts, Horizon may make a smaller setup viable for the same application. If it meets the required speed and quality, purchase, rental and total infrastructure cost may fall. That gain needs to be evaluated for each model and workload.
Data under your control
Local models for information you manage.
An organization can work with internal documents on its own infrastructure; an individual can use files on their own computer. Local inference can keep requests and responses in that environment, with access and storage defined by its operator.
Compact experts may make it possible to use existing cards or rent a smaller setup. Memory, speed, RAM and operations all enter the comparison to see whether infrastructure actually costs less for that model and workload.
To inspect a complete run, the public app offers IFEval on validated setups. Explore the app →
Related work
HHorizon RunMapCompact experts and controlled copies
Residency
Partial or complete residency, depending on configuration and executor. Campaign capacities are per layer.
Storage precision
resident-qpacked-cache-v1 documents the source in RAM and GPU cache in qpacked format with a W4A16 path. offload-qpacked uses FP16/BF16 slots reconstructed from low-bit source; these are different mechanisms.
CPU role
In the documented immutable cache, the CPU keeps the complete source in RAM and coordinates staging, admission and publication. This is the CPU’s architectural role; utilization is a separate metric.
Movement granularity
Routed expert rows, per layer; bounded temporary transport arenas.
Placement policy
Cache capacity fixed before admission. Resident experts are not necessarily active for the current token.
Overlap strategy
ADR 0009 serializes miss transactions and publishes only after completion. The contribution described here is coordination of transfer and publication.
Hardware class
The selected campaign declares an RTX 5070 on a shared Windows desktop. This hardware defines the context of recorded capacities.
Disclosed workload
OLMoE 0924, OLMoE 0125 and Qwen; 13 configurations, 12 prompts repeated in three blocks, 128 tokens per measured response.
Benchmark maturity
Controlled E3 memory and execution results. OLMoE c64 also changes executor; general limits are in the research-scope section.
Documented mechanisms and measured scope are identified separately. ADR 0009 does not authorize generalizing its P0 profile to all campaign arms.
Horizon has demonstrated its architecture in measured configurations and a complete instruction-following evaluation. Useful next tests include more hardware and models, longer requests, concurrent use and the full operating cost of a deployment.
The recorded results retain their exact model, precision, executor, hardware and workload contracts. The earlier BF16 study measures format and termination separately; factual correctness, matched quality equivalence and production operation need separate evaluation.
Answer correctness has not been graded. The instruction test checks whether stated requirements were followed, not whether every answer was factually right.
Questions for reading the research
How should the results be read?
The rates record how Horizon operated in the published configurations. A performance comparison with another runtime is a separate research question: it requires comparable checkpoint, precision, hardware, prompts and generation settings.
Which data explain the cache?
H2D is the volume transferred from the host to the GPU. The explorer shows the total recorded over measured responses. Hits, misses and per-miss cost require their own counters and timings; that volume alone does not determine those measurements.
What can I verify in the downloads?
The performance extracts support integrity and aggregate consistency checks. The IFEval bundle supports recomputing all four metrics from 541 original responses. The earlier behavioral bundle supports recalculating all 120 pair classifications. Reproducing the full experiment also requires sufficient code, environment details and raw observations. The audit block describes the shared material and code availability.
How can storage precision, execution precision, residency and data movement be coordinated without requiring the entire MoE parameter set to occupy GPU memory in expanded form?
Full residency and offloading are operating regimes of this architecture. The measurements describe how the system operates within each recorded configuration.
OLMoE 0924 instruction following in IFEval and output conformity in a separate BF16 comparison.
What this work does not claim
The fastest MoE runtime.
State-of-the-art throughput.
Universal superiority over llama.cpp, vLLM or KTransformers.
Validated scaling to frontier-size MoE models.
Each measurement belongs to its stated model, precision, executor, hardware and workload. The contribution is the implemented architecture and the evidence recorded for it.
The scope of this campaign
The performance campaign includes OLMoE 0924, OLMoE 0125 and Qwen. Each configuration uses 12 prompts in three blocks, with 128-token responses.
Short checks and historical series use different conditions and remain separate from this campaign. Each set of files identifies its execution conditions.
Unknown (UNKNOWN) means absent or unmeasured information. Unsupported (UNSUPPORTED) means the platform or dependency cannot provide the observation.
Lucas Ribeiro conceived and developed Horizon independently, without outside investment and using limited hardware. The work spans the research question, architecture, custom execution, runtime, tests, recorded experiments, documentation and public evaluation app.
The journey began with NeuronMap research and grew into its own runtime. Reused code has recorded provenance; the technical reports show the engineering decisions and the conditions of each measurement.
The work from start to finish
ResearchFramed the memory problem and the experimental question.
ArchitectureDesigned how model parts are represented, stored, moved and executed.
ExecutionBuilt the runtime and graphics-card computation routines.
TestsValidated contracts and recorded performance and instruction following.
DocumentationPublished decisions, methods, limits and inspectable files.
Public appCreated a way to prepare and follow the evaluation on another computer.
Three engineering decisions and their records
01
Memory is a budget
Problem · A step activates a few experts, but the full set of weights must remain accessible.
Engineering decision · Distinguish storage, residency and execution; declare capacity per layer and profile.
How to interpret · This decision makes the space reserved for experts explicit in each configuration.
Problem · Replacing a copy in use must preserve the authority of weights and cache readers.
Engineering decision · Keep the source in RAM immutable, prepare new copies in temporary GPU locations and make the new cache state visible all at once, after transfer completion.
How to interpret · Calculations receive a complete set of copies ready for use, while the source in RAM remains intact.
Problem · OLMoE c60 and c64 change residency and executor together.
Engineering decision · Keep the recorded results available, separate Grouped T3 and place the contract beside the measurement.
How to interpret · The reading considers both changes together. The three block rates show recorded variation; isolating the causal effect of residency requires a different experimental design.
This is the research presentation: Horizon’s architecture, technical decisions and execution records.
Technical review
Analyze Horizon with AI.
The reviewer brief brings together architecture, measurements, limits and related work. The prompt asks readers to reconstruct the complete system before assessing it and to report which sources they accessed.
I welcome conversations about applications, research, partnerships and professional opportunities. LinkedIn is the direct contact path; GitHub offers a complementary technical profile.
Lucas Ribeiro
Creator of Horizon RunMap · independent research
One person connecting research, engineering and experimentation.
I’m Lucas Ribeiro. I conceived Horizon and develop it individually, with limited resources and no institutional funding. The project brings together my work in research, runtime development and experimental analysis.
I want to bring this ability to build and investigate to AI systems challenges, working with people and organizations that value thoughtful engineering and initiative.
The performance edition contains curated extracts from the completed campaign. You can check file integrity and consistency between each extract and the aggregate dataset.
Follow the guide to download the manifest and its 19 files into a new directory, preserving names. A PowerShell command downloads the complete edition.
With Node.js 22.19.0, run inside that directory:
node verify-evidence-edition.mjs --directory .
PASS confirms 19 file hashes and agreement of 13 aggregate copies. The checker uses only local files, without installing dependencies or running models.
Verification covers file integrity and consistency of this edition’s aggregates. Experiment reproduction and recomputation from raw observations require the full material and a separate procedure. Hashes check files against the supplied manifest; publisher authentication is a distinct check.
The runtime source code remains private by the author’s decision. This presentation shares the architecture, engineering decisions and documented results. Downloads include extracts, a report, technical notes and a verifier.
All 541 original responses, per-instruction scores and the frozen evaluator are available in a separate bundle. Its verifier recomputes the four metrics on CPU.
The 240 original responses and 120 pair classifications are in a separate bundle, with a manifest, report and Python verifier that recalculates the analysis.