Residency and executor campaign

13 configurations · 39 workers · 468 measured responses · E3

PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.

OLMoE c30–c60 use PRODUCT-core scalar; c64 uses Grouped T3. This is not a residency-only ablation. Qwen uses top4-fused-v1.

ConfigurationExecutorPooled decode token/sWhole-worker VRAM GiB
OLMoE 0924 c30/64PRODUCT-core scalar23.004.11
OLMoE 0924 c40/64PRODUCT-core scalar27.694.62
OLMoE 0924 c50/64PRODUCT-core scalar31.505.15
OLMoE 0924 c60/64PRODUCT-core scalar35.015.67
OLMoE 0924 c64/64Grouped T361.265.99
OLMoE 0125 c30/64PRODUCT-core scalar23.044.15
OLMoE 0125 c40/64PRODUCT-core scalar27.124.62
OLMoE 0125 c50/64PRODUCT-core scalar31.225.15
OLMoE 0125 c60/64PRODUCT-core scalar35.445.67
OLMoE 0125 c64/64Grouped T362.076.01
Qwen c20/60top4-fused-v115.227.11
Qwen c30/60top4-fused-v117.178.08
Qwen c39/60top4-fused-v119.958.99

Pooled decode excludes first-token wait. VRAM is a device-wide whole-worker driver peak including preparation, warmup and desktop use. RSS and H2D have separate measurement scopes. See each linked result for its revision, full metric planes and limitations.

Audit material

The verifier checks hashes and curated aggregate consistency. This edition does not supply the complete raw-run bundle or independent runtime reproduction.