Residency and executor campaign
13 configurations · 39 workers · 468 measured responses · E3
PERF-REAL-v1; 12 prompts repeated in 3 blocks; 36 measured responses of 128 tokens per configuration; greedy selection, EOS suppressed; 64-token warmup excluded. Canonical initial placement, evolving LRU through fixed prompts; concurrency 1.
OLMoE c30–c60 use PRODUCT-core scalar; c64 uses Grouped T3. This is not a residency-only ablation. Qwen uses top4-fused-v1.
| Configuration | Executor | Pooled decode token/s | Whole-worker VRAM GiB |
|---|---|---|---|
| OLMoE 0924 c30/64 | PRODUCT-core scalar | 23.00 | 4.11 |
| OLMoE 0924 c40/64 | PRODUCT-core scalar | 27.69 | 4.62 |
| OLMoE 0924 c50/64 | PRODUCT-core scalar | 31.50 | 5.15 |
| OLMoE 0924 c60/64 | PRODUCT-core scalar | 35.01 | 5.67 |
| OLMoE 0924 c64/64 | Grouped T3 | 61.26 | 5.99 |
| OLMoE 0125 c30/64 | PRODUCT-core scalar | 23.04 | 4.15 |
| OLMoE 0125 c40/64 | PRODUCT-core scalar | 27.12 | 4.62 |
| OLMoE 0125 c50/64 | PRODUCT-core scalar | 31.22 | 5.15 |
| OLMoE 0125 c60/64 | PRODUCT-core scalar | 35.44 | 5.67 |
| OLMoE 0125 c64/64 | Grouped T3 | 62.07 | 6.01 |
| Qwen c20/60 | top4-fused-v1 | 15.22 | 7.11 |
| Qwen c30/60 | top4-fused-v1 | 17.17 | 8.08 |
| Qwen c39/60 | top4-fused-v1 | 19.95 | 8.99 |
Pooled decode excludes first-token wait. VRAM is a device-wide whole-worker driver peak including preparation, warmup and desktop use. RSS and H2D have separate measurement scopes. See each linked result for its revision, full metric planes and limitations.
Audit material
The verifier checks hashes and curated aggregate consistency. This edition does not supply the complete raw-run bundle or independent runtime reproduction.