Qwen3 4B → 14B
The published house pair. Prefill on 4B, decode on 14B after KV transfer. 4.3s → 1.7s first token. $11.8k → $4.6k prefill on 1M requests.
Gated on held-out quality and latency before the chute opens.
Private lockers · cacheinfer.property
Private, and we set it up. Independent benchmark leaderboards increasingly place leading open-weight models at frontier-level quality across common reasoning, coding, and knowledge tasks. We handle the GPUs and the serving. You change one line of config.
The published house pair. Prefill on 4B, decode on 14B after KV transfer. 4.3s → 1.7s first token. $11.8k → $4.6k prefill on 1M requests.
Gated on held-out quality and latency before the chute opens.
We stand up private open-weight serving. GPUs, routing, and the transfer path stay in the icehouse. You keep one config line pointed at the endpoint.
If a pair does not clear both gates, the target model uses its usual inference path. The large vault still serves; the chute simply stays shut.