Private lockers · cacheinfer.property

Serviced open-weight models

Private, and we set it up. Independent benchmark leaderboards increasingly place leading open-weight models at frontier-level quality across common reasoning, coding, and knowledge tasks. We handle the GPUs and the serving. You change one line of config.

Researched pair

Qwen3 4B → 14B

The published house pair. Prefill on 4B, decode on 14B after KV transfer. 4.3s → 1.7s first token. $11.8k → $4.6k prefill on 1M requests.

Gated on held-out quality and latency before the chute opens.

Serving

Your weights, our vault

We stand up private open-weight serving. GPUs, routing, and the transfer path stay in the icehouse. You keep one config line pointed at the endpoint.

Fallback

Ordinary prefill

If a pair does not clear both gates, the target model uses its usual inference path. The large vault still serves; the chute simply stays shut.

config.yaml · one line endpoint: https://cacheinfer.property/v1 # private bay · founder@cacheinfer.property

Open a private bay Our Research →