Warm restore for vLLM

Warm up once.
Serve in seconds.

A new vLLM replica spends minutes compiling kernels, autotuning and capturing CUDA graphs before its first token. warmctl does that once, snapshots the warm engine and restores it on any matching GPU host in seconds, at full serving speed.

new replica · 8×H100
A new GLM-5.3-Flash replica, from snapshot to serving in 37 s. A cold start of the same engine takes 650 s.
Cold start

A new replica repeats the same work. Every time.

Before a vLLM replica serves its first token, it loads and shards the weights, compiles the model, JIT-compiles and autotunes kernels for every shape, captures CUDA graphs and warms up. For a large MoE on 8 GPUs that takes more than ten minutes.

Fast storage only shortens the first step: compilers and autotuners don't care where the bytes came from. warmctl does all of that once, at snapshot time, and every replica after that starts from the warm result.

GLM-5.3-Flash · FP8 · 8×H100new replica → serving
vLLM cold start
650 s
  • start processes, NCCL, build model15%
  • load weights7%
  • compile kernels, capture CUDA graphs39%
  • autotune GEMM kernels, 967 shapes17%
  • warm up, capture more graphs23%
warmctl restore
37 s
17×faster than cold start
100–136%of stock throughput after
bitwisecanary match with the donor

Cold start and restore measured on the same node with the same image and flags. The phase split comes from a separate cold start log of the same model on 8×H100 (673 s, weights already in page cache). Restore: CRIU 8.2 s, GPU state 10.5 s, weights 13.8 s, plus staging and admission. The throughput lead comes partly from the warmup grid the snapshot carries; stock is measured right after its own cold start.

Scale

Scale to zero. Scale to many.

The warm engine lives in the snapshot, not on a node. When traffic stops, release every GPU. When it comes back, restore as many replicas as you need, each on its own node and each serving in under a minute. No warm spares kept just in case.

GLM-5.3-Flash8×H100
0replicas serving
GPUs in use0
each replica37 s to serving
statescaled to zero
GLM-5.3-Flash

From zero to serving in under 60 seconds.

Per replica, snapshot to serving with a warm page cache: GLM-5.3-Flash 37 s, Qwen3-32B up to 21 s, gpt-oss-120b up to 16 s. Replicas restore independently, so fan-out is bounded by how fast your storage serves the snapshot. Today each replica is one warmctl restore; an autoscaling operator is on the roadmap.

Edge

The whole engine. On your terms.

Most snapshot tools stop at one GPU. warmctl freezes the engine you actually run in production, sharded across GPUs with its fastest collectives on, and leaves the weights, the registry and the code in your hands.

8×H100 SXMrestore
GLM-5.3-FlashTP837 s
Qwen3-32BTP4/814–21 s
gpt-oss-120bTP2/410–16 s
4×L40S · PCIe
T-pro 2.1TP214 s
Qwen3-8BTP28–9 s
TP > 1

Multi-GPU engines, restored whole

Proven at TP2, TP4 and TP8 on 8×H100 SXM and at TP2 over PCIe on 4×L40S. The production GLM-5.3-Flash engine, FP8 with MTP across eight GPUs, serves again in 37.3–37.5 s.

8×H100 SXMafter restore
custom all-reduce✓
FlashInfer all-reduce✓
symmetric-memory all-reduce✓
all-reduce + RMSNorm fusion✓
NCCL NVLS multicast✓
4×L40S · PCIe
custom all-reduce✓
ipc: auto

Fast collectives survive

vLLM's fastest all-reduce paths share GPU memory between processes, which cuda-checkpoint alone refuses to save. Under NVIDIA cuInterpose, warmctl saves shared allocations, multicast objects and IPC handles and re-creates them at the same addresses: 1.1 s at TP8 with NVLS. Served throughput after restore: 100–136% of stock vLLM with the same flags, ITL within 2–5%.

sha256

Weights apart from the warm state

Processed weights are exported straight from vLLM's sleep-mode allocator. Every snapshot file is a blob addressed by its SHA-256, so weights two snapshots share are stored and transferred once. On restore they stage in parallel with the runtime, are SHA-verified, go back to the same GPU addresses and are checked by fingerprints.

your infrastructure
  1. snapshoton your GPUs
  2. push · pull · lsyour OCI registry
  3. verify --deepevery file hashed
  4. restoreon your GPU hosts
OCI 1.1

Bake your own. Keep your weights.

Run warmctl snapshot on your own GPUs for any model vLLM serves with sleep mode; a profile is a few lines of YAML. push, pull and ls move snapshots as OCI 1.1 artifacts through any standard registry: GHCR, Docker Hub, ECR, Harbor. Your weights never leave your infrastructure.

Apache-2.0

Open source, end to end

warmctl is Apache-2.0, and so is the NVIDIA cuInterpose build that ships beside it. Use it, change it, build it into your own platform.

release archive · linux-amd64
warmctl-linux-amd64static Go binary + vLLM extension
cuinterpose/NVIDIA cuInterpose, Apache-2.0
LICENSE · NOTICEApache-2.0
How it works

Freeze once. Ship anywhere. Thaw in seconds.

A snapshot holds the whole process tree of a running engine, the GPU state of every worker and the processed weights of every rank. Weights travel separately from the warm state, so a restore is never a disguised cold start.

01

Freeze

warmctl starts vLLM in a managed container, warms it over every declared shape, puts it to sleep and checkpoints the processes with CRIU and the GPUs with cuda-checkpoint. Weights are exported straight from vLLM's sleep-mode allocator.

$ warmctl snapshot \
    --profile glm53-flash-prod-tp8 \
    --model-dir /models/glm \
    --gpus 0-7 --output /snaps/glm
02

Ship

A snapshot is a directory with a manifest, published atomically, or an artifact in any OCI registry. verify --deep hashes every file before you trust it.

$ warmctl push /snaps/glm \
    oci://ghcr.io/you/glm:h100-tp8
03

Thaw

warmctl checks the driver and GPUs, maps the process tree back, re-attaches the GPUs, reloads weights to the same addresses, runs canaries and only then opens an OpenAI-compatible API.

$ warmctl restore /snaps/glm \
    --gpus 0-7 --listen 0.0.0.0:8000
  serving started 37.3s
Guarantees

A restore counts only if it answers like the donor did.

“Zero degradation” without checks is marketing. Every restore runs the same checks before it admits a single request.

Matching host

Before anything is mapped back, the host must match the donor: GPU model, memory, driver and topology class, and every CPU flag the donor had.

Same addresses

CUDA graphs replay raw device pointers. The GPU address layout, communicators and every buffer come back exactly where they were.

Verified weights

Weights are checked by SHA-256 while staged and by fingerprints on the GPU, before any traffic.

Canary against the donor

The restored engine must return what the donor returned, token for token. Large models like GLM-5.3 run the canary bitwise.

Measured

Restore to serving.

17 models proven on H100 and L40S so far, most of them from a YAML profile alone, with no code change.

ModelTypeGPUsRestore
GLM-5.3-FlashMoE · FP8 · MTP8×H10037 s
Qwen3-32Bdense · BF164–8×H10014–21 s
gpt-oss-120bMoE · MXFP42–4×H10010–16 s
Qwen3-30B-A3BMoE · BF161×H10017–20 s
Gemma 4 26B-A4BMoE · BF161×H10014 s
Nemotron 3.5 LightningMamba MoE · FP81×H10011 s
T-pro 2.1dense 32B · BF162×L40S14 s
Qwen3-8Bdense · BF161×H1007–10 s

Warm page cache. A range covers every run of that model across its listed GPU counts.

CLI

One binary. Eight commands.

A static Go binary with its vLLM extension built in. No sidecar repositories, no hand-written host descriptions.

snapshot
bake a warmed engine into a snapshot
restore
bring a snapshot back to serving
push · pull · ls
move snapshots through any OCI registry
verify
check a snapshot against its manifest; --deep hashes every file
stop
stop a restored instance
gc
remove blobs no recorded snapshot uses
doctor
check this host
profile
list, check, show a profile or print its vLLM argv

Any model vLLM serves with sleep mode gets a profile in a few lines. vllm: takes any vllm serve flag; image, warmup grid and canary policy come from defaults.

llama31-8b-tp1.yaml
schema: warmctl.profile/v1
id: llama31-8b-tp1
model: {id: meta-llama/Llama-3.1-8B-Instruct, revision: 0e9e39f2}
parallel: {tp: 1}
hardware: [h100-80gb-sxm]
vllm: {max-model-len: 8192, kv-cache-memory-bytes: 21474836480}

Linux x86-64 · Docker + NVIDIA runtime · driver ≥ 580 · H100 SXM, L40S

Early access

Stop paying GPUs to warm up.

warmctl is in beta. Tell us which models you serve and on which GPUs, and we will profile them first.