Warm up once.
Serve in seconds.
A new vLLM replica spends minutes compiling kernels, autotuning and capturing CUDA graphs before its first token. warmctl does that once, snapshots the warm engine and restores it on any matching GPU host in seconds, at full serving speed.
A new replica repeats the same work. Every time.
Before a vLLM replica serves its first token, it loads and shards the weights, compiles the model, JIT-compiles and autotunes kernels for every shape, captures CUDA graphs and warms up. For a large MoE on 8 GPUs that takes more than ten minutes.
Fast storage only shortens the first step: compilers and autotuners don't care where the bytes came from. warmctl does all of that once, at snapshot time, and every replica after that starts from the warm result.
- start processes, NCCL, build model15%
- load weights7%
- compile kernels, capture CUDA graphs39%
- autotune GEMM kernels, 967 shapes17%
- warm up, capture more graphs23%
Cold start and restore measured on the same node with the same image and flags. The phase split comes from a separate cold start log of the same model on 8×H100 (673 s, weights already in page cache). Restore: CRIU 8.2 s, GPU state 10.5 s, weights 13.8 s, plus staging and admission. The throughput lead comes partly from the warmup grid the snapshot carries; stock is measured right after its own cold start.
Scale to zero. Scale to many.
The warm engine lives in the snapshot, not on a node. When traffic stops, release every GPU. When it comes back, restore as many replicas as you need, each on its own node and each serving in under a minute. No warm spares kept just in case.
From zero to serving in under 60 seconds.
Per replica, snapshot to serving with a warm page cache: GLM-5.3-Flash 37 s, Qwen3-32B up to 21 s, gpt-oss-120b up to 16 s. Replicas restore independently, so fan-out is bounded by how fast your storage serves the snapshot. Today each replica is one warmctl restore; an autoscaling operator is on the roadmap.
The whole engine. On your terms.
Most snapshot tools stop at one GPU. warmctl freezes the engine you actually run in production, sharded across GPUs with its fastest collectives on, and leaves the weights, the registry and the code in your hands.
Multi-GPU engines, restored whole
Proven at TP2, TP4 and TP8 on 8×H100 SXM and at TP2 over PCIe on 4×L40S. The production GLM-5.3-Flash engine, FP8 with MTP across eight GPUs, serves again in 37.3–37.5 s.
Fast collectives survive
vLLM's fastest all-reduce paths share GPU memory between processes, which cuda-checkpoint alone refuses to save. Under NVIDIA cuInterpose, warmctl saves shared allocations, multicast objects and IPC handles and re-creates them at the same addresses: 1.1 s at TP8 with NVLS. Served throughput after restore: 100–136% of stock vLLM with the same flags, ITL within 2–5%.
Weights apart from the warm state
Processed weights are exported straight from vLLM's sleep-mode allocator. Every snapshot file is a blob addressed by its SHA-256, so weights two snapshots share are stored and transferred once. On restore they stage in parallel with the runtime, are SHA-verified, go back to the same GPU addresses and are checked by fingerprints.
- snapshoton your GPUs
- push · pull · lsyour OCI registry
- verify --deepevery file hashed
- restoreon your GPU hosts
Bake your own. Keep your weights.
Run warmctl snapshot on your own GPUs for any model vLLM serves with sleep mode; a profile is a few lines of YAML. push, pull and ls move snapshots as OCI 1.1 artifacts through any standard registry: GHCR, Docker Hub, ECR, Harbor. Your weights never leave your infrastructure.
Open source, end to end
warmctl is Apache-2.0, and so is the NVIDIA cuInterpose build that ships beside it. Use it, change it, build it into your own platform.
Freeze once. Ship anywhere. Thaw in seconds.
A snapshot holds the whole process tree of a running engine, the GPU state of every worker and the processed weights of every rank. Weights travel separately from the warm state, so a restore is never a disguised cold start.
Freeze
warmctl starts vLLM in a managed container, warms it over every declared shape, puts it to sleep and checkpoints the processes with CRIU and the GPUs with cuda-checkpoint. Weights are exported straight from vLLM's sleep-mode allocator.
$ warmctl snapshot \ --profile glm53-flash-prod-tp8 \ --model-dir /models/glm \ --gpus 0-7 --output /snaps/glm
Ship
A snapshot is a directory with a manifest, published atomically, or an artifact in any OCI registry. verify --deep hashes every file before you trust it.
$ warmctl push /snaps/glm \ oci://ghcr.io/you/glm:h100-tp8
Thaw
warmctl checks the driver and GPUs, maps the process tree back, re-attaches the GPUs, reloads weights to the same addresses, runs canaries and only then opens an OpenAI-compatible API.
$ warmctl restore /snaps/glm \ --gpus 0-7 --listen 0.0.0.0:8000 serving started 37.3s
A restore counts only if it answers like the donor did.
“Zero degradation” without checks is marketing. Every restore runs the same checks before it admits a single request.
Matching host
Before anything is mapped back, the host must match the donor: GPU model, memory, driver and topology class, and every CPU flag the donor had.
Same addresses
CUDA graphs replay raw device pointers. The GPU address layout, communicators and every buffer come back exactly where they were.
Verified weights
Weights are checked by SHA-256 while staged and by fingerprints on the GPU, before any traffic.
Canary against the donor
The restored engine must return what the donor returned, token for token. Large models like GLM-5.3 run the canary bitwise.
Restore to serving.
17 models proven on H100 and L40S so far, most of them from a YAML profile alone, with no code change.
| Model | Type | GPUs | Restore |
|---|---|---|---|
| GLM-5.3-Flash | MoE · FP8 · MTP | 8×H100 | 37 s |
| Qwen3-32B | dense · BF16 | 4–8×H100 | 14–21 s |
| gpt-oss-120b | MoE · MXFP4 | 2–4×H100 | 10–16 s |
| Qwen3-30B-A3B | MoE · BF16 | 1×H100 | 17–20 s |
| Gemma 4 26B-A4B | MoE · BF16 | 1×H100 | 14 s |
| Nemotron 3.5 Lightning | Mamba MoE · FP8 | 1×H100 | 11 s |
| T-pro 2.1 | dense 32B · BF16 | 2×L40S | 14 s |
| Qwen3-8B | dense · BF16 | 1×H100 | 7–10 s |
Warm page cache. A range covers every run of that model across its listed GPU counts.
One binary. Eight commands.
A static Go binary with its vLLM extension built in. No sidecar repositories, no hand-written host descriptions.
- snapshot
- bake a warmed engine into a snapshot
- restore
- bring a snapshot back to serving
- push · pull · ls
- move snapshots through any OCI registry
- verify
- check a snapshot against its manifest; --deep hashes every file
- stop
- stop a restored instance
- gc
- remove blobs no recorded snapshot uses
- doctor
- check this host
- profile
- list, check, show a profile or print its vLLM argv
Any model vLLM serves with sleep mode gets a profile in a few lines. vllm: takes any vllm serve flag; image, warmup grid and canary policy come from defaults.
schema: warmctl.profile/v1 id: llama31-8b-tp1 model: {id: meta-llama/Llama-3.1-8B-Instruct, revision: 0e9e39f2} parallel: {tp: 1} hardware: [h100-80gb-sxm] vllm: {max-model-len: 8192, kv-cache-memory-bytes: 21474836480}
Linux x86-64 · Docker + NVIDIA runtime · driver ≥ 580 · H100 SXM, L40S
Stop paying GPUs to warm up.
warmctl is in beta. Tell us which models you serve and on which GPUs, and we will profile them first.