thaw-vllm
The fork primitive for LLM inference. Snapshot a running session — weights + KV cache + scheduler state — and hydrate it into N divergent children that skip prefill. For RL rollouts, parallel coding agents, agent branching. Supports vLLM and SGLang.
thaw-vllm has been downloaded 10,361 times in total on PyPI, including 432 in the last 30 days. The latest version is 0.5.1, released Apr 23, 2026. It is distributed under the Apache-2.0 license.
Version0.5.1
Downloads
10.36k
LicenseApache-2.0
AuthorNils Matteson
Downloads
Weekly, last 90d.
Includes CI traffic.
VersionsTotal0.*
Range
View
Granularity
Group by
CI traffic
Stack: OffCI: Included2 / 17 series
Selected total3.59k
10.4kAll-time
432Last 30 days
6Last 24 h
<0.01/sPer second
Sponsored
Sponsorships keep pepy free to read
Version distribution
Share of downloads by released version. Computed over the last quarter.
- 0124.6%
0.6.0
442 downloadsDownloads44224.6% - 027.9%
0.5.1
142 downloadsDownloads1427.9% - 036.4%
0.5.0
114 downloadsDownloads1146.4% - 045.1%
0.3.2
91 downloadsDownloads915.1% - 054.9%
0.3.1
88 downloadsDownloads884.9% - 064.8%
0.4.0
87 downloadsDownloads874.8% - 074.8%
0.2.0
87 downloadsDownloads874.8% - 084.8%
0.3.0
86 downloadsDownloads864.8% - 0936.7%
Other
658 downloadsDownloads65836.7%
Guess the next day
Thirteen recent days of thaw-vllm downloads. Drag the green handle on the right to guess where day fourteen lands.