thaw-vllm
The fork primitive for LLM inference. Snapshot a running session — weights + KV cache + scheduler state — and hydrate it into N divergent children that skip prefill. For RL rollouts, parallel coding agents, agent branching. Supports vLLM and SGLang.
thaw-vllm has been downloaded 9,929 times in total on PyPI, including 605 in the last 30 days. The latest version is 0.5.1, released Apr 23, 2026. It is distributed under the Apache-2.0 license.
Version0.5.1
Downloads
9.93k
LicenseApache-2.0
AuthorNils Matteson
Downloads
Weekly, last 90d.
Includes CI traffic.
VersionsTotal0.*
Range
View
Granularity
Group by
CI traffic
Stack: OffCI: Included2 / 17 series
Selected total6.24k
9.9kAll-time
605Last 30 days
1Last 24 h
<0.01/sPer second
Sponsored
Sponsorships keep pepy free to read
Version distribution
Share of downloads by released version. Computed over the last quarter.
- 0123.4%
0.6.0
729 downloadsDownloads72923.4% - 0213.0%
0.5.1
406 downloadsDownloads40613.0% - 038.0%
0.5.0
249 downloadsDownloads2498.0% - 046.0%
0.4.0
187 downloadsDownloads1876.0% - 055.3%
0.3.3
164 downloadsDownloads1645.3% - 065.1%
0.3.2
158 downloadsDownloads1585.1% - 074.7%
0.3.1
146 downloadsDownloads1464.7% - 084.2%
0.2.1
130 downloadsDownloads1304.2% - 0930.5%
Other
953 downloadsDownloads95330.5%
Guess the next day
Thirteen recent days of thaw-vllm downloads. Drag the green handle on the right to guess where day fourteen lands.