thaw-vllm

The fork primitive for LLM inference. Snapshot a running session — weights + KV cache + scheduler state — and hydrate it into N divergent children that skip prefill. For RL rollouts, parallel coding agents, agent branching. Supports vLLM and SGLang.

thaw-vllm has been downloaded 10,361 times in total on PyPI, including 432 in the last 30 days. The latest version is 0.5.1, released Apr 23, 2026. It is distributed under the Apache-2.0 license.

Version0.5.1
Downloads
10.36k
LicenseApache-2.0
AuthorNils Matteson
UpdatedApr 23, 2026

Downloads

Weekly, last 90d.
Includes CI traffic.

VersionsTotal0.*
Range
View
Granularity
Group by
CI traffic
Stack: OffCI: Included2 / 17 series
Selected total3.59k
10.4kAll-time
432Last 30 days
6Last 24 h
<0.01/sPer second

Version distribution

Share of downloads by released version. Computed over the last quarter.

  • 01

    0.6.0

    442 downloads
    24.6%
  • 02

    0.5.1

    142 downloads
    7.9%
  • 03

    0.5.0

    114 downloads
    6.4%
  • 04

    0.3.2

    91 downloads
    5.1%
  • 05

    0.3.1

    88 downloads
    4.9%
  • 06

    0.4.0

    87 downloads
    4.8%
  • 07

    0.2.0

    87 downloads
    4.8%
  • 08

    0.3.0

    86 downloads
    4.8%
  • 09

    Other

    658 downloads
    36.7%

Guess the next day

Thirteen recent days of thaw-vllm downloads. Drag the green handle on the right to guess where day fourteen lands.

TRUTH65TUEWEDTHUFRISATSUNMONTUEWEDTHUFRISATSUN
    thaw-vllm · 10.4k downloads on PyPI