thaw-vllm

The fork primitive for LLM inference. Snapshot a running session — weights + KV cache + scheduler state — and hydrate it into N divergent children that skip prefill. For RL rollouts, parallel coding agents, agent branching. Supports vLLM and SGLang.

thaw-vllm has been downloaded 9,929 times in total on PyPI, including 605 in the last 30 days. The latest version is 0.5.1, released Apr 23, 2026. It is distributed under the Apache-2.0 license.

Version0.5.1
Downloads
9.93k
LicenseApache-2.0
AuthorNils Matteson
UpdatedApr 23, 2026

Downloads

Weekly, last 90d.
Includes CI traffic.

VersionsTotal0.*
Range
View
Granularity
Group by
CI traffic
Stack: OffCI: Included2 / 17 series
Selected total6.24k
9.9kAll-time
605Last 30 days
1Last 24 h
<0.01/sPer second

Version distribution

Share of downloads by released version. Computed over the last quarter.

  • 01

    0.6.0

    729 downloads
    23.4%
  • 02

    0.5.1

    406 downloads
    13.0%
  • 03

    0.5.0

    249 downloads
    8.0%
  • 04

    0.4.0

    187 downloads
    6.0%
  • 05

    0.3.3

    164 downloads
    5.3%
  • 06

    0.3.2

    158 downloads
    5.1%
  • 07

    0.3.1

    146 downloads
    4.7%
  • 08

    0.2.1

    130 downloads
    4.2%
  • 09

    Other

    953 downloads
    30.5%

Guess the next day

Thirteen recent days of thaw-vllm downloads. Drag the green handle on the right to guess where day fourteen lands.

TRUTH162FRISATSUNMONTUEWEDTHUFRISATSUNMONTUETHU