dfloat11

0.5.0
27.27k

DFloat11: Fast and memory-efficient GPU inference for losslessly compressed LLMs and diffusion models