Quark Support for HuggingFace Diffusers and SVDQuant

Diffusion models are heavy on memory and compute: a single text-to-image call runs a large transformer or UNet dozens of times. Quantization — storing weights (and sometimes activations) in low precision — is one of the most effective ways to cut both the memory footprint and the latency of these models.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论