kvcache-ai/ktransformers



A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
A Flexible Framework for Experiencing Cutting-edge LLM Inference/Fine-tune Optimizations🎯 Overview | 🚀 Inference | 🎓 SFT | 🔥 Citation | 🚀 Roadmap(2026Q2)
🎯 Overview
KTransformers is a research project focused on efficient inference and fine-tuning of large language models through CPU-GPU heterogeneous computing. The project now exposes two user-facing capabilities from the kt-kernel source tree: Inference and SFT.
🔥 Updates
June 21, 2026: MiniMax-M3 Day0 Support! ( Tutorial )
June 17, 2026: GLM-5.2 Day0 Support! ( Tutorial )
May 6, 2026: KTransformers at GOSIM Paris 2026 — "Agentic AI on Edge" track. We'll present KT's inference performance on consumer hardware.
May 02, 2026: DeepSeek-V4-Fl…