"Strata" for GLM5.3 Flash is here for some! Project Maya

I stumbled across this as I was currently having glm5.3 flash run only around 10tok/s basically unusable. With Maya i am running 30 tok/s now. GLM5.3 feels even with this low quant like a much more enjoyable model than qwen 3.8 flash next so far. Still testing . This is very exciting allowing many to run GLM5.3 flash. github.com/mw00/project-maya Project Maya A 321-billion-parameter AI model on your own GPU GLM-5.3-Flash, private and fast · one NVIDIA GPU or up to 16 · AMD and Windows (experimental) · chat, pictures, OpenAI- and Anthropic-compatible API Models this size usually need a datacenter. Maya runs GLM-5.3-Flash (zai-org, MIT license) on the hardware you already have: 321 B parameters, about 18 B active per token, and a context of up to 1 M tokens. Its engine is built for this one model. It keeps the experts your conversation uses most on the GPU, the next ones in RAM and the rest on your NVMe SSD, and moves them as you chat. It measures your GPUs, CPU, RAM and SSD and tunes itself to them. Nothing leaves your machine. Fast: up to 118 tokens/s on four RTX 4090s, 30+ on a single RTX 5090. Real numbers from real users below. Close to the original: Maya's own quants keep 97.7-99.2% of the FP8 model's zero-shot accuracy. Plugs into your tools: OpenAI- and Anthropic-compatible API with tool calls, ready for coding agents. Everything in the box: a chat and a live monitor in the browser, pictures, thinking levels, a one-command setup.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论