vllm-ascend updates (GLM5.3-flash) on dual 310p Ascend cards

Alright, I learend a few valuable lessons since I posted about a week ago, and I've made more strides with the DDR4x Ascend 96GB 310P cards I purchased, so here we go, meet the new friendly and less verbose me. GLM-5.3-Flash is a 320-billion-parameter MoE model, with about 18 billion active per token. My checkpoint matches its 45-layer, 288-expert, top-8 routing configuration. Its a W2/W3/W4 quantization reduces weight storage; the parameter count stays the same. One of the lessons I had from working qwen-flash-next initially was that one of the expensive parts of running an experiment is reloading the model eachtime. Well now I can hotswap new modules/CANNs/server changes etc while keeping the core model weights resident. This has allowed me to accomplish in 1 week for GLM what might have taken 2 or 3 additional. When I started on GLM, the math was incoherent and I rented a nvidia rig on Shadeform to extract the "golden math" and used that to calibrate what I was doing for Ascend. Once that clicked, out of the gate I was getting roughly 8 seconds per generated token. Now, after numerous optimizations, I am currently getting about 8-9 tok/s. Cold-prefill is lengthy, when you submit a cold kilocode prompt, it does take a bit because a lot of additional prompt context is supplied behind the scenes. Subsequent prompts run quicker. As far as context, with the current implemetnation, the last verified configuration was 311,040 tokens per request, including prompt and generated output , processed in 640-token prefill chunks , with up to 4 concurrent requests . Third Ascend Card Arrived Friday I spent a day and a half optimizing my kernels for n-asscend cards, and got qwen3.8-flash-next going pretty well, which supports the image processing well now too by the way. I also implemented YaRN support for 1M context windows. qwen-flash-next is running great now This had been a ton of fun, but unfortunately my Nvidia RTX 6000 Pro sucked its power cable into the fan and long story short, it disapeared from the PCIe bus. In order to test it, since the threadripper buildout with the Ascend cards has a power supply with the 2x8 pin PCIe 12V 600 watts that card takes, I put the RTX card there and decided to start working on Ascend Sidecard. preview.redd.it/lvhcu2zffvuh1.png Ascend Sidecard I've been having this idea because I think one of things that drew me into the Ascend architecture was the storage capcity of the cores. You are getting almsot as much storage as an RTX 6000 for about 14% of the price. The idea for Ascend Sidecard is to split the model execution between CUDA (or other accelerators) and offload some of the model weights and processing to Ascend cards. Other Notes Thankfully codex is pretty solid at working through these fused kernels and improvements to keep memory transfers/casts to a minimum and eliminate redundant maths. Claude refuses to have anything to do with even discussing the Ascend architecture. Also I didn't mean to post my draft yet, so I'll have to provide other updates on Ascend Sidecard later on.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论