Got my Ascent GX10 two days ago, ran REAP-pruned NVFP4 DeepSeek-V4-Flash on a single Spark, and it stays consistent at long context

Got my Ascent GX10 two days ago and spent the last couple of days pushing a REAP-pruned NVFP4 DeepSeek-V4-Flash setup on a single Spark by patching the eugr/spark-vllm-docker image. Credit where it’s due: the REAPs were done by 0xSero . I’m just the person who wired it up, validated it, and pushed it through the machine. The main thing I wanted to check was long-context consistency, and the interesting part is how steady the throughput stays as context scales up. I also vibecoded a Grafana dashboard in Herm

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论