Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0: decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter) decode at 258,794: 42.9 to 45.0 prefill at 1,004,581: 790 to 937 tok/s , 21.2 to 17.9 minutes cold prefill at 258,794: 1,086 to 1,114 tok/s Conditions: Ryzen AI Max+ 395, 128 GB. The 262k and 1M rows are one cold request each at the 1M configuration ( HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576 ), greedy, 64 tokens, the rates the response's timings report. The 32k row is the standard ten-prompt served mean and did not change. A follow-up turn over the prompt cache at 1M reaches its first token in about 0.55 s; the numbers above are the cold path. To run it at 1M: add -e HALOGEN_ROPE_YARN=4 -e HALOGEN_CTX=1048576 to the README's podman line; it needs the 128 GB box. Release notes and the full table: github.com/peonist-ai/halogen-flash-server If you have a 1M sweep of your own, I would like to see it rerun on 0.12.0. Thanks for all your support, especially huggingface.co/nightvich

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论