What speeds are everyone getting with deepseek v4 flash 0731?

What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context window of 128000, -ub/-b at 4096, “q8” unsloth’s lossless quant

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论