Adaptive speculative decoding: picking draft lengths at runtime

A follow-on to the economics of speculative decoding, we run the inference
lab simulator on MTP & DFlash drafters with real acceptance data, and find
out whether adaptively choosing the draft length is worth it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论