Up to 3.2x Faster Inference with LFM2.5-DSpark
This one is ready to use now onwards, as its PR got merged today mentioned by u/jacek2023 But don't use DSpark GGUFs(testing versions) from that PR. Use the official GGUFs by them. You could find them on model cards. Anyway sharing the table below. Draft (GGUF) Target (GGUF) LFM2.5-1.2B-Instruct-DSpark-GGUF LFM2.5-1.2B-Instruct-GGUF LFM2.5-2.6B-DSpark-GGUF LFM2.5-2.6B-GGUF LFM2.5-8B-A1B-DSpark-GGUF LFM2.5-8B-A1B-GGUF --------- Never tried speculative decoding on Mobile. I use PocketPal & ChatterUI. Any idea how to run these on Mobile?
评论
?
参与讨论