Notes on MTP, EAGLE-3 and DFlash

Speculative decoding is a huge speedup for local inference of LLMs in llama-cpp (and similar programs). Recently I have tried to adjust the settings for MTP, EAGLE-3 and DFlash to get higher tokens per second on my NVidia 3060 RTX 12 GB GPU.
MTP, EAGLE-3, DFlash
For larger models like Qwen 3.6 35B and Gemma 4 26B, I have found that using a lower spec-draft-n-max setting (for llama-cpp ) is beneficial. I have found 1 works better than 2 or 3 (or 0) for MPT, EAGLE-3 and DFlash on these larger models.
I compared DFlash with EAGLE-3 in Gemma 4 26B, and found that with a spec-draft-n-max of 1, DFlash was slightly faster than EAGLE-3. For consumer GPUs like the 3060, it seems beneficial to reduce the speculation being done, and use smaller speculative models (the DFlash model I have is smaller than the EAGLE-3 one).
Overall, I found that MTP speculative decoding improved the tokens per second in my coding model (Qwen 3.6 35B AutoRound) from 24 to 39 tokens per second, which is a 38% speedup. It is important on some GPUs to use a lower value for spec-draft-n-max, even setting it to 1.