Need maybe say "Use llama.cpp"
So I tried that miracle engine everyone is talking about. Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation: Can you help with the following problem? So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing). The thinking trace: We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated? ... 10k tokens later it degrades to: Need maybe maybe include "Use 'for code, list constraints'." Need maybe maybe include "Use 'for code, list requirements'." The same exact model in llama.cpp does produce a coherent answer without a doom loop.