At what point does local begin to make sense, not just for privacy, but also cost?
With Codex and Claude code limits getting nerfed heavily, I'm beginning to question when local inference begins to make sense. 24/7 access, no usage limits, no surprises, no silent model labotomies etc, you could pump a lot of tokens. It won't match frontier, but Qwen3.8 looks interesting and the cost of mac studios makes local seemingly make sense (if you're a 2+ 20X sub). Anyone else considering local inference?
评论
?
参与讨论