If You Already Pay for an LLM Service, Running Local Embeddings and Rerankers Feels More Useful Than Running Local LLMs

preview.redd.it/v0xtn3jdu9ch1.png preview.redd.it/vjxiucsdu9ch1.png This post was originally written in Korean, then polished and translated into English using ChatGPT. I do run llama.cpp locally on a Tesla P40, but as someone who already pays for ChatGPT Pro, I was gradually losing the practical reason to keep running local LLMs like Qwen 3.6

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论