localllama (Reddit)
r/LocalLLaMA,约 72.8 万成员的社区,专注于在本地硬件上运行大语言模型,涵盖模型选择、GPU 配置、量化技术、Ollama 与 llama.cpp 等工具,以及隐私优先的 AI 工作流。
Least to most expensive (Somewhat modern) GPU's with 32gb of vram (Under $1600) Based on ebay listings
I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.
Best local model for Blender and game dev?
I have been looking at some local models, but even the smartest ones like GLM5.3 Flash and DSv4 Flash have a hard time creating coherent models in Blender and placing them logically in game engines. I...
Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)
We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next. Victoria (coding and agents) Qwen3.8-Flash-Next cut down by 44% using a paper / technique called ...
peculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B
I have a long list of my own domain specific eval questions that I run to validate which models I can rely on: coding, coding (numpy/pandas), data analytics decision making, local RAG, and voice assis...
If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option?
Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to...
Index-Translate: 150 text languages, plus document translation, multilingual subtitles and dubbing
I'm part of the BiliBili Index LLM team. We're sharing Index-Translate and its companion models for translating text, documents, and videos. Index-Translate supports 150 text languages, with 2B, 9B, a...
Update: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark
Hi folks, I've spent the last couple of months experimenting with Strix Halo and previously I released a proof of concept I called llama-halo-hybrid. I've continued updating it and it now performs ver...
Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started
Framework Desktop Framework Desktop DIY Edition (AMD Ryzen™ AI Max 400 Series) 192GB
What's going on with DGX Spark? Price up $2k in 1 week?
I need to buy 2x DGX Spark but can't find a seller. Any hope or idea where I can get 2? The price seems to have soared. Local Microcenter had 25+ then 0 the next day. They pulling stock?
Another Ling model comes out, same receipt, 2 weeks free to use, then open source. Chinese labs do contribute a lot to open source community
Ling-3.1-flash: ~560B total params, ~25B active/token, up to 1M-token context. Across work, coding & healthcare: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professiona...
What was the mafia meeting of the tech criminals about at the White House?
Or should I assume it was mainly about trying to stop/ban/regulate China and open models,
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?
I've been closely following the rise of the Strata inference engine and as someone with 28GB VRAM and 32GB RAM I'm itching to buy 32GB more RAM just to use Flash Next. But of course before I make such...
Gliner2.5-Decide (Jev style model)
I was perusing huggingface trending and was surprised that this hadn't been posted to localllama. Looks like it's been out for about a week.
i would like to learn deeply about fine-tuning local models before burning money
There's so many new techniques like RL, RL LoRA, QLoRA, CPT LoRA. I believe i would have a usecase for them, but i don't know where to learn, youtube is filled with bad quality tutorials if i just sea...
Bartowski/AtomicChat 的 Ornith 1.5 35B A3B + sharp 模板,还是 Tiel Coder 35B A3B?
这篇 Reddit 帖讨论了 Peculiar Ragdoll 的 Tiel Coder 35B A3B 模型,指出它本质上是 Ornith 1.5 35B A3B 加上自家"编码优化"的 imatrix 量化和 sharp 聊天模板。 发帖人质疑:既然 Bartowski/AtomicChat 是已验证的量化器,与其直接用 Tiel Coder 权重,不如拿 Ornith 原权重让 Bartowski 量化后自己加上 sharp 模板,效果可能一样甚至更好。 核心问题是哪种量化器对编码场景更强,属于 LocalLLaMA 社区常见的技术选型讨论,信息密度中等,面向本地大模型量化部署的垂直读者有参考价值,但缺少实测数据或明确结论。
更新:Yandex/AliceAI 80B-A3B 微调进展
一位 Reddit 用户在 r/LocalLLaMA 更新了对 Yandex/AliceAI 80B-A3B 模型的微调进展:已完成约 40%,贴出了取自每步最后一个 micro 的 loss 曲线(波动已说明原因),并提到用了 Gemini 3.8 Flash 辅助。 作者还挂了一个 Cloudflare 临时域名直播训练过程,并链接到原帖。属于社区一手训练日志,有 live streaming 和 loss 细节,对关注本地大模型微调的读者有追更价值,但信息量有限、缺少方法论深度。
我们刚刚在 Hugging Face 开源了全球最快的 WebGPU 内核,用于本地 AI
Hugging Face 开源了 200+ 常见 ML 算子的 WebGPU 内核集合,声称是"全球最快"的 WebGPU 内核,可在浏览器中完全本地运行 AI 推理。 项目同步推进上游合并至 Transformers.js、ONNX Runtime Web、LiteRT.js 等框架。对关注本地推理、浏览器端 AI 部署的开发者是明确利好,信息增量在于算子覆盖广度与性能定位,以及多框架上游化路线图。 "最快"表述偏营销,实际基准需自行验证。