localllama (Reddit)

RSS: https://reddit.com/r/localllama/hot/.rss
r/LocalLLaMA,约 72.8 万成员的社区,专注于在本地硬件上运行大语言模型,涵盖模型选择、GPU 配置、量化技术、Ollama 与 llama.cpp 等工具,以及隐私优先的 AI 工作流。

Best local model for Blender and game dev?

I have been looking at some local models, but even the smartest ones like GLM5.3 Flash and DSv4 Flash have a hard time creating coherent models in Blender and placing them logically in game engines. I...
评论点赞收藏4 小时前

Gliner2.5-Decide (Jev style model)

I was perusing huggingface trending and was surprised that this hadn't been posted to localllama. Looks like it's been out for about a week.
评论点赞收藏11 小时前

Bartowski/AtomicChat 的 Ornith 1.5 35B A3B + sharp 模板,还是 Tiel Coder 35B A3B?

这篇 Reddit 帖讨论了 Peculiar Ragdoll 的 Tiel Coder 35B A3B 模型,指出它本质上是 Ornith 1.5 35B A3B 加上自家"编码优化"的 imatrix 量化和 sharp 聊天模板。 发帖人质疑:既然 Bartowski/AtomicChat 是已验证的量化器,与其直接用 Tiel Coder 权重,不如拿 Ornith 原权重让 Bartowski 量化后自己加上 sharp 模板,效果可能一样甚至更好。 核心问题是哪种量化器对编码场景更强,属于 LocalLLaMA 社区常见的技术选型讨论,信息密度中等,面向本地大模型量化部署的垂直读者有参考价值,但缺少实测数据或明确结论。
评论点赞收藏11 小时前

更新:Yandex/AliceAI 80B-A3B 微调进展

一位 Reddit 用户在 r/LocalLLaMA 更新了对 Yandex/AliceAI 80B-A3B 模型的微调进展:已完成约 40%,贴出了取自每步最后一个 micro 的 loss 曲线(波动已说明原因),并提到用了 Gemini 3.8 Flash 辅助。 作者还挂了一个 Cloudflare 临时域名直播训练过程,并链接到原帖。属于社区一手训练日志,有 live streaming 和 loss 细节,对关注本地大模型微调的读者有追更价值,但信息量有限、缺少方法论深度。
评论点赞收藏11 小时前

我们刚刚在 Hugging Face 开源了全球最快的 WebGPU 内核,用于本地 AI

Hugging Face 开源了 200+ 常见 ML 算子的 WebGPU 内核集合,声称是"全球最快"的 WebGPU 内核,可在浏览器中完全本地运行 AI 推理。 项目同步推进上游合并至 Transformers.js、ONNX Runtime Web、LiteRT.js 等框架。对关注本地推理、浏览器端 AI 部署的开发者是明确利好,信息增量在于算子覆盖广度与性能定位,以及多框架上游化路线图。 "最快"表述偏营销,实际基准需自行验证。
评论点赞收藏12 小时前

登录芦苇

登录后关注作者、收藏内容和参与讨论。