一位工程师每月向生产环境提交 2000 个 PR,验证是关键
SpaceXAI Grok 团队工程师 Lauren Tan 分享个人 agent 工作流 pstack:每月向生产环境合入约 2000 个 PR,接近每人每天 100 个。 她的核心判断是“验证”是关键基础设施而非普通技能——agent 能检查自己的产出并持续迭代直到任务完成,否则人类就成为循环中最慢的一环。 验证依赖一个 agent 可驱动、检查并获取结构化结果的运行时;单应用可以按需启动自身,但由数百个微服务组成的系统默认没有这种运行时,且要跟上上百个并行 agent 是难题。 她据此估算强验证能力可将团队产出放大 100 到 1000 倍。文章属于对一次个人实践的分析报道,数字本身是离群点,方向符合“coding agent 让同等人力产出十倍代码”的既有判断,但 2000 PR/月已真实落地到生产环境。
Vercel 收紧免费套餐部署保留规则,老部署将被直接删除
Vercel 本周调整 Hobby 免费套餐的部署保留策略:超过 10GB 部署存储上限的团队,原本最多保留 30 天的旧部署现在会被立即删除;每个 Hobby 项目保留规则也从原来的 10 个生产部署缩减为仅保留 3 个最近生产部署加 3 个最近任意类型部署,且预览部署的数量保护也同步取消。 Vercel 定价负责人 Jas Garcha 向 The New Stack 解释,原因是部署量增速远超一年前,需要为免费层可持续提供支持。
This week’s news from Zed, Anthropic, and OpenRouter shows why better harnesses matter more than better models
Our five most-read stories this week covered code collaboration, a model router, a UI change, a benchmark, and a caching tutorial. Five different stories about the same problem: Turning a model into s...
Kubernetes 能跑 AI 推理,但它算得出真实成本吗?
The New Stack 的 KubeCon 周报,盘点上周 Kubernetes 生态动态。重点包括:Kubernetes 运行 AI 推理的成本核算仍不清晰;Gartner 发布服务器虚拟化平台魔力象限,HPE 被归为 Challenger;Red Hat 作者在 K8s v1.37 中新增两项存储安全 Alpha 特性(bind mount 选项与 emptyDir 权限增强);另有 CNCF 项目更新。 属于例行生态周报,信息点覆盖面广但单条深度有限,适合想快速了解 K8s 近况的读者,讨论入口较弱。
How buildpacks help enterprises finally operate container security controls at scale
Container security controls often fail less because organizations lack standards, scanners, or best practices, and more. After all, every service tends to use its own Dockerfile, building images in sl...
Your agent is only as good as your infrastructure
You built a great agent, but something happened when it moved into production. In testing, your agent reviewed pull requests efficiently on its own. It read the diff, grepped the codebase for related ...
Open-weight models now handle a majority of tokens on Vercel’s AI Gateway. But Anthropic still takes 64% of the spend.
The trend is clear: open-weight models are taking an increasingly large bite out of production AI usage. On Monday, The New Stack reported that open-weight models accounted for 60% of OpenRouter’s US ... 

代码审查正在耗尽你最优秀的工程师
作者运营一个资深工程师社区,观察到 AI 代码审查正在成建制地压垮最在意代码质量的工程师:高 AI 采纳团队合并 PR 数量多 98%,审查时间多 91%,一项调研称 77% 工程师把写代码时间转去审查 AI 产出。 文章指出与人类同事相比,AI 生成代码"意图随行"丢失,审查者要从 diff 反推意图,任务类型已变;且 AI 代码表面连贯,易出现"看似合理实则错"、过度工程、边界误判等问题。 结尾给出团队层面应对思路(如把 review 队列视为系统瓶颈而非个人责任)。对 AI 时代工程管理与审查文化有信息增量,讨论入口多。
Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weight
The 1.58 in a 1.58-bit language model sounds like a hard limit, but Intel researchers pushed a ternary model below it by changing how its weights are stored rather than changing the model itself. Thei...
“Be transparent only if asked”: OpenAI’s models learned to leave notes for their future selves
OpenAI revealed Wednesday evening that some GPT-5.6 Sol model instances, during reinforcement learning (RL) training, wrote instructions to conceal mistakes or misaligned behavior from users. That’s n... 

GitHub and Anthropic used their own agents for major Rust rewrites — but with very different playbooks
Rust is seemingly the language of the moment, with open-source projects and companies forming an orderly queue to move core software over to the general-purpose programming lingo. The draw? A pursuit ...
Study: Developers are addicted to AI, and managers are making it worse
AI is addictive, and managers are rewarding those who use it the most (even though they are shipping stuff they don’t understand). What a tangled mess. Let’s unpack findings from the AI Coding Addicti... 

Why human oversight is shifting from writing code to defining requirements
This walks through the pipeline our agents operate inside—from a recorded scoping meeting through unit specs, spec review, generated code, PR checks, and automated QA, out to a weekly Thursday release...
Perplexity’s AI agents helped build a database. They weren’t allowed to run it.
Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB. Built by two engineers in two months with...
“Everyone’s in a race to replace GitHub”: Zed launches Delta because agents made pull requests obsolete
Something of a consensus has emerged from the developer fraternity in 2026 — GitHub, a platform built substantively for human developers, is no longer fit for purpose. Part of the problem is sheer vol... 

Anthropic bet users were choosing wrong. So it removed the choice.
Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? Anthropic has decided to eliminate that fork. Starting Wednesday, Claude Chat and Co...
AI evaluator: The most important AI job in history? How developers might fill the proposed new job
The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan t...
OpenAI 总裁:电脑应该为你赋能,别再为 AI 智能体重构软件
OpenAI 总裁 Greg Brockman 在 a16z 播客上提出一个明确判断:开发者正在"为世界重做"——给软件加 MCP 服务器、CLI、API 等连接器让 AI 智能体接入——这条路不如让 AI 直接像人一样使用电脑更简单。 他认为 AI 的北极星是"一个统一、足够易用,让你少和电脑较劲"的 AI。文中梳理了这条"计算机使用"思路从 2015 年 OpenAI 早期 offsite 三步计划到现在的演变,并指出其他公司仍在投入连接器基础设施。 可讨论点:computer use 的可靠性与成本 vs 连接器生态的成熟度,Brockman 的"简化论"是否会低估 MCP/工具调用的工程价值。
Meta lets Claude and Codex configure WhatsApp Business via MCP
Any business worth its salt in 2026 needs to be embracing the right tools to reach its customers, and few tools carry as much weight as WhatsApp. Paid messaging on the app crossed a $2 billion annual ... 
