Amazon Science homepage

RSS: https://www.amazon.science/index.rss
亚马逊以客户为中心的科学方法。获取 AI 与机器学习创新的最新动态,包括职位、论文、会议与活动等。

SOP-Bench:评估AI代理执行真实商业流程的新基准

Amazon Science发布SOP-Bench,一个覆盖12个商业领域、含2000+任务的AI代理评估基准,已在KDD 2026发表。每个任务配备工具接口和标准答案,测试代理在真实SOP中的执行能力,而非仅靠文本生成得分。现有基准多只测单一能力,SOP-Bench填补了真实流程中多工具协调、歧义理解和错误恢复的评估空白。
评论点赞收藏4 天前

十年数学确定性:亚马逊自动化推理组的回顾与展望

亚马逊科学团队回顾其自动化推理组(ARG)十年历程:从2016年提出用数学逻辑证明AWS系统正确性,到如今Tiros、Zelkova等工具已支撑Amazon Inspector、IAM Access Analyzer、Bedrock Guardrails等数百万客户日常使用的服务。团队还证明了Nitro Confidentiality Engine等核心基础设施的正确性,并将Lean等证明助手与语言模型结合以扩展形式化验证规模。
评论点赞收藏14 天前

AWS Trainium Frontier竞赛:在专用AI芯片上协同设计模型与内核

AWS推出Trainium Frontier竞赛,让参赛者在Trainium2芯片上从零训练语言模型,探索硬件原生架构设计空间。竞赛分两阶段:第一阶段单芯片30分钟,以验证bits-per-byte评分;第二阶段进入Top10后使用完整服务器,增加CORE推理能力评估。最终得分是训练效率与推理能力的50/50综合。
评论点赞收藏15 天前

Why Triso Matters

Amazon is investing in next-generation nuclear technology to meet the rising energy demands of AI infrastructure and cloud computing, and at the heart of that technology are tristructural isotropic (T...
评论点赞收藏48 天前

Real-world grounding in agentic AI

The year 2026 marks a definitive shift in the AI landscape: we have moved from models that simply know to agents that do. Foundation models (FMs) — large Transformer models pretrained with massive dat...
评论点赞收藏78 天前

Ground truth is a process, not a dataset

Today, the key challenge in AI isn’t only how to build better models; it’s how to build evaluation systems that can keep up. Search-augmented AI systems can now produce deep research reports — long, p...
评论点赞收藏83 天前

Making LLMs faster without sacrificing accuracy

Large language models (LLMs) keep getting bigger and better. But the cost of running them — generating text, answering questions, powering real-time applications — is scaling up, too. Obviously, model...
评论点赞收藏102 天前

登录芦苇

登录后关注作者、收藏内容和参与讨论。