Not translated yet — showing the Chinese originals.
Industry & Practice · Tagged “论文”
Research & Benchmarks//7 min
加州大学伯克利等研究者报告,八个主流 Agent 基准都能在不真正解题的情况下拿到接近满分。采购和发布不该再把榜单分数当成能力本身。
Research & Benchmarks//8 min
Sierra 与普林斯顿的 CC BY 论文把「构建 Agent」本身做成任务。53 道题、四个领域里,最强配置 Claude Opus 5 加 Claude Code 只通过 23.9% 的评测仿真,专家参考上限是 82.2%。
Research & Benchmarks//6 min
Artificial Analysis 介绍的 Terminal-Bench 2.1 仍是 89 道硬任务,但修了环境和指令,让分数反映 Agent 能力而不是环境缺口。评测用 Terminus 2,每题三次取 pass@1。