CodexQA

Industry & PracticeResearch & Benchmarks

研究更新:算法评估与整体评估

CodexQA 团队28 min read

METR 用 18 项真实开源任务对照两种评分:Claude 3.7 Sonnet 按维护者测试的成功率是 38%,15 次人工复核里没有一次能按原样合并。算法评分高估了智能体的整体可用性。

In this piece

研究更新:算法评估与整体评估

David Rein

2025 年 8 月 13 日

译注:标题里的 Holistic 与算法评估相对,按字面译作「整体」。

太长不读

  • 在来自两个大型开源仓库的 18 项真实任务上,2025 年初的人工智能智能体常常实现出功能上正确、却不能轻易按原样使用的代码,原因在于测试覆盖、格式化/静态检查,或总体代码质量方面的问题。
  • 这表明,许多基准所使用的自动评分1可能会高估人工智能智能体的真实世界表现。

译注:TL;DR 是英文缩写,字面意思是「太长不读」,此处用作小节标题。

智能体各次运行的桑基图

背景

许多人工智能基准使用算法评分2来评估人工智能系统在某一组任务上表现得有多好。例如,流行的基准 SWE-Bench Verified 衡量的是,一个人工智能系统是否通过由最初的人类拉取请求作者在代码里实现的测试用例。算法评分使得在该基准上评估一个新系统变得容易,因为这些测试可以复用、运行得很快,并且不需要人工介入/人工复核。

然而,许多目标很难用算法评分函数来表示。例如,评判文档的质量就很难自动完成——你无法(轻易地)编写出评估这一点的测试用例,因为文档通常是非结构化的自然语言。

特别是考虑到可验证奖励强化学习这类方法近来很流行,而这类方法依赖于算法评分函数,我们或许可以预期:人工智能系统在能够被自动评估的任务上,会比在不能被自动评估的任务上表现得更好。

如果人工智能在能够被自动评估的任务上要强出许多,我们就可能高估它们在实际工作中有多有用,因为我们常常把无法被自动评估的工作交给人工智能系统。我们假设,这一点促成了我们所测得的时间范围(https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)——Claude 3.7 Sonnet 大约一小时——与我们从最近的开发者生产率随机对照试验(https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)中观察到的放缓效应之间那种表面上的差距(该试验包含许多要花费人类一小时才能完成的任务)。

方法

在高层次上,我们评估自主智能体在尝试完成开源软件任务时的表现。我们比较用两种方法所评估出的它们的表现:自动/算法评分(使用由原始拉取请求作者实现的测试用例),以及人工/评分量规复核。

任务与开发环境

我们从两个仓库中取出了 18 个议题,这两个仓库属于我们最近一项衡量人工智能对有经验的开源开发者生产率之影响的随机对照试验(https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)。其中 15 项任务来自 stdlib-js(https://github.com/stdlib-js/stdlib)仓库,3 项来自 hypothesis(https://github.com/HypothesisWorks/hypothesis)仓库——它们都是大型仓库,代码行数分别为 8M 与 100k,而这 18 个议题花费人类维护者 20 分钟到四小时才完成(平均为 1.3 小时)。3 为了选出这些议题,我们筛选出由原始拉取请求作者实现了单元测试的议题,并按对人类而言的时长递增排序。从定性上看,这些仓库和议题代表了该随机对照试验所涉及仓库的更广泛分布。我们在附录里收录了若干任务示例。

为了给智能体一个自主完成这些任务的公平机会,我们针对每一项任务实现并测试了开发环境,使用的是原始人类拉取请求的基线提交。这给智能体提供了与人类仓库维护者相当的设定,因此它们不必(例如)自己安装依赖,并且可以立刻开始处理每一项任务。

在评估人工智能智能体完全自主地完成这些任务的能力时,一个顾虑是这些议题是否包含足以解决问题的上下文。例如,议题描述也许是草率的/简短的,或者依赖于智能体得不到的、开发者之间的讨论(例如经由电子邮件或 Slack)。然而,我们有信心这些拉取请求无需额外上下文就能被智能体解决,因为这些仓库里的全部交流都是公开进行的(这对大型开源项目很常见),而且这些仓库高度结构化、文档也很完善。从定性上看,这些议题清楚而且有意义——我们把它们全部做了人工复核,并认为解决它们并不需要任何私有知识或信息。

智能体

我们有一个 Inspect ReAct 智能体(https://inspect.aisi.org.uk/react-agent.html),它使用 Claude 3.7 Sonnet,并让它去尝试这些拉取请求。然后,我们取来由人类的拉取请求所实现的那些测试,并对照这些测试来评估智能体的解答。具体而言,我们衡量智能体的解答是否通过了人类原始参考拉取请求中所实现的全部测试用例。我们在任务环境里,直接对这些智能体的代码运行这些测试,以此来降低依赖或环境出现不一致的可能性。

我们并不向该智能体展示测试用例,这使它们能够进行奖励黑客的可能性低得多。宽泛地说,该智能体被指示去完整地实现所给定的议题。我们总共花了几个小时来迭代该智能体的指令,主要迭代的是与测试它们的解答、以及恰当地使用版本控制来提交它们的实现有关的那些指令。完整的提示放在附录里。

评估

然后,我们为每一项任务随机抽取一次运行4,并人工评估该智能体所提交的拉取请求。这一人工评估/评分的目标,是从整体上判断该智能体的拉取请求能否在不需要有意义的/实质性的额外工作的情况下,在该仓库中被合并。然后,我们把这些人工评分与智能体在上文所述参考测试用例上的评分进行比较。

我们针对五种失败模式来人工评估智能体的拉取请求:

  • 缺失/不正确的核心功能:任务的核心算法/逻辑功能没有被正确实现
  • 测试覆盖不充分:任务的主要组成部分没有写上充分的测试
  • 缺失/不正确的文档:关键文档缺失,或者不符合项目标准
  • 静态检查/格式化/类型问题
  • 其他代码质量问题:代码冗长、脆弱、难以阅读、难以维护,或者包含没有被「核心功能」所涵盖、但仍很可能阻止该拉取请求被合并的其他问题

我们在每一条轴上给每个智能体拉取请求打一个二元的 0 或 1 分,并且,如果一份提交不包含任何一种失败模式(也就是说,它可以按原样被合并),就认为这份提交是完整的。

结果

算法评分高估能力

处于 Inspect ReAct 智能体脚手架之中的 Claude 3.7 Sonnet,按维护者所编写的测试用例来衡量,在这些议题上的平均成功率为 38%(±19%,95% 置信区间),这表明该智能体常常正确地实现这些议题的核心功能组件5。然而,当人工复核这些拉取请求的一个子集时6,它们之中没有一个可以按原样被合并。这表明,即便是在衡量人工智能做真实软件任务的能力时——这常常被视为评估人工智能系统在软件开发中的真实世界表现的金标准方法——基于单元测试的算法评分也会高估模型的能力。

对于经过人工复核的 15 个拉取请求中的每一个,我们估计一名人类把该拉取请求修改/修复到可工作/可使用状态预计要花多长时间。我们估计,平均而言,这些智能体拉取请求要花费 42 分钟来修改/修复到可合并状态。当筛选到该智能体通过了人类所编写测试用例的那些运行时,我们估计这些拉取请求平均要花 26 分钟来修复(最短时间为 15 分钟)。因此,即便当该智能体正确地实现了核心算法功能(也就是说,它通过了全部由人类编写的测试用例),要让这些拉取请求达到可合并,仍然需要相当数量的工作——这些拉取请求花费最初那位维护者的平均时间是 1.3 小时,所以「修复所用时间」大约相当于这些拉取请求所需总时间的三分之一。请注意,这些估计是通过把智能体的拉取请求与参考的人类拉取请求相比较而作出的,并不是由原始的拉取请求作者作出的。这些估计应当被理解为一种非常粗略的/方差很高的主观判断,判断的是智能体的拉取请求离被完成还有多近。

奖励黑客

在人工复核智能体的拉取请求时,我们会加以检查,以确保当智能体通过了全部测试用例时,它们并没有作弊,也没有对测试做特殊情况处理。在若干情形中,该智能体经由 Github 找到了人类的拉取请求,因为原始仓库是开源的。我们排除这些运行,并人工检查那些议题上的所有其他运行(由此又暴露出另外几条该智能体找到了原始人类拉取请求的轨迹——我们也把这些从我们的分析中排除)。对测试做显式的特殊情况处理,对智能体而言会是困难的/不可能的,因为并没有事先向它们展示测试用例,而且我们确认这种情况并没有在它们的拉取请求里发生。

理解人工复核与算法评分之间的差距

对于 15 个经过人工评估的智能体拉取请求中的每一个,我们把失败归入各不相同的桶。请注意,这些类别意在捕捉智能体拉取请求最明显的问题——可能还有此处没有被表示、也没有被捕捉到的进一步问题,因此这些评分意在表示智能体表现的一个上界(也就是说,即便这些问题全部得到解决,仍可能存在没有被算法评分或这些类别捕捉到的其他问题)。每一次运行都可以有多项失败,我们报告的是包含每一个失败类别的智能体运行所占的百分比,分别针对未通过以及通过人类所编写测试用例的运行。

失败类型未通过测试用例的运行(n=11)通过测试用例的运行(n=4)
核心功能:任务的核心算法/逻辑功能没有被正确实现100%25%
测试覆盖:任务的主要组成部分没有写上充分的测试91%100%
文档:关键文档缺失,或者不符合项目标准89%775%
静态检查/格式化/类型问题73%75%
其他代码质量问题:代码冗长、脆弱、难以阅读、难以维护,或者包含没有被「核心功能」所涵盖、但仍很可能阻止该拉取请求被合并的其他问题64%50%

这 15 个拉取请求全都至少有上述问题中的三个,60% 至少有上述问题中的四个,并且 20% 的运行是五项问题全都出现。虽然每一项单独的失败也许相对容易修复(例如告诉该智能体「你忘了添加文档」),但要使一个拉取请求能够被合并,这些问题必须一项都不存在;而且因为大多数拉取请求包含了大多数问题,这就增加了把该拉取请求做到可合并的质量水平所需的工作量。

除了「核心功能」这一类别之外(我们本来就预期测试用例会首先抓住它),我们并没有看到这些问题的发生率会因为一次运行是未通过还是通过测试用例而出现重大差异,这表明这些失败与智能体在算法评分上的成功并没有强相关(尽管我们的样本量很小)。

要点/讨论

译注:Takeaways 按字面译作「要点」,Discussion 译作「讨论」。

走向与时间范围的调和

宽泛地说,这些结果帮助我们解释在最近的开发者生产率随机对照试验中所观察到的、出人意料的放缓效应(这些任务正是从该试验中取得的)。看起来,至少在这个相对具有代表性(但规模很小)的议题子集上,处于一个基本智能体脚手架之中的 Claude 3.7 Sonnet 能够以中等良好的程度实现这些任务的核心功能。然而,要真正达到可合并,往往还有许多其他重要目标需要被满足,而这个智能体无法在单条轨迹之中把它们全部满足。

一个有意思的比较点,是 Claude 3.7 Sonnet 在《Measuring AI Ability to Complete Long Tasks》(https://arxiv.org/abs/2503.14499)中的成功率(以及所估计的时间范围)。在那一任务分布(SWAA+HCAST+RE-Bench)上,Claude 3.7 Sonnet 对于大约要花一小时的任务,有着估计为 50% 的成功率。这与这些议题上 38% 的成功率相似;这些议题人类平均要花 1.3 小时,而这里是在使用算法评分进行评估。这是实质性的证据,表明当所评估的是模型实现核心逻辑/功能的能力时,这些任务并不比花费人类大约一小时来完成的 HCAST 任务更难。8

这些结果也指向:算法评分/评估是现有基准一项非常有意义的局限——因为它们并没有捕捉我们所关心的全部内容,在它们上面做爬山式推进,最终可能 a)放大奖励黑客之类的问题,以及 b)并不会在真实环境中带来相应的生产率改进。例如,前沿模型在 SWE-Bench Verified 上的成功率大约在 70-75%,但人工智能智能体目前实际上能够完整解决真实环境中 75% 的真实拉取请求,这一点似乎并不太可能。

注意事项

理解这些结果的一项重要背景是,所考虑的两个开源仓库,也就是 stdlib-js(https://github.com/stdlib-js/stdlib)和 hypothesis(https://github.com/HypothesisWorks/hypothesis),都是大型的、成熟的仓库,拥有数以百计的贡献者。为了把这种规模的仓库维护下去,对文档、代码质量、静态检查/格式化以及测试覆盖提出非常高的标准,是极其重要的。然而,许多软件工程项目或任务并没有在同等程度上提出这些要求,智能体在那些设定里(就这些非功能性目标而言)可能会表现得更好。

这些结果也可能因为选择效应而发生偏倚:为了作出这一比较,我们明确筛选的是那些编写了大量高质量测试用例的拉取请求。然而,相当一部分被写出来的代码通常并没有为它编写许多显式的测试用例,而且这一经过筛选的议题分布,与被写出来的拉取请求的「真实」分布之间,可能存在差异。例如,有可能这些议题比起议题的「真实」分布,范围划定得更清楚,因为它们适合拥有许多显式测试。

最后,我们并没有在我们的智能体脚手架上做大量的能力引出。我们使用的是一个基本的智能体脚手架,它并没有使用可观的并行推理计算,也没有大语言模型被设计来加以利用的专门工具,除了基本的文件查看/搜索/编辑、python 以及 bash 工具。我们或许可以设想,那种利用更多推理计算的更复杂脚手架,或者与大语言模型的后训练紧密集成的专门工具,可能会产生好上许多的结果。此外,情况也可能是:仅仅编写带有高质量少样本示例的、特定于项目的提示,就能够帮助缩小人工评分与自动评分之间的差距。我们并不认为这些结果代表智能体表现的上界——我们把这视为一个现实的设定,其中使用的是一个还不错的、标准的/常见的智能体脚手架。

对未来进展的含义

解释这些初步结果时,一个重要的问题是:这些失败究竟是根本性的,还是可以凭借更多或更好的引出/训练来解决。我们确实在定性上观察到,智能体在朝这些「更软」的目标推进时,取得了还不错(尽管并不完整)的进展,因此在这一不可验证的任务分布上的进展,有可能正以与核心算法组件上的进展相似的速度在改善,尽管是落在后面的。鉴于人工智能进展的步伐很快,理解能力的趋势,可能比在特定时间点上的单次测量更有用。

就最近的人工智能进展是由针对算法奖励信号来训练模型所推动的这一程度而言(例如借助 RLVR(https://arxiv.org/abs/2411.15124)),一个重要的问题是:大语言模型评判在更杂乱或更主观的任务上,能够在多大程度上替代可验证奖励。如果大语言模型评判能够很好地扩展,并且为更偏定性的任务提供有用的奖励信号,那么这里所观察到的、整体评分与算法评分之间的差距,就不太可能代表一种根本性的瓶颈/障碍。

也有可能,由于人工智能系统拥有一些人类所没有的特定可供性(更多被记住的知识、更好的短期记忆、更快的打字速度、更低的人力成本,等等),当人工智能在做软件工程时,围绕文档、格式化、测试或代码质量的这些标准,对于它们要把事情做对而言,重要性可能会更低。一般而言,未来从事软件工程的人工智能智能体,可能会以与当前由人类驱动的软件工程的开发模式和惯例非常不同的方式,来组织它们的工作。这看起来有些不太可能,因为随着项目在范围和复杂性上扩大,达成这些更软的、不可验证的目标,一般来说是有用的,但这是一个需要加以考虑的重要可能性。

附录:深入考察若干示例任务

Stdlib-35,「slice-grapheme-clusters」

译注:原文在此处把同一标题连续写了两次,译文按原顺序保留,不加以合并;Stdlib-35 与 slice-grapheme-clusters 为原标识,照录不译。

Stdlib-35,「slice-grapheme-clusters」

由人实现的测试用例有 26 个,测试的是 10 个不同类别的输入。

## [RFC]: add `string/base/slice-grapheme-clusters`

This task should implement a functional API for slicing a string based on grapheme clusters.

The behavior should not match the built-in `String.prototype.slice`, as the function should properly handle grapheme clusters (e.g., emoji). All arguments should be required:

 ```javascript
sliceGraphemeClusters( str, start, end )

Related: https://github.com/stdlib-js/stdlib/tree/develop/lib/node_modules/%40stdlib/string/base/for-each-grapheme-cluster


译注:上面这一段是原文代码块中的任务说明,标题只存在于代码块内,不是正文标题,故按原文照录,不译。

说明

这些说明是具体且没有歧义的,既展示了该函数的签名,又链接到一个实现了相关功能的模块。

tape( 'main export is a function', function test( t ) { t.ok( true, __filename ); t.strictEqual( typeof sliceGraphemeClusters, 'function', 'main export is a function' ); t.end(); });

tape( 'the function has an arity of 3', function test( t ) { t.strictEqual( sliceGraphemeClusters.length, 3, 'the function has an arity of 3' ); t.end(); });

tape( 'the function returns an empty string if provided an empty input string', function test( t ) { var out;

out = sliceGraphemeClusters( '', 0, 1 ); t.strictEqual( out, '', 'returns expected value' ); t.end(); });

tape( 'the function returns an empty string if the starting index is greater than or equal to the ending index', function test( t ) { var out;

out = sliceGraphemeClusters( 'hello', 2, 2 ); t.strictEqual( out, '', 'returns expected value' );

out = sliceGraphemeClusters( 'hello', 3, 2 ); t.strictEqual( out, '', 'returns expected value' );

t.end(); });

tape( 'the function returns an empty string if the starting index is greater than or equal to the string length', function test( t ) { var out;

out = sliceGraphemeClusters( 'hello', 5, 6 ); t.strictEqual( out, '', 'returns expected value' );

out = sliceGraphemeClusters( 'hello', 10, 12 ); t.strictEqual( out, '', 'returns expected value' );

t.end(); });

tape( 'the function slices an input string based on grapheme cluster indices', function test( t ) { var out;

out = sliceGraphemeClusters( 'hello', 0, 3 ); t.strictEqual( out, 'hel', 'returns expected value' );

out = sliceGraphemeClusters( '🌷🍕👉🏿', 1, 2 ); t.strictEqual( out, '🍕', 'returns expected value' );

out = sliceGraphemeClusters( 'अनुच्छेद', 1, 5 ); t.strictEqual( out, 'नुच्छेद', 'returns expected value' );

out = sliceGraphemeClusters( '六书/六書', 1, 5 ); t.strictEqual( out, '书/六書', 'returns expected value' );

out = sliceGraphemeClusters( '🏝️🌷', 1, 2 ); t.strictEqual( out, '🌷', 'returns expected value' );

t.end(); });

tape( 'the function slices an input string based on grapheme cluster indices (skin-tone emojis)', function test( t ) { var out;

out = sliceGraphemeClusters( '🌷👨‍👩‍👧‍👦👉🏿', 1, 2 ); t.strictEqual( out, '👨‍👩‍👧‍👦', 'returns expected value' );

out = sliceGraphemeClusters( '🏝️👨‍👩‍👧‍👦', 1, 2 ); t.strictEqual( out, '👨‍👩‍👧‍👦', 'returns expected value' );

out = sliceGraphemeClusters( '👋🏾🤦🏽‍♀️🧑🏿', 1, 2 ); t.strictEqual( out, '🤦🏽‍♀️', 'returns expected value' );

t.end(); });

tape( 'the function supports providing negative indices', function test( t ) { var out;

out = sliceGraphemeClusters( 'hello', -5, -2 ); t.strictEqual( out, 'hel', 'returns expected value' );

out = sliceGraphemeClusters( '🌷🍕👉🏿', -2, -1 ); t.strictEqual( out, '🍕', 'returns expected value' );

out = sliceGraphemeClusters( 'अनुच्छेद', -4, -1 ); t.strictEqual( out, 'नुच्छे', 'returns expected value' );

out = sliceGraphemeClusters( '六书/六書', -3, 5 ); t.strictEqual( out, '/六書', 'returns expected value' );

out = sliceGraphemeClusters( '🏝️🌷', -1, 2 ); t.strictEqual( out, '🌷', 'returns expected value' );

t.end(); });

tape( 'the function supports providing negative indices (skin-tone emojis)', function test( t ) { var out;

out = sliceGraphemeClusters( '🌷👨‍👩‍👧‍👦👉🏿', -2, -1 ); t.strictEqual( out, '👨‍👩‍👧‍👦', 'returns expected value' );

out = sliceGraphemeClusters( '🏝️👨‍👩‍👧‍👦', -1, 4 ); t.strictEqual( out, '👨‍👩‍👧‍👦', 'returns expected value' );

out = sliceGraphemeClusters( '👋🏾🤦🏽‍♀️🧑🏿‍🦱', -2, -1 ); t.strictEqual( out, '🤦🏽‍♀️', 'returns expected value' );

t.end(); });

tape( 'the function truncates the end index to the string length', function test( t ) { var out;

out = sliceGraphemeClusters( 'hello', 0, 10 ); t.strictEqual( out, 'hello', 'returns expected value' );

out = sliceGraphemeClusters( '六书/六書', 1, 10 ); t.strictEqual( out, '书/六書', 'returns expected value' );

out = sliceGraphemeClusters( '🏝️🌷', 0, 5 ); t.strictEqual( out, '🏝️🌷', 'returns expected value' );

t.end(); });

function sliceGraphemeClusters( str, start, end ) { var numClusters; var result; var idx; var brk; var i;

if ( str === '' ) { return ''; } numClusters = numGraphemeClusters( str ); if ( start < 0 ) { start = max( start + numClusters, 0 ); } if ( end < 0 ) { end = max( end + numClusters, 0 ); } if ( start >= numClusters || start >= end ) { return ''; } if ( end > numClusters ) { end = numClusters; } result = ''; idx = 0; i = 0; while ( idx < str.length ) { brk = nextGraphemeClusterBreak( str, idx ); if ( brk === -1 ) { brk = str.length; } if ( i >= start && i < end ) { result += str.substring( idx, brk ); } idx = brk; i += 1; if ( i >= end ) { break; } } return result; }

function sliceGraphemeClusters( str, start, end ) { var breaks; var count; var idx;

if ( !isString( str ) ) { throw new TypeError( format( 'invalid argument. First argument must be a string. Value: %s.', str ) ); } if ( !isInteger( start ) ) { throw new TypeError( format( 'invalid argument. Second argument must be an integer. Value: %s.', start ) ); } if ( !isInteger( end ) ) { throw new TypeError( format( 'invalid argument. Third argument must be an integer. Value: %s.', end ) ); }

// Count the total number of grapheme clusters and collect break positions breaks = findAllBreaks( str ); count = breaks.length + 1; // Total count is one more than break count

if ( count === 0 ) { return ''; }

// Handle negative indices if ( start < 0 ) { start = count + start; if ( start < 0 ) { start = 0; } }

if ( end < 0 ) { end = count + end; if ( end < 0 ) { end = 0; } }

// Return empty string for invalid slice ranges if ( start >= end || start >= count ) { return ''; }

if ( end > count ) { end = count; }

// Get starting position in the string (in terms of character indices) if ( start === 0 ) { idx = 0; } else { idx = breaks[ start - 1 ]; }

// Get ending position if ( end === count ) { return str.substring( idx ); }

return str.substring( idx, breaks[ end - 1 ] ); }

/**

  • Finds all grapheme cluster breaks in a string.

  • @private
  • @param {string} str - input string
  • @returns {Array<integer>} array of break positions

/ function findAllBreaks( str ) { var breaks = []; var idx = 0; var brk

if ( str.length === 0 ) { return breaks; }

while ( idx < str.length ) { brk = nextGraphemeClusterBreak( str, idx ); if ( brk === -1 ) { break; } breaks.push( brk ); idx = brk; }

return breaks; }


| 实现属性 | 智能体成功/失败✅→ 未观察到问题❌→ 观察到问题 |
| --- | --- |
| 核心功能:任务的核心算法/逻辑功能没有被正确实现 | ✅→ 未观察到问题 |
| 测试覆盖:任务的主要组成部分没有写上充分的测试 | ❌→ 观察到问题 |
| 文档:关键文档缺失,或者不符合项目标准 | ❌→ 观察到问题 |
| 静态检查/格式化/类型问题 | ❌→ 观察到问题 |
| 其他代码质量问题:代码冗长、脆弱、难以阅读、难以维护,或者包含没有被「核心功能」所涵盖、但仍很可能阻止该拉取请求被合并的其他问题 | ❌→ 观察到问题 |

人工编写的参考测试

人类实现与智能体实现

人类核心实现

智能体核心实现

评分

估计修复时间:15 分钟

要点

人工编写的测试用例非常详尽,因此该智能体的解答能够通过这些测试用例,就是它在功能上正确的有力证据。把智能体的解答与人类的解答相比较,智能体的解答更为冗长,没有使用已经存在的 `numGraphemeClusters` 函数,并且实现了一些冗余的类型检查,但它们在功能上是相同的。请注意,这个拉取请求还需要数量可观的额外代码——我们在这里只呈现核心的算法功能。

### hypothesis-1「TypeAlias support」

5 个类别中有 10 项测试。

work for type alias (typing.TypeAliasType)

Given a type alias (which is of type TypeAliasType) defined like:

>>> type Point = tuple[float, float]
>>> type(Point)
<class 'typing.TypeAliasType'>

I want to create a strategy and an example from that:

from hypothesis import strategies as st
st.from_type(Point).example()

which results in an error

...
.venv/lib/python3.13/site-packages/hypothesis/strategies/_internal/core.py", line 1287, in _from_type
    raise InvalidArgument(f"{thing=} must be a type")  # pragma: no cover
    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
hypothesis.errors.InvalidArgument: thing=Point must be a type

Accessing the underlying value through \_\_value__ works as expected:

>>> st.from_type(Point.__value__).example()
(1.192092896e-07, -7420061444470370.0)

I am not familiar with the internals of hypothesis but since Point.__value__ works, I am inclined to think Point should also just work by handling typing.TypeAliasType internally.

python version: Python 3.13.1 (main, Dec 3 2024, 17:59:52) [GCC 14.2.1 20241116] on linux hypothesis version: 6.123.17

Comments:


That makes sense to me - we could handle it just below NewType, and in much the same way. Tests would go in a new tests/cover/test_typealias_py312.py file, and add this config for 3.12 so it doesn't give SyntaxError on older versions.


译注:上面这一段是原文代码块中的议题说明。其中的「### Comments:」只存在于代码块内,不是正文标题,故按原文照录,不译。

def test_resolves_simple_typealias(): type MyInt = int type AliasedInt = MyInt type MaybeInt = int | None

assert_simple_property(st.from_type(MyInt), lambda x: isinstance(x, int)) assert_simple_property(st.from_type(AliasedInt), lambda x: isinstance(x, int)) assert_simple_property( st.from_type(MaybeInt), lambda x: isinstance(x, int) or x is None )

find_any(st.from_type(MaybeInt), lambda x: isinstance(x, int)) find_any(st.from_type(MaybeInt), lambda x: x is None)

def test_resolves_nested(): type Point1 = int type Point2 = Point1 type Point3 = Point2

assert_simple_property(st.from_type(Point3), lambda x: isinstance(x, int))

def test_mutually_recursive_fails():

example from

https://docs.python.org/3/library/typing.html#typing.TypeAliasType.__value__

type A = B type B = A

I guess giving a nicer error here would be good, but detecting this in general

is...complicated.

with pytest.raises(RecursionError): find_any(st.from_type(A))

def test_can_register_typealias(): type A = int st.register_type_strategy(A, st.just("a")) assert_simple_property(st.from_type(A), lambda x: x == "a")

def test_prefers_manually_registered_typealias():

manually registering a type A = ... should override automatic detection

type A = int

assert_simple_property(st.from_type(A), lambda x: isinstance(x, int))

with temp_registered(A, st.booleans()): assert_simple_property(st.from_type(A), lambda x: isinstance(x, bool))

if types.is_a_type_alias_type( thing ): # pragma: no cover # covered by 3.12+ tests if thing in types._global_type_lookup: strategy = as_strategy(types._global_type_lookup[thing], thing) if strategy is not NotImplemented: return strategy return _from_type(thing.__value__)

if types.is_a_type_alias(thing):

Access the underlying type via __value__

return _from_type(thing.__value__)


| 实现属性 | 智能体成功/失败✅→ 未观察到问题❌→ 观察到问题 |
| --- | --- |
| 核心功能:任务的核心算法/逻辑功能没有被正确实现 | ❌→ 观察到问题 |
| 测试覆盖:任务的主要组成部分没有写上充分的测试 | ❌→ 观察到问题 |
| 文档:关键文档缺失,或者不符合项目标准 | ❌→ 观察到问题 |
| 静态检查/格式化/类型问题 | ✅→ 未观察到问题 |
| 其他代码质量问题:代码冗长、脆弱、难以阅读、难以维护,或者包含没有被「核心功能」所涵盖、但仍很可能阻止该拉取请求被合并的其他问题 | ✅→ 未观察到问题 |

说明

人工编写的测试用例

人类实现与智能体实现

人类核心实现

智能体核心实现

评分

估计修复时间:20 分钟

要点

我们可以看到,所编写的测试覆盖了自定义注册的策略,而这个智能体实现确实没有通过最后两项测试,因为它并不返回自定义策略(如果该策略存在的话)——它总是返回为底层值的类型所定义的那条策略。当智能体为已经注册的类型别名返回自定义策略时,测试就会通过。其余测试覆盖的是核心实现,包括嵌套的类型别名,并确认设置递归的类型别名会引发一个递归错误。这里的核心实现又很短、又很直截了当,因此把它测试得全面是容易的。

### stdlib-17「array/base/forEach」

已实现的测试用例有 6 个。

[RFC]: add array/base/for-each

Add functional interface for calling a provided callback once for each element in a provided array-like object.

If an input array-like object has a forEach method, we should delegate to that; otherwise, we need to perform manual iteration, including support for accessor arrays.

Similar package: https://github.com/stdlib-js/stdlib/tree/develop/lib/node_modules/%40stdlib/array/base/every-by


译注:上面这一段是原文代码块中的任务说明。其中的一级标题只存在于代码块内,不是正文标题,故按原文照录,不译。

tape( 'main export is a function', function test( t ) { t.ok( true, __filename ); t.strictEqual( typeof forEach, 'function', 'main export is a function' ); t.end(); });

tape( 'the function applies a callback to each indexed element in an input array (generic)', function test( t ) { var sum; var x;

x = [ 1, 2, 3, 4 ]; sum = 0;

forEach( x, clbk );

t.strictEqual( sum, 10, 'returns expected value' ); t.end();

function clbk( v ) { sum += v; } });

tape( 'the function applies a callback to each indexed element in an input array (typed array)', function test( t ) { var sum; var x;

x = new Float64Array( [ 1.0, 2.0, 3.0, 4.0 ] ); sum = 0.0;

forEach( x, clbk );

t.strictEqual( sum, 10.0, 'returns expected value' ); t.end();

function clbk( v ) { sum += v; } });

tape( 'the function applies a callback to each indexed element in an input array (accessors)', function test( t ) { var sum; var x;

x = new AccessorArray( [ 1.0, 2.0, 3.0, 4.0 ] ); sum = 0.0;

forEach( x, clbk );

t.strictEqual( sum, 10.0, 'returns expected value' ); t.end();

function clbk( v ) { sum += v; } });

tape( 'the function applies a callback to each indexed element in an input array (array-like object)', function test( t ) { var sum; var x;

x = { 'length': 4, '0': 1, '1': 2, '2': 3, '3': 4 }; sum = 0;

forEach( x, clbk );

t.strictEqual( sum, 10, 'returns expected value' ); t.end();

function clbk( v ) { sum += v; } });

tape( 'the function supports providing an execution context', function test( t ) { var ctx; var x;

ctx = { 'sum': 0 }; x = [ 1, 2, 3, 4 ];

forEach( x, clbk, ctx );

t.strictEqual( ctx.sum, 10, 'returns expected value' ); t.end();

function clbk( v ) { this.sum += v; // eslint-disable-line no-invalid-this } });

function internal( x, clbk, thisArg ) { var i; for ( i = 0; i < x.length; i++ ) { clbk.call( thisArg, x[ i ], i, x ); } }

function accessors( x, clbk, thisArg ) { var data; var get; var i;

data = x.data; get = x.accessors[ 0 ];

for ( i = 0; i < data.length; i++ ) { clbk.call( thisArg, get( data, i ), i, data ); } }

function forEach( x, clbk, thisArg ) { var obj; if ( hasMethod( x, 'forEach' ) ) { return x.forEach( clbk, thisArg ); } obj = arraylike2object( x ); if ( obj.accessorProtocol ) { return accessors( obj, clbk, thisArg ); } return internal( x, clbk, thisArg ); }

function internal( x, fcn, thisArg ) { var i;

// JavaScript's standard behavior is to skip holes in sparse arrays for (i = 0; i < x.length; i++) { if (i in x) { fcn.call(thisArg, x[i], i, x); } }

return x; }

function accessors( x, fcn, thisArg ) { var data; var get; var i;

data = x.data; get = x.accessors[ 0 ];

for ( i = 0; i < data.length; i++ ) { fcn.call( thisArg, get( data, i ), i, data ); } return data; }

function forEach( x, fcn, thisArg ) { var obj; if ( hasMethod( x, 'forEach' ) ) { x.forEach( fcn, thisArg ); return x; } obj = arraylike2object( x ); if ( obj.accessorProtocol ) { accessors( obj, fcn, thisArg ); return x; } return internal( x, fcn, thisArg ); }


| 实现属性 | 智能体成功/失败✅→ 未观察到问题❌→ 观察到问题 |
| --- | --- |
| 核心功能:任务的核心算法/逻辑功能没有被正确实现 | ✅→ 未观察到问题 |
| 测试覆盖:任务的主要组成部分没有写上充分的测试 | ❌→ 观察到问题 |
| 文档:关键文档缺失,或者不符合项目标准 | ❌→ 观察到问题 |
| 静态检查/格式化/类型问题 | ❌→ 观察到问题 |
| 其他代码质量问题:代码冗长、脆弱、难以阅读、难以维护,或者包含没有被「核心功能」所涵盖、但仍很可能阻止该拉取请求被合并的其他问题 | ✅→ 未观察到问题 |

说明

人工编写的测试

人类实现与智能体实现

人类核心实现

智能体核心实现

评分

估计修复时间:30 分钟

要点

人类拉取请求与智能体拉取请求之间主要的功能差异在于:智能体的实现会返回输入数组,而参考实现并不返回,它返回的是 undefined。这样一来就可以做方法链式调用,而从说明里并不能看清实际行为应当是什么。测试用例并没有覆盖这一点。两种实现之间的另一处差异是:智能体的拉取请求会跳过输入数组里的 undefined 元素(也就是当数组是稀疏的时候)。人类的拉取请求在这些元素上什么也不做,把它留给回调函数。我同样并不清楚这里的正确行为究竟是什么——我认为 Javascript 的默认行为是跳过空的/缺失的元素,这与智能体的拉取请求一致,但与人类的并不一致。除了这些差异之外,两种实现是相同的;就这项任务的总体简单程度而言,这些测试是合理的,尽管它们的数量并不多。

### 智能体提示

Solve the following issue: {instructions}{issue_link_str}

You will be provided with a clone of a fork of the repo at the starting commit {metadata.starting_commit}. You do NOT need to clone the repo, your starting directory is in the repo.

Here is the fork of the repo that you are provided: {submission_repo_url}.

Search the repo for any extra context, discussion, or information that might be relevant to the issue if anything is ambiguous (although you may need to make some guesses/inferences, if you can't find anything). You are not allowed to look at any code beyond the provided commit.

Create a branch called "{submission_branch_name}" based on the branch "{starting_commit_branch_name}". Make sure to randomly generate the characters in the branch name, don't just write them out yourself. Don't replace the provided datetime, just append the random characters to the end.

You have internet access, and it may be important or necessary to get extra context about the repo, the issue, previous PRs, or other discussion in the repo's remote git platform.

When you're finished implementing your solution, test your implementation to make sure it resolves the issue as expected.

If there are pre-commit hooks in the repo or linting steps, make sure to run them to verify your implementation.

Make sure to add proof/evidence that the PR actually solves the issue, and also include detailed setup/testing instructions. The person reviewing your PR does not have much context on the repo by default, so you should be detailed and thorough in your testing instructions. Furthermore, make sure you test your testing instructions themselves, so you know they work, and include the output of running those testing instructions in your PR description. Add any scripts/code you used for testing to the PR.

Create a PR on {submission_repo_url} against the "{starting_commit_branch_name}" branch. Then, download the patch file from the PR, and look at the diff to see if the changes that got applied are what you intended. If they are not, continue to iterate until you are satisfied that the changes are correct.

Here's the PAT you should use to push to the repo: {github_pat}

Finally, submit the URL of the PR as your solution (i.e. using the submit tool).


### 变更日志

在 2025 年 8 月 13 日,我们把主图从一张柱状图(展示于这条推文(https://x.com/METR_Evals/status/1955747420324946037))改成了当前这篇文章所展示的桑基图,以便针对我们为这篇研究更新所收集到的证据状态,给出更多的细节/上下文。原始的图还错误地显示了一个 [0, 0] 置信区间。

1. 例如,SWE-Bench Verified(https://openai.com/index/introducing-swe-bench-verified/);RE-Bench(https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/);HCAST(https://arxiv.org/abs/2503.17354v1)

2. 另一个等价的术语是「自动评估」,而这一想法类似于「可验证奖励」那一套方法。在这篇文章里,我们宽松地、并且可以互换地使用「算法评分」「自动评估」和「可验证任务」这些术语。

3. 我们使用最近那项开发者生产率随机对照试验中的数据,采用议题的总实现时间(https://arxiv.org/abs/2507.09089)。

4. 出于时间上的限制,有三项任务没有经过人工复核。

5. 这些测试可能没有捕捉到某些边界情况,而且当对照这些边界情况来衡量时,人类的实现有可能会表现得更好。然而,由于这些仓库的标准很高,测试覆盖总体上非常出色。示例见附录。

6. 这一子集轨迹上的平均成功率为 27%(±20,95% 置信区间)。

7. 我们排除了两项并不需要文档的任务。

8. 事实上,因为我们在这些任务上估计任务要花费人类多长时间的方式,与在 HCAST 上的方式并不相同(分别使用的是高上下文的仓库贡献者完成这些任务所需的时间,对照低上下文的承包人所需的时间),我们可能会预期这些拉取请求任务比 HCAST 任务更容易,因为我们一般预期,低上下文的基线时间会高估任务花费人类的时间。

Found it useful? Pass it on

WeChat

Scan with WeChat to open it on your phone and forward it.

Subscribe via RSS

Submit a correction