Industry & PracticeCase Studies
FlowAgent: 在 开发 者 切 走 之前 修好 提交 前 测试
Google 把修复智能体放进提交前的持续集成。195 个失败上人工评估准确率 67.18%。上线后对 2,567,829 个失败变更触发,发出 295,508 条建议,开发者应用 28,554 条。中位执行 9.85 分钟,落在开发者自己动手的中位 27.48 分钟之内。

In this piece
在开发者还没切走之前修好测试:谷歌规模的低延迟智能体程序修复
Celal Ziftci、Spencer Greene、Ray Liu、Livio Dalloro、Lorenzo Dini。Google。许可证:CC BY 4.0。arXiv:2610.07289v1,2026 年 10 月 5 日。ASE 2026 预印本。
手工修复程序失败又慢又打断人,尤其在提交前(pre-submit,变更还没进仓库)阶段,失败发生在持续集成系统里。自动程序修复(APR,Automated Program Repair,自动找出并改掉导致失败的代码或测试)借大语言模型有了很大进展,但现有强方法主要做提交后的离线修复,没有“在开发者切走之前”所需的低延迟。
本文介绍 FlowAgent,一个部署在 Google 的 AI 智能体,在持续集成里的提交前外环(outer-loop,变更已经创建、要在整个代码库上查回归的阶段)自动修复测试失败。它接进内部工具 Critique 和 Cider,用 ReAct 风格(推理和行动交替)的生成–验证循环,以及执行前、执行后的弃权过滤器,在严格延迟下保证建议质量。
案例表明它有效。对 195 个真实测试失败的人工评估里,建议正确修复的准确率是 67.18%。全公司上线后,FlowAgent 在 295,508 个变更上给出修复建议,开发者预览了 65,069 个、应用了 28,554 个。访谈认为它能给出正确修复,把自主修复智能体放进工业软件工程流程是被接受的,同时仍有挑战和机会。
预印本说明
这是第 41 届 IEEE/ACM 自动化软件工程国际会议(ASE 2026)已接收论文的预印本。
1 引言
开发者经常改代码:加功能、修缺陷、提高安全性或性能。为了不把回归带进更大的代码库,这些修改必须经过严格测试。改动导致测试失败时,开发者就要做程序修复。这又花时间又打断人,尤其在提交前的外环:失败发生在集成开发环境(IDE)之外、持续集成系统之内,例如 Google 的 TAP(Memon et al., 2017; Micco, 2012),用来在变更提交前发现全局回归。
收到失败通知后,开发者通常会调查并改代码。作者的分析显示,在 Google,从失败通知到开发者下一次改代码的中位时间是 27.48 分钟。这是一个关键窗口:自动修复如果来得及,可以提高速度、减少挫折。
为减轻手工调试,APR 受到很多关注。早期方案依赖手工写的变异规则、符号约束和静态分析。近来大语言模型把方向推向机器学习和自主智能体。SWE-Agent(Yang et al., 2024)、AutoCodeRover(Zhang et al., 2024)、Google 的 Passerine(Rondon et al., 2025)和 Meta 的 Engineering Agent(Maddila et al., 2025)用 ReAct 风格循环(Yao et al., 2023)自主浏览代码库、使用工具并生成修复。但这些方法主要做提交后流程,离线修程序,没有“在开发者切出当前流程之前实时帮忙”所需的延迟约束。
本文介绍 FlowAgent:部署在 Google 的 AI 智能体,直接在提交前外环里自动调试并修复测试失败。成功的修复,连同 AI 生成的根因和修复摘要,出现在内部开发工具里,开发者可以在习惯的流程中预览并应用。
据作者所知,和提交后的 APR 工具不同,FlowAgent 是第一个做在提交前外环、并以关键延迟要求运行、以便赶上开发者还在改当前变更的 APR 系统。
作者给出据他们所知迄今报告过的最大工业数据集上的实证评估:先人工评估,再在生产中跨多个月。讨论它为数千次失败拿到有效修复的情况,并给出采用遥测和用户访谈:开发者偏好、工具采用行为,以及对 AI 生成修复的不同信任程度。
2 手工程序修复
开发者改代码的原因很多,例如修缺陷、加功能、提高性能或安全性。代码在一个或多个文件里改完后,开发者创建一个变更(change,对代码仓库的一次改动,评审后原子提交)。
2.1 程序失败阶段
失败可以是构建失败(代码编不过),或测试失败(能构建,但有测试失败)。本文只处理测试失败。
一个变更里,如果配套测试失败,可能要改程序,也可能要改测试。修复发生在两个阶段:
- 提交前:变更提交之前修,通常就在这次变更里,不影响其他开发者。
- 提交后:破坏性变更已经进仓库,接着影响所有要改这段代码的人,因为即使不再加新改动,测试也是失败的。
本文关注提交前:有问题的代码或测试还没进仓库。
提交前的自动测试,按变更生命周期又分两条:
- 内环(inner-loop):在 IDE 里改代码。开发者或 AI 编码智能体(例如 Antigravity)跑与改动相关的单元测试,在变更创建前就给早期反馈。这些测试通常紧挨着被改的代码,例如同一目录或父目录。
- 外环:本地改动完成、变更已创建。这一阶段看全局正确性:在整个单体仓库(monorepo,全公司代码在一个仓库里)里找出并执行所有依赖测试,确认变更不会把回归带进 Google 更广的代码库。
本文关注外环。下面是开发者在外环修复流程里用到的工具。
2.2 外环提交前的开发者工具
Cider:Google 内部 IDE(Nguyen et al., 2025)。开发者通常在 Cider 里开始改代码,用更多工具迭代,并在 Cider 和其他工具之间切换来收尾。
Critique:Google 内部基于网页的代码评审系统(Sadowski et al., 2018)。开发者在 Cider 改完代码后创建变更。在把变更送给其他人评审之前,自动工具和回归测试会在改过的版本上跑,用发现项(finding,标在相关代码或整次变更上的通知)和建议修复帮助作者。
TAP:测试自动化平台(Testing Automation Platform),是 Google 跑单元测试的持续集成系统(Memon et al., 2017; Micco, 2012)。对一次变更,TAP 找出单体仓库里所有依赖这次变更的测试并执行,以便在可以提交之前发现潜在回归。Critique 里的一条 TAP 测试失败发现项见图 1。
图 1:测试失败发现项出现在 Critique 里。
开发者通常先在 Cider 写代码;准备收尾时创建变更并在 Critique 里跑分析,这时 TAP 可能发现由这次改动引起的失败测试。收到通知后,开发者在 Cider 里调查,手工或借助工具,修改变更或测试本身。
作者分析了 2025 年 9 月、跨 30 天、在 Google 提交的代码变更,以理解开发者面对失败测试时的行为。他们看修复尝试延迟:从 TAP 对某一版本报告测试失败,到开发者再做任何编辑并把变更推进到下一版本的墙钟时间,见图 2。
图 2:修复尝试延迟分布。p10 为 2.37 分钟,p25 为 6.72 分钟,p50 为 27.48 分钟。纵轴为分钟,百分位从 0 到 100。
修复尝试延迟是帮助开发者的机会窗口:在他们自己继续改变更、去修测试之前,可以提供根因信息或建议修复。p25 是 6.72 分钟,p50 是 27.48 分钟。
3 用 FlowAgent 做自动程序修复
近来自主 AI 编码智能体在 APR 上表现很好。为了在程序修复上帮助开发者,作者提出 FlowAgent:接进 Google 开发工具、自动调试并修复提交前外环测试失败的自主智能体。
FlowAgent 和它所接入的工具是 Google 内部的,但它的逻辑和流程可以用业界标准开发工具复现。
高层架构见图 3。变更创建并在 Critique 触发测试后,FlowAgent 监听 TAP 的测试失败通知。失败发生时,不需要开发者参与:它分析变更,检查是否通过支持的过滤器,跑 AI 智能体修代码,检查修复是否通过安全过滤器,最后把修复发到 Critique,供开发者在 Critique 或 Cider 里预览和应用。
图 3:用 FlowAgent 自动修复程序的系统总览。TAP 测试失败进入发布/订阅队列,经执行前过滤器,由编排器读取测试日志和变更信息,进入 ReAct 循环,验证测试通过,摘要修复,再经执行后过滤器,把修复送到 Critique 和 Cider 的开发者。
3.1 执行前过滤器
FlowAgent 用执行前过滤器忽略某些带失败的变更,以避免不可行或不必要的修复尝试,也避免无关或低质量建议造成困惑。
下面这些过滤器检查变更或失败测试的特征,依据是机构知识、对一组失败的人工评估和观察,以及生产中的开发者反馈。
| 过滤器 | 跳过条件 |
|---|---|
| 构建失败 | 失败来自构建,不是测试 |
| 已经失败 | 仓库里本来就失败的测试 |
| 不稳定(flaky) | 已知不稳定测试:同一代码版本、没有新改动也会时过时不过(Fowler, 2011; Micco, 2016; Ziftci and Cavalcanti, 2020) |
| 自动生成 | 由自动工具创建的变更上的失败测试 |
| 过期 | 开发者已经又改过的变更上的失败测试 |
| 不支持的文件 | 涉及某些文件类型,例如图像和二进制文件 |
| 文件太多 | 变更超过 100 个文件 |
| 行数太多 | 变更超过 1000 行 |
一次变更上的测试失败通过全部过滤器后,才触发 FlowAgent。
3.2 智能体修复
通过执行前过滤器后,在沙箱里运行的混合系统被触发。它既有预定步骤,也有完全由智能体决定的步骤,见图 3。设置参数是:
- LLM = Gemini 2.5 Pro 的内部版本,在 Google 内部代码和数据上微调
- temperature = 0.1,使动作大体确定,便于调试和复现
- top p = 0.95,留出创造空间,同时排除极不可能的结果
编排器(orchestrator,按固定顺序启动各步的程序)先让智能体读测试日志,并读变更信息:描述、修改过的文件列表,以及 diff 格式的修改内容。
这些信息进入清单 4(a) 所示提示下的 ReAct 风格生成–验证循环。智能体可以按构建和测试反馈迭代修改修复,并有:
- 工具:读/改/删文件,搜索代码,构建,读构建状态和日志,执行测试,读测试状态和日志
- 最多 100 次工具调用,避免运行和调用过长
- 最长运行 30 分钟,与第 2.2 节的修复尝试延迟对齐,避免跑太久,以便赶上开发者还在写代码
若循环成功并产出修复 diff,再做最终验证:要么确认智能体在最后一轮验证里执行过测试,要么再执行一次测试并检查结果。这一步必要,因为大语言模型会编造;没有最后的确定性验证,就可能把无效修复展示给开发者。
验证成功后,把上面生成–验证循环的整段轨迹交给 LLM,只调用一次,按清单 4(b) 的提示摘要失败根因和修复。
最后,建议修复及其摘要进入执行后过滤器。
清单 4:FlowAgent 使用的提示。
(a)ReAct 循环的系统提示模板,原文如下:
The following change has failing pre-submit tests.
The change information is provided below:
=== Change description ===
{change_description}
=== Change diff ===
{change_diff}
Your task is to analyze the provided error logs from failing tests (paying close attention to +/- diffs indicating discrepancies), and the content of the relevant test file(s).
Based on this analysis, you must output only a list of specific modifications required to fix the failing tests, along with their target file names.
Your Analysis Steps (Internal Process):
- Identify Test Failures & Discrepancies: Parse the error logs to pinpoint failing tests, error messages, stack traces. Crucially, analyze any + (expected) / - (actual) diffs within the logs, as these often directly indicate necessary corrections in test expectations or data.
- Contextualize with Code: Analyze the test files (present in the error logs) to understand the structure, setup, assertions, and logic of the failing test(s) and related code.
- Determine Required Modifications: Synthesize all information to identify the exact changes needed. This might involve updating expected values in tests based on log diffs. Do not add any comments or whitespaces.
Your Output:
- Your output must consist exclusively of the modification instructions and target files in diff format.
- Do not include any explanations, analysis summary, greetings, apologies, or conversational text.
Fix the following tests that produced errors:
=== Test name: {test_name} ===
{test_error}(b)摘要失败根因和建议修复的提示模板,原文如下:
You are an expert software engineer. You are given a transcript of an AI agent’s thoughts and actions to repair a test failure. Please provide a concise summary of what caused the test failure and the steps taken by the agent to repair the failure.
Try your best to summarize the steps taken by the agent to fix the test. Focus on the specific actions taken and the thought process behind them. Do not simply copy the transcript of repair steps. When possible it’s best to include artifacts like file names, method headers, variables, etc.
The resulting synopsis should be at most 5 sentences and suitable for a user interface, but ideally it can be one or two. If appropriate, presenting the synopsis in bullet points can be helpful.
It is imperative that the summary follows HTML formatting, so make sure to use HTML as your only formatting style. Do not use any other formatting. The summary should be as compact as possible, this means the HTML should be written in a single line.
=== TRANSCRIPT ===
{agent_transcript}3.3 执行后过滤器
执行后过滤器用来避免把没有用的修复建议发给开发者。
| 过滤器 | 含义 |
|---|---|
| 过期 | 智能体还在处理失败版本时,开发者已经又改了变更,修复作废 |
| 已提交 | 开发者已经提交变更,修复作废 |
通过这些过滤器后,修复就可以展示。
3.4 把修复展示给开发者
开发者预览和应用修复的环境偏好不同。有人更愿意用 Critique,因为修复先发到那里;有人更愿意回到 Cider,因为他们习惯在那里改代码。为提高采用,FlowAgent 同时接进 Critique 和 Cider。
修复作为 TAP 测试失败通知的一部分显示在 Critique 里(见图 1),而不是另开一条发现项。这样开发者能把修复和失败对应起来,也不会被太多发现项淹没:每次变更会跑数百项分析,可能已经有好几条发现项。
图 5 中,修复包含失败根因摘要,让开发者先判断诊断是否正确,再决定是否查看并应用建议修复。另外提供在 Cider 里预览的链接,以及在 Critique 里预览并应用的按钮。
图 5:FlowAgent 的发现项出现在 Critique,带根因摘要和指向 Cider 的链接。
开发者可以在 Critique 里点按钮预览修复(图 6)并就地应用。
图 6:在 Critique 中展示的建议修复。
更复杂、跨多行或多文件的修复,或者还想再手工改,开发者可以在 Cider 里预览并应用(图 7)。
图 7:在 Cider 中展示的建议修复。
4 评估与讨论
这一节说明如何评估系统、案例研究的结果,以及关于可用性的访谈。
4.1 人工评估
作者先做案例研究:在来自 Google 37 个不同团队的、随机选出的 195 个 TAP 测试失败上跑 FlowAgent,见表 1。请 3 名各有至少五年软件开发经验的专家评估建议修复(类似图 6),报告修复是否符合变更意图和这次失败。然后开会过一遍报告,对齐分歧,得到每条修复的最终一致判断。
按专家报告,195 个失败里有 131 个给出了正确修复,成功率 67.18%。其余 64 个不正确或并非完全正确,下面用例子讨论。
表 1:人工评估统计。
| 项 | 数 |
|---|---|
| 拥有这些 TAP 失败的团队数 | 37 |
| 做评估的开发者数 | 3 |
| 评估的失败数 | 195 |
| 准确修复数 | 131(67.18%) |
4.2 生产采用与用户行为
人工评估之后,FlowAgent 对所有带 TAP 测试失败的变更自动运行,共 2,567,829 个,并在 Google 代码仓库上展示建议修复。从 2025 年 10 月起持续至今。统计见表 2。
表 2:FlowAgent 在 Critique 和 Cider 上的执行与使用。箭头表示相对上一行的保留比例。
| 项 | 数 |
|---|---|
| 带失败测试的变更 | 2,567,829 |
| 被执行前过滤器丢弃 | 1,785,955 |
| FlowAgent 尝试修复 | 781,874(降至 30.45%) |
| FlowAgent 得到修复 | 421,818(降至 53.95%) |
| 被执行后过滤器丢弃 | 126,310 |
| 发出修复建议的变更 | 295,508(降至 70.06%) |
| 开发者预览了修复 | 65,069(降至 22.02%) |
| 开发者应用了修复 | 28,554(降至 43.88%) |
| 在 Critique 预览 | 55,029 |
| 在 Critique 应用 | 20,826(37.84%) |
| 在 Cider 预览 | 21,279 |
| 在 Cider 应用 | 12,702(59.69%) |
#### 4.2.1 系统触发
系统在 2,567,829 个带失败测试的变更上触发。1,785,955 个被执行前过滤器跳过。这一阶段跳过很多,因为 FlowAgent 针对的是人写的、可行的修复尝试。这也指出以后可以修更广的变更,包括非人撰写的和更大的变更。
#### 4.2.2 修复尝试
表 3:FlowAgent 尝试修复 781,874 个变更,作者 36,479 人;每个变更平均 10.57 个文件、中位数 6 个;全部变更上有 917 种不同文件扩展名(例如 .py);每个变更失败测试平均 16.49 个、中位数 2 个。
表 3:FlowAgent 尝试修复的代码变更。
| 项 | 数 |
|---|---|
| 尝试修复的变更 | 781,874 |
| 这些变更的不同作者 | 36,479 |
| 每变更文件数均值 | 10.57 |
| 每变更文件数中位数 | 6 |
| 不同文件扩展名 | 917 |
| 每变更失败测试均值 | 16.49 |
| 每变更失败测试中位数 | 2 |
| 每次执行的工具调用均值 | 21.37 |
| 每次执行的 token 均值 | 424,149 |
| 迄今 token 总量 | 6430 亿 |
每次执行平均 21.37 次工具调用,平均 424,149 个 token,迄今合计 6430 亿 token。
图 8:执行延迟。从 TAP 失败通知到拿到成功修复,p10 为 4.78 分钟,p50 为 9.85 分钟,p90 为 25.83 分钟。
p50 的 9.85 分钟和 p90 的 25.83 分钟,落在第 2.2 节中位修复尝试延迟 27.48 分钟之内。
FlowAgent 为 421,818 个变更成功生成建议修复。不成功是因为给执行加了上限,避免跑太久——用户可能已经收到失败通知并自己动手修。作者根据内部实验预测:这一阶段相当一部分失败尝试,用研究之后发布的更强语言模型,可能能够产出修复。
#### 4.2.3 成功但被丢掉的修复建议
126,310 个成功修复被执行后过滤器丢掉,因为开发者已经修改或提交了变更,智能体的修复作废。
有些修复在开发者改过变更后仍可能相关,因为他们改的可能是无关文件。但这类修复可能造成困惑,并与开发者的改动冲突,因此执行后过滤器把它们整条丢掉。
尽管 FlowAgent 按低延迟来优化,仍有不可忽略的一部分成功修复被执行后过滤器丢掉。这说明还要再降延迟,才能赶上开发者改变更的流程。
#### 4.2.4 建议的修复
执行后过滤之后,295,508 条修复发到 Critique。在 Critique 和 Cider 上,65,069 条被预览,28,554 条被开发者接受。
作者手工分析了一部分“预览了但没应用、后来由开发者自己修好”的失败测试,看到不用建议修复有几个原因,例子见图 9。
第一,修复能让测试通过,但有时无关或错误,见图 9(a) 和 9(b)。这类修复不该建议给开发者,长期会损害信任;应在智能体架构里再加一步验证。作者把这留作后续工作。
第二,修复是对的,但不是开发者想要的样子,见图 9(c)。这时开发者多半在预览时已经知道该改哪里,然后自己改。
还有 FlowAgent 建议撤销变更里的修改,等于取消开发者的修改意图。这类建议不该发给开发者,也被记为智能体上要做的后续改进。
最后,有时修复是对的,开发者仍自己把同样的改动敲进去。作者推测这些人没意识到可以在 Critique 里应用,或者用的编辑器不是 Cider,发现项里的链接用不上,或其他原因,其中一部分在第 4.3 节讨论。
图 9:用户没有应用的建议修复例子。(a)给测试加上 @Ignore。(b)用 try/catch 吞掉断言。(c)修得接近但不精确:应当断言 "== 3",而不是 "> 2"。
#### 4.2.5 应用修复时的开发者偏好
Critique 里预览 55,029、应用 20,826;Cider 里预览 21,279、应用 12,702。
有些建议修复在 Critique 和 Cider 都被预览过。结合第 4.3 节访谈:开发者有时在 Critique 预览,再到 Cider 应用,以便随后继续改。
另外,一次变更可能有多条修复建议:开发者把变更改到满意的过程中,测试可能在多个版本上失败。因此表 2 里 Critique 与 Cider 应用数之和,大于“开发者应用了修复”的变更数。
应用/预览比在 Critique 是 37.84%,在 Cider 是 59.69%。开发者似乎更愿意在 Cider 里查看和应用。作者认为原因包括:还想在建议修复上再改,以及习惯在 IDE 里看 diff。下一节访谈支持这一点。
4.3 用户访谈
作者与 9 名开发者做了 45 分钟当面访谈。他们来自 9 个不同组,在 Google 至少工作 6 个月,并在 Critique 或 Cider 里预览和/或应用过建议修复。下面是摘录。这一节按参与者的回答归纳主要收获。
#### 4.3.1 对 FlowAgent 的看法
先问:不由自己触发、就自动收到修复建议,感觉如何。P-6 和 P-7 说喜欢它接进他们已经习惯的开发流程。
再问他们猜 FlowAgent 怎么工作。所有人都说它很可能是基于 AI 的产品,因为这些修复建议不像非 AI 工具能给出来的,名字也暗示用了 AI。
接着问这会不会造成对建议修复的预期偏差。P-1 概括了大家的说法:建议从哪来不重要,只要相关且准确。
- [P-7]:我先去做别的任务,然后这个我从没见过的叫 FlowAgent 的东西就把测试给我修好了。
- [P-6]:这个流程很好。它和很多 Critique 流程一致,所以不用想这是哪种评审,做法总是一样的。
- [P-1]:我不在乎建议是什么、谁、从哪来,只要相关且准确。
#### 4.3.2 发现项里的根因摘要
关于根因摘要,P-4、P-5、P-6 说读发现项里的根因分析就能比较容易明白失败原因。
- [P-4]:它把问题解释得相当清楚。谢谢做了这个。以后肯定会用。
- [P-5]:它把结论抽出来放在消息顶部……扫一眼就知道错误是什么;如果没有这个,要挖到真正的错误会更久。
- [P-6]:这信息有用;通常得进测试日志,现在信息都放在前面,省时间;不用进日志、按错误过滤、也不用把那些全扫一遍,有摘要;可以跳过找问题的那一步。
#### 4.3.3 预览并应用了的修复
P-1、P-7、P-8 说修复很有用,尤其是重复、枯燥的改动,例如改 mock(测试里代替真实依赖的假对象),以及和重构有关的改动。
- [P-7]:你们是不是在 Critique 上做了 FlowAgent?这东西太好了。它修好了那次变更。我改了些代码,测试失败,因为它用了一堆 mock。我讨厌改这种测试。
- [P-1]:FlowAgent 修好复杂测试(尤其是很多 mock)的体验,是我对 LLM 智能体最大的惊喜之一。
- [P-8]:谢谢开发和维护这个对开发者速度很有希望、很有用的工具。第一次用就让我兴奋:可以把翻多份文档和过去已解决缺陷、找简单参数改动来修测试这种脏活交出去。
#### 4.3.4 预览了但没有应用的修复
遥测显示,有时开发者预览了却没有应用。作者问了原因。
P-1、P-2、P-3 说 FlowAgent 有时给出无关或错误的修复。
- [P-1]:修复不准确,我不想给列表排序。
- [P-2]:它建议加上并不需要的东西,所以什么也没修好。
- [P-3]:不准确,我不认为该去掉那个反斜杠。
P-2 和 P-3 说发现项有时来得太晚,他们已经开始自己手工修。如果及时看到,会觉得有用。
- [P-2]:说实话,坏掉的东西让我烦了一整天。我手工调过这个问题,跨这么多层找根因很麻烦。我可以确认你们智能体的输出是对的;如果在我手工翻日志之前看到,会省很多时间。
- [P-3]:那条发现项确实找出了问题,或至少其中一个。它肯定会有帮助。
还有几位开发者说,他们有时预览建议修复,但不在 Critique 或 Cider 里应用,而是自己在 IDE 里敲,因为提议不完全符合他们的代码风格,他们想要稍有不同的编辑,类似图 9。
P-6 说自己手工应用是为了省时间,并避免 Critique 流程和他们想在 IDE 里做的改动冲突。
- [P-6]:我有第二块显示器,看我在哪工作也许是第三块……但 Cider 已经在那边开着。所以我不是去点 Show Fix、应用修复……我只是把鼠标移过去改我的代码,以免陷入冲突或什么奇怪的情况。
P-9 说,即使自己修了失败,有一个智能体独立给反馈仍然有用,因为它对变更意图和由此产生的失败有自己的上下文。
- [P-9]:我喜欢 Critique 里另有一个智能体也插话,因为它的上下文是新的。
尽管系统提示里有明确指示,FlowAgent 有时仍建议撤销变更或把失败测试注释掉,见图 9。几位开发者说这类建议让人沮丧,访谈和内部论坛上还有人拿 FlowAgent 开玩笑。这是关键的用户信任问题。作者正在图 3 的流程里加一步:在把建议发给开发者之前,验证建议修复不是在删测试或撤销开发者的变更。
#### 4.3.5 可用性反馈
访谈中,几位参与者希望有一个剩余时间估计,以便决定是等 FlowAgent 的修复建议,还是自己修。
- [P-4]:如果有个指示说它正在跑,我就可以说“好,它在试着修一些东西”,也许去做点别的再回来。
这个功能没能马上加上,因为 Critique 不支持。作者观察到开发者正在适应和 AI 智能体一起决定自己的等待行为。他们正在加通知:告诉用户 FlowAgent 正在处理其变更上的失败。
5 对有效性的威胁
人工评估:由三名有经验的开发者完成。但他们并不拥有失败测试的生产代码或测试代码,报告 FlowAgent 修复是否准确时可能出错。另外,这些失败是在 Google 随机选的,不一定代表全部失败。
过滤器:执行前过滤器基于早期一小群 FlowAgent 用户的反馈。相当一批变更因这些过滤器被跳过,评估结果未必推广到具有那些特征的变更。
开发者偏差:FlowAgent 给 Google 全部合格代码变更建议修复。但开发者对在开发中使用 AI 和 AI 辅助的正负判断不同。评估预览和应用时没有控制这些变量,可能引入选择偏差,并影响结果的准确性和有效性。访谈对象是自愿反馈的变更作者,这是自选群体;没参加调查和访谈的人可能有不同看法,从而改变这些定性结论。
开发基础设施:案例发生在 Google 特定的开发环境和基础设施里。虽然内部工具在业界都有对应产品,发现未必能推广到其他组织。
LLM:研究用的是 Gemini 2.5 Pro 的内部版本。流行的大模型在编码等任务上表现好,但结果推广到其他 LLM 可能有限。另外 FlowAgent 依赖给 Gemini 的特定提示模板。措辞、结构或指示顺序的小改动,都可能导致表现不同或变差。提示工程对细小变化或未来模型更新是否稳健,是对发现可推广性的潜在威胁。
测试:测试由开发者编写,不同组、不同系统的质量和细节可能差很多,这可能影响 FlowAgent 在不同测试上的准确率,以及案例结果。
6 相关工作
APR 文献很多:先是非机器学习方法,然后机器学习,再到基于 LLM 的技术,最近是基于 LLM 智能体的方法。
非机器学习:把 APR 当成带手工变异规则和修复模式的搜索,或从人写的补丁得到半自动变换规则;用符号约束抽出修复;用静态分析先定位再修;或用同一代码库里相似但正确的代码替换有问题的代码。除功能修复外,还有针对语法错误、性能问题、安全漏洞、类型错误和构建损坏的方法。
机器学习:早期用机器学习在候选修复里排序和选择。后来用机器学习生成修复,包括神经机器翻译把有问题的代码映射到修好的代码、树变换预测、针对更具体缺陷类型的神经模型,以及为特定修复训练模型。SelfAPR 用扰动被修复程序的先前版本得到训练样本,迫使模型抓住项目特有的知识。
基于 LLM:近来更常见的是用通用 LLM,而不是为任务单独训练。AlphaRepair 是零样本方法,把 APR 当成填空,不需要额外训练数据。FitRepair 在此之上加入“整形手术假设”:修缺陷所需的代码通常已经在同一代码库里。若干技术用一次提示,把问题描述或错误信息发给模型,模型返回提议的代码修复。更新的技术是迭代:多次提示,并把上一次修复尝试的反馈传给下一轮。ChatRepair 是对话式方法,和工程师对话,逐步收集信息、提出可能的补丁,并随时间生成更好的补丁。ITER 把部分补丁改进到看起来合理且正确。Agentless 用生成的测试和预定义的三阶段流水线(定位、修复、验证),以很低的成本达到当时的强结果,超过此前的开源技术。
基于智能体:LLM 自主规划并借助工具完成修复所需的任务。FixAgent 用专门的 LLM 智能体(“Tester”和“Debugger”),通过提示链做端到端故障定位和修复。LLM4FL 用两个 LLM 智能体迭代分析代码,给可疑方法排序。SWE-Agent 实现 ReAct 风格的开放循环,用一组 API 提供终端访问。RepairAgent 在“用工具生成并执行代码”的先前工作上,给出受状态机约束的智能体,不允许某些工具使用序列,这一点和 SWE-Agent 不同。AutoCodeRover 利用包括类和方法定义在内的程序元数据。SpecRover 在其上加入对期望行为的自然语言规格,并用一个审阅潜在补丁的智能体来引导修复智能体。CodeR 用“经理”智能体协调写成任务图的子任务。MarsCode Agent 提出动态、迭代的 APR,用多个子智能体的生成–验证流水线。OpenHands(曾用名 OpenDevin)是受 Devin 启发的开源智能体平台,为软件工程和其他领域提供智能体基础。Google 的 Passerine 是 ReAct 风格智能体,工具包括基于测试套件的度量计算器做验证,已部署到生产,据报告在修复机器生成的缺陷上成功率为原先的三倍;最近又加上过滤不太可能被人接受的补丁的技术。Meta 的 Engineering Agent 也实现 ReAct 风格循环,用迭代循环和验证工具自动修失败的单元测试。延伸工作 MetaMateCR 为评审意见提供 AI 辅助修复。
本文的 APR 结合上面基于 LLM 和基于智能体的做法。但先前工作聚焦提交后:在已有代码上离线修,没有严格的低延迟要求。据作者所知,本文是第一个在提交前开发流程里、失败一发生就自动触发 APR、延迟要求对开发者体验至关重要,并分别报告提交前流程的工业规模结果和学习的工作。
APR 基准:社区做了若干公开和私有基准。HumanEvalFix 覆盖三个软件工程下游任务的翻译基准,包括程序修复。SWE-Bench 来自 12 个流行开源项目的 GitHub issue,含 2294 个 Python 缺陷和修复;另有 300 对的 SWE-Bench-Lite;SWE-Bench Verified 是人工标注为“可解”的子集,再分“易”和“难”。SWE-Bench 已成为许多 APR 技术的标准评测基准。DebugEval 定义若干软件工程任务,包括代码修复,以跨语言评估 LLM。DebugBench 是大量 LeetCode 题,人工注入缺陷,用来比较 LLM 的 APR。SWE-smith 是大规模生成软件工程任务的流水线,支持自动合成破坏现有测试的代码变更。SWE-Lancer 包含 1400 多个自由职业软件工程任务。最近,Google 的 Passerine 用类似 SWE-Bench 的内部基准,Meta 的 Engineering Agent 用内部机器生成任务基准。本文的内部基准由 Google 的真实程序失败组成,见第 4 节。由于第 8 节所述限制,不能公开这个基准数据集。
7 结论
本文给出 FlowAgent:部署在 Google 的自主 AI 编码智能体,在提交前外环自动调试并修复测试失败。它直接接进 Critique 和 Cider,在开发者从改变更的流程里切走之前给出建议修复。
为了在低延迟下保证建议质量,系统使用执行前和执行后过滤器,以及由 LLM 驱动的迭代生成–验证循环。
评估表明这一做法在工业规模上有效、可落地。3 名专家的人工评估中,FlowAgent 对所评测试失败的 67.18% 给出了准确修复。2025 年 10 月生产部署之后,系统建议了 295,508 条修复,开发者接受并应用了 28,554 条。
用户访谈给出采用和行为上的观察。开发者报告,在诊断和修复复杂测试失败时节省了时间、减少了挫折。挑战仍在:要把执行延迟优化到开发者切换上下文之前,并减少“把测试注释掉”这类没有帮助的建议。
这些发现说明,把基于 LLM 的自主智能体放进对延迟敏感的工业开发流程里做 APR,是有潜力的。
8 后续工作
案例研究中看到若干改进机会。
第一,改进 FlowAgent 删除测试和撤销代码的行为,避免让开发者沮丧。
第二,用更新的语言模型提高延迟和成功率,从而给出更多修复。
另外,增加一项功能:告诉开发者 FlowAgent 正在修他们变更上的失败测试,以便他们可以选择等待。
最后,把 FlowAgent 扩展到 AI 撰写的变更,因为业界正更多转向 AI 生成的代码变更。
数据可用性声明
支持本研究发现的数据归 Google 专有,不能公开。由于严格的保密政策、商业敏感性和保密协议,作者无法提供公开的复现包或数据集。
致谢
感谢 Gaurav Johari 在 Google 调查 APR 领域的工作;感谢 Adrian Berding、Ahmed Omran、Allen Edwards、Assaf Raman、Bernhard Konrad、Chris Lewis、Daotong(Daniel) Dai、Gaurav Johari、James Wilson、Jeyavaishnavi Muralikumar、John Penix、Mariana Stariolo、Nic Cornejo、Peter Spragins、Rob Dryke、Rose Rodrigues、Sara Toth、Sruthi Santhosh Nair、Will Findley、Yang Liu、Yurun Shen 在 FlowAgent 上的工作;感谢 Elaine Thai、Laurie Pham、Maria Arguello、Ryan McGarry 协助做 FlowAgent 使用情况的用户访谈。
Google Gemini 被用来生成本文的若干部分,具体是若干 LaTeX 表格,以及 Google Colab 里的若干图。
参考文献
[1] Johannes Bader, Andrew Scott, Michael Pradel, and Satish Chandra. 2019. Getafix: Learning to fix bugs automatically.Proceedings of the ACM on Programming Languages3, OOPSLA (2019), 1–27. doi:10.1145/3360585 [2] Earl T Barr, Yuriy Brun, Premkumar Devanbu, Mark Harman, and Federica Sarro. 2014. The plastic surgery hypothesis. InProceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering. 306–
- doi:10.1145/2635868.2635898
[3] Rohan Bavishi, Hiroaki Yoshida, and Mukul R Prasad. 2019. Phoenix: Automated data-driven synthesis of repairs for static analysis violations. InProceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 613–624. doi:10.1145/ 3338906.3338952 [4] Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. 2024. RepairAgent: An autonomous, LLM-based agent for program repair. InProceedings of the 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). doi:10.1109/ICSE55347.2025.00157 [5] José Cambronero, Michele Tufano, Sherry Shi, Renyao Wei, Grant Uy, Runxiang Cheng, Chin-Jung Liu, Shiying Pan, Satish Chandra, and Pat Rondon. 2025. Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair. InProceedings of the IEEE/ACM 48th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP ’26). 118–129. doi:10. 1145/3786583.3786858 [6] Dong Chen, Shaoxin Lin, Muhan Zeng, Daoguang Zan, Jian-Gang Wang, Anton Cheshkov, Jun Sun, Hao Yu, Guoliang Dong, Artem Aliev, et al. 2024. CodeR: Issue resolving with multi-agent and task graphs.arXiv preprint arXiv:2406.01304 (2024). doi:10.48550/arXiv.2406.01304 [7] Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Martin Monperrus. 2019. Sequencer: Sequence-to-sequence learning for end-to-end program repair.IEEE Transactions on Software Engineer- ing47, 9 (2019), 1943–1959. doi:10.1109/TSE.2019.2940179 [8] Yiu Wai Chow, Luca Di Grazia, and Michael Pradel. 2024. PyTy: Repairing static type errors in Python. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13. doi:10.1145/3597503.3639184 [9] Cognition. 2024. Introducing Devin, the first AI software engineer. https: //www.cognition.ai/blog/introducing-devin [10] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261(2025). doi:10.48550/arXiv.2507.06261 [11] Martin Fowler. 2011.Eradicating Non-Determinism in Tests. https://martinfowler. com/articles/nonDeterminism.html [Online; Accessed 2026-03-25]. [12] Alexander Frömmgen, Jacob Austin, Peter Choy, Nimesh Ghelani, Lera Kharatyan, Gabriela Surita, Elena Khrapko, Pascal Lamblin, Pierre-Antoine Man- zagol, Marcus Revaj, et al. 2024. Resolving Code Review Comments with Machine Learning. InProceedings of the 46th International Conference on Software Engi- neering: Software Engineering in Practice. 204–215. doi:10.1145/3639477.3639746 [13] Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided Language Models. InProceedings of the 40th International Conference on Machine Learning. 10764–
- doi:10.48550/arXiv.2211.10435
[14] Google Inc. 2026. Google Antigravity. https://antigravity.google. Online; Accessed 2026-03-25. [15] Google Inc. 2026. Google Colab. https://colab.research.google.com. Online; Accessed 2026-03-25. [16] Google Inc. 2026. Google Gemini: Google’s AI Assistant. https://gemini.google. com. Online; Accessed 2026-03-25. [17] Rahul Gupta, Aditya Kanade, and Shirish Shevade. 2019. Deep reinforcement learning for syntactic error repair in student programs. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 930–937. doi:10.1609/aaai.v33i01. 3301930 [18] Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade. 2017. Deepfix: Fixing common c language errors by deep learning. InProceedings of the aaai conference on artificial intelligence, Vol. 31. doi:10.1609/aaai.v31i1.10742 [19] Jacob Harer, Onur Ozdemir, Tomo Lazovich, Christopher Reale, Rebecca Russell, Louis Kim, et al. 2018. Learning to repair software vulnerabilities with generative adversarial networks.Advances in neural information processing systems31 (2018). doi:10.48550/arXiv.1805.07475 [20] Dániel Hidvégi, Kamyar Etemadi, S Bobadilla, and Martin Monperrus. 2024. Cigar: Cost-efficient program repair with LLMs.arXiv preprint arXiv:2402.06598 (2024). doi:10.48550/arXiv.2402.06598 [21] James W. Hunt and M. Douglas McIlroy. 1976.An algorithm for differential file comparison. Technical Report Computing Science Technical Report 41. Bell Laboratories, Murray Hill, NJ. http://www.cs.dartmouth.edu/~doug/diff.pdf [22] Naman Jain, Shubham Gandhi, Atharv Sonwane, Aditya Kanade, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma. 2023. Staticfixer: From static analysis to static repair.arXiv preprint arXiv:2307.12465 (2023). doi:10.48550/arXiv.2307.12465 [23] Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan. 2023. Impact of code language models on automated program repair. InProceedings of the IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1430–1442. doi:10.1109/ICSE48619.2023.00125 [24] Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. InThe Twelfth International Conference on Learning Representations. doi:10.48550/arXiv.2310.06770 [25] Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench Lite. https://www.swebench. com/lite.html [Online; Accessed 2026-03-25]. [26] Harshit Joshi, José P Cambronero Sánchez, Sumit Gulwani, Vu Le, Gust Ver- bruggen, and Ivan Radicek. 2023. Repair is nearly generation: Multilingual program repair with LLMs. InProceedings of the AAAI Conference on Artificial Intelligence. 5131–5140. doi:10.1609/aaai.v37i4.25642 [27] Sungmin Kang, Bei Chen, Shin Yoo, and Jian-Guang Lou. 2023. Explainable automated debugging via large language model-driven scientific debugging. Empirical Software Engineering(2023). doi:10.1007/s10664-024-10594-x [28] Yalin Ke, Kathryn T Stolee, Claire Le Goues, and Yuriy Brun. 2015. Repairing programs with semantic code search (t). In2015 30th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 295–306. doi:10.1109/ ASE.2015.60 [29] Dongsun Kim, Jaechang Nam, Jaewoo Song, and Sunghun Kim. 2013. Automatic patch generation learned from human-written patches. In2013 35th international conference on software engineering (ICSE). IEEE, 802–811. doi:10.1109/ICSE.2013. 6606626 [30] Leslie Lamport. 1994.LaTeX: A Document Preparation System(2nd ed.). Addison- Wesley. [31] Xuan Bach D Le, David Lo, and Claire Le Goues. 2016. History driven program repair. In2016 IEEE 23rd international conference on software analysis, evolution, and reengineering (SANER), Vol. 1. IEEE, 213–224. doi:10.1109/SANER.2016.76 [32] Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2011. Genprog: A generic method for automatic software repair.Ieee transactions on software engineering38, 1 (2011), 54–72. doi:10.1109/TSE.2011.104 [33] Claire Le Goues, Michael Pradel, and Abhik Roychoudhury. 2019. Automated program repair.Commun. ACM62, 12 (2019), 56–65. doi:10.1145/3318162 [34] Cheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang, Zhouruixing Zhu, Lingming Zhang, and Michael R Lyu. 2025. Unidebugger: Hierarchical multi- agent framework for unified software debugging. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 18248–18277. doi:10.18653/v1/2025.emnlp-main.921 [35] LeetCode. 2026. LeetCode - LeetCode - The World’s Leading Online Programming Learning Platform. https://leetcode.com [Online; Accessed 2026-03-25]. [36] Yi Li, Shaohua Wang, and Tien N Nguyen. 2020. Dlfix: Context-based code transformation learning for automated program repair. InProceedings of the ACM/IEEE 42nd international conference on software engineering. 602–614. doi:10. 1145/3377811.3380345 [37] Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F Bissyandé. 2019. TBar: Revisiting template-based automated program repair. InProceedings of the 28th ACM SIGSOFT international symposium on software testing and analysis. 31–42. doi:10.1145/3293882.3330577 [38] Yizhou Liu, Pengfei Gao, Xinchen Wang, Jie Liu, Yexuan Shi, Zhao Zhang, and Chao Peng. 2024. Marscode agent: Ai-native automated bug fixing.arXiv preprint arXiv:2409.00899(2024). doi:10.48550/arXiv.2409.00899 [39] Yu Liu, Sergey Mechtaev, Pavle Subotić, and Abhik Roychoudhury. 2023. Pro- gram repair guided by datalog-defined static analysis. InProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1216–1228. doi:10.1145/3611643.3616363 [40] Fan Long and Martin Rinard. 2016. Automatic patch generation by learning correct code. InProceedings of the 43rd annual ACM SIGPLAN-SIGACT symposium on principles of programming languages. 298–312. doi:10.1145/2837614.2837617 [41] Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li, Moshi Wei, and Lin Tan. 2020. Coconut: combining context-aware neural translation models using ensemble for program repair. InProceedings of the 29th ACM SIGSOFT international symposium on software testing and analysis. 101–114. doi:10.1145/ 3395363.3397369 [42] Chandra Maddila, Negar Ghorbani, James Saindon, Parth Thakkar, Vijayaragha- van Murali, Rui Abreu, Jingyue Shen, Brian Zhou, Nachiappan Nagappan, and Peter C Rigby. 2026.AI-Assisted Fixes to Code Review Comments at Scale. As- sociation for Computing Machinery, New York, NY, USA, 439–449. https: //doi.org/10.1145/3803437.3805217 [43] Chandra Maddila, Adam Tait, Claire Chang, Daniel Cheng, Nauman Ahmad, Vijayaraghavan Murali, Marshall Roch, Arnaud Avondet, Aaron Meltzer, Victor Montalvao, et al. 2025. Agentic Program Repair from Test Failures at Scale: A Neuro-symbolic approach with static analysis and test execution feedback.IEEE Transactions on Software Engineering(2025). doi:10.1109/TSE.2026.3696849 [44] David H. Maister. 1985. The Psychology of Waiting Lines. InThe Service En- counter: Managing Employee/Customer Interaction in Service Businesses, John A. Czepiel, Michael R. Solomon, and Carol F. Surprenant (Eds.). Lexington Books, Lexington, MA, 113–123. [45] Yacine Majdoub and Eya Ben Charrada. 2024. Debugging with open-source large language models: An evaluation. InProceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement. 510–516. doi:10.1145/3674805.3690758 [46] Petros Maniatis and Daniel Tarlow. 2023.Large sequence models for software development activities. https://research.google/blog/large-sequence-models-for- software-development-activities [Online; Accessed 2026-03-25]. [47] Sergey Mechtaev, Jooyong Yi, and Abhik Roychoudhury. 2016. Angelix: Scalable multiline program patch synthesis via symbolic analysis. InProceedings of the 38th international conference on software engineering. 691–701. doi:10.1145/2884781. 2884807 [48] Atif Memon, Zebao Gao, Bao Nguyen, Sanjeev Dhanda, Eric Nickell, Rob Siem- borski, and John Micco. 2017. Taming Google-scale continuous testing. In 2017 IEEE/ACM 39th International Conference on Software Engineering: Software Engineering in Practice Track (ICSE-SEIP). IEEE, 233–242. doi:10.1109/ICSE- SEIP.2017.16 [49] J Micco. 2012. Tools for continuous integration at Google scale. https://www. youtube.com/watch?v=KH2_sB1A6lA.Google Tech Talk, Google Inc(2012). On- line; Accessed 2026-03-25. [50] John Micco. 2016.Flaky Tests at Google and How We Mitigate Them. https: //testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html [On- line; Accessed 2026-03-25]. [51] Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke.
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance
Software Engineering?arXiv preprint arXiv:2502.12115(2025). doi:10.48550/ arXiv.2502.12115 [52] Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre. 2023. Octopack: Instruction tuning code large language mod- els. InNeurIPS 2023 workshop on instruction tuning and instruction following. doi:10.48550/arXiv.2308.07124 [53] Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chan- dra. 2013. Semfix: Program repair via semantic analysis. In2013 35th International Conference on Software Engineering (ICSE). IEEE, 772–781. doi:10.1109/ICSE.2013. 6606623 [54] Vincent Nguyen, Guilherme Herzog, José Cambronero, Marcus Revaj, Aditya Kini, Alexander Frömmgen, and Maxim Tabachnyk. 2025. Smart Paste: Automat- ically Fixing Copy/Paste for Google Developers. InProceedings of the IEEE/ACM 48th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). doi:10.1145/3786583.3786859 [55] Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez. 2023. Go- rilla: Large language model connected with massive APIs.Advances in Neural Information Processing Systems(2023). doi:10.52202/079017-4020 [56] Md Nakhla Rafi, Dong Jae Kim, Tse-Hsun Chen, and Shaowei Wang. 2024. A multi-agent approach to fault localization via graph-based retrieval and reflexion. arXiv preprint arXiv:2409.13642(2024). doi:10.48550/arXiv.2409.13642 [57] Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, and Satish Chandra. 2025. Evaluating agent-based program repair at Google. InProceedings of the 47th International Conference on Software Engineering: Software Engineering in Practice Track. doi:10.1109/ICSE- SEIP66354.2025.00038 [58] Haifeng Ruan, Yuntong Zhang, and Abhik Roychoudhury. 2024. SpecRover: Code intent extraction via LLMs. InProceedings of the 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). doi:10.1109/ICSE55347.2025.00080 [59] Caitlin Sadowski, Emma Söderberg, Luke Church, Michal Sipko, and Alberto Bacchelli. 2018. Modern code review: a case study at Google. InProceedings of the 40th International Conference on Software Engineering: Software Engineering in Practice. ACM, 181–190. doi:10.1145/3183519.3183525 [60] Georgios Sakkas, Madeline Endres, Philip J Guo, Westley Weimer, and Ranjit Jhala. 2022. Seq2parse: neurosymbolic parse error repair.Proceedings of the ACM on Programming Languages6, OOPSLA2 (2022), 1180–1206. doi:10.1145/3563330 [61] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems36 (2023), 68539–68551. doi:10.52202/ 075280-2997 [62] Daniel Tarlow, Subhodeep Moitra, Andrew Rice, Zimin Chen, Pierre-Antoine Manzagol, Charles Sutton, and Edward Aftandilian. 2020. Learning to fix build errors with graph2diff neural networks. InProceedings of the IEEE/ACM 42nd international conference on software engineering workshops. 19–20. doi:10.1145/ 3387940.3392181 [63] Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019. On learning meaningful code changes via neural machine translation. In2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE). IEEE, 25–36. doi:10.1109/ICSE.2019.00021 [64] Rijnard Van Tonder and Claire Le Goues. 2018. Static automated program repair for heap properties. InProceedings of the 40th International Conference on Software Engineering. 151–162. doi:10.1145/3180155.3180250 [65] Marko Vasic, Aditya Kanade, Petros Maniatis, David Bieber, and Rishabh Singh.
- Neural Program Repair by Jointly Learning to Localize and Repair.CoRR
abs/1904.01720 (2019). arXiv:1904.01720 doi:10.48550/arXiv.1904.01720 [66] Ke Wang, Rishabh Singh, and Zhendong Su. 2018. Search, align, and repair: data-driven feedback generation for introductory programming exercises. In Proceedings of the 39th ACM SIGPLAN conference on programming language design and implementation. 481–495. doi:10.1145/3192366.3192384 [67] Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. 2025. OpenHands: An open platform for AI software developers as generalist agents. InInternational Conference on Learning Representations, Vol. 2025. 65882–65919. doi:10.48550/ arXiv.2407.16741 [68] Kari Edison Watkins, Brian Ferris, Alan Borning, G Scott Rutherford, and David Layton. 2011. Where Is My Bus? Impact of mobile real-time information on the perceived and actual wait time of transit riders.Transportation Research Part A: Policy and Practice45, 8 (2011), 839–848. doi:10.1016/j.tra.2011.06.010 [69] Wikipedia contributors. 2026. diff - Wikipedia. https://en.wikipedia.org/wiki/ Diff [Online; Accessed 2026-03-25]. [70] Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2025. Demystifying llm-based software engineering agents.Proceedings of the ACM on Software Engineering2, FSE (2025), 801–824. doi:10.1145/3715754 [71] Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023. Automated program repair in the era of large pre-trained language models. InProceedings of the IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1482–1494. doi:10.1109/ICSE48619.2023.00129 [72] Chunqiu Steven Xia and Lingming Zhang. 2022. Less training, more repairing please: revisiting automated program repair via zero-shot learning. InProceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 959–971. doi:10.1145/3540250.3549101 [73] Chunqiu Steven Xia and Lingming Zhang. 2024. Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 819–831. doi:10.1145/3650212.3680323 [74] Jifeng Xuan, Matias Martinez, Favio Demarco, Maxime Clement, Sebastian Lame- las Marcote, Thomas Durieux, Daniel Le Berre, and Martin Monperrus. 2016. Nopol: Automatic repair of conditional statement bugs in java programs.IEEE Transactions on Software Engineering43, 1 (2016), 34–55. doi:10.1109/TSE.2016. 2560811 [75] Deheng Yang, Xiaoguang Mao, Liqian Chen, Xuezheng Xu, Yan Lei, David Lo, and Jiayu He. 2022. Transplantfix: Graph differencing-based code transplantation for automated program repair. InProceedings of the 37th IEEE/ACM International Con- ference on Automated Software Engineering. 1–13. doi:10.1145/3551349.3556893 [76] John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-Agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems37 (2024), 50528–50652. doi:10.52202/079017-1601 [77] John Yang, Kilian Lieret, Carlos E Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang. 2025. SWE-Smith: Scaling data for software engineering agents.Advances in Neural Information Processing Systems(2025). doi:10.52202/085713-3239 [78] Weiqing Yang, Hanbin Wang, Zhenghao Liu, Xinze Li, Yukun Yan, Shuo Wang, Yu Gu, Minghe Yu, Zhiyuan Liu, and Ge Yu. 2025. COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis. InFindings of the Association for Computational Linguistics: NAACL
- Association for Computational Linguistics, 2570–2585. doi:10.18653/v1/
2025.findings-naacl.139 [79] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InThe Eleventh International Conference on Learning Representations. doi:10.48550/arXiv.2210.03629 [80] He Ye, Matias Martinez, Xiapu Luo, Tao Zhang, and Martin Monperrus. 2022. SelfAPR: Self-supervised program repair with test execution diagnostics. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (ASE). ACM, 92:1–92:13. doi:10.1145/3551349.3556926 [81] He Ye, Matias Martinez, and Martin Monperrus. 2022. Neural program repair with execution-based backpropagation. InProceedings of the 44th international conference on software engineering. 1506–1518. doi:10.1145/3510003.3510222 [82] He Ye and Martin Monperrus. 2024. Iter: Iterative neural repair for multi-location patches. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE). doi:10.1145/3597503.3623337 [83] Tingting Yu and Michael Pradel. 2018. Pinpointing and repairing performance bottlenecks in concurrent programs.Empirical Software Engineering23, 5 (2018), 3034–3071. doi:10.1007/s10664-017-9578-1 [84] Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024. Autocoderover: Autonomous program improvement. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 1592–
- doi:10.1145/3650212.3680384
[85] Qihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang, Kang Yuan, Yingfei Xiong, and Lu Zhang. 2021. A syntax-guided edit decoder for neural program repair. InProceedings of the 29th ACM joint meeting on European software engineering conference and symposium on the foundations of software engineering. 341–353. doi:10.1145/3468264.3468544 [86] Celal Ziftci and Diego Cavalcanti. 2020. De-flake your tests: Automatically locating root causes of flaky tests in code at Google. In2020 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 736–745. doi:10. 1109/ICSME46990.2020.00083
出处:Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini,Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale,2026-10-05,https://arxiv.org/abs/2610.07289,CC BY 4.0。
Found it useful? Pass it on
Scan with WeChat to open it on your phone and forward it.