Not translated yet — showing the Chinese originals.
Industry & Practice · Tagged “评估器”
Research & Benchmarks//8 min
华为云 AgentArts 把评测集当成考卷:至少 30 到 50 条,正向、边界、对抗、安全大约按 50%、20%、20%、10% 配置。防幻觉必须单列 context,actual_output 不能提前写进试卷。
Research & Benchmarks//6 min
华为云 AgentArts 预置 50 多个评估器。只拿正确性满分不能上线:订票答对却没调 API、事实对但风格不合、回答正确却带偏见,都要靠结果、调用链和体验三类评估器互相制约。
Research & Benchmarks//11 min
阿里云 AgentLoop 在 UGC 游戏 Agent 上用 3 天走完观测、分层评估、根因和回归。事实层按正确项除以正确加错误项打分,一条样本 16 项里 9 项错。回写后的新分数没有公布。