构建模型应用评估与回归集
截至 2026-10-09,AI 把这项工作做到 L2 部分自动化。7 条更新涉及它,最强的证据:仅来自发布方。
- 级别
- L2部分自动化
- 更新
- 7
- 公司
- 5Microsoft · Tencent · OpenAI
- 最强证据
- T3仅来自发布方
是什么把它推到这里
涉及这项工作的全部更新,最新的在前。
- 2026-10-06MicrosoftGitHub releases ReviewBench for AI code review agents
GitHub introduces ReviewBench, an open benchmark for AI code review agents, built from analysis of 103.9M pull requests with 219 PRs across 19 languages.
- 2026-09-22TencentTencent Hunyuan introduces WebCraftBench for agent web-app evaluation
Tencent Hunyuan announced WebCraftBench, a benchmark where agents use a live app, with coverage-guided exploration and scoring of aesthetics and usability.
- 2026-09-15OpenAICognition's Devin uses GPT-6 Astra for test generation
OpenAI Devs states that GPT-6 Astra helps Cognition's Devin back up 'it works' with tests before the team ships.
- 2026-09-14OpenAIPerplexity engineer uses GPT-6 Astra in Codex for end-to-end testing
An engineer at Perplexity reports using GPT-6 Astra in Codex to build test harnesses and mock third-party API responses for end-to-end testing.
- 2026-09-10GoogleGoogle Research introduces ToolGrad framework for tool-use datasets
Google Research introduced ToolGrad, an efficient framework for generating tool-use datasets that generates ground-truth tool-use chains before prompts, achieving a nearly 100% pass rate and improved LLM tool-use performance.
- 2026-09-03AlibabaQwen introduces E-Commerce Bench benchmark
Alibaba's Qwen team presents E-Commerce Bench, a benchmark for long-horizon autonomous business operations where agents run online stores for 365 days with ¥100,000 starting capital.
- 2026-08-26MicrosoftMicrosoft Foundry adds agent cost optimization via routing, caching and evaluation
Microsoft Foundry describes four ways to optimize agent costs through smarter routing, caching and evaluation.