← Software testing and defect analysis
Design test scenarios and test data
As of 2026-10-09, AI does this work at L2 Partial automation. 3 updates bear on it. Strongest evidence: From the publisher only.
- Level
- L2Partial automation
- Updates
- 3
- Companies
- 2Tencent · OpenAI
- Strongest evidence
- T3From the publisher only
L0ManualL1AssistedL2Partial automationL3Conditional automationL4High automationL5Full automation
What moved it
Every update that bears on this work, newest first.
- 2026-09-22TencentTencent Hunyuan introduces WebCraftBench for agent web-app evaluation
Tencent Hunyuan announced WebCraftBench, a benchmark where agents use a live app, with coverage-guided exploration and scoring of aesthetics and usability.
- 2026-09-15OpenAICognition's Devin uses GPT-6 Astra for test generation
OpenAI Devs states that GPT-6 Astra helps Cognition's Devin back up 'it works' with tests before the team ships.
- 2026-09-14OpenAIPerplexity engineer uses GPT-6 Astra in Codex for end-to-end testing
An engineer at Perplexity reports using GPT-6 Astra in Codex to build test harnesses and mock third-party API responses for end-to-end testing.