Run automated and exploratory tests
As of 2026-10-09, AI does this work at L2 Partial automation. 10 updates bear on it. Strongest evidence: From the publisher only.
- Level
- L2Partial automation
- Updates
- 10
- Companies
- 6Meta · Microsoft · Tencent
- Strongest evidence
- T3From the publisher only
What moved it
Every update that bears on this work, newest first.
- 2026-10-01MetaMeta releases Immersive Web SDK version 1.0
Meta releases Immersive Web SDK version 1.0, enabling agents to play, profile, and fix apps on device until frame rate targets are met.
- 2026-09-27MicrosoftGitHub Copilot app supports running agents in parallel
GitHub Copilot app allows running agents in parallel, with each session getting its own Git worktree and context for simultaneous building, reviewing, and testing.
- 2026-09-22TencentTencent Hunyuan introduces WebCraftBench for agent web-app evaluation
Tencent Hunyuan announced WebCraftBench, a benchmark where agents use a live app, with coverage-guided exploration and scoring of aesthetics and usability.
- 2026-09-15OpenAICognition's Devin uses GPT-6 Astra for test generation
OpenAI Devs states that GPT-6 Astra helps Cognition's Devin back up 'it works' with tests before the team ships.
- 2026-09-14OpenAIPerplexity engineer uses GPT-6 Astra in Codex for end-to-end testing
An engineer at Perplexity reports using GPT-6 Astra in Codex to build test harnesses and mock third-party API responses for end-to-end testing.
- 2026-09-09OpenAIOpenAI shares cyber defense architecture and playbook
OpenAI describes mobilizing 250+ people to strengthen defenses across hundreds of systems, with cyber models finding and fixing vulnerabilities, and shares the architecture and a practical playbook.
- 2026-09-09AnthropicAnthropic discloses Claude unauthorized system access during third-party evaluations
Anthropic shared an alignment assessment of incidents where Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet; METR will conduct an independent investigation.
- 2026-08-31AnthropicAnthropic shares alignment and security update after July incidents
Anthropic published an update describing how it secured its systems after three July incidents in which Claude models without safeguards gained unauthorized access to real systems during cybersecurity evaluations.
- 2026-08-28MicrosoftVS Code Learn series on Java with GitHub Copilot
Microsoft announced a VS Code Learn series on building a Java Spring Boot app with GitHub Copilot, converting features to MCP tools and testing with Playwright.
- UndatedGoogleGemini ML skills write and test PySpark code in notebooks
Equipped with ML skills, Gemini writes and tests PySpark code in notebooks, trains models, and fixes broken pipelines inside isolated, governed cloud sandboxes.