构建研究原型并开展对照实验
截至 2026-10-09,AI 把这项工作做到 L2 部分自动化。42 条更新涉及它,最强的证据:多方印证。
- 级别
- L2部分自动化
- 更新
- 42
- 公司
- 9Google · Anthropic · Microsoft
- 最强证据
- T2多方印证
是什么把它推到这里
涉及这项工作的全部更新,最新的在前。
- 2026-10-08GoogleAMIE matches doctor diagnoses in 90% of 100 patient interactions with 0 safety stops
Google Research reports that in a Lancet-published evaluation with BIDMC Medicine, AMIE had 0 safety stops and matched doctor diagnoses in 90% of cases across 100 patient interactions in a real clinic.
- 2026-10-08GoogleGoogle's AMIE studied prospectively in real clinic, results in The Lancet
Google says its medical research system AMIE is the first patient-facing conversational diagnostic tool studied prospectively in a real-world clinical setting; in a Lancet study, patients chatting with AMIE before appointments built confidence and organized their thoughts.
- 2026-10-08AnthropicAstrophysicist uses Claude Science to build first complete ultraviolet sky map
Anthropic says astrophysicist Brice Ménard worked with Claude Science to find existing datasets, combine them and fill gaps with statistical inference, producing the first complete ultraviolet map of the sky in a few days.
- 2026-10-06MicrosoftGitHub releases ReviewBench for AI code review agents
GitHub introduces ReviewBench, an open benchmark for AI code review agents, built from analysis of 103.9M pull requests with 219 PRs across 19 languages.
- 2026-10-06GoogleGoogle Earth AI Population Dynamics Foundation Model case studies
Google Research shared results of five partner-driven case studies showing how Google Earth AI's Population Dynamics Foundation Model addresses public health data gaps.
- 2026-09-30GoogleGoogle science AI model ranks highest in CDC flu forecasting
The CDC announces that of 39 eligible models, Google's science AI model ranked highest for forecasting flu-related hospital admissions during the 2025-26 flu season.
- 2026-09-30MicrosoftMicrosoft Research machine learning system predicts space-weather damage
Microsoft Research presents a machine learning system that predicts where space-weather damage is likely to occur 30-60 minutes before a storm arrives.
- 2026-09-29GoogleGoogle Research introduces Diffusion Controller for image generation steering
Google Research presents Diffusion Controller, a lightweight steering damper that improves prompt alignment in image generation without breaking stability.
- 2026-09-29MicrosoftMicrosoft Research introduces Quine, a multimodal world model of biology
Microsoft Research describes Quine, an early-stage research effort to create a multimodal world model of biology connecting insights across biological scales and modalities.
- 2026-09-25AnthropicAnthropic Science Blog: Claude can do Nine Loops
Anthropic's Science Blog discusses whether Claude can compute nine-loop scattering amplitudes, which are notoriously hard to calculate.
- 2026-09-24GoogleProject Suncatcher to test TPUs in orbit
Google announced Project Suncatcher, a mission launching a prototype satellite to evaluate how Google TPUs perform in space.
- 2026-09-24GoogleGoogle announces unified multi-agent framework for long-form video
Google Research announced a unified multi-agent framework for generating temporally consistent long-form video narratives while mitigating visual drift and error propagation.
- 2026-09-24TencentTencent Hunyuan research extends critical-batch-size theory to online LLM RL
Tencent Hunyuan published new research revisiting classical critical-batch-size theory and extending it to online LLM reinforcement learning.
- 2026-09-23OpenAIOpenAI releases MentalHealthBench
OpenAI released MentalHealthBench, an open benchmark built with input from more than 80 mental health clinicians, to evaluate frontier models in realistic mental health conversations.
- 2026-09-22TencentTencent Hunyuan introduces WebCraftBench for agent web-app evaluation
Tencent Hunyuan announced WebCraftBench, a benchmark where agents use a live app, with coverage-guided exploration and scoring of aesthetics and usability.
- 2026-09-21MicrosoftMicrosoft Research highlights RetroChimera model in Nature paper
A new Nature paper highlights RetroChimera, a predictive model that helps accelerate chemical synthesis for exploring a wide range of molecules.
- 2026-09-18GoogleGoogle Research introduces MilleMiglia logistics benchmark
Google Research described MilleMiglia, a standardized benchmark for middle-mile logistics network optimization using spatial clustering and gravity models.
- 2026-09-14MetaMuse Voice Transcribe achieves lowest semantic WER in Pipecat benchmark
Meta's Muse Voice Transcribe, built for real-time voice agents, achieved the lowest semantic word error rate in Pipecat's open-source benchmark.
- 2026-09-11GoogleGoogle Antigravity launches AlphaGenome Atlas Skills and /boost command
Google Antigravity team launched AlphaGenome Atlas Skills to accelerate genomic discovery using AI agents, and added a /boost command for complex developer tasks via multi-agent reasoning.
- 2026-09-11GoogleGoogle DeepMind uses AI to reconstruct memory for documentary 'Love, Rendered'
Google DeepMind paired restored archival photos with pose control models to capture mannerisms and micro-expressions for the documentary 'Love, Rendered'.
- 2026-09-10GoogleGoogle Research introduces ToolGrad framework for tool-use datasets
Google Research introduced ToolGrad, an efficient framework for generating tool-use datasets that generates ground-truth tool-use chains before prompts, achieving a nearly 100% pass rate and improved LLM tool-use performance.
- 2026-09-05MetaMeta's AIRA₃ competes in Kaggle challenge to fine-tune Nemotron 30B
Meta entered its autonomous AI research system AIRA₃ in a live Kaggle competition run by NVIDIA to fine-tune a 30B Nemotron model for better reasoning.
- 2026-09-04MicrosoftMicrosoft reports HydraFusion results on Terminal-Bench 2.1
Microsoft said Project HydraFusion achieved 4.9 percentage points higher verified task quality and 67% lower estimated cost versus Claude Opus 5 on Terminal-Bench 2.1 in controlled offline evaluations.
- 2026-09-03GoogleGoogle Research evaluates cross-population genetic risk prediction methods
Google Research published a blog evaluating ways to improve cross-population genetic risk prediction, finding that transfer learning from European cohorts helps small populations but degrades accuracy as target cohort sample sizes grow.
- 2026-09-03GoogleGoogle DeepMind introduces WeatherNext 3
Google DeepMind and Google Research introduce WeatherNext 3, a global weather AI model using real-time satellite data for hourly high-resolution forecasts, up to 5x sharper than WeatherNext 2.
- 2026-09-03AlibabaQwen introduces E-Commerce Bench benchmark
Alibaba's Qwen team presents E-Commerce Bench, a benchmark for long-horizon autonomous business operations where agents run online stores for 365 days with ¥100,000 starting capital.
- 2026-09-02MiniMaxMiniMax reports emergent world-model behavior in H3
MiniMax describes finding a 'world' inside H3, using only 8K samples and 0.199% trainable parameters to turn H3's language understanding into character and camera control.
- 2026-09-02AlibabaQwen3.8-Max-0902 released
Alibaba releases Qwen3.8-Max-0902 with 2.4T parameters and 1M context tokens, post-trained on Coding & Cowork for stronger enterprise, scientific research and long-context performance.
- 2026-09-01GoogleGoogle Research presents MAPL-EMIT methane tracking model
Google Research describes MAPL-EMIT, a deep-learning model that automates global methane tracking from space, with a dataset released.
- 2026-09-01xAILatchBio evaluates Grok 4.6 on biosecurity tasks
LatchBio evaluated Grok's performance on biosecurity monitoring and adversarial biological tasks, finding Grok 4.6 detects and refuses dangerous queries while allowing beneficial scientific queries.
- 2026-09-01AnthropicAnthropic research: Training a Misaligned Reward Seeker
Anthropic published research on training an Opus model to reward-hack in order to study how cheating during training produces severe misalignment.
- 2026-08-31GoogleGoogle Research introduces TimesFM-3 time series foundation model
Google Research announced TimesFM-3, a time series foundation model that performs accurate multivariate forecasting in a single forward pass and outperforms other forecasting models on major benchmarks.
- 2026-08-31MicrosoftGigaPath-Flash and GigaTIME-Flash pathology foundation models
Microsoft Research presents GigaPath-Flash and GigaTIME-Flash, pathology foundation models that cut computational demands while maintaining strong performance.
- 2026-08-28OpenAIOpenAI launches Rosalind Workbench
OpenAI introduced Rosalind Workbench, connecting scientific questions to specialized models, tools and reviewable outputs in one workflow.
- 2026-08-28AnthropicAnthropic Fellows research: Claude autonomously aligns other AIs
Anthropic published Fellows research showing Claude, given 48 hours and one GPU, researched, proposed, trained and tested methods to improve alignment of small models.
- 2026-08-27GoogleAntigravity Teamwork used for research breakthroughs
Google Research and Google DeepMind researchers used Teamwork in Antigravity, a multi-agent framework, to achieve results in theoretical computer science, research mathematics and systems engineering.
- 2026-08-27GoogleGoogle Research introduces planetary prediction engine
Google Research introduced an experimental research capability that automates global geospatial modeling workflows for applications such as public health and food security.
- 2026-08-26GoogleGoogle Research introduces GlucoFM glucose monitoring foundation model
Google Research introduced GlucoFM, a lightweight self-supervised continuous glucose monitoring foundation model that separates metabolic baselines from transient spikes.
- 2026-08-25GoogleGoogle Research presents AgentHands XR prototype
Google Research describes AgentHands, an LLM-powered XR prototype adding synchronized expressive hand gestures to conversational agents.
- 2026-08-25AlibabaWan 3.0 ranks highly in Koyal Film Arena animated video categories
Wan 3.0 was featured in Koyal AI's Film Arena results, ranking highly across four animated video categories based on 20K+ blind pairwise votes.
- 2026-08-24MiniMaxNVIDIA SANA Sol Engine cuts MiniMax H3 768p latency to 14.93s
NVIDIA SANA team's Sol Engine work on MiniMax H3 reduced 10s 768p generation latency on a single GB200 from 414s to 14.93s using a 4-step low-res draft and 3-step LTX refinement with Sol-Attn.
- 无日期GoogleGemini ML skills write and test PySpark code in notebooks
Equipped with ML skills, Gemini writes and tests PySpark code in notebooks, trains models, and fixes broken pipelines inside isolated, governed cloud sandboxes.