# Oct 5–11, 2026 | AI turned in 722 mathematics papers and passed 9.4% of its legal tasks

> What AI took over this week, what it changed, and what people are arguing about. Mathematical proofs, clinic intake and live dashboards were taken; legal documents were not; call center jobs began shrinking 4% a year. Every item links to its source.

Published 2026-10-11 · https://surviagi.com/articles/what-ai-took-over-2026-10-11

![What AI took over this week: the weekly board and a seven-week timeline](https://surviagi.com/figures/what-ai-took-over-2026-10-11/board.en.gif)

This week we read more than 3,700 posts from 276 accounts and kept about 100 things that actually happened. The board above shows the six that drew the most attention, ordered by views of the original post. Three sections follow: what was taken over, what it changed, and what people are arguing about.

## 1. What AI took over this week

**1. Software engineers: the answer is no longer a paragraph. It is an interface.** OpenAI rolled out [GPT-6 and Intelligent UI](https://x.com/OpenAI/status/2107894997538525580) to everyone ([the work it touches](https://surviagi.com/updates/ev-07edf2ac8d908387bb9c55ed50ca89870ebcdc85)): a ChatGPT answer can be an interface you click and operate, or a small tool made on the spot.

**2. Mathematicians: three hours, one mathematics paper.** OpenAI [released a broad range of mathematical results](https://x.com/OpenAI/status/2107596713791767021) ([the work it touches](https://surviagi.com/updates/ev-bc8366531732dcd185541ab4dab93a1e5705d43b)) produced by an internal model that has not been released.

- As described by accounts that read the release: [722 manuscripts in 372 families of results](https://x.com/kimmonismus/status/2107597320028065793), drawn from about 4,000 research problems, at an average of about three hours of compute per result.
- Many of the proofs [have been translated into Lean](https://x.com/WesRoth/status/2107682329187778998), so a machine can check them.
- OpenAI's Mark Chen says the run represents [a decade of mathematical progress in a week](https://x.com/markchen90/status/2107876922147668198). Noam Brown says models have [surpassed top human experts on some research problems](https://x.com/polynoamial/status/2107947189184164056), and also that their capabilities remain jagged.

**3. Designers: one sentence, one live dashboard.** Anthropic's [Claude Dashboards and Claude Motion](https://x.com/claudeai/status/2108271552991252810) entered beta ([the work it touches](https://surviagi.com/updates/ev-3fa4f4c1b8795b7450ba8be42bf122c016ebaf38)): data becomes a live dashboard, an idea becomes an animated explainer. In the same week Gemini in Google Docs began [turning a written brief into a presentation deck](https://x.com/GoogleWorkspace/status/2107818978722648568).

**4. Call center agents: after fifteen years of growth, the jobs shrink 4% a year.** a16z [posted a chart](https://x.com/a16z/status/2108663228029112702): call center jobs grew about 4% a year for 15 years and are now shrinking about 4% a year. The data is from Revelio Labs.

**5. Lawyers: it handed in the paper and passed 9.4%. This one was not taken.** Artificial Analysis and Harvey [moved their legal agent benchmark to v1.1](https://x.com/ArtificialAnlys/status/2108264572545310824): a task counts only when the deliverables meet every rubric criterion and contain no material misstatement.

- The leader, Grok 4.7, scores 9.4%.
- Without that check Muse Spark 1.3 would lead at 26.7%. With it, 8.9% remains. Of three results that looked like passes, two carried an error that would mislead a reader.

**6. Doctors: it takes the history first, then the doctor takes over.** Google's diagnostic system AMIE [saw 100 patients in a real clinic](https://x.com/GoogleResearch/status/2108326125676192100) ([the work it touches](https://surviagi.com/updates/ev-e006788115400035bf3a498d84fd276a3d13d37e)): 0 safety stops, and its diagnosis matched the doctor's in 90% of cases, published in The Lancet. Patients talk to it first and [a physician takes over](https://x.com/Google/status/2108324514442461225).

**The ones that were taken have something in common: when the work is done, someone or something can tell whether it is right. A proof can be handed to a machine, a diagnosis is reviewed by a doctor, a dashboard is right or wrong at a glance. A wrong date in a legal document reads as smoothly as a right one.**

## 2. What it changed

**Mathematicians.**

- Ethan Mollick relayed [first-hand accounts](https://x.com/emollick/status/2107964764316209326) from mathematicians: problems solved in ways no person would use. Two days later he [wrote](https://x.com/emollick/status/2108251560752857474) that at least some of the proofs had set off rapid iterative advances from many collaborators.
- Epoch AI [reports](https://x.com/EpochAIResearch/status/2108669023986860322) that in 3 of the 18 math subfields it tracks, more than half of arXiv papers acknowledge using AI; in differential geometry the share went from about 8% in July to about 57% in September.

**Call centers.**

- US Bureau of Labor Statistics figures, [as compiled by one reader](https://x.com/Kaiyc12/status/2108801848903938284): telephone call center payrolls fell from 322,200 in September 2025 to 307,200 in August 2026, down 4.7%. He adds that this shows headcount fell and does not prove AI caused it.
- Revelio's research also finds that across 12 middle-income countries, [hiring slowed more than departures did](https://x.com/VizierPrime/status/2108916070354735272). The jobs are shrinking through fewer hires.
- A reading in the other direction: a post [citing a study of 5,172 call center agents](https://x.com/_Setlur/status/2108402561246249416) says AI made the least skilled 36% more productive, while top performers declined slightly.

**Everyone else.**

- McKinsey: the US may have more jobs in 2035 than today, yet [roughly 11 million workers may need to change occupations](https://x.com/McKinsey_MGI/status/2108211068430332111).
- Google Research announced a three-month [randomized controlled trial with practicing patent attorneys](https://x.com/GoogleResearch/status/2107929855988048269), testing both short-term productivity and longer-term skill building.

**The number of jobs is falling, and among those who stay the weakest gain the most. The two readings describe two stages of the same thing.**

## 3. What people are arguing about

Each of these questions gained new claims this week. You can vote on each one, once per quarter.

- [Will AI replace most mathematicians?](https://surviagi.com/topics/replacement-mathematician-74b7e010) Kent Beck asks: [we prove theorems, it proves them better, so who are we now?](https://x.com/KentBeck/status/2108559701550170124) Another answer is that mathematicians [move on to explaining the proofs AI produces](https://x.com/TheStalwart/status/2108220325854851451).
- [Will AI replace most software engineers?](https://surviagi.com/topics/replacement-software-engineer-cf74e6b5) The question with the most claims. Most accounts say far fewer will be needed, but none of the claims is a measurement.
- [Will AI replace most lawyers?](https://surviagi.com/topics/replacement-lawyer-cd6623ca) Most claims say they stay and do harder work. This week's 9.4% is the first reading on the question.
- [Are call center jobs growing or shrinking?](https://surviagi.com/topics/replacement-call-center-worker-46b310c0) Many accounts say shrinking, but nearly all of them repeat the same chart.
- [Will AI replace most sales representatives?](https://surviagi.com/topics/replacement-sales-representative-1bb369ff) There are claims on both sides.
- [Will AI agents bypass Airbnb and book directly?](https://surviagi.com/topics/viability-airbnb-388d97e1) A question that first appeared this week.

## You

What follows is this issue's judgment and guesswork, not data.

**One: look at your own work and ask who would notice if you got it wrong, and how fast.** Work that a machine or the next step can check quickly gets taken first; that is how mathematics arrived where it is. Work that cannot be checked is safe for now, but not because AI cannot do it. It is safe because nobody dares to trust it.

**Two: if your work cannot be checked, your value is moving from doing to vouching.** A lawyer's signature, a doctor's review: what is paid for is someone answering for the result.

**Three: do not read "it delivered" as "it got it right".** On legal tasks, two in three of the results that looked like passes carried an error. Before you hand work over, have a check you trust.

**A guess: the next field to see "a decade in a week" will be another one where right and wrong can be decided automatically.** Chip verification, compilers and cryptographic protocols are candidates.

---

*Data as of 02:55 UTC on October 11, 2026. View counts are the original posts' readings when collected. Figures about the mathematics release come from accounts describing it; OpenAI's own text and repository are authoritative. Other figures come from the poster's own statement. "Work touched by AI" is a machine mapping that no person has reviewed. For the work each update touches, see [all updates](https://surviagi.com/updates); for every question, see [topics](https://surviagi.com/topics).*
