SurviAGI
← All articles

Oct 5–11, 2026 | AI turned in 722 mathematics papers and passed 9.4% of its legal tasks

What AI took over this week: the weekly board and a seven-week timeline

This week we read more than 3,700 posts from 276 accounts and kept about 100 things that actually happened. The board above shows the six that drew the most attention, ordered by views of the original post. Three sections follow: what was taken over, what it changed, and what people are arguing about.

1. What AI took over this week

1. Software engineers: the answer is no longer a paragraph. It is an interface. OpenAI rolled out GPT-6 and Intelligent UI to everyone (the work it touches): a ChatGPT answer can be an interface you click and operate, or a small tool made on the spot.

2. Mathematicians: three hours, one mathematics paper. OpenAI released a broad range of mathematical results (the work it touches) produced by an internal model that has not been released.

3. Designers: one sentence, one live dashboard. Anthropic's Claude Dashboards and Claude Motion entered beta (the work it touches): data becomes a live dashboard, an idea becomes an animated explainer. In the same week Gemini in Google Docs began turning a written brief into a presentation deck.

4. Call center agents: after fifteen years of growth, the jobs shrink 4% a year. a16z posted a chart: call center jobs grew about 4% a year for 15 years and are now shrinking about 4% a year. The data is from Revelio Labs.

5. Lawyers: it handed in the paper and passed 9.4%. This one was not taken. Artificial Analysis and Harvey moved their legal agent benchmark to v1.1: a task counts only when the deliverables meet every rubric criterion and contain no material misstatement.

  • The leader, Grok 4.7, scores 9.4%.
  • Without that check Muse Spark 1.3 would lead at 26.7%. With it, 8.9% remains. Of three results that looked like passes, two carried an error that would mislead a reader.

6. Doctors: it takes the history first, then the doctor takes over. Google's diagnostic system AMIE saw 100 patients in a real clinic (the work it touches): 0 safety stops, and its diagnosis matched the doctor's in 90% of cases, published in The Lancet. Patients talk to it first and a physician takes over.

The ones that were taken have something in common: when the work is done, someone or something can tell whether it is right. A proof can be handed to a machine, a diagnosis is reviewed by a doctor, a dashboard is right or wrong at a glance. A wrong date in a legal document reads as smoothly as a right one.

2. What it changed

Mathematicians.

  • Ethan Mollick relayed first-hand accounts from mathematicians: problems solved in ways no person would use. Two days later he wrote that at least some of the proofs had set off rapid iterative advances from many collaborators.
  • Epoch AI reports that in 3 of the 18 math subfields it tracks, more than half of arXiv papers acknowledge using AI; in differential geometry the share went from about 8% in July to about 57% in September.

Call centers.

  • US Bureau of Labor Statistics figures, as compiled by one reader: telephone call center payrolls fell from 322,200 in September 2025 to 307,200 in August 2026, down 4.7%. He adds that this shows headcount fell and does not prove AI caused it.
  • Revelio's research also finds that across 12 middle-income countries, hiring slowed more than departures did. The jobs are shrinking through fewer hires.
  • A reading in the other direction: a post citing a study of 5,172 call center agents says AI made the least skilled 36% more productive, while top performers declined slightly.

Everyone else.

The number of jobs is falling, and among those who stay the weakest gain the most. The two readings describe two stages of the same thing.

3. What people are arguing about

Each of these questions gained new claims this week. You can vote on each one, once per quarter.

You

What follows is this issue's judgment and guesswork, not data.

One: look at your own work and ask who would notice if you got it wrong, and how fast. Work that a machine or the next step can check quickly gets taken first; that is how mathematics arrived where it is. Work that cannot be checked is safe for now, but not because AI cannot do it. It is safe because nobody dares to trust it.

Two: if your work cannot be checked, your value is moving from doing to vouching. A lawyer's signature, a doctor's review: what is paid for is someone answering for the result.

Three: do not read "it delivered" as "it got it right". On legal tasks, two in three of the results that looked like passes carried an error. Before you hand work over, have a check you trust.

A guess: the next field to see "a decade in a week" will be another one where right and wrong can be decided automatically. Chip verification, compilers and cryptographic protocols are candidates.


Data as of 02:55 UTC on October 11, 2026. View counts are the original posts' readings when collected. Figures about the mathematics release come from accounts describing it; OpenAI's own text and repository are authoritative. Other figures come from the poster's own statement. "Work touched by AI" is a machine mapping that no person has reviewed. For the work each update touches, see [all updates](/updates); for every question, see [topics](/topics).