SurviAGI
← All articles

Sep 29–Oct 7, 2026 | Judgment now costs two cents

Between September 29 and October 7, 2026 we read just over 3,400 posts and kept about 240 things that actually happened. Eleven companies announced 68 of them; the rest came from evaluation groups and individuals. Each is a news item on its own. Together they make five threads.

1. Decision models became a category in one week

Perplexity also posted a bill: its model cleared the Elite Four and the Champion in Pokémon FireRed in 137 API calls for 2.8 cents.

Eight days, four vendors: one open-sourced, one cut in half, one free, and now one you can run locally. A new category went straight into a price war.

2. Evaluation moved from answering questions to doing a long job

A model can write a 17,000-line proof and still not keep a vending machine in business. Whether a long job can be handed over depends on whether anyone catches it when it goes wrong.

3. People started handing whole chores to agents

The supporting pieces appeared the same week:

What gets handed over first is not the job but the chores. And payment companies and retailers have started building roads for agents.

4. Large models running on your own machine

The biggest thing on this thread did not come from a company. Strata, an open-source inference engine written by one developer, put Qwen3.8-Flash-Next, a model of more than a hundred billion parameters, onto ordinary gaming PCs.

  • Its author @coldniko wrote that a cheap old DDR3 machine with 64 or 128 GB of memory and a GPU of 8 GB or more runs it at over 70 tokens per second. A day later he raised output speed from 64.5 to 76.4 tokens per second.
  • @Yacamochi_db measured the 87 GB model at 120 to 150 tokens per second on a single RTX 5090.
  • @rS_alonewolf measured 100 to 140 on an RTX 5090 and 50 to 70 on an RTX 5070 Ti.
  • @daluoseo measured 96 tokens per second with 12 GB of GPU memory.
  • @chikitleung reached 38 on an 8 GB RTX 4060 Ti.

Among the posts our search returned, 39 accounts published results from their own machines, most of them writing in Japanese or Chinese.

Apple silicon moved too:

Not everything worked: @Gimenorum asked a more heavily compressed version to fix the same problem five times in a row, and it did not.

In one week dozens of people ran the same large model on their own old graphics cards. The bar dropped from a server to a gaming PC. But seven of nine setups leaking data says that running and being safe to rely on are still two different things.

5. Open-weight models arrived in a cluster

  • Mistral released a preview of Mistral Large 4: one trillion parameters, 49 billion active at a time, weights due at the end of October. It is listed on OpenRouter at 68 cents per million input tokens.
  • Google released EmbeddingGemma 2 (the work it bears on), a 740-million-parameter model for on-device use.
  • Cristóbal Valenzuela announced Praxis-1, an open-weight action model for robots.
  • Perplexity's decision model is open source as well.
  • A new non-profit, Trillium Labs, launched to build open training recipes for frontier models.

In one week, from a trillion-parameter model down to an embedding model for phones, each came in a version you can download.

You

What follows is this issue's judgment and guesswork, not data.

One: if your product sells "helping people choose", this week it stopped being a product. Routing, triage, first-pass screening, recommendation, review: at two cents these are a feature. The question is no longer whether your judgment is accurate. It is what you hold besides judgment: the data, the channel, or the liability when it goes wrong.

Two: the hours you spend on chores each week will disappear before your actual job does. Not because the models suddenly got smarter, but because someone has started building roads for agents. Check-ins, expenses, second-hand listings, forms: what they share is that a mistake can be undone.

Three: do not hand over a whole long job yet. This week's best model only partly managed to run a vending machine. What you can hand over safely is work where a mistake is visible and fixable. The test is simple: if it gets this wrong, how long before you notice?

A guess: voice is the next thing to fall to a few cents. Microsoft's voice models appeared on several platforms within days of release, which is usually how a price war starts. If your business assumes speech transcription is expensive, redo the arithmetic now.


Data through 10:32 UTC on October 7, 2026. Every figure is as stated by the account that posted it; follow the links to the original posts. For the work each update bears on and the level it reached, see [all updates](/updates) and [occupations](/occupations).