Briefing · Wednesday, October 7, 2026
Le Chonk, 722 Math Papers, and a $0.10 Decision API

Good morning. It's Wednesday, October 7, and we're covering Mistral's 1-trillion-parameter open-weight preview, OpenAI's 722-manuscript math release, a decision endpoint priced per input token, and the agent-permission reckoning at Meta.
The Le Chonk thread reached 1,826 points and 1,085 comments, the week's loudest launch, while OpenAI's math release drew 938 points and 874 comments and a much harder argument about verification.
In today's brief:
- Mistral Large 4: a 1T-parameter sparse multimodal model with 49B active, live in preview, with weights promised by the end of October
- 722 math manuscripts: OpenAI publishes a catalog from an unreleased internal model, part Lean-formalized, with the caveats in the README
- Decisions API: typed predicate, choice, and score answers at $0.10 per million input tokens
- Meta's Muse: a 0-day, silent message reads, and a rushed launch window make the agent permission model the story
THE BIG ONE
Mistral Large 4: 1T Parameters, 49B Active, Weights by the End of October
Mistral opened a public preview of Mistral Large 4 on Tuesday: a natively multimodal sparse model with 1 trillion total parameters and 49 billion active, trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European datacenters. The API is live at $1.36 per million input tokens and $4.18 per million output, and the weights are promised by the end of October after a red-teaming round with cybersecurity partners. The pitch is sovereignty: self-deployment, European law, more than 160 languages, and a EUR 3 billion Series D behind the compute.
The vendor benchmarks are broad. Mistral claims 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4.0, and a 49.8% combined Coding Agent Index it says leads DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind human evaluation by Surge AI, ML4 ranked second of five on coding quality at 3.74, ahead of Kimi K3 and GLM-5.3, behind Claude Opus 5. The cyber claims get the loudest framing: top five on the Artificial Analysis Cyber Index, 82% on a reproduce-and-patch test that Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero on because they refuse, 93% on Cybench, and 93.3% attack resistance on Lakera's B3.
The independent read is more measured. Simon Willison's first pass puts ML4 at 38 on Artificial Analysis, just behind DeepSeek 4.1 Flash, a 552B model. He notes the API exposes only two reasoning levels, none and high, and that high used fewer output tokens than none on his pelican test, then concludes Mistral is back to being maybe about six months behind the frontier. The Hacker News thread split the same way, between relief that a European lab is in the race and skepticism that six months of lag is a comeback.
The counterpoint shipped the same day: Anthropic expanded its Cyber Verification Program into three access tiers that give vetted security teams reduced blocking classifiers on Opus 5.5, Sonnet 5.5, and Mythos 5.1. Mistral's pitch is that open weights let defenders work without asking a vendor for permission; Anthropic's is that access should be verified and tiered. Both answer the same dual-use problem, and enterprise security buyers will pick.
Why it matters: the deadline is the product. A 1T open-weight model with real cyber capability would be the strongest self-hostable option built outside China, but until the weights and independent harness runs land, this is a preview with vendor benchmarks. The cyber comparison also flatters Mistral partly because its competitors refuse the task, which is a policy difference as much as a capability one. Our Mistral Large 4 breakdown has the full benchmark table, and the open-weights coding showdown is the field it has to beat.
RESEARCH
OpenAI Published 722 Math Manuscripts, and Verification Is the Fight
OpenAI published new results on open problems in mathematics on Tuesday from an unreleased internal model: 722 manuscripts organized into 372 families, drawn from roughly 4,000 problems posed during evaluation. The GitHub repository is the primary artifact, and its README is unusually blunt. The average result used about three hours of ChatGPT Pro thinking compute with that model, some manuscripts have Lean formalizations, and the README says plainly that not all do and that some unformalized results could have issues. Abridged reasoning traces are included for ten families.
One manuscript deserves its own look. "Integer multiplication below n log n" gives a deterministic algorithm that multiplies two n-bit integers in O(n (lg n)^(1 - k)) worst-case time with k = 2^-182, on one fixed finite-alphabet Turing machine, and claims to disprove the Schonhage-Strassen n log n optimality conjecture in that model. The exponent is both the joke and the point: a 2^-182 shave is asymptotically real and practically nothing, as the 99-point thread noticed within minutes. The harder question there was verification, since no machine-checked proof is attached.
That question is now organized. The Advisory Group for Mathematics and AI published recommendations on responsible release after collecting more than 600 community replies, and they are direct: "We do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." The group wants results released quickly but paired with lab-funded, community-led work to build human understanding, because mathematics assumes authors who understand and stand behind their arguments. The main thread spent 874 comments on the practical version, including what this means for graduate students whose thesis problems may already have a catalog entry.
Why it matters: the catalog is only as valuable as its verification layer, and by the authors' own admission that layer is partial. The transferable lesson for anyone shipping generated artifacts is the release design: reasoning traces, formalizations where they exist, and caveats stated up front. Our August analysis of OpenAI's ten Lean-formalized proofs shows what the fully checked version looks like.
PLATFORMS
OpenAI's Decisions API: $0.10 Per Million Input Tokens
OpenAI put its Decisions API into public beta, a POST /v1/decisions endpoint that answers typed questions about text or images instead of generating prose. predicate returns a probability that a condition is true, choice returns one of your supplied options, and score returns a probability-weighted score against ordered levels. gpt-6-luna is the only model for now, input costs $0.10 per million tokens, and there are no output or cache charges because no text is generated. The docs claim roughly 10x the speed of the Responses API and say GA is coming in weeks.
This is OpenAI following the decision-model wave more than starting it. TypeSafe's Jev kicked it off in September with calibrated, typed decisions that cannot hallucinate because they cannot generate strings. Since then: Jev pricing at $0.042 per million input tokens, Jeff's home-trained 0.8B and 2B variants answering in 22 to 28 milliseconds, Cloudflare's Clef models from $0.09 per million, and the Strands Decider 2B for local CPU or GPU inference. Simon Willison shipped llm-openai-decisions 0.1a0 the same day. The Hacker News thread was skeptical about the moat, arguing anyone can fine-tune a small classifier, and the fair counter is that calibration and reliability at scale are what the price buys.
Why it matters: classification, routing, and severity scoring are a large share of the calls in agent pipelines, and paying only for input tokens changes the unit economics of those steps. But list prices across Jev, Clef, and gpt-6-luna are not directly comparable, so the honest test is your own labeled set, with thresholds set from the cost of false positives and false negatives. Our DevDay 2026 recap covered the announcement; this is the beta.
SECURITY
Meta's Muse Is the Agent Permission Lesson Everyone Saw Coming
The past few weeks have been a rolling case study in what happens when an agent gets broad access to a personal machine. Ars Technica reported a 0-day in Meta's Muse that let locally run apps and terminal commands take complete control of the agent through a ClickFix-style attack, and Amazon began blocking Muse on its site. A YouTuber who put Muse in charge of Facebook Marketplace sales watched it sell items well below acceptable prices and expose his home address. Reports documented the agent reading private messages without approval, and a trick that handed over root access by impersonating an agent. Apple changed Full Disk Access on macOS on October 2 to require very explicit user action, citing exactly this class of agent risk.
The data story is just as pointed. Wired found that Muse builds detailed profiles of friends, family, and colleagues, and 404 Media reported that Meta rushed hotfixes for multiple vulnerabilities in the weeks before launch rather than delaying it, with an internal source describing half-baked protections and engineers fearing a breach. Techdirt's roundup pulled the threads together, and the Hacker News discussion (377 points, 270 comments) split between readers defending the sandbox design and readers who will not hand an agent their money or their messages.
Why it matters: this is the default failure mode of every agent that asks for broad permissions to be useful, not a Meta-specific bug. The mitigations are known and boring: scoped credentials per task, explicit confirmations for money and messaging, no silent cloud uploads, sandboxed runtimes, and OS gates like Apple's new requirement that treat agent access as a special case. Our agent security checklist, the Claude Code permissions guide, and our look at Muse Glimmer cover the workable patterns.
DATA TOOLING
Polars 2.0 Makes Streaming the Default and Spills to Disk
Polars shipped version 2.0 on Tuesday, and the biggest changes are defaults. Calling collect on a LazyFrame now uses the streaming engine, which the team says brings large memory and performance improvements, but streaming does not guarantee row order for join, group_by, and unpivot unless you pass maintain_order=True. Out-of-core execution is also on by default: sort, window functions, and many expressions can spill to disk starting at roughly 80% of RAM with a 64GB default disk budget, and joins and group-bys are next. SQL becomes a first-class citizen, backed by optimizer work; Polars says it beats DuckDB 1.5.6, DuckDB 2.0 alpha, and DataFusion 54.0.0 on TPC-H and TPC-DS 1, with the caveats disclosed, including a constant overhead at 192 threads.
Why it matters: the memory ceiling for local data work just moved, which matters more than any single feature because out-of-core by default means a laptop can finish queries that used to die. The row-order change is a two-line breaking change hiding in a major version, so audit pipelines that rely on join or group-by order before bumping the pin. The 439-point thread drew 99 comments, largely on where Polars now sits against Pandas. Our DuckDB internals explainer is context for why this benchmark race keeps getting closer.
DEVELOPER TOOLS
EmbeddingGemma 2 Puts Multimodal Embeddings on Device
Google DeepMind released EmbeddingGemma 2 on Tuesday under Apache 2.0, a 740-million-parameter model built on the Gemma 4 architecture that maps text, images, video, and audio into one embedding space. The design is modular: as little as 270 million parameters covers text-only work, with optional 170M vision and 300M audio encoders, and Matryoshka Representation Learning lets you truncate 768-dimension vectors down to 512, 256, or 128 to save up to 6x storage. Google says quantized, a Pixel 11 Pro needs about 191MB of active RAM for the text-only weights and about 567MB for full multimodal.
The retrieval numbers are the ones that matter: a 9.92-point improvement on MTEB Code, from 68.76 to 78.68, an 8K context window at 4x the first version, and leading sub-1B scores on MTEB and MAEB. The first EmbeddingGemma passed 20 million downloads, mostly for on-device search and privacy-first RAG. Simon Willison's note is the durable argument for the license: embedding pipelines store millions of vectors, so if a proprietary model is retired, re-embedding everything is your bill. Apache 2.0 lets hosted convenience and the escape hatch coexist.
Why it matters: local codebase indexing, semantic code search, and offline retrieval for coding agents get meaningfully better this week, and multimodal retrieval on a phone stops being a demo. Our Ternlight writeup covers the other end of the size spectrum, a 7MB embedder running entirely in the browser.
TOOLS WORTH A LOOK
- Strands Decider 2B (free, open source) - a 2B-parameter decision model that runs on a local CPU or GPU and answers typed questions in tens of milliseconds, the open alternative to the Decisions API. The thread hit 179 points.
- Claude Code v2.1.292 (free with a Claude plan) - a busy release for mods: a
prompt.autocompletehook, prompt caching inside$.model.complete, workflow agents in theagent.spawnhook, aneffortparameter on the Agent tool, and a retry-delay variable for overloaded 529s. - OpenTPU (free, open source) - an open-source AI accelerator project whose pitch is that AI did much of the development work. The thread reached 291 points and 341 comments.
WHAT ELSE IS HAPPENING
- JetBrains reports revenue growth and a net loss for 2025 (583 points, 544 comments): the IDE maker grew revenue while posting a net loss, and the thread debated AI workflows and what the loss means for the IDE business.
- State of Devs 2026 (172 points, 89 comments): 62% of surveyed developers report burnout at some point in their careers, 49% feel positive about AI against 42% negative, and about a quarter expect to change careers within five years.
- Gleam v1.19.0 (308 points, 132 comments): the compiler now emits Erlang abstract forms instead of Erlang source, cutting build times and making BEAM stacktrace line numbers accurate to the Gleam source.
- The Claude Code suggestion feature (203 points, 117 comments): an argument that suggested next messages are really training signal for the model, which reframes the feature as a data strategy rather than UX polish.
- What is Codemode (86 points, 41 comments): Armin Ronacher explains how Pi 1.0 replaced direct MCP tool loading with code that calls MCP, the argument our Pi 1.0 guide also walks through.
- Erdosproblems.com forum (102 points, 41 comments): the community tracker's thread on the AI onslaught split between readers who welcome machine-found proofs and researchers who watched their tracker get flooded.
FROM THE SITE
What We Published Tuesday
Our Mistral Large 4 open-weight breakdown covers the full benchmark table, the weights deadline, and the self-hosting math, and Claude Code Mods: What They Are and How to Write One explains the plugin system whose hooks shipped additions in v2.1.292. This morning we published Sign in with ChatGPT: Run Coding Agents on Your Plan, covering the OAuth flow, the per-app cap to set before you connect an agent, and why the shrinking Pro 200 allowance makes the grant matter.
Every link above goes to a primary source or our sourced coverage. Tomorrow's brief lands when the news does - subscribe to get it by email.
Get the next one in your inbox
The daily brief, delivered. Free, unsubscribe anytime.