Skip to main content
Watch: I Asked Claude to Build Me a Business

AI MODELS

112 items

111 posts, 1 guide

Blog
Grok 4.6: xAI's Agent-Focused Update Matches GPT-5.6 Sol at the Same $2/$6 Price

xAI shipped Grok 4.6 on August 12, 2026: it matches GPT-5.6 Sol on the AA Intelligence Index (61), beats it on CursorBench 3.2, and keeps Grok 4.5's $2/$6 per million token pricing. Available in Cursor and Grok Build today, and in OpenCode as opencode/grok-4.6.

Blog
LFM2.5-VL-3B: Liquid AI's 3B Vision Model Reads Screens, Grounds Objects, and Calls Tools on a Laptop

Liquid AI released LFM2.5-VL-3B on August 12, 2026: a 3.1B open-weights vision-language model that averages 80.7 on ScreenSpot-v2, doubles ToolSandbox to 59.5, and decodes at 228 tokens/s on an M5 Max in about 3 GB of memory. Here is what shipped, the benchmark caveats, and how to run it.

Blog
Cactus Needle 2: The 14MB Agentic LLM That Runs on a Raspberry Pi 5

Cactus open-sourced Needle 2, a 45M-parameter agentic LLM in a single 14MB binary that runs a full tool-calling session in 28MB of RAM. 500 tok/s on a Raspberry Pi 5, ESP32-S3 class parts, Apache 2.0. Here is what the benchmarks actually show.

Blog
Encrypted Chain-of-Thought Is Not Private: New Paper Decodes Reasoning Traces From Anthropic, OpenAI, and Google APIs

A new arXiv paper shows the encrypted reasoning blocks that Anthropic, OpenAI, and Google return to API clients can be replayed into weaker models from the same provider and transcribed verbatim. The authors decoded 315,320 blocks from public repositories and recovered 367 PII artifacts and 182 credentials.

Blog
Muse Glimmer 30B: Meta's Open-Weight Local Agent Model, Benchmarks, and Hardware Reality

Meta open-sourced Muse Glimmer, a 30B Apache 2.0 multimodal agent model that runs in a 24GB envelope at up to 233 tok/s. MCP Atlas 75.5, SWE-Bench Verified 76.0, 131K context. Here is what the numbers actually say.

Blog
Grok Imagine Image 2.0 Ships: xAI's Typography-Aware Image Model Is Already on Vercel's AI Gateway

xAI released Grok Imagine Image 2.0 on August 7 as the new Quality Mode on grok.com and mobile, ranked second worldwide on both text-to-image and image-editing leaderboards. A 2.0 preview build is already callable through Vercel's AI Gateway with the AI SDK, before xAI's own API access goes live.

Blog
DeepMind Open-Sources WeatherNext Cyclones After a Nature-Verified Breakthrough

WeatherNext Cyclones adds a full day of lead time to tropical cyclone forecasts - roughly a decade of meteorological progress - and now the weights, code, and data feeds are public. What the paper actually shows and how to run it.

Blog
Kimi K3 Is GA in GitHub Copilot: Pricing, Rollout, and What It Means for Model Choice

GitHub made Kimi K3 generally available in Copilot on August 6 at $3/$15 per million tokens, hosted on Fireworks AI. It is off by default for Business and Enterprise, the rollout was paused mid-day by a GitHub Actions incident, and it changes the price/quality calculus in the model picker.

Blog
Meta Ships Muse Code and Muse Spark 1.2: A Terminal Agent With a 12x Cheaper Contributor Tier

Meta released Muse Code, a terminal coding agent, and Muse Spark 1.2 on August 5, 2026. The model co-trains with the harness, logs every call to a replay-safe event log, and offers a $0.10/$0.20 contributor tier if Meta may train on your data.

Blog
OpenAI Retunes GPT-5.6 Sol in ChatGPT and Makes Luna the Free Tier Default

GPT-5.6 Sol gets a chat-focused retune with 68% fewer factual errors in OpenAI's internal eval, a new effort slider, and GPT-5.6 Luna becomes the default model for Free and Go users with unlimited text chats. What the API did not change and why the split matters.

Blog
LFM2.5-2.6B: Liquid AI's On-Device Agent Model Runs at 220 Tokens/s in Under 2.5 GB

Liquid AI shipped LFM2.5-2.6B on August 4, 2026: a 2.6B open-weight model trained for agentic work inside real harnesses, decoding at 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen CPU. Here is how it was trained, what the benchmarks say, and how to run it.

Blog
Qwen 3.8 Max Ships: 2.4T MoE, 1M Context, $2/$6 per MTok, Open Weights Next Week

Alibaba released Qwen 3.8 Max on August 3, 2026 - a 2.4T-parameter MoE with 95B active per token, a 1M context window, and $2/$6 per million tokens on QwenCloud. It leads PaperBench at 93.0, and the weights open next week.

Blog
Gemini 2.5 Pro and Gemini 3 Flash Deprecated in GitHub Copilot: What to Switch To

GitHub deprecated Gemini 2.5 Pro and Gemini 3 Flash in every Copilot surface on July 31, 2026. The suggested replacements are Gemini 3.1 Pro and Gemini 3.6 Flash. Here is what changed, what it costs, and how to migrate cleanly.

Blog
Your AI Session Is No Longer Yours: How Providers Seal Reasoning, Search, and Subagent State

An analysis from the Earendil team behind the Pi harness documents how OpenAI, Anthropic, and Google now return provider-sealed state instead of portable transcripts - encrypted reasoning blobs, opaque compaction, hidden subagent messages. The five tests and seven rules for session portability, and why session lock-in matters more than model lock-in.

Blog
Budget AI Coding Models Compared September 2026: GPT-6 Luna vs V4.1-Flash vs Gemini 3.5 Flash vs Haiku 4.5

The cheap coding tier repriced again: GPT-6 Luna opened at $0.10/$0.50 and took the floor from DeepSeek, whose V4.1-Flash cut rates to $0.15/$0.60 off-peak. Gemini 3.5 Flash and Claude Haiku 4.5 hold the hosted middle. Prices verified September 26, 2026.

Blog
Claude Mythos Preview Explained: Anthropic's Gated Frontier Model and Project Glasswing

Claude Mythos Preview is the model that found thousands of zero-days, and you could not buy it. Here is what it is, who got access through Project Glasswing, what it actually found, and where the model line went after it retired.

Blog
DeepSeek V4 Flash 0731: The Budget Tier Just Overtook Pro Preview on Agent Benchmarks

DeepSeek re-post-trained V4 Flash into an agent workhorse: Terminal Bench 82.7, DeepSWE 54.4, native Responses API, and first-party Codex support - all at $0.14/$0.28 per million tokens. What changed, what the numbers actually mean, and how to wire it up today.

Blog
DeepSeek V4 Flash 0731: The Official Release, Benchmarks, and How to Run It in OpenCode

DeepSeek shipped the official V4 Flash release on July 31, 2026. The re-post-trained 0731 build beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million tokens. Here is what changed and how to run it through OpenCode today.

Blog
Inkling-Small: Thinking Machines Ships a 12B-Active Open Model That Beats Its Big Sibling on Agent Work

Inkling-Small is a 276B-parameter MoE with 12B active per token, Apache 2.0, and open weights. It beats the 975B Inkling on SWEBench Verified (80.2), HLE (31.6), and tool use at a quarter of the size and a third of the output price.

Blog
MiniMax H3: An Omni-Modal Video Model With Native Audio, 2K Output, and Open Weights Coming

MiniMax launched H3, an omni-modal generation model that takes text, image, video, and audio input and outputs 2K video with native stereo sound at 0.80 CNY per second. Open weights are promised in the coming days.

PreviousPage 2 of 6Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever