Magnitude: A Self-Tuning Inference Engine Bids to Be Your Agent's Local Backend

TL;DR
Magnitude launched as an Apache-2.0 inference engine that compiles and tunes kernels on your own hardware, claims up to 2x faster decode than llama.cpp, and wires itself into OpenCode, Codex and Claude Code.
Last updated: October 1, 2026.
Magnitude, an Apache-2.0 inference engine for local agents, launched on Hacker News on September 30 with a specific claim: instead of shipping kernels precompiled for broad hardware classes, it compiles and tunes kernels on your device before a model runs, then serves open-weight models to the coding agent you already use. The launch post reached 178 points and 87 comments, and the GitHub repo showed about 6,000 stars and 401 forks at the time of writing. The same day, its post topped r/LocalLLaMA.
That lands in the middle of the site's runtime question: which engine should sit under a coding agent. Magnitude bets the generalists leave performance on the table.
What actually shipped#
Magnitude ships as a desktop app for macOS, Linux, and Windows with the magnitude CLI bundled, running on Apple Silicon, NVIDIA, AMD, or CPU only. The architecture choices in the launch post and the docs are the substance:
- On-device compilation and tuning. Kernel parameters are tuned on your chip before a model runs, about a minute per model, once.
- Dynamic memory allocation. Only weights are reserved up front; session memory grows with context and frees when agents stop. The launch claims 27 to 28 percent less per-agent memory.
- Hybrid paged attention. Concurrent sessions share prefix caches, borrowing an idea from SGLang-style serving while tuning for single-session speed.
- Kernels for popular open-weight families only. It specializes instead of generalizing, so the catalog is narrower than llama.cpp's.
The headline benchmark, run against llama.cpp with Qwen 3.6 35B A3B (4-bit) at 64k context, no speculative decoding:
| Hardware | Decode (tok/s) | Prefill (tok/s) | Per-agent memory |
|---|---|---|---|
| Mac M4 Pro 48GB, Metal | 30 -> 57 (+92%) | 466 -> 507 (+9%) | -28% |
| DGX Spark, CUDA | 49 -> 58 (+19%) | 2,033 -> 2,507 (+23%) | -27% |
The decode win is Metal-first; the prefill win is CUDA-first. These are vendor numbers, and the thread spends most of its energy on exactly that, which matters because trying the product is free.
The Monday move#
Magnitude does not ask you to replace your agent. It exposes an OpenAI-compatible endpoint and writes the config into your existing harness. These commands are copied verbatim from the CLI reference and the OpenCode integration page, not run on this machine, so treat them as the documented path rather than a test result:
# from docs.magnitude.dev/reference
magnitude hardware
magnitude catalog recommendations --preference balanced --limit 10
magnitude catalog pull <model-id>
magnitude connections add opencode --set-model <model-id>
# then, from your project folder
opencode --model magnitude/<MODEL_ID>
The connection writes ~/.config/opencode/opencode.json and points it at http://ADDRESS:10100/inference/v1; the docs show the raw config for machines Connections cannot reach. Supported harnesses are Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, with anything else on the OpenAI-compatible API. Test it with a model your hardware already handles and compare per-turn latency, not single-stream tokens per second; our KV caching guide explains why prefix reuse often matters more to how an agent feels than a decode benchmark.
Where it breaks, per the thread: GPUs detected as multiple devices, pessimistic fit estimates, hardware detection failures on Windows, no small models for low-VRAM machines, and prefill regressions against Apple-specific engines. It is a 0.2.x project and reads like one.
Who wins, who loses#
The direct losers, if the tuning advantage survives more hardware, are the precompiled generalists: Ollama and LM Studio compete on breadth, llama.cpp on reach. Magnitude argues those are the wrong optimizations for an agent session, and the Apple-specific engines are the real benchmark on Metal. The most squeezed pitch is a runtime that is only a local server.
The second-order effect is that magnitude connections add writes agent configs and can install a Magnitude skill into the harness. If engines own model routing inside harnesses, switching cost moves from the API shape, which is trivially swappable, to the integration layer: skills, model IDs, per-harness defaults. That is why the easiest engine for a harness to delegate to beats the engine with the best single benchmark. It also explains the business model that came up on HN: the engine is free, and the founders said they plan to sell hybrid local and cloud inference billed per token.
What people are actually saying#
The Hacker News thread is unusually technical for a Launch HN, and it does not hand the team a clean win.
- Agreement. The on-device autotuner is what people respect. One commenter running 11 Macs approved of the design: kernel-level search, tuning budget split by measured time share, winners cached per device and toolchain.
- The benchmark criticism. The sharpest question was why llama.cpp is the baseline. One commenter called beating llama.cpp "a low bar" and listed what actually matters for agents: speculative decoding, KV cache size at 100k-plus context, and degradation as context grows. Another asked for MLX numbers and noted how much llama.cpp varies with configuration.
- Practitioner reports are mixed. An M5 Max user measured llama.cpp b10853+ roughly 2x faster than Magnitude 0.2.1 in prefill and decode, and a two-GPU NVIDIA user saw llama.cpp 20 to 30 percent faster at decode. On the other side, an M5 Pro user beat oMLX on decode (82.8 vs 76.5 tok/s) while losing on prefill (709 vs 1,843), and a 64GB RTX 5080 user found the app capping recommendations at 9B while a 35B MoE ran at 90-plus tokens per second.
- The counter-case. The bugs are real: one user stuck at "Assessing Models" with no downloads, one Windows machine where hardware was not detected, and fit scores that look calibrated against one Mac configuration. The founders acknowledged a Metal 4 matmul gap on M5-generation chips, a mature answer that is also an admission that the fastest hardware is not yet the best supported.
On the r/LocalLLaMA thread, the framing was an LM Studio or Unsloth Desktop alternative that optimizes itself, and the top-of-day placement suggests appetite for another serious local runtime. The disagreement matches HN: the idea is right, the 2x claim needs independent verification on more hardware.
The read#
Magnitude is worth trying if you already run open models against a coding agent and accept early-user risk. The integration path is the differentiator: one command writes your harness config and the runtime fades into the background. It is not yet the default local backend, and the community numbers say so. But the direction is clear: as agents become the main consumer of local models, the runtime that wins will tune for the machine under it and get out of the way of the agent on top.
Continue Reading#
- Ollama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents - the broader decision this launch competes for
- LM Studio Bionic: Local AI Agents Without the Cloud - the closed-source incumbent in the same niche
- KV Caching for Transformer Inference - why prefix caches and memory layout decide agent latency
- vLLM vs TGI vs SGLang: Which Inference Server to Self-Host - the server-side engines Magnitude borrows from
- OpenCode Developer Guide - the harness used in the integration example
Sources#
- GitHub: magnitudedev/magnitude - launch README, benchmarks, star and fork counts, checked October 1, 2026
- Magnitude docs: Quick start, CLI reference, and OpenCode integration - commands copied verbatim
- Hacker News: Launch HN: Magnitude (YC S25) - launch text and commenter reports, fetched October 1, 2026
- r/LocalLLaMA: Open source inference engine thread - top-of-day community signal, October 1, 2026
- Google Trends: blocked for this run; no trend numbers claimed.
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on AI coding tools
Ollama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents
A fair, sourced comparison of the four runtimes developers reach for when they want a coding agent talking to a model on their own hardware instead of an API: Ollama's convenience, LM Studio's GUI, vLLM's throughput, and llama.cpp's control. What each is actually for, and which to pick.
10 min readLM Studio Bionic: A Local-First AI Agent for Open Models
LM Studio launches Bionic, a standalone agent harness for open models with local inference, voice input, and zero data retention cloud options.
6 min readKV Caching: A Practical Guide to Optimizing Transformer Inference
How KV caching speeds up LLM inference - the math, the code, the memory tradeoffs, and when it stops helping. Every dev running local models hits this wall.
11 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








