Skip to main content
Watch: I Asked Claude to Build Me a Business

LOCAL AI

21 items

16 posts, 5 tools

Blog
Hermes Agent Guide (2026): What It Is, Local Setup With Ollama, Skills, and Hermes vs OpenClaw

Hermes Agent is Nous Research's MIT-licensed, self-improving AI agent: it writes and refines its own skills, keeps memory across sessions, and talks to you from the terminal or Telegram, Discord, and Slack. Here is how to run it fully locally with Ollama, how its skills system works, and how it compares with OpenClaw.

Blog
GitHub Copilot for JetBrains Gains Persistent Memory and Ollama BYOK

The August 11 JetBrains plugin release adds Copilot memory across chat sessions, Ollama as a bring-your-own-key provider, and enterprise managed settings for MCP access and permission bypass. Here is what each feature actually does and why the IDE just became the control point for agent tooling.

Blog
Cactus Needle 2: The 14MB Agentic LLM That Runs on a Raspberry Pi 5

Cactus open-sourced Needle 2, a 45M-parameter agentic LLM in a single 14MB binary that runs a full tool-calling session in 28MB of RAM. 500 tok/s on a Raspberry Pi 5, ESP32-S3 class parts, Apache 2.0. Here is what the benchmarks actually show.

Blog
Muse Glimmer 30B: Meta's Open-Weight Local Agent Model, Benchmarks, and Hardware Reality

Meta open-sourced Muse Glimmer, a 30B Apache 2.0 multimodal agent model that runs in a 24GB envelope at up to 233 tok/s. MCP Atlas 75.5, SWE-Bench Verified 76.0, 131K context. Here is what the numbers actually say.

Blog
Running Gemma 4 26B at 5 Tokens/Sec on a 13-Year-Old Xeon With No GPU

A developer got Google's Gemma 4 26B running on 2013 Xeon hardware for under $300. The fix for a silent MoE bug is now upstream - here's what it means for local inference.

Blog
Colibri: Running GLM 5.2 on a 32GB Laptop with Disk Streaming and Expert Offloading

A solo developer built a 1,300-line C inference engine that runs the 744B GLM 5.2 model on consumer hardware by streaming routed experts from disk. Here's how it works.

Blog
Kokoro: Local, CPU-Friendly TTS That Actually Sounds Good

An 82M parameter text-to-speech model that runs on CPU and produces high-quality speech across multiple languages - no cloud APIs or GPU required.

Blog
Program-as-Weights Turns Prompts Into Local Fuzzy Functions

The Program-as-Weights paper is a useful signal for developers: some LLM calls may move from per-request API prompts into compact local artifacts that behave like reusable fuzzy functions.

Blog
GLM-5.2 Local Deployment: Running Z.ai's 744B Model on Consumer Hardware

Unsloth's dynamic quantization makes GLM-5.2 runnable on a 256GB Mac or a 24GB GPU with CPU offloading. Here is the hardware math, the quantization tradeoffs, and what the HN community learned from actually running it.

Blog
DiffusionGemma: Google Bets Diffusion Can Make Text Generation 4x Faster

Google released DiffusionGemma today, a 26B MoE open model that generates entire 256-token blocks in parallel instead of one token at a time. Here is what that means for latency, local inference, and the post-autoregressive landscape.

Blog
What Is Cline? The Open-Source AI Coding Tool That Runs in VS Code

Cline is a free, open-source VS Code extension that brings autonomous AI coding to your editor, and a common Cursor and Copilot alternative for developers who want to bring their own model. It works with local models or cloud APIs, handles multi-file changes, and runs terminal commands without proprietary lock-in.

Blog
Client-Side Tool Calling Is the Privacy Pattern AI Apps Need

A Show HN PDF form demo points at a bigger architecture shift: keep sensitive documents local, expose narrow browser tools to the model, and make AI assistance inspectable.

Tool
Ollama

The easiest way to run LLMs locally. One command to pull and run any model. OpenAI-compatible API. 52M+ monthly downloads. Supports GGUF, Safetensors, and custom Modelfiles.

Tool
LM Studio

Desktop app for discovering, downloading, and running local LLMs. Clean chat UI, OpenAI-compatible API server, and automatic GPU detection. MLX engine optimized for Apple Silicon.

Tool
Jan

Open-source ChatGPT alternative that runs 100% offline. Desktop app with local models, cloud API connections, custom assistants, and MCP integration. AGPLv3 licensed.

Tool
GPT4All

Private local AI chatbot by Nomic. 250K+ monthly users, 65K GitHub stars. LocalDocs feature lets you chat with your own files. Runs on Windows, macOS, and Linux.

Tool
LocalAI

Open-source OpenAI API replacement. Runs LLMs, vision, voice, image, and video models on any hardware - no GPU required. 35+ backends. Distributed mode for scaling.

Blog
DeepSeek R1 and V3: The Developer's Guide to Open-Source AI

DeepSeek's R1 and V3 models deliver frontier-level performance under an MIT license. Here's how to use them through the API, run them locally with Ollama, and decide when they beat closed-source alternatives.

Blog
Llama 4: The Complete Developer's Guide to Meta's Open Source Models

Meta's Llama 4 family brings mixture-of-experts to open source with Scout and Maverick. Here's how to run them locally, access them through APIs, and decide when they beat the competition.

Blog
NVIDIA Nemotron Nano 9B V2: Local AI That Punches Up

NVIDIA's Nemotron Nano 9B V2 delivers something rare: a small language model that doesn't trade capability for speed. This 9B parameter model outperforms Qwen 3B across instruction following, math,...

Page 1 of 2Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever