Skip to main content
Watch: I Asked Claude to Build Me a Business

OPEN SOURCE

134 items

107 posts, 27 tools

Blog
Cloudflare OS: The Open Source Agent Workspace That Treats Apps Like Files

On August 5 Cloudflare open sourced Cloudflare OS, the agent workspace it has run internally since May: capability-based Gatekeepers instead of ambient MCP access, apps as private per-user instances, and approvals that simulate outcomes so agents never stall. A concrete blueprint for the company-wide agent platform.

Blog
LFM2.5-2.6B: Liquid AI's On-Device Agent Model Runs at 220 Tokens/s in Under 2.5 GB

Liquid AI shipped LFM2.5-2.6B on August 4, 2026: a 2.6B open-weight model trained for agentic work inside real harnesses, decoding at 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen CPU. Here is how it was trained, what the benchmarks say, and how to run it.

Blog
Prime Agent: A Self-Improving Coding Harness Where Everything Is Python

Prime Intellect open-sourced Prime Agent on August 5, 2026. It gives the model exactly one tool - a persistent IPython kernel - and lets the harness rewrite its own prompts, skills, memory, and sub-agents mid-run. Here is how it works, what the benchmarks actually show, a full provider and model guide, and an honest comparison to Claude Code, Codex, OpenCode, OpenClaw, Hermes, and Pi.

Blog
Mistral Shieldstral: A 3B Open-Weight Policy-Adaptive Moderation Model That Beats Models 7x Its Size

Shieldstral is a 3B-parameter Apache 2.0 multimodal safety classifier that takes your moderation policy as a plain-language question at inference time, scores content 0-1 in a single forward pass, and runs on one 16GB GPU. It beats 12B-20B guard models on text safety and sets state of the art on multimodal benchmarks.

Blog
Vibe Maintenance: A Practical Workflow for AI-Generated Pull Requests

Steve Yegge's response to AI-generated pull requests suggests a better maintainer workflow: automate triage, repair good ideas, and keep human taste at the boundary.

Blog
The RipGrep Musl Segfault That Led to a One-Line Linux Kernel Patch

A ripgrep musl binary crashing during very-large searches turned out to be a suspected Linux 7.0 kernel race - a thread's own store vanishing mid-function. The reporter's instrumentation pinned it, and a kernel-hardening maintainer posted a one-line fix candidate for testing.

Blog
Your AI Session Is No Longer Yours: How Providers Seal Reasoning, Search, and Subagent State

An analysis from the Earendil team behind the Pi harness documents how OpenAI, Anthropic, and Google now return provider-sealed state instead of portable transcripts - encrypted reasoning blobs, opaque compaction, hidden subagent messages. The five tests and seven rules for session portability, and why session lock-in matters more than model lock-in.

Blog
Coding Agents Almost Never Read Open Source Contribution Rules: RepoComplianceBench Study

A new 106-issue benchmark across 49 repositories finds frontier coding agents rarely retrieve AI contribution rules on their own - and never refuse to contribute in AI-banned repositories, no matter the prompt. Disclosure and verification can be fixed; bans cannot.

Blog
DeepSeek V4 Flash 0731: The Official Release, Benchmarks, and How to Run It in OpenCode

DeepSeek shipped the official V4 Flash release on July 31, 2026. The re-post-trained 0731 build beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million tokens. Here is what changed and how to run it through OpenCode today.

Blog
GitHub Case-Folds 480TB of Code at >45 GiB/s: The Branchless Casefold Crate

GitHub open-sourced casefold, a Rust crate that folds the case of every byte Blackbird indexes at memory bandwidth. The counterintuitive trick: delete the early-exit, kill the branches, and fold Unicode as byte arithmetic.

Blog
Inkling-Small: Thinking Machines Ships a 12B-Active Open Model That Beats Its Big Sibling on Agent Work

Inkling-Small is a 276B-parameter MoE with 12B active per token, Apache 2.0, and open weights. It beats the 975B Inkling on SWEBench Verified (80.2), HLE (31.6), and tool use at a quarter of the size and a third of the output price.

Blog
MiniMax H3: An Omni-Modal Video Model With Native Audio, 2K Output, and Open Weights Coming

MiniMax launched H3, an omni-modal generation model that takes text, image, video, and audio input and outputs 2K video with native stereo sound at 0.80 CNY per second. Open weights are promised in the coming days.

Blog
Buzz by Block: The Open-Source Workspace Where Humans and AI Agents Build Together

A companion guide to the Buzz video: Block's open-source Nostr relay workspace where humans and AI agents share the same rooms, with agent-first CLI, git integration, and workflows. Here is what it does and where it fits in the agentic dev stack.

Blog
Superlogical: Mitchell Hashimoto's Multiplexer Company

Mitchell Hashimoto launched Superlogical, a company building a terminal multiplexer for all work: local dev, remote access, and agents. HN reaction, analysis.

Blog
Buzz: Block's Agent-Native Messaging Layer on Nostr

Block open-sourced Buzz, a team workspace where agents are cryptographic identities instead of bot tokens. Every message is a signed Nostr event, the relay is yours to run, and the CLI is JSON in, JSON out.

Blog
TurboFieldfare: Running Gemma 4 26B in 2 GB of RAM on Any M-Series Mac

TurboFieldfare is a custom Swift and Metal inference engine that runs Google's 26B-parameter Gemma 4 MoE model in roughly 2 GB of RAM on any Apple Silicon Mac, including 8 GB base models.

Blog
PGSimCity: A 3D Interactive City That Visualizes How PostgreSQL Works

PGSimCity is an explorable 3D city that models PostgreSQL internals - shared buffers, WAL, autovacuum, checkpoints, and replication. Built with three.js and TypeScript, it hit #1 on HN with 682 points.

Blog
Debian Debates LLM Usage: Four Proposals, One Fork in the Road

Debian is voting on four proposals to regulate LLM-generated contributions - from an outright ban to full acceptance. The HN discussion reveals the fault lines in open source's biggest AI policy debate yet.

Blog
Ruff v0.16.0: 413 Default Rules, Markdown Formatting, and What Zero-Config Linting Means for Python

Ruff v0.16.0 ships 413 default rules (up from 59), Markdown code-block formatting, and a new ruff: ignore system. Here is what changed, what HN is saying, and why zero-config linting matters more with AI coding agents.

Blog
How My Images Are Dithered - Simulating Halftone Printing with ImageMagick

A technical deep dive into AM halftoning with ImageMagick hit the HN front page at 195 points. We break down the technique, the HN debate on dithering vs halftoning, and why this matters for developers.

PreviousPage 2 of 7Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever