4 items
4 posts
DeepSeek shipped experimental vision for V4 Flash as deepseek-v4-flash-vision-exp. JPEG, PNG, GIF, and WebP; three input methods; 384 tokens per image. Here is the API contract and how to run it in OpenCode today.
Liquid AI released LFM2.5-VL-3B on August 12, 2026: a 3.1B open-weights vision-language model that averages 80.7 on ScreenSpot-v2, doubles ToolSandbox to 59.5, and decodes at 228 tokens/s on an M5 Max in about 3 GB of memory. Here is what shipped, the benchmark caveats, and how to run it.
How to ship Claude's vision API in production. OCR, charts, UI audits, real cost numbers, TypeScript SDK code, and the gotchas that bite at 100k images a month.
NVIDIA's Nemotron Nano 2 VL delivers vision-language capabilities at a fraction of the computational cost. This 12-billion-parameter open-source model processes videos, analyzes documents, and reas...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.