Skip to main content
Watch: I Asked Claude to Build Me a Business

MULTIMODAL

8 items

6 posts, 2 tools

Blog
Gemini's Agentic Video Understanding Cuts Video Tokens by 88%: How It Works and What Breaks

Google shipped agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite: the model decides which frames, audio, and transcripts to inspect instead of swallowing video at a fixed frame rate. Verified numbers: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy on video benchmarks. Here is what changed and where the agentic loop still leaks.

Blog
Gemini Omni 1.1 Flash Goes GA: Scene Extension, Keyframe Control, and 4K in the Gemini API

Google made Gemini Omni 1.1 Flash generally available today: 10-second scene-extension context, first and last frame interpolation, 360p drafts at a third of the cost, and 4K upscaling. Verified pricing: about $0.10 per second of 720p video.

Blog
Gemini Robotics ER 2: Video-Feeding Embodied Reasoning Model Opens to All Developers

Google DeepMind's Gemini Robotics ER 2 is now publicly available via the Gemini API. It watches live video feeds to track task progress, orchestrates VLA models as tools, and coordinates multiple robots. The numbers: 57.4% progress classification, 91.3% moment finding at 0.96s offset.

Blog
MiniMax H3: An Omni-Modal Video Model With Native Audio, 2K Output, and Open Weights Coming

MiniMax launched H3, an omni-modal generation model that takes text, image, video, and audio input and outputs 2K video with native stereo sound at 0.80 CNY per second. Open weights are promised in the coming days.

Blog
FLUX 3: Black Forest Labs Ships a Unified Multimodal Foundation Model for Image, Video, Audio, and Robotics

Black Forest Labs released FLUX 3, a single multimodal model trained jointly on images, video, and audio that also drives robots on Audi production lines. Here is what it does, how it works, and how to try it.

Blog
Claude Vision API: Image Analysis At Production Scale

How to ship Claude's vision API in production. OCR, charts, UI audits, real cost numbers, TypeScript SDK code, and the gotchas that bite at 100k images a month.

Tool
LocalAI

Open-source OpenAI API replacement. Runs LLMs, vision, voice, image, and video models on any hardware - no GPU required. 35+ backends. Distributed mode for scaling.

Tool
Gemini

Google's frontier model family. Gemini 2.5 Pro has 1M token context and top-tier coding benchmarks. Gemini 3 Pro pushes reasoning further. Free tier via AI Studio.

AI Development Stack

Get Smarter About AI Dev

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.

One email per weekReal code, not theoryFree forever