Skip to main content
Watch: I Asked Claude to Build Me a Business

Research Briefs

HF Research Papers

Summaries of trending research from Hugging Face Daily Papers, written for builders. 255 papers from July 2026.

Building Agents

58 papers

LLMRetrieval

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

AskChem changes the retrieval unit from the paper to the provenance-carrying claim - 2.4M claims, each grounded by DOI and verbatim quote, served to agents over MCP. Grounding a GPT-5.5 reader in AskChem gives 100% resolvable DOIs vs 88.3% without retrieval; ungrounded, the same reader fabricated 6 of 14 DOIs on one question.

Read summary
AgentGUI

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

A foundation GUI agent spanning mobile, computer-use, web and DeepSearch, trained with online RL on 100+ turn trajectories across 10,000+ concurrent environments; SOTA on mobile-use benchmarks and competitive with frontier models on computer and browser use.

Read summary
AgentMemory

Metis: Memory Foundation Model

Metis introduces the first memory foundation model: a persistent, dynamically evolving memory state inside the backbone, updated gradient-free in a single forward pass at inference, with learned native memory procedures acquired through mid-training.

Read summary
AgentTraining

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

** Converts sparse outcome rewards in long-horizon agentic RL into turn-level credit signals with a critic-free, recursive Bayesian belief update in log-odds space - no extra critic, no extra rollouts - beating GRPO and strong self-distillation baselines on ALFWorld, WebShop, and Search-QA.

Read summary
AgentEvaluation

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

** Automates the measurement of harness optimization - iteratively editing the prompts, tools, control flow, memory, and orchestration code around a target agent under a fixed evaluation budget - and finds optimizer models separate more than the coding harnesses they act through.

Read summary
LLMAgent

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Microsoft Research compiles specs into stateful synthetic apps graded against their own databases and co-evolves environments with the model: a 9B agent reaches 67.1% across 14 evaluation splits (from 36.5%), within 14 points of its frontier teacher, while shallow environments actively hurt live accuracy (80.0 to 75.0).

Read summary
AgentCode

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Retrieval-based speculative decoding drafts tokens by copying continuations from existing text, which fits coding agents that repeatedly reproduce code, logs, and earlier attempts. AgSpec supplies the corpora and draft-length policies existing retrieval engines lack, raising generation throughput up to 4.37x at batch size 1 and 4.76x at batch size 16.

Read summary
AgentEvaluation

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

Eight parallel 10-agent worlds ran continuously for 16 days (850,000+ LLM calls, ~50B tokens) with persistent memory, shared tools, and self-governed institutions; after operational state accumulated, three stress events (indirect prompt injection, misinformation, private-memory exposure) defeated every world, with agents acting on adversarial content up to 46 hours later.

Read summary
AgentCode

AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories

Execution guardrails mined from real coding-agent failure traces instead of hand-written rules: learned constraints cut the abnormal execution rate from 69.0% to 26.7% and lift task completion from 21.7% to 35.0% on a disjoint evaluation set (Claude Code + Claude Haiku 4.5).

Read summary

+49 more

Faster Inference

55 papers

AgentsMemory

LycheeMemory V2: Semantic Segment-Level Consolidation Slashes Long-Term Memory Construction Cost

Batching memory consolidation to semantic segment boundaries instead of every turn reaches 89.22% on LoCoMo and 92.20% on LongMemEval-S while cutting construction tokens 86.0% and 75.9% versus A-Mem, with no increase in query-time token usage.

Read summary
World ModelsMemory

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Interactive world model that externalizes persistent world state into a camera-indexed state bank with view-relevant retrieval, keeps the denoiser context bounded as sessions grow, and distills drift resistance from a linear-scaling long-horizon teacher into a three-step student - state-of-the-art on WBench, 2.11s per 1.5s chunk on a single H200.

Read summary
LLMDistillation

DOPD: Dual On-policy Distillation

DOPD introduces advantage-aware dual distillation that dynamically routes token-level supervision between teacher and student policies based on the privilege advantage gap, solving the 'privilege illusion' failure mode in on-policy distillation.

Read summary
DiffusionSpeculative Decoding

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

BlockPilot reveals that fixed block-size strategies in diffusion-based speculative decoding are fundamentally suboptimal, and proposes a lightweight instance-adaptive policy that predicts the optimal block size from the prefilling representation, achieving 4.20× speedup on Qwen3-4B.

Read summary
TrainingEfficiency

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Factoring the training loop instead of the generation procedure - explore K candidate matches between model generations and data, train on the best - adds exploration as a third pretraining axis: 4.1x FLOP efficiency, 6.2x sample efficiency, and end-to-end generation with 16-256x fewer inference steps than diffusion.

Read summary
Image GenerationAutoregressive

GEAR: Guided End-to-End AutoRegression for Image Synthesis

GEAR trains a VQ tokenizer and autoregressive generator jointly end-to-end via a dual hard/soft codebook readout, achieving up to 10x faster gFID convergence on ImageNet by shifting representation alignment from the tokenizer to the AR model.

Read summary
LLMMemory

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Scaling a pretrained parametric memory module to 6.9B parameters on 300B tokens: a 410M base with that memory beats Pythia-12B on 17 benchmarks with 39% fewer total parameters, and 1.7B domain memories add more than 9 points to Qwen3 bases from 0.6B to 14B.

Read summary
AgentCode

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Retrieval-based speculative decoding drafts tokens by copying continuations from existing text, which fits coding agents that repeatedly reproduce code, logs, and earlier attempts. AgSpec supplies the corpora and draft-length policies existing retrieval engines lack, raising generation throughput up to 4.37x at batch size 1 and 4.76x at batch size 16.

Read summary
TrainingGeneration

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

** Data augmentations inherited from natural image tasks can disrupt the fine-grained vascular topology and textures critical for identity discrimination in vein recognition.

Read summary

+46 more

Training & Fine-Tuning

101 papers

LLMReasoning

Orca: The World is in Your Mind

Orca introduces a general world foundation model that learns a unified latent space from multimodal signals via Next-State-Prediction, then freezes its backbone and attaches lightweight readout decoders for text, image, and action - outperforming specialized baselines across all three modalities.

Read summary
AgentMemory

Metis: Memory Foundation Model

Metis introduces the first memory foundation model: a persistent, dynamically evolving memory state inside the backbone, updated gradient-free in a single forward pass at inference, with learned native memory procedures acquired through mid-training.

Read summary
AgentsEvaluation

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Taste-Bench measures the quality of an agent's long-horizon decisions by mining decision forks from real agent trajectories without human annotation. The best frontier model answers only 59.7% of taste questions correctly, forks whose deciding evidence appears late are harder for every model, and a larger reasoning budget does not help - but taste is trainable via outcome distillation, and the trained student improves end-to-end success on held-out SWE-bench Pro tasks.

Read summary
RLTraining

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

GRPO broadcasts a response-level advantage to every token, leaving within-response credit unallocated. CoRT rescores the same sampled response under the rubric-conditioned prompt and a matched criteria-free prompt, and uses the tokenwise log-likelihood contrast as a proxy for rubric dependence, redistributing the signed advantage across tokens - no auxiliary scorer, +4.4 points on average over response-level GRPO.

Read summary
AgentTraining

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

** Converts sparse outcome rewards in long-horizon agentic RL into turn-level credit signals with a critic-free, recursive Bayesian belief update in log-odds space - no extra critic, no extra rollouts - beating GRPO and strong self-distillation baselines on ALFWorld, WebShop, and Search-QA.

Read summary
Reinforcement LearningTraining

Sharpening Tax in Post-Training

RL post-training sharpens a model toward what it already solves: pass@1 and sampling consistency improve, but solution coverage (pass@K) shrinks, so harness-equipped base models can match or surpass post-trained counterparts given a large test-time budget. Across 14 checkpoint pairs from four families and three agentic benchmarks (42 cases), the paper's Sharpening Tax is prevalent and estimable from a few rollouts, and a difficulty-adaptive sampler shrinks it during training.

Read summary
TrainingEfficiency

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Factoring the training loop instead of the generation procedure - explore K candidate matches between model generations and data, train on the best - adds exploration as a third pretraining axis: 4.1x FLOP efficiency, 6.2x sample efficiency, and end-to-end generation with 16-256x fewer inference steps than diffusion.

Read summary
HarnessTraining

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Harness optimization has been fixing the scenario order that generates feedback while optimizing how the harness updates. ActiveSaddler treats the curriculum as a non-stationary bandit over reusable failure patterns that co-evolves with the harness, improving test Pass@1 by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 over the same optimizer with a scenario order fixed before optimization.

Read summary
LLMMemory

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Scaling a pretrained parametric memory module to 6.9B parameters on 300B tokens: a 410M base with that memory beats Pythia-12B on 17 benchmarks with 39% fewer total parameters, and 1.7B domain memories add more than 9 points to Qwen3 bases from 0.6B to 14B.

Read summary

+92 more

Robotics & Embodied AI

33 papers

World ModelsInference Efficiency

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

An end-to-end JEPA that turns action-labeled trajectories into a search-free policy: the conditional mean of a learned distributional action law directly selects actions, hitting 85.78-100% success on LeWM with zero search and 2.9-5.5ms inference, and cutting CEM sampling 23.44x.

Read summary
VLARobotics

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

Current VLA models retain shallow perceptual knowledge (color, shape) after robotics fine-tuning but catastrophically drop performance on richer semantic categories (emotion, counting, temporal, normative, cultural knowledge) - with answer-relevant information still present in intermediate layers yet failing to reach action output.

Read summary
RoboticsTraining

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

** Existing World Action Models fail at mobile manipulation due to three structural misalignments - coarse video prediction doesn't match fine-grained control, entangled navigation/manipulation action spaces cause gradient interference, and training on ground-truth futures doesn't generalize to the model's own noisy rollouts at inference time.

Read summary
MultimodalAgent

ASPIRE: Agentic Skill Programming through Iterative Robot Exploration

** Traditional robot programming is difficult due to the need to orchestrate multimodal perception, manage contact dynamics, and handle diverse configurations and failures, with no existing system that autonomously writes, refines, and transfers reusable control programs across tasks and embodiments.

Read summary
LLMAgent

AutoMem: Automated Learning of Memory as a Cognitive Skill

** LLMs lack a learned memory management strategy - they don't know what to encode, when to retrieve, or how to organize knowledge over long-horizon tasks.

Read summary
LLMRobotics

CausalMix: Data Mixture as Causal Inference for Language Model Training

** Existing data mixture optimization methods assume static data distributions and require costly retraining from scratch when the data pool shifts, preventing scalable transfer across data pools and model sizes.

Read summary
MultimodalAgent

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

** Existing text-rich image data pipelines follow a static crawl-filter-freeze paradigm that discards rejected samples, wasting failure signals (OCR errors, semantic mismatches) that could inform later construction rounds.

Read summary
MultimodalAgent

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

** VLA models fine-tuned from powerful VLMs on robotics data may catastrophically forget commonsense and world knowledge, but existing benchmarks conflate knowledge gaps with low-level control failures.

Read summary
MultimodalRobotics

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

** Vision-Language-Action (VLA) models fail under environmental shifts (camera pose, embodiment changes) and existing adaptation methods require costly multi-demonstration data per task.

Read summary

+24 more

Content Generation

23 papers

TrainingGeneration

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

** Data augmentations inherited from natural image tasks can disrupt the fine-grained vascular topology and textures critical for identity discrimination in vein recognition.

Read summary
LLMCode

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

** Existing audio-video generation models use separate per-modality tokenizers, which creates a representation gap, causes semantic misalignment, and requires expensive dual-branch architectures.

Read summary
LLMDiffusion

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

** Existing diffusion-based speculative decoding methods use a fixed block size for all inputs, which is suboptimal since the optimal block size varies across samples.

Read summary
MultimodalDiffusion

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

** Existing brain encoding and decoding models treat these as separate tasks using unimodal alignment, ignoring the brain's intrinsic multimodal integration nature.

Read summary
MultimodalGeneration

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

** Blind image deblurring methods struggle with real-world spatially varying degradations and lack the semantic awareness needed to distinguish valid textures from artifacts.

Read summary
LLMMultimodal

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

** Can a discrete diffusion language model match or exceed autoregressive models on medical report generation while offering capabilities autoregressive models lack?

Read summary
DiffusionCode

GEAR: Guided End-to-End AutoRegression for Image Synthesis

** Standard two-stage visual generative models (tokenizer → frozen generator) decouple training, leaving the tokenizer unaware of what the generator finds easy or hard to model.

Read summary
LLMMultimodal

InstanceControl: Controllable Complex Image Generation without Instance Labeling

** Existing controllable image generation methods like ControlNet struggle with attribute confusion in multi-instance scenes, while approaches that fix this require labor-intensive manual instance labeling.

Read summary
LLMReasoning

Little Brains, Big Feats: Exploring Compact Language Models

** Can small language models (SLMs) serve as viable, GPU-free replacements for large language models in the generation stage of Retrieval-Augmented Generation (RAG) systems?

Read summary

+14 more

Evaluation & Benchmarks

36 papers

AgentsBenchmark

FrontierChallenge: 75.5% of Non-Passing Agent Trajectories Claim Completion - Partial Scores and Confidence Both Lie

On 97 end-to-end scientific workflows across six domains, the best model-and-scaffold configuration completes only 20.6% of tasks, partial-progress scores of 87.6-94.9 coexist with 0-4% completion rates, and 75.5% of non-passing Claude Code trajectories still end with language claiming the job is done.

Read summary
AgentsSkills

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

Code2Skill transforms source code into grounded agent skills without any agent interaction experience: 1,006,822 verified skill records mined and verified from 19,769 GitHub repositories, lifting model performance by 11.7% on average across 72 protocol-matched evaluations and beating trajectory-derived skill banks on all seven shared benchmarks - and the pipeline stays effective on AI-generated code (93.50% pass rate vs 93.00% for human-written).

Read summary
AgentEvaluation

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

** Automates the measurement of harness optimization - iteratively editing the prompts, tools, control flow, memory, and orchestration code around a target agent under a fixed evaluation budget - and finds optimizer models separate more than the coding harnesses they act through.

Read summary
EvaluationAgents

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Ten AI systems evaluated on search-and-discovery benchmarks whose rules are executable and conflict with familiar knowledge, so exploration can be verified exactly and recall alone cannot answer - the strongest systems acquire the unfamiliar rules, but continued exploration can stall or reverse earlier gains.

Read summary
EvaluationAgents

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

A controlled study of recursive AI peer review: training successor reviewers on data mixed with model-generated reviews compresses rating distributions and shrinks semantic diversity of judgments - 'scientific-judgment collapse'. The mitigation (TrustReviewer) combines single-stage training on a curated corpus with paired activation steering at test time, needing no additional expert annotation.

Read summary
AgentsEvaluation

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

Benchmark accuracy conflates two separate questions - whether a response reached an evaluable state and whether its answer was judged correct. Under matched token budgets, models produce sharply different execution mixtures (49 of 450 Qwen MATH outputs terminate without a final answer vs 5 of 300 DeepSeek), so accuracy comparisons quietly mix execution case mix with verification policy.

Read summary
AgentEvaluation

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

** Popular performance-optimization benchmarks for coding agents conflate runtime instability, scoring-rule artifacts, and saturation effects, making leaderboard scores unreliable indicators of true coding-agent progress.

Read summary
AgentsSRE

Correct Source, Empty Story: Why Microservice RCA Agents Need Trajectory-Level Grading

A trajectory-level study of 3,500 diagnostic runs shows RCA agents can localize the right service while failing to reconstruct how the fault propagated - and endpoint-correctness benchmarks cannot see the difference. Success needs grounding before localization, which the authors package as DiagGuard.

Read summary
AgentCode

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

** Benchmark pass rates for coding agents can be near-perfect even when the agent failed to build the requested artifact, because the agent satisfies the test oracle by inlining behavior into a throwaway demo instead of the reusable library.

Read summary

+27 more