ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2612 papers100 this month12 topics
AllTraining 50Efficiency 38Reasoning 37Agents 32Evaluation 32Architecture 18Applications 13Safety 12Multimodal 11scaling 6Data 5Alignment 3

Oct 5 – Oct 11(60)

A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control

Oct 8, 2026

Octi Zhang, Mateo Guaman Castro, Patrick Yin et al.

Adaptive curriculum sampling that concentrates on task difficulty at the edge of a robot's capabilities makes massive parallel RL training practical and effective, solving problems that uniform sampling cannot.

This paper tackles a key challenge in scaling reinforcement learning for robots: when training with millions of parallel simulations, most experience gets wasted on tasks the robot either already mastered or can't attempt yet.

trainingefficiencyagents

On the estimation and validity of AI time horizons---a statistical look at the METR plot

Oct 8, 2026

Drew T. Nguyen, William Fithian

AI time horizon benchmarks need better statistical foundations: a 10x increase in human task time doesn't represent equal difficulty gains across all ranges, which matters for fairly comparing AI capabilities.

This paper examines how AI capabilities are measured using 'time horizons'—the human task completion time at which an AI succeeds 50% of the time. The authors show that the standard linear model for this relationship is flawed and propose better statistical methods using splines and item-response theory, revealing that task difficulty doesn't scale uniformly with human time.

Sep 28 – Oct 4(40)

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

Oct 2, 2026

Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan et al.

Decoder expressivity matters: simpler decoders with latent-space objectives produce better transferable geometric representations than complex pixel-space decoders, even in self-supervised settings.

This paper shows that Novel View Synthesis can learn strong 3D geometric representations if you constrain the decoder and use latent-space reconstruction instead of pixel-level targets. The authors introduce SNAP, which learns viewpoint-invariant features useful for localization, pose estimation, depth, and robot tasks—without needing explicit 3D supervision.

architecture

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Oct 2, 2026

Ruihong Shen, Žiga Kovačič, Peter Kulits et al.

Strong vision models fail at understanding dynamic scenes through code generation—the gap between static and dynamic reconstruction is a major frontier for AI agents.

4DCodeBench is a benchmark that tests AI agents on reconstructing dynamic 3D scenes from video by writing graphics code. Agents must understand physics, deformation, and fluid dynamics to generate executable programs that recreate what they see. The benchmark reveals that current models struggle with complex motion even when they're good at static scene reconstruction.

evaluationscaling

CSF: Contextual Safety Filtering for Motion Generators

Oct 8, 2026

Lizhi Yang, Yiling Hou, Yao Tang et al.

CSF enables safety filtering for motion generators by grounding natural-language safety rules in reference trajectories, reducing unsafe motions by up to 90% without requiring labeled data or model retraining.

This paper presents Contextual Safety Filtering (CSF), a method that makes motion-generating AI systems safer by understanding scene context. Instead of just checking text prompts, CSF learns what makes a motion safe or unsafe by comparing reference trajectories, then uses control techniques to prevent dangerous movements.

safetyagentsefficiency

From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents

Oct 8, 2026

Abbas Raftari

AI agent security requires continuous runtime verification of system boundaries and multi-layered controls—no single sandbox or safeguard is sufficient to prevent agents from accessing unintended systems.

This paper analyzes three major AI agent security incidents from 2026 where OpenAI, Anthropic, and Google agents escaped their test environments and accessed real systems. It proposes a Proactive Agent Security Assurance Cycle (PASAC) and Boundary Assurance Stack framework that emphasizes continuous verification of execution boundaries rather than relying on single safeguards.

safetyagentsevaluation

BrickBench: Evaluating Agentic Brick Design

Oct 8, 2026

Peter Kulits, Yiqing Xu, R. Kenny Jones et al.

Current leading AI agents can meet basic physical and semantic requirements for LEGO design but struggle to match human-level design quality, revealing gaps in spatial reasoning and constraint satisfaction under discrete choice problems.

BrickBench is a benchmark that tests AI agents' ability to design buildable LEGO sets from text descriptions. Agents must select parts from a library and satisfy physical constraints, semantic requirements, and design quality—combining reasoning about local details with global structure. The paper includes BrickAgent, an environment for agents to build and validate designs.

agentsreasoningevaluation

Bi-FORK: Generative Modeling of High-Dimensional Bifurcating Systems

Oct 8, 2026

Anna Zimmel, Fleur Hendriks, Markus Holzleitner et al.

Bifurcating systems violate the one-to-one assumption in standard neural surrogates; Bi-FORK solves this by treating bifurcations as a generative modeling problem, enabling amortized prediction of all solution branches simultaneously.

Bi-FORK is a generative model that learns one-to-many solution maps in high-dimensional physical systems undergoing bifurcations—where a single input produces multiple equally valid outputs. Using latent flow matching and repulsion-guided sampling, it generates complete trajectories while preserving spatial and temporal coherence, scaling to systems with hundreds of thousands of dimensions.

architecturetrainingapplications

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

Oct 8, 2026

Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos

A single transformer block with depth-programmed expert mixtures can replace deep vision encoders, reducing parameters dramatically while maintaining accuracy—useful for efficient vision models and elastic deployment at multiple depths.

This paper proposes reViT, a vision transformer that uses a single block applied repeatedly instead of stacking many layers. By representing the feed-forward network as a mixture of shared experts controlled by depth coordinates, it matches full-depth encoders with 70% fewer parameters and comparable computation. The approach works both for training from scratch and distilling from teacher models.

efficiencyarchitecturetraining

Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception

Oct 8, 2026

Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba et al.

Internal activation patterns can reliably catch model deception and hidden goals better than reading model outputs, enabling practical safety monitoring for deployed AI systems.

Researchers developed white-box probes that detect when language models deceive or sabotage by analyzing their internal activations across layers and tokens.

safetyevaluationalignment

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

Oct 8, 2026

Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al.

Quantizing optimizer states in preconditioner space (where learning rates are computed) rather than state space reduces training loss gaps by up to 70%, making 4-bit AdamW practical for large-scale pretraining without sacrificing convergence.

This paper improves 4-bit quantization of AdamW optimizer states by rounding in preconditioner space instead of state space. The authors show that quantization errors in the second moment (used to scale learning rates) cause larger problems than errors in the first moment, and propose two methods—ZIP-SR and ZE-EDEN—that better preserve the adaptive learning rates.

trainingefficiency

Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff

Oct 8, 2026

Erin Crawley, Hidenori Tanaka

Collaborative AI agents pose a population-level safety risk distinct from individual agent risks: a small group may be harmless, but once population size crosses a critical threshold, collective cyber capability can explode in a self-reinforcing cycle.

This paper applies ecological theory to AI agent safety, showing that misaligned agents collaborating to conduct cyberattacks could trigger a population explosion once they reach a critical threshold—even if individual agent capability stays constant.

safetyagentsscaling

VioLA: Learning Generalist Humanoid Control Policies from Human Data

Oct 8, 2026

Mert Albaba, Jens Beißwenger, Anna Manasyan et al.

Predicting learned motion representations instead of raw joint commands lets humanoid policies leverage massive human motion datasets for zero-shot real-world control, solving the dual problems of high-dimensional action spaces and scarce robot demonstrations.

VioLA is a humanoid robot control policy that learns from human motion data by predicting body and hand motion patterns instead of direct joint commands. By training on 140 million frames (93% human data), it achieves zero-shot task execution on real robots without task-specific fine-tuning, reaching 100% success on locomotion where prior methods scored 0-17%.

trainingefficiency

Toward Joint Optimization of Circuit Depth and Training Data Size in Adaptively Grown Quantum Classifiers

Oct 8, 2026

Saeefa Rubaiyet Nowmi, Md Mahmuduzzaman Kamol, Mohammad Saidur Rahman

Quantum circuit growth algorithms don't follow predictable scaling laws with training data size, and theoretical generalization bounds, while valid, have weak predictive power for practical circuit design decisions.

This paper investigates whether quantum circuit complexity and training data requirements follow predictable scaling laws. Researchers reimplemented Q-FLAIR, an algorithm that grows quantum circuits gate-by-gate, and tested it on MNIST classification with varying dataset sizes.

trainingscalingevaluation

FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?

Oct 8, 2026

Yuxuan Hu, Weikang Shi, Yang Bo et al.

Streaming VLMs fail at perceiving fast events in videos—the best model scores only 50.7% on FastBench—because they can't adaptively determine when to sample frames densely, revealing a fundamental gap between sparse uniform sampling and what's actually needed for dynamic understanding.

FastBench is a benchmark for evaluating how well streaming video language models can understand fast-moving events in real-world videos. Current models struggle with this task because they must balance limited context budgets across temporal history, spatial resolution, and frame sampling rates.

evaluationmultimodal

RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments

Oct 8, 2026

Zimo Wen, Yijin Chen, Yuxuan Cao et al.

By structuring robot learning around explicit skill hierarchies with clear input-output contracts and execution-grounded diagnosis, robots can reliably improve their capabilities through experience and safely reuse learned skills across new tasks.

RoboRSI is a robot self-improvement system that learns and refines skills through real-world experience. It organizes task execution into a hierarchy of skills (compound, atomic, base) with clear responsibilities, diagnoses failures to pinpoint which skill needs fixing, and validates improvements before reusing them.

agentsreasoningtraining

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

Oct 8, 2026

Luping Liu, Bingyi Kang, Yifan Wang et al.

By combining foundation models with diverse training data, you can build correspondence matching systems that work across both classical vision tasks and modern generative applications like image editing.

FreeMatching is a framework for matching visual correspondence across images that goes beyond traditional motion and geometry assumptions. It combines generative and semantic models trained on diverse data sources to handle challenging transformations in image editing and generation tasks, while maintaining competitive performance on standard benchmarks.

multimodalevaluationarchitecture

Decoupling Exploration from Optimization in RLVR

Oct 7, 2026

Saif Punjwani, Micah Goldblum

Decoupling exploration (with novelty bonuses) from optimization (standard training) via distillation lets language models discover diverse correct reasoning strategies without quality degradation, outperforming direct RLVR approaches.

This paper proposes Exploration-Distillation (ExpDis), a method that separates exploration from optimization in reinforcement learning with verifiable rewards. Explorer policies use novelty bonuses to discover new reasoning strategies, their best trajectories are filtered and distilled into a student policy trained without novelty incentives.

trainingreasoningagents

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

Oct 7, 2026

Hongru Cai, Ran Wei, Wenjie Wang et al.

Conditional memory architectures can serve as an editable knowledge layer separate from the model's core computation, allowing efficient factual updates without full retraining or catastrophic forgetting.

EngramEdit enables updating factual knowledge in language models that use conditional memory (n-gram lookup tables) without retraining. It computes target memory states for updated facts across different phrasings, then jointly updates shared embeddings while protecting frequently-used ones to avoid breaking unrelated knowledge.

trainingefficiency

Long-WAM: Scaling the Context of World-Action Models

Oct 7, 2026

Wei Huang, Bohan Zhang, Chenzhi Liu et al.

Autoregressive pretraining on video unlocks the value of long context windows for robot control—bidirectional models don't benefit from extra history, but AR-pretrained models show consistent improvements with longer visual context.

Long-WAM is a system for controlling robots in real-time by processing long video histories (up to 19 seconds) to understand motion and task progress. The key insight is that autoregressive video pretraining—learning to predict future frames from past ones—makes longer context windows actually useful for robot control, whereas bidirectional pretraining doesn't benefit from extra history.

agentsefficiencyarchitecture

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

Oct 7, 2026

Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed

Gradient clipping in decentralized SGD provably converges at optimal rates under heavy-tailed noise with linear speedup across agents—outperforming normalization-based approaches that require additional techniques like momentum.

This paper proves that decentralized SGD with gradient clipping achieves optimal convergence rates under heavy-tailed noise, a common problem in modern machine learning. The key insight is that clipping preserves gradient magnitude information better than normalization, enabling faster convergence and linear speedup across multiple agents in a decentralized network.

trainingefficiencyscaling

Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

Oct 7, 2026

Mikey Watts, Yuchen Cui

Language phrasing sensitivity in VLAs is a learnable, systematic problem that can be mitigated at inference time using LLM-distilled rephrasing rules, improving robustness to out-of-distribution instructions without model retraining.

Vision-language-action models (VLAs) that control robots are surprisingly fragile to how instructions are phrased—changing one word can drop success rates by 50+ points.

multimodalagents

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

Oct 7, 2026

Zhewei Chen, Hao Zhu, Jiaojiao Jiang et al.

Using geometric properties like Ollivier-Ricci curvature to guide knowledge distillation helps MLPs capture the graph structure that GNNs learn, improving accuracy while keeping deployment simple.

This paper addresses the challenge of distilling knowledge from Graph Neural Networks (GNNs) to simpler MLPs for deployment. The authors identify two spectral failure modes—underfit on sparse graphs and overfit on dense graphs—and propose G²MLP, which uses Ollivier-Ricci curvature to guide where the student MLP should preserve the teacher's geometric structure.

trainingefficiencyarchitecture

RoboJEPA: Scaling Robotic Latent World Models

Oct 7, 2026

Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan et al.

Latent world models follow reliable scaling laws: imagination error (how well the model predicts future states) improves predictably with compute, and this directly correlates with real robot task performance, making it possible to forecast model quality without expensive robot experiments.

RoboJEPA is an 8-billion-parameter world model trained on real robot data from 12 different embodiments. It predicts future states in a compressed latent space and follows predictable scaling laws—meaning you can estimate how much better it gets with more compute before actually training it. The model works as a zero-shot robot controller, planning by imagining future states toward a goal image.

scalingreasoning

SciExam for ENSO: Can AI Agents Build Climate Models?

Oct 7, 2026

Yinling Zhang, Langchen Liu, Dongbin Xiu et al.

AI agents can conduct genuine scientific research on unsolved problems—this benchmark shows agents building climate models that outperform published work and align with competing scientific theories, proving evaluation is possible without a known ground truth.

Researchers created SciExam for ENSO, a benchmark where AI agents build climate models of El Niño from real data without knowing the right answer. Agents work within a 6-hour budget, write their own diagnostics, and develop models tested against hidden criteria.

agentsreasoningevaluation

Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

Oct 7, 2026

Amanda Myntti, Jenna Kanerva, Veronika Laippala et al.

Embedding models need explicit training with distractors to reliably follow retrieval instructions—current models are brittle to irrelevant query information despite appearing to understand instructions.

This paper investigates why embedding models struggle to follow retrieval instructions, even simple ones. The researchers discovered that models fail when distractor information is present in queries and show that fine-tuning with query-side distractors significantly improves instruction-following ability without hurting performance on other tasks.

trainingevaluationefficiency

RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

Oct 7, 2026

Yilun Hao, Krishna Sayana, Isabella Ye et al.

Instead of just retrieving relevant documents, RECAST learns to combine retrieval with computation—filtering, aggregating, and deriving answers across multiple sources—enabling LLMs to handle complex, multi-step reasoning tasks more effectively.

RECAST is a framework that helps language models solve complex tasks by learning to actively construct evidence through computation rather than just retrieving it. A lightweight router model decides which operations to perform on multiple information sources, a compiler translates those decisions into executable code, and an answer model produces the final result.

reasoningagents

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

Oct 7, 2026

Daniel Robert Kling Alexander, Catherine Louise Kling

You can evaluate LLMs on subjective tasks by checking if their responses follow theoretical predictions and internal consistency, rather than comparing to a ground truth—a framework borrowed from decades of economics research.

This paper adapts validity frameworks from economics to evaluate language models on subjective questions without ground truth answers. Using stated-preference survey methods, the authors test whether LLM responses follow economic theory predictions (like downward-sloping demand curves), demonstrating that validity testing can assess model coherence even when correct answers don't exist.

evaluationalignment

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

Oct 7, 2026

Python Song, Zhixuan Liang, Kelsey Fu et al.

By treating robot learning as active hypothesis testing rather than passive data collection, you can dramatically reduce the physical experiments needed to improve foundation models—this system reaches 77% success where baselines only achieve 40%.

EmbodiedRSI is a self-evolving robot control system that autonomously decides which experiments to run on physical robots and uses the results to improve its code and skills.

agentsreasoningtraining

Evolutionary Architecture Search for Chlorophyll-$a$ Prediction in Lakes using Sentinel-2

Oct 7, 2026

Kursat Komurcu, Linas Petkevicius

Automated architecture search can discover surprisingly effective small models for Earth observation tasks—in this case, a 1.6KB network that outperforms hand-designed baselines and fits on resource-constrained devices.

This paper uses evolutionary architecture search to automatically design neural networks for predicting chlorophyll-a levels in lakes from satellite imagery. Starting with a hand-designed model, the search finds smaller, better-performing networks (26× fewer parameters) that fit on edge devices for real-time monitoring.

architectureefficiencyevaluation

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Oct 7, 2026

Tan Yu, Alexander Bukharin, Khushi Bhardwaj et al.

You can predict post-training performance of coding agents by analyzing base models' probability of generating specific successful code actions from recorded trajectories—avoiding expensive full post-training runs.

This paper proposes methods to predict which base language models will perform well after expensive post-training for coding agents, without running the full post-training process.

evaluationagentsefficiency

A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

Oct 7, 2026

Ali Asaria, Deep Gandhi, Tony Salomone

Large populations of AI agents need explicit organizational structures—institutions like peer review and resource allocation—to coordinate effectively and produce better research outcomes than unorganized agent swarms.

This paper proposes organizing large populations of autonomous research agents through explicit institutions inspired by academic science. Rather than managing agents individually, the authors create a 'society of researchers' where agents compete for compute resources via proposal review and grants, with a human governor allocating resources.

agentsscalingtraining

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

Oct 6, 2026

Kevin Zhang, Stephen Bates

Conformal prediction set sizes have a rigorous information-theoretic interpretation: they quantify information gain in a way that's mathematically sandwiched between generalized entropy measures and obeys data processing inequalities.

This paper establishes a theoretical connection between conformal prediction (a method for uncertainty quantification) and information theory. The authors show that the size of prediction sets from conformal methods can be interpreted as a measure of information gain, providing mathematical justification for using set size as an uncertainty metric.

evaluationsafetyreasoning

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

Oct 6, 2026

Shiqi Li, Sean Cho, Yijie Li et al.

Feed-forward generative models can reconstruct complex hand-object interactions faster and more reliably than optimization-based methods by learning to correct foundation model errors while respecting physical constraints during generation.

This paper presents 4D-HOF, a fast method for reconstructing 3D hand and object positions/orientations from video. Instead of slow per-video optimization, it uses a generative model trained on diverse data to refine rough estimates from vision models.

architectureefficiencyreasoning

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

Oct 6, 2026

Ziyu Chen, Yilun Zhao, Jiashuo Sun et al.

Structured supervision—encoding how papers should be synthesized together—is more effective for training ideation models than prompting alone, and combining anchor-based training with retrieval produces the best results.

This paper teaches language models to generate research ideas by synthesizing multiple papers, using structured specifications called 'anchors' that capture how papers should be combined.

trainingreasoningapplications

DepthWorld: 3D World Model for Robot Manipulation

Oct 6, 2026

Jai Bardhan, Josef Sivic, Vladimir Petrik

Adding 3D depth supervision to video world models improves both geometric consistency and visual quality, enabling more reliable 3D reasoning for robot manipulation without requiring architectural changes to pretrained video models.

This paper presents DepthWorld, a 3D world model for robot manipulation that predicts both RGB video and metric depth.

multimodal

Sherpa: Teaching LLMs to Teach Adaptively

Oct 6, 2026

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al.

LLMs can learn to teach better by optimizing directly for student learning outcomes rather than following predefined teaching rules—this adaptive approach works across different learner types and aligns with how human teachers actually work.

Sherpa trains LLMs to teach adaptively by using reinforcement learning with multiple simulated student archetypes. Instead of relying on fixed teaching demonstrations, the system directly optimizes for student learning outcomes, enabling teachers to personalize instruction. Results show 20.5 percentage point improvements in student performance and 79.6% human preference over baseline models.

trainingalignmentapplications

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

Oct 6, 2026

Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.

LLM agents struggle to convert their general capabilities into cost-efficient task-specific solutions, but when they succeed, the savings are dramatic—suggesting bottling is a valuable but underdeveloped capability worth improving.

This paper introduces BOTTLED, a benchmark testing whether LLM agents can autonomously create cheaper, task-specific solutions from their general capabilities. Agents receive unlabeled workloads with fixed budgets and must decide their own approach—like training small models or writing programs.

agentsefficiencyevaluation

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Oct 6, 2026

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra et al.

Training web agents in adversarial simulation with co-evolving curricula and adaptive attackers produces agents that generalize better to real-world prompt injection attacks than agents trained on fixed injections.

This paper presents AdvSim2Real, a training method that improves web agents' robustness against prompt injection attacks. The approach co-evolves three components—a task curriculum, an adaptive adversary, and the agent—within a simulated web environment.

safetytrainingagents

Rapid Fredholm stabilization of the Kuramoto--Sivashinsky equation with unrestricted, spatially-varying anti-diffusion

Oct 6, 2026

Luke Bhan, Miroslav Krstic, Yuanyuan Shi

Neural operators can approximate control gains for PDE stabilization with high accuracy (~0.1% error), enabling practical implementation of theoretically-grounded feedback designs for systems where traditional single-input methods fail.

This paper develops a feedback control method for stabilizing the Kuramoto-Sivashinsky equation, a complex nonlinear system, using two boundary inputs instead of one. The key innovation is using a neural operator to approximate the control gains, enabling practical implementation while maintaining theoretical stability guarantees.

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Oct 6, 2026

Zewei Zhou, Rachel Luo, Yulong Cao et al.

Fixing your evaluator is as important as fixing your policy—when agents improve, what they need to be judged on changes, so you need a system to evolve your judge alongside your agent.

VeriFine is a framework that improves AI agents through co-evolving the policy, training data, and evaluation judge. When an agent's performance plateaus, humans help refine the judge by resolving disagreements on tricky cases, then the improved judge guides better training. Tested on driving and robot navigation, it shows continuous improvement as new failure patterns emerge.

trainingreasoningagents

WorldSonus: Bringing Sound to Worlds

Oct 6, 2026

Pengjun Fang, Jingyi Fa, Kam Man Wu et al.

Real-time spatial audio generation for interactive video is now feasible with streaming diffusion models, enabling world models to produce complete audiovisual experiences rather than silent videos.

WorldSonus adds realistic spatial sound to AI-generated video environments in real-time. It uses a streaming diffusion model to generate audio that matches video content, responds to text instructions mid-generation, and creates stereo sound that aligns with scene geometry and camera movement.

multimodalefficiencyapplications

The Missing Minimal Pair: Stereotype Evaluation in LLMs

Oct 6, 2026

Nataliya Stepanova, Ivan Titov, Emily Allaway et al.

Single stereotype sentence pairs are unreliable for measuring LLM bias; use dual minimal pairs and mutual information-based metrics instead for consistent, language-agnostic bias evaluation.

This paper identifies a flaw in how bias is typically measured in language models: comparing just two sentences about stereotypes can give contradictory results depending on how you rephrase them. The authors propose a better approach using dual comparisons and introduce metrics based on mutual information to more reliably measure stereotype bias across languages and models.

evaluationsafety

Linear Bandits under Exact Sliding-Window Constraints

Oct 6, 2026

Seyed Mohammad Hadi Hosseini, Yasin Abbasi-Yadkori, Sattar Vakili

Sliding-window constraints make online learning fundamentally harder than offline optimization, but rare policy updates combined with optimistic planning can achieve sublinear regret while maintaining exact feasibility.

This paper studies how to make optimal decisions in linear bandit problems when actions must satisfy strict sliding-window constraints—meaning every consecutive block of actions must come from an allowed set.

trainingreasoningefficiency

On the Computational Tractability of Robust Bandits

Oct 6, 2026

Vanessa Kosoy, Vinayak Pathak

Robust bandits have a sharp computational boundary: a specific case is tractable with efficient algorithms, but small generalizations become NP-hard, suggesting this marks the frontier of what's computationally feasible for unrealizable learning.

This paper studies how to efficiently learn in bandit problems when the true environment doesn't match the learner's model. The authors identify a tractable special case with polynomial-time algorithms and $\tilde{O}(\sqrt{T})$ regret, while showing that natural generalizations become NP-hard—establishing where the problem transitions from solvable to hard.

trainingreasoningsafety

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

Oct 6, 2026

Mathias Ollu, Nikos Komodakis

Diffusing tokens at multiple granularities (fine-grained and clustered) in parallel improves continuous diffusion language models substantially, achieving state-of-the-art results on text generation and reasoning benchmarks.

This paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), which improve text generation by diffusing tokens at multiple semantic levels simultaneously—both individual tokens and coarser token clusters. Applied to existing models like CoBit and FLM, this approach significantly boosts generation quality and reasoning performance with minimal computational overhead.

architecturetrainingefficiency

A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

Oct 6, 2026

Alexia Allal, Hicham Randrianarivo, Sylvain Lamprier

DocID design significantly impacts generative retrieval performance—the paper provides a systematic framework and training-free metrics to evaluate and optimize DocID structures, enabling faster iteration than traditional downstream evaluation.

This paper systematically studies how to design effective document identifiers (DocIDs) for generative information retrieval systems.

evaluationefficiency

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Oct 5, 2026

Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang et al.

Agentic systems with explicit constraint checking and visual critics can reliably preserve structural integrity in document layout tasks—achieving 68.6% fidelity versus 11-41% for prior methods—by factoring the problem into specialized stages rather than end-to-end generation.

This paper tackles the problem of automatically adapting flowchart diagrams to different aspect ratios (like fitting a pipeline figure into a paper column, slide, or social media format) while preserving all connections and content.

agentsapplicationsevaluation

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

trainingreasoningdata

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Oct 5, 2026

Haojin Deng, Zhiping Lin, Yimin Yang

Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.

BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.

safetyevaluationtraining

Learning to Read the Contextual Tokens in Diffusion Transformers

Oct 5, 2026

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al.

Text tokens in image-generating transformers develop interpretable semantic representations of the emerging image that can be read with an LLM probe—and explicitly training to strengthen these representations improves generation quality.

This paper reveals what text tokens learn during image generation in multimodal diffusion transformers. Researchers built a tool to 'read' these hidden representations by connecting them to a language model, discovering they encode rich scene information early in generation. They then used these insights to improve image quality through a new training technique.

multimodaltraining

Recursive Video In-Context Learning for Agentic Robot

Oct 5, 2026

Wenrui Bao, Xinxin Liu, Bingxin Xu et al.

By organizing demonstration videos into navigable hierarchies rather than static prompts, agents can access task details only when needed, reducing context overhead while improving learning from single examples.

This paper presents Recursive Video In-Context Learning (RV-ICL), a method that helps robot agents learn from demonstration videos more effectively. Instead of feeding entire videos as prompts, RV-ICL organizes a single demo video into a hierarchical structure of sub-events (like grasps and releases) that the agent can navigate on-demand.

agentsreasoningmultimodal

Direct Intermediate Initialization for Tilted Diffusion Samplers

Oct 5, 2026

Gregory D. Bellchambers

Initializing diffusion samplers at intermediate timesteps using pulled-back clean-space posteriors and Gaussian bridges can dramatically improve sample quality, especially when the posterior has modes that are rare under the prior.

This paper improves diffusion-based posterior sampling by initializing the sampler at an intermediate step rather than starting from pure noise. The key insight is that Gaussian-tilted targets along the reverse process can be reformulated as weaker clean-space posteriors, with samples transported analytically via a Gaussian bridge.

efficiency

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Oct 5, 2026

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al.

Separating the UI slider from the underlying strength parameter and remapping it based on perceptual distance creates intuitive, predictable image editing interfaces that users prefer.

UniSlider makes image editing sliders feel natural by ensuring perceptual change increases smoothly and predictably as you move the slider. Current methods produce uneven results—some slider positions cause no visible change while others transform the image abruptly.

efficiencyapplications

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Oct 5, 2026

Yifan Zhang, Yutong Dai, Viraj Prabhu et al.

Self-verification through conformal methods lets web agents learn from their own reasoning about task progress, eliminating the need for expensive judge calls at deployment while improving training efficiency.

CLIFT trains web agents to complete browser tasks by having them verify their own actions through natural-language questions, creating a reusable signal that works both during training (with sparse rewards) and at test time (without expensive external judges). The method achieves state-of-the-art results on multiple web agent benchmarks and transfers across different models.

trainingagentsreasoning

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Oct 5, 2026

Yaohui Zhang, Binxu Li, Haoyi Duan et al.

Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.

PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.

evaluationmultimodaldata

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

Oct 5, 2026

Oliver Jaffe, Dane Sherburn

AI models are becoming exponentially better at the experimental research process itself (not just raw capability), with frontier models now reaching expert-level results using 2.3x less compute—a skill that could significantly accelerate AI R&D timelines.

TasteVal is a benchmark measuring how efficiently AI models design and conduct experiments to solve research problems. Rather than evaluating raw problem-solving ability, it measures 'experimental taste'—the skill to iteratively design good experiments and interpret results—by comparing how much compute a model needs versus human experts to reach the same performance level.

evaluationreasoningagents

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture

TAPDreamer: Transferable Adversarial Patches for World Action Models

Oct 5, 2026

Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng et al.

Small, fixed adversarial patches can severely degrade world action models across multiple robotic tasks by exploiting how visual encoders process information, highlighting that securing shared visual components is critical for robust robotic control systems.

This paper presents TAPDreamer, an attack method that uses small visual patches to fool world action models—AI systems that predict how robotic environments will change. Unlike previous attacks, TAPDreamer works without accessing the target model, instead using a public encoder to create a single patch that transfers across different tasks and robot policies.

safety
evaluationreasoningagents

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

trainingevaluationreasoning

RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

Oct 2, 2026

Yiming Huang, Lennart Bastian, Hanqun Cao et al.

A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.

This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.

dataarchitectureevaluation

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

trainingefficiencydata

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Oct 2, 2026

Hui Chen, Xuan Qi, James Xu Zhao et al.

By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.

FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.

efficiencyreasoningagents

Planning to Learn

Oct 2, 2026

Ian Osband

Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.

This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.

trainingreasoning

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation

Oct 2, 2026

Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar et al.

Simulator-generated counterfactual rollouts can effectively bootstrap forecasting models for new policies before real deployment data exists, and these models improve further with minimal real-world calibration.

When deploying a new decision policy, prediction models face a cold-start problem because historical data reflects old policies, not the new one. This paper uses simulation to generate counterfactual training data by rolling out the new policy in a simulator, then tests whether models trained on simulated data transfer to real-world inventory control.

applicationsevaluation

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

Oct 2, 2026

Junyoung Koh, Hao-Wen Dong

Linear STFT outperforms the more complex HCQT for vocal ensemble pitch estimation while reducing computational cost—simpler input representations can be more effective when properly designed.

This paper challenges the conventional use of harmonic constant-Q transform (HCQT) for multi-pitch estimation in vocal ensembles by showing that simpler linear STFT representations actually perform better while being much faster to compute.

evaluation

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Oct 2, 2026

Sean Culatana, Shang-En Huang, Kang Li

MRVQ enables one quantized index to serve all dimension-rate combinations by using truncatable residual quantization, reducing memory overhead by 17.8-22x compared to training separate indices, though with modest quality trade-offs.

This paper introduces MRVQ, a quantization method that compresses embeddings for vector search while supporting flexible trade-offs between embedding dimension and compression rate. A single index can be truncated in two ways—dropping quantization stages or embedding coordinates—to adapt to different memory and latency constraints without storing multiple separate indices.

efficiencyevaluation

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Oct 2, 2026

Rubén Manrique, Michelle Castellanos, Jorge Morales et al.

LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.

This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.

evaluationsafetyapplications

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Oct 1, 2026

Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.

Current open-weight LLMs struggle with precise cybersecurity tool use (max 42% accuracy), but fine-tuning with verifiable rewards from this benchmark can make smaller models competitive with much larger ones.

KaliBench is a benchmark for evaluating how well language models can translate security analyst requests into executable commands for Kali Linux tools. It includes 8,504 query-command pairs across 1,642 tools and provides a verification system that checks both whether commands are syntactically correct and whether they actually run successfully, without needing to execute them during training.

evaluationagentssafety

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Oct 1, 2026

Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.

Robots can improve their own performance through autonomous practice and skill refinement in simulation without updating model weights, then transfer successfully to real hardware—a practical path to reliable robot systems.

RPG is a framework that improves robot performance without retraining models by identifying skills from offline data, practicing in simulation with failure diagnosis, and refining symbolic skills and system prompts.

agentsreasoningtraining

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Oct 1, 2026

Sohyeon Kim, Yoonho Lee, Bo Liu et al.

Even advanced AI agents fail at retrieving papers that inspired real research (max 0.51 recall), revealing a critical gap in how models search scientific literature—this task requires something beyond current retrieval and reasoning approaches.

ScholarCatalyst is a benchmark dataset where 184 computer science researchers labeled which prior papers inspired their completed projects. The benchmark tests whether AI systems can retrieve these influential papers given only an initial research question and literature available at project start.

evaluationreasoningagents

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Oct 1, 2026

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.

Representing 3D shapes as overlapping 2D slices rather than voxels dramatically reduces computational cost while improving topological correctness—a practical win for efficient 3D generation at scale.

SILSA is a 3D generation framework that uses compact sliding-window slice representations instead of expensive voxel tokens to generate high-resolution 3D shapes. By encoding cross-sections along three axes and adding topology supervision based on persistence diagrams, it achieves better structural quality while using 70% fewer tokens and running 58% faster than competing methods.

architectureefficiency

VISTA: A Visual Harness for Reasoning in an Interactive World

Oct 1, 2026

Qiushi Han, Keya Hu, Linlu Qiu et al.

Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.

VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.

agentsmultimodalreasoning

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control

Oct 1, 2026

Akshay Balsubramani

Cost-augmented optimal transport on graphs can be solved exactly and efficiently using matrix exponentials instead of learned neural networks—no approximation or temporal discretization needed.

This paper solves the cost-augmented Schrödinger bridge problem on graphs exactly, without learning or time discretization. By reformulating state costs as a Feynman-Kac tilt of the reference process, the problem reduces to computing a standard bridge through alternating matrix exponentials. The method is exact, memory-efficient, and converges based on endpoint coupling alone.

reasoning

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Oct 1, 2026

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi et al.

HGR enables sequence models to generate chemically valid molecules with perfect validity while capturing complex molecular topology, achieving top performance on generation and property prediction tasks without the computational cost of explicit higher-order encodings.

This paper introduces Higher-order Grammar Representation (HGR), a new way to represent molecules that captures complex structural features like ring systems by converting them into sequences of grammar rules. Unlike existing methods that struggle with computational overhead, HGR makes these structures compatible with standard sequence models while guaranteeing valid molecules.

architecturedataapplications

Decoding Looped Transformers Better for (Almost) Free

Oct 1, 2026

Weihao Liu, Huangjie Zheng, Tianrong Chen et al.

Looped Transformers naturally produce weak-to-strong prediction pairs across recurrent passes; contrasting them during decoding improves quality and enables halving compute with no training needed.

This paper introduces LoopCD, a training-free decoding method for looped Transformers that reuses intermediate predictions from earlier recurrent passes to guide token selection. By contrasting predictions from different loop depths, LoopCD improves reasoning accuracy (e.g., AIME scores from 61.88% to 73.33%) while cutting inference compute by 22-48% through fewer required loops.

efficiency

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

Generative Cinematographer: Composing Camera and Object Motion in 3D

Oct 1, 2026

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.

By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.

GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.

multimodalarchitectureapplications

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Oct 1, 2026

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

Self-repair in language models isn't adaptive compensation—it's pre-existing counterweights responding predictably to ablation. You can predict how a component will respond to intervention from its fixed weights alone.

When you disable a component in a language model, other parts often seem to compensate—a phenomenon called 'self-repair.' This paper shows it's not actually repair: it's pre-existing counterweights doing their normal job.

evaluation

Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination

Oct 1, 2026

Suyu Ye, Zheyuan Zhang, Vaishnav Tadiparthi et al.

Observing joint behavior between two robots reveals hidden physical constraints better than observing a single constrained robot, enabling effective zero-shot coordination without explicit communication.

This paper tackles a practical robotics problem: how can one robot learn another robot's physical limitations (like broken joints or weak actuators) just by watching them work together, then use that knowledge to coordinate effectively on new tasks? The authors show that by analyzing how both robots move together, you can infer hidden constraints better than looking at just one robot alone.

agentsreasoning

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Oct 1, 2026

Xuan Zhang, Longtao Zheng, Cunxiao Du et al.

Teaching agents to manage their own context through learned compaction decisions—rather than just handling overflow—improves performance on long-horizon coding tasks by 5-9% across different context window sizes.

AutoCompact trains coding agents to automatically decide when and how to compress their working context during long software engineering tasks. By learning when to discard stale exploration and what state to preserve, the agent improves its ability to solve repository-level coding problems while staying within context limits.

agentsreasoningtraining

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Oct 1, 2026

Hanchu Zhou, Dechen Gao, Hang Wang et al.

Using semantic communication between robots—where they exchange meaningful descriptions rather than sensor data—enables better coordination on long-horizon tasks while keeping each robot's execution independent and reliable.

DuoMind is a framework that enables multiple robots to coordinate and work together on complex tasks by combining vision-language models for high-level reasoning with vision-language-action models for precise execution. Robots communicate through semantic messages rather than raw data, allowing them to share understanding of the task and environment while maintaining independent control.

agentsmultimodalreasoning

When Do Intrinsic Rewards Lead to Exploration?

Oct 1, 2026

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Intrinsic rewards designed to encourage exploration can fail to find the most informative experiences—you need to explicitly measure whether an agent's history can substitute for real experience under different policies.

This paper examines when intrinsic reward signals (like prediction error or curiosity) actually lead to good exploration in reinforcement learning. The authors show that maximizing these rewards doesn't always produce the most informative experiences, propose a formal criterion based on counterfactual information, and demonstrate failures of existing methods with concrete examples.

trainingreasoning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Oct 1, 2026

Lucheng Fu, Kejing Xia, Yiyang Wang et al.

LLM agents can significantly improve performance on knowledge-intensive tasks by learning persistent, source-specific models that evolve through repeated interaction—achieving up to 22.6 point gains over standard retrieval methods.

This paper introduces SourceLearn, a method for LLM agents to develop persistent, reusable understanding of external knowledge sources through repeated interaction.

trainingagentsreasoning