ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2612 papers59 this month12 topics
AllTraining 50Efficiency 38Reasoning 37Agents 32Evaluation 32Architecture 18Applications 13Safety 12Multimodal 11scaling 6Data 5Alignment 3

Oct 5 – Oct 11(29)

A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control

Oct 8, 2026

Octi Zhang, Mateo Guaman Castro, Patrick Yin et al.

Adaptive curriculum sampling that concentrates on task difficulty at the edge of a robot's capabilities makes massive parallel RL training practical and effective, solving problems that uniform sampling cannot.

This paper tackles a key challenge in scaling reinforcement learning for robots: when training with millions of parallel simulations, most experience gets wasted on tasks the robot either already mastered or can't attempt yet.

trainingefficiencyagents

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

Oct 8, 2026

Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos

A single transformer block with depth-programmed expert mixtures can replace deep vision encoders, reducing parameters dramatically while maintaining accuracy—useful for efficient vision models and elastic deployment at multiple depths.

This paper proposes reViT, a vision transformer that uses a single block applied repeatedly instead of stacking many layers. By representing the feed-forward network as a mixture of shared experts controlled by depth coordinates, it matches full-depth encoders with 70% fewer parameters and comparable computation. The approach works both for training from scratch and distilling from teacher models.

Sep 28 – Oct 4(48)

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

trainingevaluationreasoning

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

Sep 21 – Sep 27(19)

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Sep 25, 2026

Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe et al.

Training language models to predict their confidence in intermediate reasoning steps—using only self-supervised learning—makes them generate shorter reasoning traces at inference time without any explicit length penalties or early-stopping mechanisms.

This paper shows that reasoning models can generate shorter, more efficient reasoning traces by learning to predict their own confidence in answers—without explicitly optimizing for length.

trainingefficiencyreasoning

First-Order Stationarity of Reverse Diffusions

Sep 25, 2026

Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Diffusion models have provable convergence guarantees similar to optimization algorithms—reverse diffusions contract divergence exponentially fast, and discrete samplers achieve measurable stationarity bounds that don't depend on data convexity.

This paper connects optimization theory to diffusion models by proving that reverse-time diffusion processes contract Fisher divergence at exponential rates under strong convexity conditions. The authors also establish first-order stationarity bounds for practical discrete samplers, showing how optimization guarantees translate to sampling quality without requiring global convexity.

Sep 14 – Sep 20(4)

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Sep 18, 2026

Hongyang Du, Lan Yan, Christian Flores et al.

Procedural memory—a continuously updated library of natural-language design skills—enables frozen frontier models to improve at complex agentic tasks by learning from execution failures without model retraining or human annotation.

This paper shows how a frozen AI model can continuously improve at graphic design by building and refining a library of reusable design procedures from real user projects. Without updating the model's weights or using human labels, the system learns 139 design skills from 1,406 real briefs, improving success rates from 73% to 99% by accumulating new procedures and fixing failed ones.

agentstrainingapplications

Cross-sector generalization of accident-process role classification in occupational accident narratives

Sep 18, 2026

Aho Yapi, Pierre Latouche, Arnaud Guillin et al.

Fine-tuned language models can generalize accident report classification across different industries without retraining, enabling scalable occupational safety analysis across sectors.

This paper develops an automated system to classify key information in occupational accident reports (work situations, unsafe conditions, events, consequences) and tests whether models trained on construction-sector narratives can work across different industries like metallurgy and chemistry.

efficiencyarchitecturetraining

Bi-FORK: Generative Modeling of High-Dimensional Bifurcating Systems

Oct 8, 2026

Anna Zimmel, Fleur Hendriks, Markus Holzleitner et al.

Bifurcating systems violate the one-to-one assumption in standard neural surrogates; Bi-FORK solves this by treating bifurcations as a generative modeling problem, enabling amortized prediction of all solution branches simultaneously.

Bi-FORK is a generative model that learns one-to-many solution maps in high-dimensional physical systems undergoing bifurcations—where a single input produces multiple equally valid outputs. Using latent flow matching and repulsion-guided sampling, it generates complete trajectories while preserving spatial and temporal coherence, scaling to systems with hundreds of thousands of dimensions.

architecturetrainingapplications

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

Oct 8, 2026

Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al.

Quantizing optimizer states in preconditioner space (where learning rates are computed) rather than state space reduces training loss gaps by up to 70%, making 4-bit AdamW practical for large-scale pretraining without sacrificing convergence.

This paper improves 4-bit quantization of AdamW optimizer states by rounding in preconditioner space instead of state space. The authors show that quantization errors in the second moment (used to scale learning rates) cause larger problems than errors in the first moment, and propose two methods—ZIP-SR and ZE-EDEN—that better preserve the adaptive learning rates.

trainingefficiency

VioLA: Learning Generalist Humanoid Control Policies from Human Data

Oct 8, 2026

Mert Albaba, Jens Beißwenger, Anna Manasyan et al.

Predicting learned motion representations instead of raw joint commands lets humanoid policies leverage massive human motion datasets for zero-shot real-world control, solving the dual problems of high-dimensional action spaces and scarce robot demonstrations.

VioLA is a humanoid robot control policy that learns from human motion data by predicting body and hand motion patterns instead of direct joint commands. By training on 140 million frames (93% human data), it achieves zero-shot task execution on real robots without task-specific fine-tuning, reaching 100% success on locomotion where prior methods scored 0-17%.

trainingefficiency

Toward Joint Optimization of Circuit Depth and Training Data Size in Adaptively Grown Quantum Classifiers

Oct 8, 2026

Saeefa Rubaiyet Nowmi, Md Mahmuduzzaman Kamol, Mohammad Saidur Rahman

Quantum circuit growth algorithms don't follow predictable scaling laws with training data size, and theoretical generalization bounds, while valid, have weak predictive power for practical circuit design decisions.

This paper investigates whether quantum circuit complexity and training data requirements follow predictable scaling laws. Researchers reimplemented Q-FLAIR, an algorithm that grows quantum circuits gate-by-gate, and tested it on MNIST classification with varying dataset sizes.

trainingscalingevaluation

RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments

Oct 8, 2026

Zimo Wen, Yijin Chen, Yuxuan Cao et al.

By structuring robot learning around explicit skill hierarchies with clear input-output contracts and execution-grounded diagnosis, robots can reliably improve their capabilities through experience and safely reuse learned skills across new tasks.

RoboRSI is a robot self-improvement system that learns and refines skills through real-world experience. It organizes task execution into a hierarchy of skills (compound, atomic, base) with clear responsibilities, diagnoses failures to pinpoint which skill needs fixing, and validates improvements before reusing them.

agentsreasoningtraining

Decoupling Exploration from Optimization in RLVR

Oct 7, 2026

Saif Punjwani, Micah Goldblum

Decoupling exploration (with novelty bonuses) from optimization (standard training) via distillation lets language models discover diverse correct reasoning strategies without quality degradation, outperforming direct RLVR approaches.

This paper proposes Exploration-Distillation (ExpDis), a method that separates exploration from optimization in reinforcement learning with verifiable rewards. Explorer policies use novelty bonuses to discover new reasoning strategies, their best trajectories are filtered and distilled into a student policy trained without novelty incentives.

trainingreasoningagents

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

Oct 7, 2026

Hongru Cai, Ran Wei, Wenjie Wang et al.

Conditional memory architectures can serve as an editable knowledge layer separate from the model's core computation, allowing efficient factual updates without full retraining or catastrophic forgetting.

EngramEdit enables updating factual knowledge in language models that use conditional memory (n-gram lookup tables) without retraining. It computes target memory states for updated facts across different phrasings, then jointly updates shared embeddings while protecting frequently-used ones to avoid breaking unrelated knowledge.

trainingefficiency

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

Oct 7, 2026

Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed

Gradient clipping in decentralized SGD provably converges at optimal rates under heavy-tailed noise with linear speedup across agents—outperforming normalization-based approaches that require additional techniques like momentum.

This paper proves that decentralized SGD with gradient clipping achieves optimal convergence rates under heavy-tailed noise, a common problem in modern machine learning. The key insight is that clipping preserves gradient magnitude information better than normalization, enabling faster convergence and linear speedup across multiple agents in a decentralized network.

trainingefficiencyscaling

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

Oct 7, 2026

Zhewei Chen, Hao Zhu, Jiaojiao Jiang et al.

Using geometric properties like Ollivier-Ricci curvature to guide knowledge distillation helps MLPs capture the graph structure that GNNs learn, improving accuracy while keeping deployment simple.

This paper addresses the challenge of distilling knowledge from Graph Neural Networks (GNNs) to simpler MLPs for deployment. The authors identify two spectral failure modes—underfit on sparse graphs and overfit on dense graphs—and propose G²MLP, which uses Ollivier-Ricci curvature to guide where the student MLP should preserve the teacher's geometric structure.

trainingefficiencyarchitecture

Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

Oct 7, 2026

Amanda Myntti, Jenna Kanerva, Veronika Laippala et al.

Embedding models need explicit training with distractors to reliably follow retrieval instructions—current models are brittle to irrelevant query information despite appearing to understand instructions.

This paper investigates why embedding models struggle to follow retrieval instructions, even simple ones. The researchers discovered that models fail when distractor information is present in queries and show that fine-tuning with query-side distractors significantly improves instruction-following ability without hurting performance on other tasks.

trainingevaluationefficiency

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

Oct 7, 2026

Python Song, Zhixuan Liang, Kelsey Fu et al.

By treating robot learning as active hypothesis testing rather than passive data collection, you can dramatically reduce the physical experiments needed to improve foundation models—this system reaches 77% success where baselines only achieve 40%.

EmbodiedRSI is a self-evolving robot control system that autonomously decides which experiments to run on physical robots and uses the results to improve its code and skills.

agentsreasoningtraining

A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

Oct 7, 2026

Ali Asaria, Deep Gandhi, Tony Salomone

Large populations of AI agents need explicit organizational structures—institutions like peer review and resource allocation—to coordinate effectively and produce better research outcomes than unorganized agent swarms.

This paper proposes organizing large populations of autonomous research agents through explicit institutions inspired by academic science. Rather than managing agents individually, the authors create a 'society of researchers' where agents compete for compute resources via proposal review and grants, with a human governor allocating resources.

agentsscalingtraining

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

Oct 6, 2026

Ziyu Chen, Yilun Zhao, Jiashuo Sun et al.

Structured supervision—encoding how papers should be synthesized together—is more effective for training ideation models than prompting alone, and combining anchor-based training with retrieval produces the best results.

This paper teaches language models to generate research ideas by synthesizing multiple papers, using structured specifications called 'anchors' that capture how papers should be combined.

trainingreasoningapplications

Sherpa: Teaching LLMs to Teach Adaptively

Oct 6, 2026

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al.

LLMs can learn to teach better by optimizing directly for student learning outcomes rather than following predefined teaching rules—this adaptive approach works across different learner types and aligns with how human teachers actually work.

Sherpa trains LLMs to teach adaptively by using reinforcement learning with multiple simulated student archetypes. Instead of relying on fixed teaching demonstrations, the system directly optimizes for student learning outcomes, enabling teachers to personalize instruction. Results show 20.5 percentage point improvements in student performance and 79.6% human preference over baseline models.

trainingalignmentapplications

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Oct 6, 2026

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra et al.

Training web agents in adversarial simulation with co-evolving curricula and adaptive attackers produces agents that generalize better to real-world prompt injection attacks than agents trained on fixed injections.

This paper presents AdvSim2Real, a training method that improves web agents' robustness against prompt injection attacks. The approach co-evolves three components—a task curriculum, an adaptive adversary, and the agent—within a simulated web environment.

safetytrainingagents

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Oct 6, 2026

Zewei Zhou, Rachel Luo, Yulong Cao et al.

Fixing your evaluator is as important as fixing your policy—when agents improve, what they need to be judged on changes, so you need a system to evolve your judge alongside your agent.

VeriFine is a framework that improves AI agents through co-evolving the policy, training data, and evaluation judge. When an agent's performance plateaus, humans help refine the judge by resolving disagreements on tricky cases, then the improved judge guides better training. Tested on driving and robot navigation, it shows continuous improvement as new failure patterns emerge.

trainingreasoningagents

Linear Bandits under Exact Sliding-Window Constraints

Oct 6, 2026

Seyed Mohammad Hadi Hosseini, Yasin Abbasi-Yadkori, Sattar Vakili

Sliding-window constraints make online learning fundamentally harder than offline optimization, but rare policy updates combined with optimistic planning can achieve sublinear regret while maintaining exact feasibility.

This paper studies how to make optimal decisions in linear bandit problems when actions must satisfy strict sliding-window constraints—meaning every consecutive block of actions must come from an allowed set.

trainingreasoningefficiency

On the Computational Tractability of Robust Bandits

Oct 6, 2026

Vanessa Kosoy, Vinayak Pathak

Robust bandits have a sharp computational boundary: a specific case is tractable with efficient algorithms, but small generalizations become NP-hard, suggesting this marks the frontier of what's computationally feasible for unrealizable learning.

This paper studies how to efficiently learn in bandit problems when the true environment doesn't match the learner's model. The authors identify a tractable special case with polynomial-time algorithms and $\tilde{O}(\sqrt{T})$ regret, while showing that natural generalizations become NP-hard—establishing where the problem transitions from solvable to hard.

trainingreasoningsafety

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

Oct 6, 2026

Mathias Ollu, Nikos Komodakis

Diffusing tokens at multiple granularities (fine-grained and clustered) in parallel improves continuous diffusion language models substantially, achieving state-of-the-art results on text generation and reasoning benchmarks.

This paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), which improve text generation by diffusing tokens at multiple semantic levels simultaneously—both individual tokens and coarser token clusters. Applied to existing models like CoBit and FLM, this approach significantly boosts generation quality and reasoning performance with minimal computational overhead.

architecturetrainingefficiency

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

trainingreasoningdata

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Oct 5, 2026

Haojin Deng, Zhiping Lin, Yimin Yang

Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.

BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.

safetyevaluationtraining

Learning to Read the Contextual Tokens in Diffusion Transformers

Oct 5, 2026

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al.

Text tokens in image-generating transformers develop interpretable semantic representations of the emerging image that can be read with an LLM probe—and explicitly training to strengthen these representations improves generation quality.

This paper reveals what text tokens learn during image generation in multimodal diffusion transformers. Researchers built a tool to 'read' these hidden representations by connecting them to a language model, discovering they encode rich scene information early in generation. They then used these insights to improve image quality through a new training technique.

multimodaltraining

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Oct 5, 2026

Yifan Zhang, Yutong Dai, Viraj Prabhu et al.

Self-verification through conformal methods lets web agents learn from their own reasoning about task progress, eliminating the need for expensive judge calls at deployment while improving training efficiency.

CLIFT trains web agents to complete browser tasks by having them verify their own actions through natural-language questions, creating a reusable signal that works both during training (with sparse rewards) and at test time (without expensive external judges). The method achieves state-of-the-art results on multiple web agent benchmarks and transfers across different models.

trainingagentsreasoning

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture
trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

trainingefficiencydata

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

Planning to Learn

Oct 2, 2026

Ian Osband

Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.

This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.

trainingreasoning

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Oct 1, 2026

Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.

Robots can improve their own performance through autonomous practice and skill refinement in simulation without updating model weights, then transfer successfully to real hardware—a practical path to reliable robot systems.

RPG is a framework that improves robot performance without retraining models by identifying skills from offline data, practicing in simulation with failure diagnosis, and refining symbolic skills and system prompts.

agentsreasoningtraining

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Oct 1, 2026

Xuan Zhang, Longtao Zheng, Cunxiao Du et al.

Teaching agents to manage their own context through learned compaction decisions—rather than just handling overflow—improves performance on long-horizon coding tasks by 5-9% across different context window sizes.

AutoCompact trains coding agents to automatically decide when and how to compress their working context during long software engineering tasks. By learning when to discard stale exploration and what state to preserve, the agent improves its ability to solve repository-level coding problems while staying within context limits.

agentsreasoningtraining

When Do Intrinsic Rewards Lead to Exploration?

Oct 1, 2026

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Intrinsic rewards designed to encourage exploration can fail to find the most informative experiences—you need to explicitly measure whether an agent's history can substitute for real experience under different policies.

This paper examines when intrinsic reward signals (like prediction error or curiosity) actually lead to good exploration in reinforcement learning. The authors show that maximizing these rewards doesn't always produce the most informative experiences, propose a formal criterion based on counterfactual information, and demonstrate failures of existing methods with concrete examples.

trainingreasoning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Oct 1, 2026

Lucheng Fu, Kejing Xia, Yiyang Wang et al.

LLM agents can significantly improve performance on knowledge-intensive tasks by learning persistent, source-specific models that evolve through repeated interaction—achieving up to 22.6 point gains over standard retrieval methods.

This paper introduces SourceLearn, a method for LLM agents to develop persistent, reusable understanding of external knowledge sources through repeated interaction.

trainingagentsreasoning

Faynt: Scaling and Optimizing Policies for Competitive Melee

Oct 1, 2026

Ali Janati, Nikita Kuzmin, Rohit Swamy et al.

Scaling, architecture choices, and post-training curricula matter more than model size alone—a smaller, optimized policy outperforms a larger pretrained one, suggesting careful design beats raw parameter count for complex game-playing tasks.

Faynt introduces transformer-based AI policies for Super Smash Bros. Melee that control all 26 characters with a single model. A 10M-parameter version wins 98.4% of same-character matches against existing AI opponents and beats a zero-delay competitor, using techniques like supervised pretraining on 840k human replays, curriculum learning, and distillation from a larger 75M model.

trainingreasoningagents

Finetuning with Sampling: SFT Learns Better Than You Think

Oct 1, 2026

Aayush Karan, Sitan Chen, Yilun Du

SFT isn't inherently worse than RL for posttraining—the gap comes from data distribution mismatch. By reshaping data to be more on-policy before training, SFT can generalize better and forget less than strong RL baselines.

This paper shows that supervised finetuning (SFT) can match or beat reinforcement learning for posttraining if you transform the training data first. The authors use an MCMC sampling algorithm to gradually shift off-policy expert demonstrations toward on-policy trajectories that a reference model can actually learn from.

trainingdata

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Oct 1, 2026

Shivam Kumar, Nabarun Deb

Discrete diffusions can sample from structured categorical data with provable sample complexity guarantees, and weight-sharing neural networks learn scores more efficiently than fully-connected ones for this task.

This paper develops theoretical guarantees for sampling from high-dimensional categorical distributions using discrete diffusion models.

trainingreasoningscaling

Local Support Learning

Oct 1, 2026

Assaf Ben-Kish, Akarsh Kumar, James Glass et al.

You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.

This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.

trainingefficiencyalignment

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Oct 1, 2026

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al.

Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.

This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.

trainingmultimodaldata

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Oct 1, 2026

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

HAO enables stable reinforcement learning under FHE encryption by preventing polynomial approximation errors from accumulating—achieving 0% boundary violations versus 83.8% for unprotected baselines, making privacy-preserving RL practically viable.

This paper solves a critical problem in privacy-preserving reinforcement learning: when you encrypt data with Fully Homomorphic Encryption (FHE) for cloud computation, you must replace nonlinear operations with polynomial approximations, which causes training to diverge.

safetyefficiencytraining

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Oct 1, 2026

Arman Behnam, Binghui Wang

Current memory systems can't measure the value of memories that are never retrieved.

Memory-augmented language models struggle to identify which memories are actually useful because some memories are never retrieved, making their value impossible to measure.

reasoningevaluationtraining

Semifactual Credit-Augmented Policy Optimization

Sep 30, 2026

Junshu Pan, Zhizhang Fu, Shulin Huang et al.

Token-level stability under prompt variations is a useful training signal for improving reasoning in LLMs—you can boost performance by penalizing tokens that change meaning when irrelevant prompt details change.

This paper identifies that language models trained with reinforcement learning are sensitive to irrelevant prompt changes, even when the problem stays the same. The authors propose SCAPO, an improved training method that assigns credit to individual tokens based on how stable they are under these prompt variations.

trainingreasoningalignment

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Sep 30, 2026

Razan El Mais, Ali Chehab, Ibrahim Issa et al.

When fine-tuning LLMs with differential privacy, untying input/output embeddings outperforms the standard weight-tied design and enables 60% memory savings—suggesting privacy-preserving training requires rethinking standard model architectures.

This paper investigates weight tying (sharing parameters between input and output embeddings) in large language models trained with differential privacy. The authors find that untying embeddings actually improves performance under DP-SGD, achieving up to 4.74% accuracy gains, while also enabling more memory-efficient privacy techniques.

safetyefficiencytraining

Turbo Harness: Instance-Adaptive Harness Optimization

Sep 30, 2026

Tunyu Zhang, Hao Wang, Kai Xu et al.

Adapting execution harnesses to individual task instances—rather than using a single global harness—consistently improves agent performance, and this adaptation can be automated by learning from previous optimization runs.

This paper introduces Turbo Harness, a system that automatically customizes AI agent execution frameworks (harnesses) for individual tasks by learning from past optimization runs. Instead of using one fixed harness for all tasks, it generates task-specific modifications that improve agent performance across diverse domains like interactive tasks, coding, and long-horizon planning.

agentstrainingefficiency

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Sep 30, 2026

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang et al.

By separating semantic planning from physical execution and using failure evidence to guide targeted capability improvements, robots can achieve 4x better performance on long-horizon manipulation tasks compared to frozen policies.

DynaHarness is a system that improves robot manipulation by coupling semantic reasoning with physical execution monitoring. It uses a two-level architecture where a 'slow brain' plans high-level actions and a 'fast brain' grounds and monitors execution, refusing unsafe actions and requesting replans when needed. The system learns from failures to improve reusable capabilities.

agentsreasoningtraining

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

Yi Pan, Haocheng Xi, Kan Zhu et al.

You can compress linear attention's recurrent state to 8-bit without significant quality loss by quantizing only at window boundaries and preserving outliers as special tokens—enabling faster inference on long sequences.

LeapQuant reduces the inference cost of linear attention models by quantizing their recurrent state to 8-bit precision while maintaining accuracy. It uses per-window quantization to limit error buildup and compensator tokens to handle outliers, achieving 2-3.7x speedups on real hardware.

efficiencytraining

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Sep 29, 2026

Dor Tirosh, Ido Amos, Mor Geva

By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.

This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.

architecturetrainingreasoning

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Sep 29, 2026

Cheng Qian, Kunlun Zhu, Beibin Li et al.

AI systems can improve other AI systems' performance by learning to build better execution environments—a form of test-time optimization that's reusable across tasks without modifying model weights.

This paper studies how an AI system (Builder) can learn to design better execution environments for another AI system (Target) without changing either model's weights. The Builder learns reusable principles called Meta-Skills from feedback on development tasks, then applies these to construct better environments for new tasks. Results show significant performance improvements across benchmarks.

agentsreasoningtraining

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Sep 29, 2026

Rishabh Agrawal, Hejie Cui, Shasha Li et al.

Selective learning from feedback—keeping corrections that significantly change model behavior while filtering out those that don't—improves advisor performance and generalization to new tasks and model families.

This paper presents AdviSD, a method for training small AI advisors that guide frozen large language models through natural-language feedback. The key innovation is selectively learning from corrections based on how much they actually change the executor's behavior, avoiding learning from corrections that don't meaningfully affect outcomes.

trainingreasoningefficiency

Telescopic Language Models

Sep 28, 2026

Zhilin Guo, Boqiao Zhang, Hakan Aktas et al.

You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.

This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.

trainingefficiencyarchitecture

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Sep 28, 2026

Zimo Wang, Junkun Yuan, Angtian Wang et al.

Critic error accumulation is a fundamental bottleneck in distilling video diffusion models; filtering it via projection dramatically improves sample quality without architectural changes or extra computation.

This paper improves video diffusion model distillation by fixing a key problem: critic errors that accumulate during training and degrade sample quality. PDMD uses a simple mathematical projection to filter out these errors while preserving useful learning signals, achieving better video quality with fewer computational steps—all with just a one-line code change.

efficiencytrainingevaluation

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Sep 28, 2026

Yijia Fan, Ziqi Huang, Zhongang Cai et al.

Unified models can learn to self-correct their outputs by applying RL to complete reflection loops, where both the reasoning about what's wrong and the actual image fixes improve together without needing external verifiers.

This paper presents UMM-Reflection, a method that teaches unified multimodal models to critique and fix their own image generations through reinforcement learning. Instead of just generating images once, the model can now look at what it created, identify problems, revise the image, and repeat—all within a single model.

trainingreasoningmultimodal

How to Loop MoE: Flatten the Experts, Untie the Attention

Sep 28, 2026

Shouren Wang, Chuang Ma, Mohsen Hariri et al.

To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.

This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.

architectureefficiencytraining

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Sep 28, 2026

Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.

KV-streams enables efficient scaling of agentic LLMs to longer horizons by streaming cached computations rather than recomputing them, making long-context RL training practical without sacrificing performance.

This paper introduces KV-streams, a technique that speeds up training of long-horizon agentic language models by streaming the key-value cache forward during context compaction instead of repeatedly refilling it. The method achieves 2.6-5x training speedup while maintaining performance, and shows that the streamed cache can retain information beyond the visible context window.

efficiencytrainingagents

Towards Communication-Efficient Social Intelligence in Language Agents

Sep 28, 2026

Linxiao Gong, Yijie Xu, Tianfu Wang et al.

TACT enables language agents to achieve better social outcomes while using fewer tokens and messages by having specialized teachers refine communication strategy and expression, then distilling improvements into the student agent.

This paper introduces Teacher-Assisted Communication Training (TACT), a method that helps language agents communicate more efficiently during social interactions. TACT improves how agents negotiate and coordinate by having a teacher refine both what agents say (expression) and how they say it (strategy), then distills these improvements back into the student agent.

trainingagentsefficiency

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Sep 28, 2026

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan et al.

When adapting pre-trained vision models to use linear attention, directly copy MLP weights but distill attention behavior—this simple strategy closes the performance gap between efficient and standard transformers.

This paper shows how to initialize linear Vision Transformers (efficient attention models) using weights from standard Softmax ViTs. The key insight: copy the MLP layers directly since they learn general representations, but use distillation to transfer the attention mechanism since it's operator-specific.

efficiencyarchitecturetraining

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Sep 28, 2026

Jonathan Light, Christopher Zhang Cui, Jeonghye Kim et al.

Agents can learn to act better by learning to explain their actions—training on self-generated retrospections alone improves future performance without RL, suggesting explanation is a useful learning signal for behavior improvement.

This paper shows that language model agents can improve their performance by training on self-generated explanations of their own experiences, without needing reinforcement learning or external rewards. The method, called Retrospection-Only Fine-Tuning (ROFT), has an agent attempt tasks, generate explanations of what happened, and then fine-tune on predicting those explanations.

trainingagentsreasoning

Harness Learning Enables Generalizable Test-Time Adaptation

Sep 28, 2026

Alvin Zhang, Xuecheng Liu, Zixuan Wang et al.

Language model agents can adapt to new tasks by learning to revise their execution harness (program structure) rather than their weights, enabling test-time adaptation that generalizes to unseen tasks.

This paper introduces harness learning, a method where an AI agent learns to improve its own executable program (harness) that controls how a language model makes decisions and uses tools. Instead of changing the model's weights, a separate proposer model learns to revise the harness structure based on task feedback.

agentstrainingreasoning
trainingevaluationreasoning

New LoRA Skills Should Read but Never Write

Sep 25, 2026

Zeyan Li, Panqi Yang, Qirong Guo et al.

When combining multiple LoRA adapters, the internal representation and directional coupling between them matters more than the adapter weights themselves—fixing these choices lets you add skills sequentially without degrading previous ones.

This paper solves the problem of combining multiple fine-tuned LoRA adapters into a single model without interference. The key insight is that LoRA updates have multiple equivalent forms, and the choice matters when combining adapters.

trainingefficiency

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Sep 25, 2026

Varun Reddy, Bernardo L. Sabatini, Houman Safaai

Common-mode error in direct feedback alignment causes training stalls by saturating hidden units; this can be prevented by centering batch errors or calibrating the readout baseline, enabling faster learning without changing the core algorithm.

Direct feedback alignment trains neural networks using fixed random error projections, but gets stuck learning near a baseline predictor. The paper identifies that a shared error component across inputs drives hidden units to saturation, slowing learning. Simple fixes like centering errors or adjusting the baseline readout can prevent this collapse and speed up training.

trainingefficiency

Strategically Diverse Sampling for Self-Training

Sep 25, 2026

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

For self-training, sampling diverse problem-solving strategies matters more than correctness or teacher model size—a small model trained on varied approaches beats distillation from a 235B teacher.

This paper shows that self-training works better when you sample diverse problem-solving approaches rather than just correct answers. The authors introduce GROOT (a tree-based sampling method) and Verbalized Sampling to generate strategically different solutions, and find that models trained on diverse but incorrect traces outperform those trained on correct answers from much larger models.

trainingreasoning

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Sep 24, 2026

Wenhao Li, Zhibin Wu, Chong Xiao et al.

Using LLM-generated semantics as a shared anchor point for aligning incomplete multimodal data is more robust than trying to reconstruct missing modalities or design complex fusion mechanisms.

SemMSA tackles multimodal sentiment analysis when some data is missing by using large language models to create rich semantic representations that ground all modalities together. Instead of reconstructing missing features, it aligns visual, acoustic, and text representations through spectral methods, achieving better results on standard benchmarks.

multimodalalignmenttraining

PoEM: Predicting RL Outcomes from Existing Policies

Sep 24, 2026

Kimia Hamidieh, Giannis Daras, Antonio Torralba

You can predict RL outcomes for new reward functions by combining existing trained models mathematically, avoiding the computational cost of retraining—useful when experimenting with different objectives or combining multiple goals.

PoEM predicts what a reinforcement learning model will do with a new reward function by combining existing models trained on different rewards, without running expensive RL training. The method works by finding that RL policies live in a low-rank space that can be reconstructed as a linear combination of existing policies.

trainingefficiencyreasoning

Anchored Extra-Proximal Methods: Optimal Higher-Order Methods for Monotone Inclusion Problems

Sep 24, 2026

Ruichen Jiang, TaeHo Yoon

Higher-order optimization methods can solve monotone inclusion problems with complexity O(ε^{-2/(3p-1)}), which is provably optimal and improves prior bounds by using anchored extrapolation with Taylor approximations of operators.

This paper develops optimal higher-order methods for solving monotone inclusion problems—a fundamental class of optimization problems. The authors introduce the Anchored Extra-Proximal framework that achieves better convergence rates than prior methods by combining extrapolation with proximal updates. They prove their approach is optimal up to logarithmic factors across all orders.

training

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

Sep 24, 2026

Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek et al.

Training latent dynamics models for long-horizon stability requires explicitly optimizing for multi-step rollout accuracy, not just reconstruction—this restructures the solution space in ways that conventional metrics don't capture.

This paper shows that neural surrogate models for physics simulations fail during long predictions not because of poor compression, but because they're trained only to reconstruct data. The authors introduce training techniques—including Koopman operator learning and noise injection—that restructure the latent space to support stable long-horizon forecasting.

efficiencytrainingreasoning

Intrinsic-Extrinsic Coupling in Learning Dynamics

Sep 24, 2026

Qinyou Wang

A model's internal state and external training conditions interact in complex, non-additive ways—the same intervention can help or hurt depending on what happens next, which matters for understanding continual learning and model adaptation.

This paper studies how a machine learning model's current state interacts with future training dynamics. The authors develop methods to measure and manipulate learning states—like classifier weights and historical information—to understand when interventions help or hurt performance.

trainingevaluation

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

Sep 24, 2026

Nayoung Choi, Shengjian Chen, Xiaokai Wei et al.

Optimizing search query understanding components individually with search-engine-derived rewards outperforms single end-to-end optimization, showing that understanding how each component affects downstream retrieval matters more than just matching labels.

This paper presents a reinforcement learning framework for query understanding in search systems that optimizes multiple components (like intent classification and query expansion) separately using rewards from live search engine interactions, rather than treating it as a single end-to-end problem.

trainingreasoningapplications

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

Sep 24, 2026

Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh

Token-level tolerance to transcription ambiguity in ASR training reduces word error rates by ~9.5% on average by letting models skip disputed individual characters while keeping supervision for the rest of the word.

This paper addresses a real problem in speech recognition: reference transcripts often contain ambiguous pronunciations or spellings that the audio doesn't uniquely determine.

trainingevaluation

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Sep 22, 2026

Laizhen Li, Jiarui Li, Juanjuan Zhao et al.

You can move repetitive agent control logic from expensive LLM context into persistent, reusable code—cutting inference costs by 74-99% while keeping smaller models effective on complex tasks.

This paper introduces Growing Harness, a method that automatically builds reusable agent control code from task feedback instead of asking language models to repeatedly solve the same control problems. By learning executable code that handles recurring decisions, the approach reduces LLM calls by 76-92% while maintaining or improving task success rates across different model sizes.

agentsefficiencytraining

FleXray: Universal Clinical X-ray Segmentation

Sep 22, 2026

Victor Ion Butoi, Vivek Gopalakrishnan, John V. Guttag et al.

You can train powerful medical imaging models without expensive manual annotation by simulating realistic training data from existing 3D datasets—FleXray segments full-body X-rays and works on real clinical data despite being trained entirely on synthetic images.

FleXray is a generalist AI model that segments 60 anatomical structures in clinical X-rays across the entire body. Rather than manually labeling thousands of X-rays, the researchers built a physics-based simulator that generates realistic synthetic X-rays from existing 3D CT scans, then trained the model on these simulations.

trainingdataapplications

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 22, 2026

Yuanteng Chen, Zhilei Liu, Peisong Wang et al.

On-policy distillation recovers reasoning capabilities in ultra-low-bit quantized models by training on the model's own generated outputs rather than fixed data, fixing the exposure bias problem that causes long-form reasoning to fail.

This paper tackles a critical problem in quantized language models: when you compress models to very low precision (under 3 bits), they lose the ability to do math and coding tasks because errors compound during long generation.

trainingefficiencyreasoning

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Sep 21, 2026

Lei Yang, Mengyin Liu, Jia Wang et al.

Token-level correction during annotation is significantly faster than full rewriting and produces on-policy training data that preserves the model's natural generation patterns while providing precise supervision signals.

onPanda is an interactive annotation tool that helps create training data for AI models by letting annotators correct responses token-by-token. Instead of rewriting entire outputs, annotators find the first mistake, fix it, and let the model regenerate from that point.

trainingdataalignment

Harness-Zero: Harness Distillation via Agent-as-Harness

Sep 21, 2026

Haoran Ye, Yuxing Lu, Haonan Dong et al.

You can distill specialized harness behaviors into model weights by having an intermediate agent translate between different harness action spaces during training, letting you deploy with simpler harnesses while keeping performance gains.

This paper tackles how to transfer the benefits of specialized agent harnesses (external systems that improve model-environment interaction) into model weights so they work with simpler harnesses at deployment.

trainingagentsapplications

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Sep 21, 2026

Peng Xia, Rujun Han, Zifeng Wang et al.

Automatically improving agent harnesses through constrained evolution can boost performance while staying generalizable—the key is regularizing the search process to favor reusable mechanisms over task-specific tricks.

This paper presents RRSI, a method for automatically improving LLM agent systems by evolving their harnesses (prompts, tools, memory, control flow) while avoiding overfitting to training tasks. It uses regularization techniques like edit budgets and change filtering to find improvements that generalize to new benchmarks, achieving strong gains on both in-distribution and out-of-distribution tasks.

agentstrainingefficiency

Learning Physics from an Imperfect Ancestor

Sep 21, 2026

S. Mohammad Mousavi, Teeratorn Kadeethum, Nikolaos Bouklas et al.

Neural operators don't need to be accurate to be useful—they can guide PINNs away from spurious solutions by providing the right structural prior, enabling reliable PDE solving in regimes where either method alone would fail.

This paper shows how to combine neural operators (fast but inaccurate) with physics-informed neural networks (accurate but optimization-fragile) to solve PDEs reliably. An imperfect neural operator provides a structural hint about which solution the PINN should find, while the PDE residual refines it to high accuracy.

trainingreasoning
trainingevaluationapplications

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Sep 18, 2026

Bowen Ye, Lei Li, Shicheng Li et al.

Using source code itself as the primary input, you can automatically generate thousands of high-quality RL training tasks for coding agents without relying on manual annotations or development artifacts like issues.

CodeMidas automatically creates reinforcement learning training tasks from open-source code by using AI agents to explore codebases, generate test cases, and validate tasks. This approach scales coding agent training to 5,545 diverse tasks across 23 languages, improving performance on code repair, program synthesis, and terminal tasks by 8-18%.

trainingagentsreasoning

Benchmarking World Models for Continual Learning on Compositional Tasks

Sep 18, 2026

Haoyu Zhou, Joe Watson, Anson Lei et al.

Modular world model architectures better balance knowledge reuse with avoiding catastrophic forgetting in continual learning, but the field still lacks methods that effectively retain and reuse knowledge across sequential robot tasks.

This paper creates a benchmark to test how well world models (AI systems that learn to predict environment dynamics) can learn continuously across robot tasks without forgetting previous knowledge. The key innovation is using compositional tasks—where new tasks combine elements from earlier ones—to isolate what knowledge gets reused versus forgotten.

trainingevaluationarchitecture