ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2612 papers43 this month12 topics
AllTraining 50Efficiency 38Reasoning 37Agents 32Evaluation 32Architecture 18Applications 13Safety 12Multimodal 11scaling 6Data 5Alignment 3

Oct 5 – Oct 11(23)

A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control

Oct 8, 2026

Octi Zhang, Mateo Guaman Castro, Patrick Yin et al.

Adaptive curriculum sampling that concentrates on task difficulty at the edge of a robot's capabilities makes massive parallel RL training practical and effective, solving problems that uniform sampling cannot.

This paper tackles a key challenge in scaling reinforcement learning for robots: when training with millions of parallel simulations, most experience gets wasted on tasks the robot either already mastered or can't attempt yet.

trainingefficiencyagents

CSF: Contextual Safety Filtering for Motion Generators

Oct 8, 2026

Lizhi Yang, Yiling Hou, Yao Tang et al.

CSF enables safety filtering for motion generators by grounding natural-language safety rules in reference trajectories, reducing unsafe motions by up to 90% without requiring labeled data or model retraining.

This paper presents Contextual Safety Filtering (CSF), a method that makes motion-generating AI systems safer by understanding scene context. Instead of just checking text prompts, CSF learns what makes a motion safe or unsafe by comparing reference trajectories, then uses control techniques to prevent dangerous movements.

Sep 28 – Oct 4(40)

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

Sep 21 – Sep 27(21)

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Sep 25, 2026

Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe et al.

Training language models to predict their confidence in intermediate reasoning steps—using only self-supervised learning—makes them generate shorter reasoning traces at inference time without any explicit length penalties or early-stopping mechanisms.

This paper shows that reasoning models can generate shorter, more efficient reasoning traces by learning to predict their own confidence in answers—without explicitly optimizing for length.

trainingefficiencyreasoning

New LoRA Skills Should Read but Never Write

Sep 25, 2026

Zeyan Li, Panqi Yang, Qirong Guo et al.

When combining multiple LoRA adapters, the internal representation and directional coupling between them matters more than the adapter weights themselves—fixing these choices lets you add skills sequentially without degrading previous ones.

This paper solves the problem of combining multiple fine-tuned LoRA adapters into a single model without interference. The key insight is that LoRA updates have multiple equivalent forms, and the choice matters when combining adapters.

Sep 14 – Sep 20(16)

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

Sep 18, 2026

Yiming Zhang, Jinghong Zhang, Haoran Zhao et al.

When using RAG with LLMs, blindly trusting all retrieved memories causes hallucinations; a lightweight geometric decision layer can filter unreliable memories without any learned parameters, making RAG safer and more trustworthy.

This paper introduces Memory Decision Layer (MDL), a parameter-free controller that decides whether to trust retrieved memories in RAG systems. It uses three signals—relevance, reliability, and task risk—combined through geometric operations to detect conflicting memories and prevent hallucinations, reducing errors by 56% when memories contradict each other.

reasoningsafetyefficiency

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Sep 18, 2026

Richard Zhe Wang

Attention heads need both the ability to abstain from attending and to filter noise from values—their importance shifts with model scale, suggesting future architectures should support both primitives.

This paper identifies two missing capabilities in standard softmax attention: abstention (allowing heads to output nothing instead of always producing weighted combinations) and noise filtering (suppressing interference from mixed features).

safetyagentsefficiency

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

Oct 8, 2026

Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos

A single transformer block with depth-programmed expert mixtures can replace deep vision encoders, reducing parameters dramatically while maintaining accuracy—useful for efficient vision models and elastic deployment at multiple depths.

This paper proposes reViT, a vision transformer that uses a single block applied repeatedly instead of stacking many layers. By representing the feed-forward network as a mixture of shared experts controlled by depth coordinates, it matches full-depth encoders with 70% fewer parameters and comparable computation. The approach works both for training from scratch and distilling from teacher models.

efficiencyarchitecturetraining

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

Oct 8, 2026

Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al.

Quantizing optimizer states in preconditioner space (where learning rates are computed) rather than state space reduces training loss gaps by up to 70%, making 4-bit AdamW practical for large-scale pretraining without sacrificing convergence.

This paper improves 4-bit quantization of AdamW optimizer states by rounding in preconditioner space instead of state space. The authors show that quantization errors in the second moment (used to scale learning rates) cause larger problems than errors in the first moment, and propose two methods—ZIP-SR and ZE-EDEN—that better preserve the adaptive learning rates.

trainingefficiency

VioLA: Learning Generalist Humanoid Control Policies from Human Data

Oct 8, 2026

Mert Albaba, Jens Beißwenger, Anna Manasyan et al.

Predicting learned motion representations instead of raw joint commands lets humanoid policies leverage massive human motion datasets for zero-shot real-world control, solving the dual problems of high-dimensional action spaces and scarce robot demonstrations.

VioLA is a humanoid robot control policy that learns from human motion data by predicting body and hand motion patterns instead of direct joint commands. By training on 140 million frames (93% human data), it achieves zero-shot task execution on real robots without task-specific fine-tuning, reaching 100% success on locomotion where prior methods scored 0-17%.

trainingefficiency

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

Oct 7, 2026

Hongru Cai, Ran Wei, Wenjie Wang et al.

Conditional memory architectures can serve as an editable knowledge layer separate from the model's core computation, allowing efficient factual updates without full retraining or catastrophic forgetting.

EngramEdit enables updating factual knowledge in language models that use conditional memory (n-gram lookup tables) without retraining. It computes target memory states for updated facts across different phrasings, then jointly updates shared embeddings while protecting frequently-used ones to avoid breaking unrelated knowledge.

trainingefficiency

Long-WAM: Scaling the Context of World-Action Models

Oct 7, 2026

Wei Huang, Bohan Zhang, Chenzhi Liu et al.

Autoregressive pretraining on video unlocks the value of long context windows for robot control—bidirectional models don't benefit from extra history, but AR-pretrained models show consistent improvements with longer visual context.

Long-WAM is a system for controlling robots in real-time by processing long video histories (up to 19 seconds) to understand motion and task progress. The key insight is that autoregressive video pretraining—learning to predict future frames from past ones—makes longer context windows actually useful for robot control, whereas bidirectional pretraining doesn't benefit from extra history.

agentsefficiencyarchitecture

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

Oct 7, 2026

Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed

Gradient clipping in decentralized SGD provably converges at optimal rates under heavy-tailed noise with linear speedup across agents—outperforming normalization-based approaches that require additional techniques like momentum.

This paper proves that decentralized SGD with gradient clipping achieves optimal convergence rates under heavy-tailed noise, a common problem in modern machine learning. The key insight is that clipping preserves gradient magnitude information better than normalization, enabling faster convergence and linear speedup across multiple agents in a decentralized network.

trainingefficiencyscaling

Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

Oct 7, 2026

Zhewei Chen, Hao Zhu, Jiaojiao Jiang et al.

Using geometric properties like Ollivier-Ricci curvature to guide knowledge distillation helps MLPs capture the graph structure that GNNs learn, improving accuracy while keeping deployment simple.

This paper addresses the challenge of distilling knowledge from Graph Neural Networks (GNNs) to simpler MLPs for deployment. The authors identify two spectral failure modes—underfit on sparse graphs and overfit on dense graphs—and propose G²MLP, which uses Ollivier-Ricci curvature to guide where the student MLP should preserve the teacher's geometric structure.

trainingefficiencyarchitecture

Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

Oct 7, 2026

Amanda Myntti, Jenna Kanerva, Veronika Laippala et al.

Embedding models need explicit training with distractors to reliably follow retrieval instructions—current models are brittle to irrelevant query information despite appearing to understand instructions.

This paper investigates why embedding models struggle to follow retrieval instructions, even simple ones. The researchers discovered that models fail when distractor information is present in queries and show that fine-tuning with query-side distractors significantly improves instruction-following ability without hurting performance on other tasks.

trainingevaluationefficiency

Evolutionary Architecture Search for Chlorophyll-$a$ Prediction in Lakes using Sentinel-2

Oct 7, 2026

Kursat Komurcu, Linas Petkevicius

Automated architecture search can discover surprisingly effective small models for Earth observation tasks—in this case, a 1.6KB network that outperforms hand-designed baselines and fits on resource-constrained devices.

This paper uses evolutionary architecture search to automatically design neural networks for predicting chlorophyll-a levels in lakes from satellite imagery. Starting with a hand-designed model, the search finds smaller, better-performing networks (26× fewer parameters) that fit on edge devices for real-time monitoring.

architectureefficiencyevaluation

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Oct 7, 2026

Tan Yu, Alexander Bukharin, Khushi Bhardwaj et al.

You can predict post-training performance of coding agents by analyzing base models' probability of generating specific successful code actions from recorded trajectories—avoiding expensive full post-training runs.

This paper proposes methods to predict which base language models will perform well after expensive post-training for coding agents, without running the full post-training process.

evaluationagentsefficiency

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

Oct 6, 2026

Shiqi Li, Sean Cho, Yijie Li et al.

Feed-forward generative models can reconstruct complex hand-object interactions faster and more reliably than optimization-based methods by learning to correct foundation model errors while respecting physical constraints during generation.

This paper presents 4D-HOF, a fast method for reconstructing 3D hand and object positions/orientations from video. Instead of slow per-video optimization, it uses a generative model trained on diverse data to refine rough estimates from vision models.

architectureefficiencyreasoning

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

Oct 6, 2026

Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.

LLM agents struggle to convert their general capabilities into cost-efficient task-specific solutions, but when they succeed, the savings are dramatic—suggesting bottling is a valuable but underdeveloped capability worth improving.

This paper introduces BOTTLED, a benchmark testing whether LLM agents can autonomously create cheaper, task-specific solutions from their general capabilities. Agents receive unlabeled workloads with fixed budgets and must decide their own approach—like training small models or writing programs.

agentsefficiencyevaluation

WorldSonus: Bringing Sound to Worlds

Oct 6, 2026

Pengjun Fang, Jingyi Fa, Kam Man Wu et al.

Real-time spatial audio generation for interactive video is now feasible with streaming diffusion models, enabling world models to produce complete audiovisual experiences rather than silent videos.

WorldSonus adds realistic spatial sound to AI-generated video environments in real-time. It uses a streaming diffusion model to generate audio that matches video content, responds to text instructions mid-generation, and creates stereo sound that aligns with scene geometry and camera movement.

multimodalefficiencyapplications

Linear Bandits under Exact Sliding-Window Constraints

Oct 6, 2026

Seyed Mohammad Hadi Hosseini, Yasin Abbasi-Yadkori, Sattar Vakili

Sliding-window constraints make online learning fundamentally harder than offline optimization, but rare policy updates combined with optimistic planning can achieve sublinear regret while maintaining exact feasibility.

This paper studies how to make optimal decisions in linear bandit problems when actions must satisfy strict sliding-window constraints—meaning every consecutive block of actions must come from an allowed set.

trainingreasoningefficiency

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

Oct 6, 2026

Mathias Ollu, Nikos Komodakis

Diffusing tokens at multiple granularities (fine-grained and clustered) in parallel improves continuous diffusion language models substantially, achieving state-of-the-art results on text generation and reasoning benchmarks.

This paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), which improve text generation by diffusing tokens at multiple semantic levels simultaneously—both individual tokens and coarser token clusters. Applied to existing models like CoBit and FLM, this approach significantly boosts generation quality and reasoning performance with minimal computational overhead.

architecturetrainingefficiency

A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

Oct 6, 2026

Alexia Allal, Hicham Randrianarivo, Sylvain Lamprier

DocID design significantly impacts generative retrieval performance—the paper provides a systematic framework and training-free metrics to evaluate and optimize DocID structures, enabling faster iteration than traditional downstream evaluation.

This paper systematically studies how to design effective document identifiers (DocIDs) for generative information retrieval systems.

evaluationefficiency

Direct Intermediate Initialization for Tilted Diffusion Samplers

Oct 5, 2026

Gregory D. Bellchambers

Initializing diffusion samplers at intermediate timesteps using pulled-back clean-space posteriors and Gaussian bridges can dramatically improve sample quality, especially when the posterior has modes that are rare under the prior.

This paper improves diffusion-based posterior sampling by initializing the sampler at an intermediate step rather than starting from pure noise. The key insight is that Gaussian-tilted targets along the reverse process can be reformulated as weaker clean-space posteriors, with samples transported analytically via a Gaussian bridge.

efficiency

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Oct 5, 2026

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al.

Separating the UI slider from the underlying strength parameter and remapping it based on perceptual distance creates intuitive, predictable image editing interfaces that users prefer.

UniSlider makes image editing sliders feel natural by ensuring perceptual change increases smoothly and predictably as you move the slider. Current methods produce uneven results—some slider positions cause no visible change while others transform the image abruptly.

efficiencyapplications

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture
trainingefficiencydata

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Oct 2, 2026

Hui Chen, Xuan Qi, James Xu Zhao et al.

By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.

FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.

efficiencyreasoningagents

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Oct 2, 2026

Sean Culatana, Shang-En Huang, Kang Li

MRVQ enables one quantized index to serve all dimension-rate combinations by using truncatable residual quantization, reducing memory overhead by 17.8-22x compared to training separate indices, though with modest quality trade-offs.

This paper introduces MRVQ, a quantization method that compresses embeddings for vector search while supporting flexible trade-offs between embedding dimension and compression rate. A single index can be truncated in two ways—dropping quantization stages or embedding coordinates—to adapt to different memory and latency constraints without storing multiple separate indices.

efficiencyevaluation

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Oct 1, 2026

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.

Representing 3D shapes as overlapping 2D slices rather than voxels dramatically reduces computational cost while improving topological correctness—a practical win for efficient 3D generation at scale.

SILSA is a 3D generation framework that uses compact sliding-window slice representations instead of expensive voxel tokens to generate high-resolution 3D shapes. By encoding cross-sections along three axes and adding topology supervision based on persistence diagrams, it achieves better structural quality while using 70% fewer tokens and running 58% faster than competing methods.

architectureefficiency

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

Decoding Looped Transformers Better for (Almost) Free

Oct 1, 2026

Weihao Liu, Huangjie Zheng, Tianrong Chen et al.

Looped Transformers naturally produce weak-to-strong prediction pairs across recurrent passes; contrasting them during decoding improves quality and enables halving compute with no training needed.

This paper introduces LoopCD, a training-free decoding method for looped Transformers that reuses intermediate predictions from earlier recurrent passes to guide token selection. By contrasting predictions from different loop depths, LoopCD improves reasoning accuracy (e.g., AIME scores from 61.88% to 73.33%) while cutting inference compute by 22-48% through fewer required loops.

efficiency

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Oct 1, 2026

Juan S. Santillana

Don't trust tool-use benchmarks for small models—use verbatim-reproduction checks and token-probability probes to verify genuine capability before claiming tool use works.

Small language models can appear to use tools correctly on standard benchmarks while actually just memorizing training examples. This paper shows how keyword-matching tests miss real failures, proposes cheap diagnostic checks to catch false positives, and demonstrates a targeted fix that repairs tool-use ability in a 1.1B parameter model using minimal compute.

evaluationefficiency

Local Support Learning

Oct 1, 2026

Assaf Ben-Kish, Akarsh Kumar, James Glass et al.

You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.

This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.

trainingefficiencyalignment

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Oct 1, 2026

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

HAO enables stable reinforcement learning under FHE encryption by preventing polynomial approximation errors from accumulating—achieving 0% boundary violations versus 83.8% for unprotected baselines, making privacy-preserving RL practically viable.

This paper solves a critical problem in privacy-preserving reinforcement learning: when you encrypt data with Fully Homomorphic Encryption (FHE) for cloud computation, you must replace nonlinear operations with polynomial approximations, which causes training to diverge.

safetyefficiencytraining

Image Classifiers are Efficient Self-Supervised Video Representation Learners

Sep 30, 2026

Owais Iqbal, Sudipta Sarkar, Shyam Marjit et al.

You can efficiently learn video representations by repurposing pretrained image models with clever masking strategies, avoiding expensive 3D architectures and reconstruction overhead.

VideoMSN uses standard image Vision Transformers to learn video representations without 3D models or reconstruction. It treats videos as grids of frames and masks either spatial patches or temporal frames, then aligns the two views using a Siamese loss. This achieves top results on video benchmarks while needing 32-160x fewer training epochs than prior methods.

efficiencymultimodal

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Sep 30, 2026

Razan El Mais, Ali Chehab, Ibrahim Issa et al.

When fine-tuning LLMs with differential privacy, untying input/output embeddings outperforms the standard weight-tied design and enables 60% memory savings—suggesting privacy-preserving training requires rethinking standard model architectures.

This paper investigates weight tying (sharing parameters between input and output embeddings) in large language models trained with differential privacy. The authors find that untying embeddings actually improves performance under DP-SGD, achieving up to 4.74% accuracy gains, while also enabling more memory-efficient privacy techniques.

safetyefficiencytraining

Turbo Harness: Instance-Adaptive Harness Optimization

Sep 30, 2026

Tunyu Zhang, Hao Wang, Kai Xu et al.

Adapting execution harnesses to individual task instances—rather than using a single global harness—consistently improves agent performance, and this adaptation can be automated by learning from previous optimization runs.

This paper introduces Turbo Harness, a system that automatically customizes AI agent execution frameworks (harnesses) for individual tasks by learning from past optimization runs. Instead of using one fixed harness for all tasks, it generates task-specific modifications that improve agent performance across diverse domains like interactive tasks, coding, and long-horizon planning.

agentstrainingefficiency

Scaling Laws for Looped Mixture of Experts

Sep 30, 2026

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

Looped MoE models combine two orthogonal efficiency axes: recurrence increases computational depth while sparsity expands capacity, and their scaling laws enable principled design of efficient models that match much larger dense models at the same compute budget.

This paper develops scaling laws that jointly model looped transformers (which use recurrence for computational depth) and Mixture-of-Experts (which use sparsity for capacity).

scalingefficiencyarchitecture

Looped Diffusion Transformer

Sep 30, 2026

Yong Xien Chng, Tianyi Chen, Wenwen Tong et al.

Looping computation within denoising steps is more efficient than scaling model size—you can get better images with fewer parameters by iteratively refining representations through repeated block execution.

This paper proposes Looped Diffusion Transformer, which improves text-to-image generation by repeatedly processing the same Transformer blocks within each denoising step rather than making models larger.

efficiencyarchitecturescaling

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Sep 29, 2026

Bingchen Yao, Haobo Xu, Haokun Lin et al.

Quantization errors in recurrent states don't matter equally: errors in long-lived memory and in dimensions that strongly influence outputs cause more accuracy loss, so allocating precision based on these factors enables aggressive compression without sacrificing performance.

STEPQuant is a quantization method that compresses the persistent memory states used in linear attention models. By analyzing how quantization errors affect model outputs differently across time and space, the method allocates precision strategically—giving more bits to memory that lasts longer and to state dimensions that matter most for predictions.

efficiency

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

Yi Pan, Haocheng Xi, Kan Zhu et al.

You can compress linear attention's recurrent state to 8-bit without significant quality loss by quantizing only at window boundaries and preserving outliers as special tokens—enabling faster inference on long sequences.

LeapQuant reduces the inference cost of linear attention models by quantizing their recurrent state to 8-bit precision while maintaining accuracy. It uses per-window quantization to limit error buildup and compensator tokens to handle outliers, achieving 2-3.7x speedups on real hardware.

efficiencytraining

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Sep 29, 2026

Rishabh Agrawal, Hejie Cui, Shasha Li et al.

Selective learning from feedback—keeping corrections that significantly change model behavior while filtering out those that don't—improves advisor performance and generalization to new tasks and model families.

This paper presents AdviSD, a method for training small AI advisors that guide frozen large language models through natural-language feedback. The key innovation is selectively learning from corrections based on how much they actually change the executor's behavior, avoiding learning from corrections that don't meaningfully affect outcomes.

trainingreasoningefficiency

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Sep 29, 2026

Yu Xu, Yuxin Zhang, Xiao Yang et al.

Video models need different scaling strategies than language models—splitting experts by semantic role and allowing imbalanced routing produces better video generation than forcing uniform expert usage.

This paper proposes SplitMoE, a new way to scale video generation models using a split expert architecture that avoids forcing uniform expert usage.

scalingarchitectureefficiency

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Sep 29, 2026

Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen et al.

Long-context evaluation needs to measure both accuracy and computational efficiency—the same model can be dramatically more or less efficient depending on its processing strategy, and current benchmarks don't capture this variation.

This paper introduces LongHarness Bench, a benchmark for evaluating how well language models handle long documents using different processing strategies. The benchmark tests both accuracy and efficiency, requiring models to find relevant information across scattered context and reason strategically rather than reading everything.

evaluationefficiency

Telescopic Language Models

Sep 28, 2026

Zhilin Guo, Boqiao Zhang, Hakan Aktas et al.

You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.

This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.

trainingefficiencyarchitecture

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Sep 28, 2026

Zimo Wang, Junkun Yuan, Angtian Wang et al.

Critic error accumulation is a fundamental bottleneck in distilling video diffusion models; filtering it via projection dramatically improves sample quality without architectural changes or extra computation.

This paper improves video diffusion model distillation by fixing a key problem: critic errors that accumulate during training and degrade sample quality. PDMD uses a simple mathematical projection to filter out these errors while preserving useful learning signals, achieving better video quality with fewer computational steps—all with just a one-line code change.

efficiencytrainingevaluation

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Sep 28, 2026

Chaoqian Ouyang, Ling Yue, Libin Zheng et al.

Token consumption in agentic LLM workflows is unpredictable and can vary 10x+ per task—TokenCast forecasts it accurately by tracking execution segments and context growth, improving budget planning by 14.5% on average.

TokenCast predicts how many tokens an LLM agent will consume during task execution, which varies wildly across runs due to tool use and growing context. It learns cost patterns for each execution step and updates predictions as the agent runs, enabling better budget control without extra LLM calls.

agentsefficiencyevaluation

Neural Harmonic Measure Operator

Sep 28, 2026

Jinjin He, Sinan Wang, Yuchen Sun et al.

NHMO enables fast, geometry-aware PDE solving by learning a domain-specific kernel once, then reusing it for any boundary conditions or sources—eliminating expensive retraining for each new problem.

This paper presents Neural Harmonic Measure Operator (NHMO), a neural network approach that solves elliptic PDEs (like Laplace and Poisson equations) on domains with varying shapes.

reasoningefficiencyarchitecture

How to Loop MoE: Flatten the Experts, Untie the Attention

Sep 28, 2026

Shouren Wang, Chuang Ma, Mohsen Hariri et al.

To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.

This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.

architectureefficiencytraining

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Sep 28, 2026

Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.

KV-streams enables efficient scaling of agentic LLMs to longer horizons by streaming cached computations rather than recomputing them, making long-context RL training practical without sacrificing performance.

This paper introduces KV-streams, a technique that speeds up training of long-horizon agentic language models by streaming the key-value cache forward during context compaction instead of repeatedly refilling it. The method achieves 2.6-5x training speedup while maintaining performance, and shows that the streamed cache can retain information beyond the visible context window.

efficiencytrainingagents

Towards Communication-Efficient Social Intelligence in Language Agents

Sep 28, 2026

Linxiao Gong, Yijie Xu, Tianfu Wang et al.

TACT enables language agents to achieve better social outcomes while using fewer tokens and messages by having specialized teachers refine communication strategy and expression, then distilling improvements into the student agent.

This paper introduces Teacher-Assisted Communication Training (TACT), a method that helps language agents communicate more efficiently during social interactions. TACT improves how agents negotiate and coordinate by having a teacher refine both what agents say (expression) and how they say it (strategy), then distills these improvements back into the student agent.

trainingagentsefficiency

Improving Test-Time Scaling with Adaptive Looped Transformers

Sep 28, 2026

Yichen You, Tianyu Fu, Aosong Feng et al.

Adaptive token-level iteration selection during inference can significantly improve test-time scaling efficiency—TaH2 achieves 53% better accuracy-per-compute gains than fixed looping by intelligently deciding which tokens deserve extra processing passes.

This paper improves how AI models use extra computation time during inference by introducing TaH2, which selectively applies multiple processing passes to tokens that benefit most from them. Unlike standard looped transformers that process every token multiple times, TaH2 learns which tokens need extra iterations, achieving better accuracy gains per unit of compute on math reasoning tasks.

efficiencyreasoning

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Sep 28, 2026

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan et al.

When adapting pre-trained vision models to use linear attention, directly copy MLP weights but distill attention behavior—this simple strategy closes the performance gap between efficient and standard transformers.

This paper shows how to initialize linear Vision Transformers (efficient attention models) using weights from standard Softmax ViTs. The key insight: copy the MLP layers directly since they learn general representations, but use distillation to transfer the attention mechanism since it's operator-specific.

efficiencyarchitecturetraining
trainingefficiency

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Sep 25, 2026

Varun Reddy, Bernardo L. Sabatini, Houman Safaai

Common-mode error in direct feedback alignment causes training stalls by saturating hidden units; this can be prevented by centering batch errors or calibrating the readout baseline, enabling faster learning without changing the core algorithm.

Direct feedback alignment trains neural networks using fixed random error projections, but gets stuck learning near a baseline predictor. The paper identifies that a shared error component across inputs drives hidden units to saturation, slowing learning. Simple fixes like centering errors or adjusting the baseline readout can prevent this collapse and speed up training.

trainingefficiency

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Sep 25, 2026

Quang Nguyen, Hieu Nguyen, Hien Hoang et al.

Building effective AI tutors for non-English regions requires both technical efficiency (faster inference on consumer hardware) and domain-specific learning (accumulating local knowledge from real interactions rather than relying on pre-training).

DeepEdu-v1 is an AI tutoring system for Vietnamese students that runs locally to protect data privacy and avoid hallucinations from Western-trained models. It uses two key innovations: a smarter way to handle long conversations that reduces processing time by 35%, and a self-improving system that learns from past tutoring interactions instead of requiring expensive retraining.

efficiencyapplicationsagents

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Sep 25, 2026

Kunxiong Zhu, Zhihao Shu, Hangyu Zheng et al.

Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.

EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.

efficiencymultimodalarchitecture

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

Sep 24, 2026

Jiabin Qiu, Zixuan Chen, Hongye Cao et al.

World models for planning need to preserve action-discriminative information through training, not just minimize prediction error—this simple insight significantly improves both simulation and real-world robotic control performance.

This paper shows that world models trained only to predict what actually happens often fail at model predictive control, which requires comparing different action choices.

reasoningagentsefficiency

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

Sep 24, 2026

Sudip Bhujel, Shanghao Shi, Ruiquan Huang et al.

Distributed RL agents that share only gradients—not raw data—still leak sensitive trajectory information through temporal correlations; defending against this requires sequence-aware privacy mechanisms, not just per-step protections.

This paper reveals a critical privacy vulnerability in distributed embodied AI systems. When agents send policy gradients to a server instead of raw sensor data, attackers can reconstruct the agent's complete trajectory of observations and actions by analyzing the temporal patterns in these gradients.

safetyagentsefficiency

Rolling-WAM: World Action Models with Rolling Imagination

Sep 24, 2026

Yinghua Zhou, Junjie Ye, Yiqi Zhao et al.

Distributing prediction computation across replanning cycles via a rolling noise schedule achieves 4.5x speedup in robot control latency without sacrificing task performance.

Rolling-WAM speeds up robot control by spreading the computation of predicting future actions and images across multiple planning cycles instead of doing it all at once. Instead of fully planning the entire future from scratch each time, it maintains a sliding window of partially-computed predictions at different stages, letting them gradually refine as new camera data arrives.

efficiencyagentsreasoning

PoEM: Predicting RL Outcomes from Existing Policies

Sep 24, 2026

Kimia Hamidieh, Giannis Daras, Antonio Torralba

You can predict RL outcomes for new reward functions by combining existing trained models mathematically, avoiding the computational cost of retraining—useful when experimenting with different objectives or combining multiple goals.

PoEM predicts what a reinforcement learning model will do with a new reward function by combining existing models trained on different rewards, without running expensive RL training. The method works by finding that RL policies live in a low-rank space that can be reconstructed as a linear combination of existing policies.

trainingefficiencyreasoning

Minimally Invasive Steering of Language Models

Sep 24, 2026

Taha Entesari, Jingyu Zhang, Daniel Khashabi et al.

You can adapt a frozen language model to new rewards at inference time by carefully controlling how much you perturb its hidden states—using Fisher information to measure and limit distributional changes prevents quality degradation.

This paper introduces MISVO, a technique for steering frozen language models at test time by adding vectors to hidden states while minimizing unwanted changes to output quality. Using Fisher information geometry, the method penalizes interventions that distort the token distribution, enabling efficient reward optimization without retraining the model.

efficiencyalignment

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

Sep 24, 2026

Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek et al.

Training latent dynamics models for long-horizon stability requires explicitly optimizing for multi-step rollout accuracy, not just reconstruction—this restructures the solution space in ways that conventional metrics don't capture.

This paper shows that neural surrogate models for physics simulations fail during long predictions not because of poor compression, but because they're trained only to reconstruct data. The authors introduce training techniques—including Koopman operator learning and noise injection—that restructure the latent space to support stable long-horizon forecasting.

efficiencytrainingreasoning

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Sep 24, 2026

Linghua Zhang

Decoupling VLM planning from action execution using a lightweight executor reduces model serving costs dramatically without sacrificing task performance on mobile GUI automation.

Jev-Mobile improves mobile GUI agents by separating planning from execution: a vision-language model makes high-level decisions infrequently, while a lightweight decision model (Jev) handles repeated low-level actions. This cuts inference costs by 73% and execution time by 33% while maintaining 79% task success on Android tasks.

agentsefficiencyapplications

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Sep 22, 2026

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

Diffusion LLMs can be dramatically accelerated by jointly optimizing memory I/O patterns in KV caching and using the model itself for draft-and-verify decoding, rather than treating these optimizations separately.

This paper presents Flash-dLLM, a system that speeds up diffusion language models (an alternative to traditional autoregressive LLMs) by optimizing how the model stores and reuses computation results during inference.

efficiencyarchitecture

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Sep 22, 2026

Trang Nguyen, Eulrang Cho, Bingqing Chen et al.

By compacting context through selective truncation rather than rephrasing, agents can maintain performance on long-horizon coding tasks while reducing inference costs significantly—making test-time scaling more economical.

CliffCompaction is a technique that compresses long conversation histories for AI coding agents by selectively removing less important content while keeping everything else unchanged. This cuts costs by up to 50% while maintaining performance, enabling agents to work on complex coding problems that span millions of tokens across multiple sessions.

efficiencyagentsreasoning

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Sep 22, 2026

Jennifer Williams, Dave Farris, Jeff Farris et al.

AI agents can complete inference engineering tasks locally, but production correctness is much harder—end-to-end serving tests catch failures that other tests miss, exposing a critical gap between development and deployment.

SWE-Serve is a benchmark with 53 real production tasks from SGLang that tests whether AI agents can implement inference serving features correctly—not just locally, but in production. It reveals a major gap: one-third of code changes that pass unit tests fail end-to-end serving tests, showing that current agents struggle with production-grade correctness.

evaluationagentsefficiency

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Sep 22, 2026

Laizhen Li, Jiarui Li, Juanjuan Zhao et al.

You can move repetitive agent control logic from expensive LLM context into persistent, reusable code—cutting inference costs by 74-99% while keeping smaller models effective on complex tasks.

This paper introduces Growing Harness, a method that automatically builds reusable agent control code from task feedback instead of asking language models to repeatedly solve the same control problems. By learning executable code that handles recurring decisions, the approach reduces LLM calls by 76-92% while maintaining or improving task success rates across different model sizes.

agentsefficiencytraining

Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

Sep 22, 2026

Remy Stewart, Olabode Anise, Andrew Hogan et al.

AI design tools deliver real time savings (~20%) in controlled settings, but benefits vary by user expertise—product managers gain more than professional designers, indicating task-dependent value.

Researchers tested whether AI-powered design tools (Figma Make) actually save time by having 50 designers and 50 product managers complete design tasks with and without the tool. They found about 20% faster completion times with AI assistance, especially for non-designers, suggesting these tools may let product managers do more design work themselves.

evaluationapplicationsefficiency

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 22, 2026

Yuanteng Chen, Zhilei Liu, Peisong Wang et al.

On-policy distillation recovers reasoning capabilities in ultra-low-bit quantized models by training on the model's own generated outputs rather than fixed data, fixing the exposure bias problem that causes long-form reasoning to fail.

This paper tackles a critical problem in quantized language models: when you compress models to very low precision (under 3 bits), they lose the ability to do math and coding tasks because errors compound during long generation.

trainingefficiencyreasoning

LoRA-generating hypernetworks for efficient on-device LLM generative personalization

Sep 21, 2026

Sean Augenstein, Li Ding, Jihwan Lee et al.

Hypernetworks can efficiently generate personalized LoRA adapters on mobile devices by mapping user context to weight modifications, avoiding the latency costs of context extension while remaining computationally feasible for resource-constrained devices.

This paper presents a method for personalizing language models on mobile devices by training a hypernetwork that generates customized LoRA (low-rank adaptation) weights based on a user's context.

efficiencyapplications

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Sep 21, 2026

Peng Xia, Rujun Han, Zifeng Wang et al.

Automatically improving agent harnesses through constrained evolution can boost performance while staying generalizable—the key is regularizing the search process to favor reusable mechanisms over task-specific tricks.

This paper presents RRSI, a method for automatically improving LLM agent systems by evolving their harnesses (prompts, tools, memory, control flow) while avoiding overfitting to training tasks. It uses regularization techniques like edit budgets and change filtering to find improvements that generalize to new benchmarks, achieving strong gains on both in-distribution and out-of-distribution tasks.

agentstrainingefficiency

Rare Event Estimation via Iterative Unalignment

Sep 21, 2026

Hanming Yang, Daksh Mittal, Jing Dong et al.

For safe AI deployment, you need to estimate how often catastrophic rare events occur in agent behavior. This paper provides a practical method using importance sampling with learned weight perturbations, achieving massive efficiency gains over naive approaches.

This paper tackles estimating extremely rare event probabilities in AI agent trajectories—events so uncommon that standard Monte Carlo sampling is impractical. The authors develop a new importance sampling method that tweaks a language model's weights to generate more likely rare events, using gradient-based optimization to search efficiently.

safetyevaluationefficiency
architectureefficiencytraining

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

Sep 17, 2026

Nitish Dashora, Douglas Chen, Idan Shenfeld et al.

By distilling VLM-identified task-salient information into a learned latent token during training, robots can efficiently handle long-horizon tasks at deployment time without expensive in-the-loop reasoning.

This paper introduces workspace tokens, a lightweight memory system for robotic manipulation that learns which task-relevant information to remember during training using a vision-language model, then uses this compressed memory at deployment without needing expensive VLM queries.

efficiencyagentstraining

How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?

Sep 17, 2026

Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi et al.

Pretraining neural PDE surrogates provides significant data efficiency gains (2-3x fewer samples needed), but this benefit shrinks or reverses when the target task involves different physics modeling than the pretraining source.

This paper investigates how pretraining neural networks to simulate fluid dynamics (PDE surrogates) helps when switching to new airfoil designs or physics models. The authors show that pretraining benefits depend on three factors: how much target data you have, how diverse that data is, and whether the source and target use different physics models.

trainingefficiencyevaluation

Score Centering Stabilizes Off-policy Reinforcement Learning

Sep 17, 2026

Martin Marek, Max Ryabinin

Score centering is a lightweight, composable fix for training-inference mismatch in RL that works by correcting accumulated bias rather than trying to eliminate the mismatch entirely—making it practical for large language models.

This paper identifies drift—a persistent bias that accumulates during training—as the main cause of instability when reinforcement learning models behave differently during training versus deployment. The authors propose 'score centering,' a simple mathematical correction that stabilizes training without requiring expensive changes to the inference engine.

trainingefficiencyalignment

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

Sep 17, 2026

Xin Chen, Sen Chen, Yujuan Ding et al.

By analyzing the geometric properties of diffusion-based trajectory generation, you can detect prediction uncertainty and dynamically adjust action planning horizons without retraining—improving robotic control performance by up to 8.7 percentage points.

This paper introduces GeoAAC, a method that dynamically adjusts how many steps ahead a robot should plan based on task difficulty. Instead of using a fixed planning horizon, it analyzes the geometry of the prediction process to detect when the model is uncertain, then shortens or lengthens the planning window accordingly.

reasoningefficiencyagents

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Sep 17, 2026

Hanchu Zhou, Brendan Lynch, Raman Goyal et al.

By treating vision and tactile signals at different timescales in a lightweight architecture, you can build fast tactile-aware robot controllers (11.9ms latency) that outperform larger models on real contact-rich tasks.

This paper presents Agile-WAM, a tactile world action model that predicts future robot states and actions for contact-rich manipulation tasks.

architecturemultimodalefficiency

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

Sep 17, 2026

Damiano Da Col, Maximilian Igl, Peter Karkus et al.

Using a privileged teacher (trained on high-level inputs like maps and bounding boxes) to supervise a camera-based student during closed-loop fine-tuning is 1000× more sample-efficient than direct RL post-training for autonomous driving.

OPTED improves autonomous driving policies by using a privileged teacher trained with reinforcement learning to guide a camera-based student model during closed-loop fine-tuning. This approach avoids expensive direct RL training in simulation while keeping the policy close to human demonstrations, achieving 1.6-9.5× improvements in driving performance.

trainingefficiencyapplications

dQwen3.5: Hybrid-Attention Diffusion Language Models

Sep 17, 2026

Anton Xue, Litu Rout, Aditya Akella et al.

Hybrid architectures combining attention and RNNs are surprisingly efficient starting points for diffusion language models, reaching comparable performance with 2x faster training than full-attention models.

This paper shows that hybrid-attention language models (mixing attention and RNN layers) can be efficiently adapted into diffusion language models, which generate text in any order rather than left-to-right. The dQwen3.5 models reach the same training loss in half the tokens compared to full-attention baselines, while maintaining strong performance in parallel decoding.

architecturetrainingefficiency

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Sep 17, 2026

Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.

Zero-shot sim-to-real transfer for autonomous driving is achievable by training on a consistent semantic representation in simulation and applying the same representation to real sensor data, eliminating the need for manual policy adaptation.

MILER is a reinforcement learning framework for autonomous driving that bridges simulation and real-world deployment without manual tuning. It trains policies in simulation using a semantic bird's-eye-view representation, then transfers them to real vehicles by converting camera and LiDAR data into the same representation format.

trainingapplicationsefficiency

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Sep 17, 2026

Haocheng Xi, Yiming Xie, Hexu Zhao et al.

Hybrid attention architectures that combine local Softmax with linear memory can dramatically accelerate video generation without sacrificing quality—enabling practical real-time video synthesis on modern hardware.

Video DeltaNet combines local attention with efficient linear memory to speed up video diffusion models. By mixing Softmax attention for fine details with a new Video Delta Attention mechanism for long-range context, it generates high-quality videos 14.5x faster than baseline models while maintaining visual quality.

efficiencyarchitecturemultimodal

On-Demand Attention: Language Models Know When to Recall

Sep 17, 2026

Haibo Feng, Ruiqi Liang, Hanyang Peng et al.

Pretrained models already know which tokens matter for the next prediction—you can train a lightweight module to detect this and skip expensive full-context reads, cutting inference cost without retraining the base model.

This paper shows that language models can predict when they need to read their full context history during generation. The authors introduce On-Demand Attention (ODA), which uses a small trained module to decide when to use expensive global attention versus cheaper local attention, reducing computation while maintaining quality on long documents.

efficiency

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Sep 17, 2026

Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.

Activation steering can be automated to find optimal intervention points in transformers, but practitioners should be aware that stronger steering increases vulnerability to prompt injection attacks—a critical concern for deployed agent systems.

Deep Noir automatically discovers where and how to steer LLM activations to change model behavior without retraining. Using a technique called Logit Lens to find optimal intervention points, the system improves spam detection by up to 42 percentage points across different model sizes.

safetyefficiency

RISC-V and machine learning: a survey

Sep 17, 2026

Shriman Keshri, Apparna Singh, Chinmaya Kumar Palo et al.

RISC-V offers a customizable, open alternative to proprietary chips for ML, but realizing its potential requires standardized extensions, better toolchain maturity, and specialized accelerator designs tailored to neural network workloads.

This survey examines how RISC-V, an open-source processor architecture, is being adapted for machine learning workloads. It analyzes existing implementations, software tools, and real-world applications, identifying strengths in energy efficiency and instruction extensions while highlighting challenges like fragmentation and standardization gaps.

architectureefficiency

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Sep 17, 2026

Zimu Han, Yiming Zeng, Jiyao Zhang et al.

You can improve robot manipulation models through human feedback without robot execution by using a handheld interface to detect when the policy is uncertain and identifying which parts of demonstrations are most important for learning.

This paper presents HIL-UMI, a method for improving vision-language-action robot models without needing a physical robot during training. Instead of repeatedly running the robot to collect new data, humans demonstrate tasks using a handheld interface while the system queries the current policy and intelligently decides when to collect new examples based on policy uncertainty and task progress.

trainingagentsefficiency

COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression

Sep 17, 2026

Zewen Yang, Xiaobing Dai, Zhenxiao Yin et al.

When building distributed sensor networks with incomplete data, combining state observers with cooperative Gaussian Process learning can accurately estimate both what's happening in the system and how it behaves, with proven error bounds.

This paper presents a method for distributed systems with multiple sensors to estimate both system states and unknown dynamics when only partial measurements are available. It combines observer-based estimation with online Gaussian Process regression across a network, includes a smart data collection strategy, and provides theoretical guarantees on estimation accuracy.

reasoningefficiency