Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.
Yuxuan Hu, Weikang Shi, Yang Bo et al.
Streaming VLMs fail at perceiving fast events in videos—the best model scores only 50.7% on FastBench—because they can't adaptively determine when to sample frames densely, revealing a fundamental gap between sparse uniform sampling and what's actually needed for dynamic understanding.
FastBench is a benchmark for evaluating how well streaming video language models can understand fast-moving events in real-world videos. Current models struggle with this task because they must balance limited context budgets across temporal history, spatial resolution, and frame sampling rates.
Luping Liu, Bingyi Kang, Yifan Wang et al.
By combining foundation models with diverse training data, you can build correspondence matching systems that work across both classical vision tasks and modern generative applications like image editing.
FreeMatching is a framework for matching visual correspondence across images that goes beyond traditional motion and geometry assumptions. It combines generative and semantic models trained on diverse data sources to handle challenging transformations in image editing and generation tasks, while maintaining competitive performance on standard benchmarks.
Qiushi Han, Keya Hu, Linlu Qiu et al.
Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.
VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.
Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.
By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.
GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
Most hate speech detection datasets focus on English; this dataset enables models to learn culturally-specific patterns in Mexican Spanish video content, which is essential since hate speech is deeply tied to local context and language nuances.
MexHat is a new video dataset with ~1,000 annotated clips for detecting hate speech in Mexican Spanish. It addresses the lack of non-English resources by capturing linguistic and cultural context specific to Mexico, with annotations for both general categories (offensive vs. hate speech) and fine-grained hate speech subtypes.
Kunxiong Zhu, Zhihao Shu, Hangyu Zheng et al.
Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.
EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.
Kevin Qu, Tao Sun, Massimiliano Viola et al.
By processing multiple sparse observations together rather than single images, and training with a motion-range supervision objective, the model better understands articulation without needing large labeled datasets—it generates synthetic training data procedurally.
FAMOS is a feed-forward model that predicts how articulated objects (like doors, drawers) move and which parts are movable from multiple sparse 3D views.
Ji Xie, Dewei Zhou, Xinyu Huang et al.
By combining real image supervision with pure-color anchors and a shared hex-prompt interface, you can now control exact object colors in both image generation and editing tasks with professional-grade precision.
Paint-Anything enables precise color control in AI image generation and editing by letting users specify exact colors using hex codes (like #FF5733). The system learns to understand hex values through a training dataset of 500K images with color labels and pure-color reference images, achieving much better color accuracy than existing methods.
William Zhou, Mayukha Siripuram, Xiao Yan et al.
Small deployable VLMs can identify species above chance but degrade sharply on real camera-trap images; specialized training data outweighs model scale, and all models occasionally hallucinate fake species names when uncertain.
This paper evaluates whether small vision-language models (2-8B parameters) suitable for edge deployment can accurately identify animal species from camera-trap photos.
Akshaj Gupta, Hwi Joo Park, Andrea Guzman et al.
This is the first guitar transcription system that produces both fingering positions and expressive techniques directly from audio, achieving significant improvements over prior work and handling real-world noisy recordings.
TART is a four-stage system that converts guitar audio into tablature (written notation showing which strings and frets to play). It tackles three real problems: recognizing guitar techniques like slides and bends, correctly identifying which string-fret combinations produce each note, and handling noisy real-world recordings.
Linzhan Mou, Jiahui Lei, Zhiyang Dou et al.
For the first time, a single model can animate any skeleton topology (bipedal, quadrupedal, insects, etc.) from text alone—no per-character fine-tuning or reference motions needed at inference time.
UniMate is a foundation model that generates realistic motion for any 3D character skeleton from text descriptions, without needing to retrain for each new character type. It uses a specialized neural architecture that understands skeleton structure through graph-based attention mechanisms, and was trained on a diverse dataset of 13,000+ motion sequences across different creature types.
Ji Soo Lee, Xilun Chen, Pierce Chuang et al.
Most LLMs struggle with real-world wearable health reasoning (19.6%-72.9% accuracy), revealing a significant gap between general language abilities and the specialized reasoning needed to interpret longitudinal physiological data.
WearableQA is a benchmark with 4,084 multiple-choice questions testing whether AI models can reason about real health data from wearables. It uses 500 days of measurements from 200 actual users, including heart rate, sleep, and blood tests, organized into 16 question types that test both data computation and health interpretation skills.
Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor et al.
State-of-the-art vision-language models fail at interpreting real scientific artifacts that domain experts find straightforward, highlighting a critical gap between general image understanding and specialized scientific visual reasoning needed for practical biotech applications.
VIALS is a benchmark of 161 visual question-answering tasks based on real scientific images (gel blots, microscopy, flow cytometry plots, etc.) from biotech workflows. Current vision-language models struggle with these domain-specific images despite excelling at natural images, revealing gaps in scientific reasoning that limit their usefulness in professional life sciences research.
Haonan Jia, Shichao Dong, Zenghui Sun et al.
Retrieval can guide reinforcement learning to improve vision-language models for image captioning by helping identify and correct errors, outperforming standard supervised fine-tuning approaches.
This paper proposes Re³Cap, a method that uses retrieval-guided reasoning to improve image captioning with reinforcement learning. Instead of just fine-tuning models, it retrieves similar images and captions to help identify and fix errors (hallucinations and omissions) in generated descriptions, achieving better results than supervised fine-tuning without needing extra labeled data.
Bobo Li, Hao Fei, Tianjie Ju et al.
Direct perception of raw scientific data—not just text summaries—is critical for AI systems to conduct rigorous, evidence-grounded research. OmniScientist shows that multimodal input improves all aspects of automated scientific discovery.
OmniScientist is an AI system that conducts scientific research across multiple disciplines by directly processing raw data in many formats—images, videos, audio, 3D structures, tables, and more.
Yunsung Chung, Yingshuo Liu, Abboud F. Hassan et al.
Clinical prediction improves when models treat recovery as an evolving process rather than a static snapshot—incorporating asynchronous post-procedure events and imaging can significantly boost outcome forecasting accuracy for cardiac interventions.
This paper presents a clinical AI model that predicts post-surgery outcomes in heart rhythm procedures by tracking how a patient's condition evolves over time.
Youjun Zhao, Alex Warren, Gary K. L. Tam et al.
Mirror reflection generation requires modeling two distinct challenges—semantic consistency (what reflects) and geometric accuracy (how it's arranged)—which can be addressed by distilling relational knowledge and learning spatial transformations.
MirrorWorld tackles the problem of generating realistic mirror reflections in videos by teaching diffusion models to understand what scene content should appear in mirrors and how it should be spatially arranged. The method uses semantic guidance from visual foundation models and geometric transformation learning to ensure reflections are consistent with their surroundings.
Zixuan Lan, Luzhe Sun, Matthew R. Walter et al.
VLMs have a critical weakness: they often trust learned world knowledge over what's actually shown in images, and SABRE provides a reusable framework to systematically identify and measure such failures.
SABRE is an automated pipeline that creates stress tests for vision-language models by converting task designs into images and question-answer pairs. It tests whether VLMs rely on visual evidence or learned assumptions about the world, revealing that current models struggle significantly (17.8-31.3% accuracy) when images contradict expectations.