Research Feed

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Efficiency & Inference / Multimodal By Xiaoyu Zhan 2026-08-27 7
Magpie: Real-Time World Renderer for Interactive Games

Magpie is a real-time renderer that uses foundation models to generate game visuals by processing white-box frames as a continuous denoising condition.

Multimodal / Benchmarks & Evals By Yuandong Pu 2026-08-27
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

The paper introduces PAWBench, a new evaluation framework that measures how well video generation models predict the physical outcomes of various scenarios.

Agents / Multimodal By Yuqiang Lin 2026-08-26
TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

TAU-Agent is an agentic framework designed to improve traffic anomaly detection by using retrieval-augmented generation to integrate video descriptions and object trajectories.

Multimodal / Benchmarks & Evals By Nabaraj Subedi 2026-08-26
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

PlanSightRAG leverages a vision-first multimodal framework to automate compliance checking for complex civil engineering drawings.

Robotics / Multimodal By Zaibin Zhang 2026-08-26
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

The researchers developed a vision-language-action model designed to improve how multiple robotic arms collaborate on complex tasks by using techniques that enforce role-agnostic instruction following.

Multimodal / Benchmarks & Evals By Mengdi Qin 2026-08-26
RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps

RotDroid uses a vision-language model to detect GUI rotation bugs by comparing visual states between portrait and landscape orientations.

Agents / Multimodal By Yiwen Chen 2026-08-26
Code World Model: Coding Agent as World Brain

The paper introduces a framework where a coding agent generates deterministic code to manage world state, which then guides a video model to maintain visual consistency in simulated environments.

Agents / Multimodal By Min Zeng 2026-08-26
VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

The researchers developed an agentic, iterative framework called VISA to generate high-quality training data for multimodal models by using feedback-driven loops instead of static one-pass pipelines.

Reasoning / Multimodal By Junxiang Xu 2026-08-26
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

The authors introduce VBVR-Pro, a suite designed to improve visual reasoning capabilities in models by using verifiable generative tasks.

Multimodal / Computer Vision By Muhammad Asad Ali 2026-08-25
MoTE: Mixture of Task Experts for Multi-Task Video Understanding

The paper introduces a Mixture of Task Experts architecture that uses task-specific modules within a video-language decoder to improve performance across diverse video understanding tasks.

Agents / Multimodal By Jiatong Li 2026-08-25
Molecular LLM Agents: From Architectural Design to Scientific Autonomy

The paper introduces a framework and an autonomy classification system for deploying Large Language Model agents in scientific molecular discovery workflows.

Multimodal / Benchmarks & Evals By Emre Kuru 2026-08-24
Evaluating Modern RAG: Textual, Multimodal, Dense, and Late Interaction Pipelines

The paper evaluates various retrieval-augmented generation pipelines, finding that multimodal vision-based approaches significantly outperform text-based methods despite introducing higher latency and storage costs.

Multimodal / Benchmarks & Evals By Junjie Zhou 2026-08-25 36
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

The paper introduces WeMM-Embedding, a series of multimodal models based on Qwen3.5 that achieve state-of-the-art performance on retrieval benchmarks and demonstrate consistent gains in production applications.

Multimodal / Benchmarks & Evals By Jiapeng Li 2026-08-24
IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

The paper introduces IntentQA and the X-CaVIR framework to help models infer latent human intentions in video content through cognitive context reasoning.

Benchmarks & Evals / Multimodal By Hang Wang 2026-08-24
FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

The authors introduce FinixDoc, an agentic parsing system powered by a specialized vision-language model, to improve accuracy and structural consistency in real-world financial document processing.

Multimodal / Computer Vision By Nan Duan 2026-08-24
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.

Multimodal / Computer Vision By Songchun Zhang 2026-08-24 19
EchoWM: Open and Enterable Omnimodal World Models

EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.

Multimodal / Efficiency & Inference By Zhifei Chen 2026-08-24
ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld enables interactive video generation with consistent long-term spatial memory by using an efficient chunk-based caching strategy.

Robotics / Multimodal By Siyuan Ma 2026-08-20
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

DECOWAM is a new model architecture that optimizes how legged robots coordinate whole body actions with visual environment predictions.

Efficiency & Inference / Multimodal By Yangshuai Liu 2026-08-21
Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

The researchers developed a text anchor method to enable high-speed reference image caching in diffusion transformers without sacrificing model performance.