Research Feed

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Computer Vision / Robotics By Yueen Ma 2026-08-26
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

The 4DGS-WAM model enables future video prediction by explicitly decomposing scenes into dynamic objects and a static background using 4D Gaussian Splatting.

Computer Vision / Efficiency & Inference By Yogesh Kumar 2026-08-25
Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

The paper introduces a causal state-space model for video anomaly detection that runs directly on edge hardware without needing frame buffering.

Benchmarks & Evals / Computer Vision By Hao Chen 2026-08-25
What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

The paper introduces ZID, a new evaluation metric for generative models that identifies and ranks failures in image generation where traditional metrics like FID fail.

Multimodal / Computer Vision By Muhammad Asad Ali 2026-08-25
MoTE: Mixture of Task Experts for Multi-Task Video Understanding

The paper introduces a Mixture of Task Experts architecture that uses task-specific modules within a video-language decoder to improve performance across diverse video understanding tasks.

Computer Vision / Training & Fine-Tuning By Wenxuan Shen 2026-08-25
Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

The authors introduce GameCleaner and the Game2World engine to remove distracting interface elements from game footage to improve the training of world models.

Computer Vision / Reinforcement Learning By Tong Wang 2026-08-24 3
From Generation to Simulation: How Far Are World Models from Being True Simulators?

The researchers audited 200 world model papers to determine how their capabilities measure up against the structural requirements of professional physical simulators.

Multimodal / Computer Vision By Nan Duan 2026-08-24
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.

Multimodal / Computer Vision By Songchun Zhang 2026-08-24 19
EchoWM: Open and Enterable Omnimodal World Models

EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.

Computer Vision / Training & Fine-Tuning By Marko Haralović 2026-08-21
When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning

The paper identifies that representational drift in the decoder and output layers significantly impacts out of distribution performance when fine-tuning the MedSAM medical image segmentation foundation model.

Computer Vision / Efficiency & Inference By Yunze Tong 2026-08-21 11
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

InfinityEdit uses a lightweight adapter to enable consistent, long-term video editing for continuous data streams.

Training & Fine-Tuning / Computer Vision By Yansen Han 2026-08-20
Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

The paper introduces ThermoDPO, a new training method that stabilizes generative model output by preventing reward-driven distortion of the underlying data distribution.

Computer Vision By Muhammad Asad Ali 2026-08-20
HandMvNet: Real-Time 3D Hand Pose Estimation Using Multi-View Cross-Attention Fusion

HandMvNet uses multi-view cross-attention to estimate 3D hand poses from multiple camera angles without requiring complex calibration.

Computer Vision / Efficiency & Inference By Yudong Jin 2026-08-20
4DAnyone: Create Anyone in 4D from a Casual Monocular Video

4DAnyone reconstructs high fidelity 4D human models from single casual videos by using geometric guidance and optimized multiview diffusion techniques to prevent structural drift.

Computer Vision / Benchmarks & Evals By Yigit Ekin 2026-08-20
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal

The paper introduces a new benchmark and evaluation protocol to accurately measure whether video object removal tools correctly eliminate both the object and its associated physical side effects like shadows and reflections.

Computer Vision / Efficiency & Inference By Jiayi Song 2026-08-18 21
EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

EditBridge uses a diffusion bridge framework to enable 4K image editing by reducing attention complexity to linear scaling.

Agents / Computer Vision By Rui-Huan Wang 2026-08-18
aDSL: Agentic 3D Creation via Joint Agent-Program Design

The paper introduces aDSL, a domain-specific language and agent-based system designed to reliably convert natural language instructions into functional 3D programs and geometry.

Multimodal / Computer Vision By Jinsheng Quan 2026-08-14 1
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

The paper introduces SPARGen, an instruction-conditioned multimodal generative framework that unifies 3D reconstruction, dense correspondence estimation, and spatial reasoning without task-specific prediction heads.

Computer Vision / Training & Fine-Tuning By Mohamed Abdelsamad 2026-08-14
GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

GhostPoint improves 3D object detection by training models to predict the hidden structure of objects that are partially occluded in LiDAR sensor data.

Computer Vision / Efficiency & Inference By Mahesh Reddy 2026-08-14
MagnifiQ: Patch-aware Text Guided Progressive Upscaling for High-Resolution Image Restoration

MagnifiQ uses a modular patching architecture and LLM-based text prompts to perform efficient high resolution image restoration.

Computer Vision By Mahdi Saberi 2026-08-14
UMPIRE-Net: Unrolled Magnitude-Phase Regularization Network for Accelerated MRI

The paper introduces UMPIRE-Net, an accelerated MRI reconstruction method that independently regularizes magnitude and sign components to improve performance in partial Fourier imaging.