Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The 4DGS-WAM model enables future video prediction by explicitly decomposing scenes into dynamic objects and a static background using 4D Gaussian Splatting.
The paper introduces a causal state-space model for video anomaly detection that runs directly on edge hardware without needing frame buffering.
The paper introduces ZID, a new evaluation metric for generative models that identifies and ranks failures in image generation where traditional metrics like FID fail.
The paper introduces a Mixture of Task Experts architecture that uses task-specific modules within a video-language decoder to improve performance across diverse video understanding tasks.
The authors introduce GameCleaner and the Game2World engine to remove distracting interface elements from game footage to improve the training of world models.
The researchers audited 200 world model papers to determine how their capabilities measure up against the structural requirements of professional physical simulators.
The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.
EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.
The paper identifies that representational drift in the decoder and output layers significantly impacts out of distribution performance when fine-tuning the MedSAM medical image segmentation foundation model.
InfinityEdit uses a lightweight adapter to enable consistent, long-term video editing for continuous data streams.
The paper introduces ThermoDPO, a new training method that stabilizes generative model output by preventing reward-driven distortion of the underlying data distribution.
HandMvNet uses multi-view cross-attention to estimate 3D hand poses from multiple camera angles without requiring complex calibration.
4DAnyone reconstructs high fidelity 4D human models from single casual videos by using geometric guidance and optimized multiview diffusion techniques to prevent structural drift.
The paper introduces a new benchmark and evaluation protocol to accurately measure whether video object removal tools correctly eliminate both the object and its associated physical side effects like shadows and reflections.
EditBridge uses a diffusion bridge framework to enable 4K image editing by reducing attention complexity to linear scaling.
The paper introduces aDSL, a domain-specific language and agent-based system designed to reliably convert natural language instructions into functional 3D programs and geometry.
The paper introduces SPARGen, an instruction-conditioned multimodal generative framework that unifies 3D reconstruction, dense correspondence estimation, and spatial reasoning without task-specific prediction heads.
GhostPoint improves 3D object detection by training models to predict the hidden structure of objects that are partially occluded in LiDAR sensor data.
MagnifiQ uses a modular patching architecture and LLM-based text prompts to perform efficient high resolution image restoration.
The paper introduces UMPIRE-Net, an accelerated MRI reconstruction method that independently regularizes magnitude and sign components to improve performance in partial Fourier imaging.