Research Feed Page 18
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces the MV2 dataset to test how well autonomous driving vision models generalize when generating views from different camera perspectives.
The paper introduces BavGround, a benchmark designed to evaluate how well large language models understand Bavarian regional culture and dialects.
The paper introduces CardioState-JEPA, a unified cardiac foundation model that learns a single shared representation across heterogeneous signals like electrocardiography, photoplethysmography, and phonocardiography by accounting for physiological delays.
The authors developed a clinical world model that uses longitudinal data and latent state transitions to predict long term cardiac surgery outcomes without requiring follow up imaging at inference.
The paper demonstrates a method to generate retinal images conditioned on specific clinical metadata while evaluating the gap between synthetic and real-world representations.
The paper introduces a standardized benchmark and an automated preference-based scoring model to address flaws in existing humanoid motion tracking evaluation metrics.
The paper benchmarks four foundation models against traditional supervised learning baselines across 19 movement and health monitoring tasks to determine their practical effectiveness.
The paper introduces Wasserstein Filtering, a method to recover clean data distributions by selecting a subset of samples that maximizes the distance from contaminated outliers.
The researchers created a framework to test if AI models acting as judges maintain consistent opinions when subjected to adversarial pressure or repetitive questioning.
AmalthAI is an open source, dockerized machine learning platform that enables non technical cultural heritage experts to perform dataset management, training, and inference.
The researchers developed a method that combines foundation models with Gaussian splatting to reconstruct detailed 3D scenes from limited compressive image data.
The FineX framework improves fine-grained human action recognition by fusing distinct visual and pose signals through a mixture-of-experts architecture.
Flex-Pi is a world-action model that achieves superior robot control by integrating 3D geometry and object semantics alongside traditional visual data.
The researchers developed a dual-branch neural network that separates environmental and social visual context to help robots learn appropriate actions without forgetting previous knowledge.
The TRACE model uses a Transformer architecture to simultaneously forecast which courses a student will take and what grades they will earn.
The Clear2Fog pipeline enables generation of physically realistic fog across RGB and LiDAR modalities to improve the accuracy of autonomous vehicle perception systems.
V-RAE improves video generation quality by creating compact latent spaces from frozen vision foundation models rather than relying on traditional pixel-level reconstruction.
Researchers built a benchmark and evaluation framework to test how well vision-language models interpret changes in MRI scans over time for clinical decision-making.
LongEarth-R1 enhances long duration satellite image analysis by aligning vision language models with structured temporal reasoning and reward based feedback.
MapRoute++ provides a system for removing specific visual concepts from diffusion models using input-conditioned routing to redirect target tokens toward safe surrogates.