Research Feed Page 18

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Computer Vision / Benchmarks & Evals By Sanjay Bhargav Dharavath 2026-08-12 1
MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

The paper introduces the MV2 dataset to test how well autonomous driving vision models generalize when generating views from different camera perspectives.

Benchmarks & Evals By Jophin John 2026-08-13 1
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

The paper introduces BavGround, a benchmark designed to evaluate how well large language models understand Bavarian regional culture and dialects.

Multimodal / Training & Fine-Tuning By Hamza Shafiq 2026-08-13 1
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

The paper introduces CardioState-JEPA, a unified cardiac foundation model that learns a single shared representation across heterogeneous signals like electrocardiography, photoplethysmography, and phonocardiography by accounting for physiological delays.

Computer Vision / Benchmarks & Evals By Yunsung Chung 2026-08-13
Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology

The authors developed a clinical world model that uses longitudinal data and latent state transitions to predict long term cardiac surgery outcomes without requiring follow up imaging at inference.

Computer Vision / Benchmarks & Evals By Zuzanna A. Wakefield-Skórniewska 2026-08-13
Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

The paper demonstrates a method to generate retinal images conditioned on specific clinical metadata while evaluating the gap between synthetic and real-world representations.

Robotics / Benchmarks & Evals By Dairu Liu 2026-08-13
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

The paper introduces a standardized benchmark and an automated preference-based scoring model to address flaws in existing humanoid motion tracking evaluation metrics.

Benchmarks & Evals / Training & Fine-Tuning By Alexander Bräuer 2026-08-13
Foundation models for movement data: Are they ready for prime-time?

The paper benchmarks four foundation models against traditional supervised learning baselines across 19 movement and health monitoring tasks to determine their practical effectiveness.

Safety & Alignment / Efficiency & Inference By Yikai Xu 2026-08-13
Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

The paper introduces Wasserstein Filtering, a method to recover clean data distributions by selecting a subset of samples that maximizes the distance from contaminated outliers.

Benchmarks & Evals / Agents By Justin Zhao 2026-08-12
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

The researchers created a framework to test if AI models acting as judges maintain consistent opinions when subjected to adversarial pressure or repetitive questioning.

Computer Vision / Efficiency & Inference By Christos Chatzisavvas 2026-08-13
AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage

AmalthAI is an open source, dockerized machine learning platform that enables non technical cultural heritage experts to perform dataset management, training, and inference.

Computer Vision By Yanming Yang 2026-08-13
GS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model Priors

The researchers developed a method that combines foundation models with Gaussian splatting to reconstruct detailed 3D scenes from limited compressive image data.

Computer Vision / Benchmarks & Evals By Imtiaz Ul Hassan 2026-08-13
Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

The FineX framework improves fine-grained human action recognition by fusing distinct visual and pose signals through a mixture-of-experts architecture.

Robotics / Multimodal By Ge Yan 2026-08-13
Flex-π: A Multi-Stream World-Action Model with Compute Flexibility

Flex-Pi is a world-action model that achieves superior robot control by integrating 3D geometry and object semantics alongside traditional visual data.

Robotics / Agents By Rafal Robert Karpinski 2026-08-13
Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement

The researchers developed a dual-branch neural network that separates environmental and social visual context to help robots learn appropriate actions without forgetting previous knowledge.

Benchmarks & Evals / Efficiency & Inference By Paul Savala 2026-08-13
Jointly Predicting Courses and Grades Using a Transformer-Based Model

The TRACE model uses a Transformer architecture to simultaneously forecast which courses a student will take and what grades they will earn.

Computer Vision / Benchmarks & Evals By Mohamed Ahmed Mohamed 2026-08-13
A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline

The Clear2Fog pipeline enables generation of physically realistic fog across RGB and LiDAR modalities to improve the accuracy of autonomous vehicle perception systems.

Computer Vision / Efficiency & Inference By Minghui Guo 2026-08-13
V-RAE: Rethinking Video Latent Spaces for Generation

V-RAE improves video generation quality by creating compact latent spaces from frozen vision foundation models rather than relying on traditional pixel-level reconstruction.

Multimodal / Benchmarks & Evals By Wafa Al Ghallabi 2026-08-13
How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?

Researchers built a benchmark and evaluation framework to test how well vision-language models interpret changes in MRI scans over time for clinical decision-making.

Multimodal / Benchmarks & Evals By Yupan Ding 2026-08-13
LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

LongEarth-R1 enhances long duration satellite image analysis by aligning vision language models with structured temporal reasoning and reward based feedback.

Multimodal / Safety & Alignment By Ashok Urlana 2026-08-13
MapRoute++: Surrogate-Guided Semantic Routing for Visual Concept Unlearning

MapRoute++ provides a system for removing specific visual concepts from diffusion models using input-conditioned routing to redirect target tokens toward safe surrogates.