Research Feed
Search source-linked summaries of recent AI research.
Research results
The paper introduces a framework and auditing protocol to fix automatic speech recognition systems that systematically fail low-resource, Indigenous, and non-standard language varieties.
The paper introduces a human-inspired pipeline using Object-Centric Learning and graph networks to model image composition more effectively than frozen foundation models.
Researchers developed an audio to score transcription system using a Transformer architecture trained on the new SheetSage A2S dataset to bridge the gap between classical and popular music transcription.
The researchers developed a method that enables robots to systematically cover unknown surfaces by combining real-time geometric reconstruction with ergodic trajectory generation.
The paper presents a method for robots to replicate human handwriting by learning trajectories from human demonstrations using Gaussian models.
LAWM-3D enables robots to learn 3D-aware actions by training world models on human videos using a new geometric alignment method.
The PrivacyPeek framework benchmarks how LLM-based agents acquire sensitive information beyond their necessary operational scope during task completion.
The paper evaluates whether super-resolution techniques applied to MRI scans risk erasing real small white-matter lesions or hallucinating false ones.
The paper provides a formal framework for determining when distribution classes are learnable and quantifies the exact performance cost of replacing interactive queries with static ones.
The researchers applied self-pretraining to transformer models to boost accuracy in medical time series classification without requiring external data.
CogVis introduces a modular perception and memory framework that decouples image analysis from query processing to improve speed and accuracy in open-vocabulary change detection.
The CFGPNet framework improves multispectral object detection by optimizing feature interaction and gradient flow while reducing computational overhead.
The researchers developed Skewon, an algorithm that provides an exact closed-form solution for the Stiefel Muon optimization problem.
The paper investigates whether vision language models serving as robot controllers truly rely on visual inputs or merely leverage non visual shortcuts like simulator rewards.
BendTwin adds bending stiffness to spring mass models to improve the stability and accuracy of physical reconstructions from sparse video data.
Robust-WAM introduces a method to align video-generation model latent spaces with semantic features, enabling robots to handle visual changes more reliably.
The paper introduces Prior-SG, a framework that casts scene graph generation as a probabilistic alignment problem using task-conditioned priors and visual-geometric feature fusion to handle arbitrarily structured environments.
The researchers developed a transformer architecture that estimates physical pressure during hand-object interactions using only monocular video input.
The paper introduces a touchscreen interface and hybrid control architecture that improves precision and reduces cognitive load during robotic surface interaction tasks.
RxnCLF is a new reaction foundation model that uses condensed graph structures and contrastive learning to better capture chemical transformation information for reactivity prediction.