Research Feed

Search source-linked summaries of recent AI research.

Active filters Date: August 2026
Filters

Browse by date

Topics

Resource filters

Sort options

Research results

Computer Vision / Training & Fine-Tuning By Manuel Laufer 2026-08-06
Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation

The paper introduces a synthetic data generation framework that allows AI to accurately assess patient poses for X-rays by training on generated depth images and radiographs.

Efficiency & Inference / Benchmarks & Evals By Dae-Jin Lee 2026-08-06
Learning Latent Memory States from Longitudinal Athlete Monitoring Data

The paper introduces a statistical method to represent an athlete's historical data as a reusable latent memory table to improve performance tracking and prediction.

Benchmarks & Evals / Safety & Alignment By Jay L. Cunningham 2026-08-06
Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

The paper introduces a framework and auditing protocol to fix automatic speech recognition systems that systematically fail low-resource, Indigenous, and non-standard language varieties.

Computer Vision / Efficiency & Inference By Fatemeh Behrad 2026-08-06
Learning visual representations for compositional analysis of artworks and photographs

The paper introduces a human-inspired pipeline using Object-Centric Learning and graph networks to model image composition more effectively than frozen foundation models.

Multimodal / Benchmarks & Evals By Eoin Cummins 2026-08-06
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

Researchers developed an audio to score transcription system using a Transformer architecture trained on the new SheetSage A2S dataset to bridge the gap between classical and popular music transcription.

Robotics By Stefan Schneyer 2026-08-06
ErgoSurf: Ergodic Control for the Coverage of Unknown Surfaces

The researchers developed a method that enables robots to systematically cover unknown surfaces by combining real-time geometric reconstruction with ergodic trajectory generation.

Robotics / Benchmarks & Evals By Alperen Kenan 2026-08-06
Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

The paper presents a method for robots to replicate human handwriting by learning trajectories from human demonstrations using Gaussian models.

Robotics / Computer Vision By Jiarui Yang 2026-08-06
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

LAWM-3D enables robots to learn 3D-aware actions by training world models on human videos using a new geometric alignment method.

Agents / Safety & Alignment By Mingxuan Zhang 2026-08-06
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

The PrivacyPeek framework benchmarks how LLM-based agents acquire sensitive information beyond their necessary operational scope during task completion.

Computer Vision / Efficiency & Inference By Zahra Khodakarami 2026-08-06
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?

The paper evaluates whether super-resolution techniques applied to MRI scans risk erasing real small white-matter lesions or hallucinating false ones.

Benchmarks & Evals By Zonghuan Xu 2026-08-06
Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

The paper provides a formal framework for determining when distribution classes are learnable and quantifies the exact performance cost of replacing interactive queries with static ones.

Training & Fine-Tuning / Benchmarks & Evals By Omar Coser 2026-08-06
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

The researchers applied self-pretraining to transformer models to boost accuracy in medical time series classification without requiring external data.

Computer Vision / Efficiency & Inference By Zijie Wang 2026-08-06
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?

CogVis introduces a modular perception and memory framework that decouples image analysis from query processing to improve speed and accuracy in open-vocabulary change detection.

Multimodal / Computer Vision By Nima Hatami 2026-08-06
CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection

The CFGPNet framework improves multispectral object detection by optimizing feature interaction and gradient flow while reducing computational overhead.

Training & Fine-Tuning / Efficiency & Inference By Mikhail Solonko 2026-08-06
Muon on the Stiefel Manifold Admits an Exact Closed-Form Update

The researchers developed Skewon, an algorithm that provides an exact closed-form solution for the Stiefel Muon optimization problem.

Robotics / Computer Vision By J. de Curtò 2026-08-06
Visual Grounding in Zero-Shot Vision-Language Control

The paper investigates whether vision language models serving as robot controllers truly rely on visual inputs or merely leverage non visual shortcuts like simulator rewards.

Computer Vision By Yixiong Jing 2026-08-06
BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models

BendTwin adds bending stiffness to spring mass models to improve the stability and accuracy of physical reconstructions from sparse video data.

Robotics / Multimodal By Haodong Yan 2026-08-06
Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Robust-WAM introduces a method to align video-generation model latent spaces with semantic features, enabling robots to handle visual changes more reliably.

Computer Vision / Robotics By Giorgio Tonetti 2026-08-06
Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

The paper introduces Prior-SG, a framework that casts scene graph generation as a probabilistic alignment problem using task-conditioned priors and visual-geometric feature fusion to handle arbitrarily structured environments.

Computer Vision / Robotics By Subin Jeon 2026-08-06
HOPE: Hand-Object Pressure Estimation from Monocular Videos

The researchers developed a transformer architecture that estimates physical pressure during hand-object interactions using only monocular video input.