Research Feed Page 20

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Multimodal / Benchmarks & Evals By Zongyun Zhang 2026-08-13
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

The researchers introduced the Edit2TikZ benchmark and a curriculum learning process to help models accurately edit scientific figures using TikZ code.

Benchmarks & Evals / Robotics By Edward Holmberg 2026-08-13
A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms

The authors created a browser-native digital test range using WebAssembly to benchmark ocean-glider planning algorithms through standardized kinematic mission simulation.

Multimodal / Efficiency & Inference By Jiaqian Li 2026-08-13
When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

The paper presents a framework to select the most efficient intervention method for multimodal models by diagnosing specific task characteristics instead of testing every approach.

Computer Vision / Benchmarks & Evals By Jisoo Jeong 2026-08-13
SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation

SNM-VFI uses a hybrid approach combining structural flow estimation with generative diffusion models to synthesize smoother, high-quality intermediate video frames.

Robotics / Efficiency & Inference By Shivam Vats 2026-08-13
Deliberate Practice: Learning Robot Skills under a Budget

The paper introduces a method called Deliberate Practice to optimally distribute a fixed training budget across various robot skills to improve overall performance.

Benchmarks & Evals By Valentin Noël 2026-08-13
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

The paper demonstrates that how you select measurement tokens for sparse autoencoder evaluation significantly alters results and proposes a shared reporting protocol to address this bias.

Robotics By James Zhao 2026-08-13
NestDex: Nested Policy Learning with Copilot Assisted Teleoperation for Dexterous Manipulation

NestDex improves dexterous robot manipulation by combining pre-trained hand skill policies with operator teleoperation using an action-compressing variational autoencoder.

Multimodal / Benchmarks & Evals By AlayaWorld Team 2026-08-13
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

AlayaWorld improves long horizon video generation by replacing traditional depth warping with a streaming 3D point cache for better geometric consistency.

Agents / Efficiency & Inference By Aimilios Hadjiliasi 2026-08-13
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

The paper explores using Small Language Models on NVIDIA Jetson Orin NX hardware to handle cognitive tasks like service routing and memory management for virtual agents.

Computer Vision / Agents By Daniel Perkins 2026-08-13
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

The paper introduces ARMDIL, a system that uses an MLLM router to dynamically assign images to specialized vision backbones to improve classification accuracy.

Robotics / Computer Vision By Zheyu Zhuang 2026-08-13
Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning

The paper introduces a Seeker module that learns to focus robot vision on relevant spatial regions, significantly increasing success rates in complex environments.

Multimodal / Computer Vision By Yi-Chung Chen 2026-08-13
TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

TraVEL improves driving video retrieval by training embedding models to prioritize ego-vehicle movement patterns over static visual shortcuts.

Reinforcement Learning / Benchmarks & Evals By Joyjeet Singh 2026-08-13 1
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

The paper demonstrates that latent world models often fail at long horizon planning because their training objectives do not align with the latent information already present in their internal representations.

Reinforcement Learning / Training & Fine-Tuning By Guibin Zhang 2026-08-13
Latent On-Policy Self-Distillation

The researchers developed a method to replace hand-designed improvement rules with a system that learns to generate its own contextual guidance for model training.

Computer Vision / Training & Fine-Tuning By Qing Zhao 2026-08-13
Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy

The authors introduce Lipschitz-regularized object detection to fix instability caused by functional mismatches when chaining image restoration with detection models.

Computer Vision / Training & Fine-Tuning By Ebenezer Tarubinga 2026-08-13
CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

The paper introduces CW-BASS v2 to fix pseudo-labeling errors that occur when high-performance foundation models become overly confident and biased during training.

Computer Vision / Multimodal By Vsevolod Skorokhodov 2026-08-13
PixSDS: Why Latent SDS Makes Noisy Pixels

PixSDS introduces a method to repair color artifacts and high frequency noise in latent based image generation by forcing pixel space consistency during the optimization process.

Multimodal / Benchmarks & Evals By Peng Li 2026-08-13
UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

UniTraffic-Agent improves traffic video reasoning by using a shared event interpretation approach to handle multi-question queries and sparse visual data.

Reasoning / Benchmarks & Evals By Yicheng Bao 2026-08-13
SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

The authors introduce SPARED, a reasoning-based detector that uses adversarial image editing to train models to identify synthetic content without relying on provenance shortcuts.

Safety & Alignment By Katherine Van Koevering 2026-08-13
It's How You Ask: Gender-Associated Linguistic Bias in LLMs

The paper investigates whether prompts containing linguistic features associated with women negatively impact the quality of responses generated by large language models.