Research Feed Page 39

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Computer Vision / Training & Fine-Tuning By Vladan Stojnić 2026-08-05 9
Invisible Shortcuts: Why Vision Encoders Know Your Camera

Researchers discovered that deep vision models unintentionally learn invisible camera metadata as shortcuts, which degrades their performance when image distribution shifts.

Agents / Benchmarks & Evals By Fanzhe Meng 2026-08-06
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

CalibForge uses automated adversarial feedback from software solvers to ensure that training tasks for LLM agents are neither too simple nor impossible to solve.

Robotics / Multimodal By Junfeng Li 2026-08-06
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

DyPES-VLA unifies robot control by using shared dynamics priors learned from video to enable action generation across different robot types without manual alignment.

Training & Fine-Tuning / Reasoning By ZhiYan Hou 2026-08-06
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

The researchers developed a method called DASH that dynamically adjusts how reasoning models learn from their own outputs to produce more accurate results.

Reasoning / Training & Fine-Tuning By Xinye Wang 2026-08-06
RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

The paper introduces RP-OPSD, a method that improves how language models transfer English reasoning skills to low-resource languages by selectively applying privileged distillation based on reasoning-pivot signals.

Agents / Efficiency & Inference By Nossa Iyamu 2026-08-06 15
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

The paper introduces a deterministic method to compile raw screen activity into structured, auditable memory frames for computer-use agents.

Robotics / Efficiency & Inference By Yuhao Pan 2026-08-05 20
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

The researchers developed a method to predict future wrist camera observations to improve how robots perform fine-grained physical tasks.

Safety & Alignment / Benchmarks & Evals By Arya Labroo 2026-08-06
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

The researchers developed a method using Concept Activation Vectors to identify if Transformer based speaking assessment systems rely on irrelevant speaker attributes rather than proficiency.

Benchmarks & Evals / Multimodal By Qifeng Zhang 2026-08-06 36
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

The authors introduce GST-Bench to evaluate and improve how vision-language models maintain consistent spatial understanding across long, continuous video streams.

Agents / Safety & Alignment By Praphul Chandra 2026-08-06
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

The paper presents a mechanism for managing AI agent deployment and compute resources through a stakeholder voting model that prioritizes broad support over aggregate wealth.

Agents / Computer Vision By Chunchao Guo 2026-08-05 50
WorldClaw: Agentic 3D Open-World Generation at Scale

WorldClaw uses an agentic pipeline to transform text prompts into spatially consistent and editable 3D environments.

Computer Vision / Robotics By Jacek Komorowski 2026-08-06
UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

The paper introduces UQ-Loc, a method that adds uncertainty estimates to LiDAR scene coordinate regression to improve localization robustness.

Agents / Reinforcement Learning By Zi-Han Wang 2026-08-06 74
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

AgentOPSD introduces a recursive self-distillation method to provide granular credit assignment for multi-turn agentic tasks by analyzing turn-level evidence.

Multimodal / Reasoning By Yifan Shen 2026-08-06 31
ChronoVision: Temporal Reasoning via Latent State Reconstruction

ChronoVision introduces a visual-focused training framework to help multimodal large language models track and reason about continuous changes in images.

Agents / Benchmarks & Evals By Jiale Han 2026-08-06 27
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

The paper provides a systems blueprint for constructing generative economic simulators that use heterogeneous agents to model market interactions and institutional dynamics.

Multimodal / Efficiency & Inference By Aniri 2026-08-05 9
OPD-V: Visual On-Policy Self-Distillation with Modality Balance

OPD-V improves multimodal model performance and reduces latency by using visual-based self-distillation to balance how the model uses image and text data.

Agents / Benchmarks & Evals By Noam Koren 2026-08-06
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

The authors introduce a reference-free framework that uses LLM judges to automatically assess the quality and consistency of task-oriented conversational agent benchmarks.

Training & Fine-Tuning By Yi He 2026-08-06
SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models

SkillTFM enables tabular foundation models to adapt to new tasks and data distributions by dynamically retrieving and evolving skills from a pre-verified skill bank.

Multimodal / Benchmarks & Evals By Sarvesh Baskar 2026-08-06
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

The paper demonstrates that current video-language models fail to accurately track event counts and sequences when event frequency and load increase.

Benchmarks & Evals By Omid Bazgir 2026-08-06
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

The researchers developed a method to increase the clinical realism of synthetic datasets for AI agents while maintaining operational utility through constrained optimization.