Research Feed Page 39
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Researchers discovered that deep vision models unintentionally learn invisible camera metadata as shortcuts, which degrades their performance when image distribution shifts.
CalibForge uses automated adversarial feedback from software solvers to ensure that training tasks for LLM agents are neither too simple nor impossible to solve.
DyPES-VLA unifies robot control by using shared dynamics priors learned from video to enable action generation across different robot types without manual alignment.
The researchers developed a method called DASH that dynamically adjusts how reasoning models learn from their own outputs to produce more accurate results.
The paper introduces RP-OPSD, a method that improves how language models transfer English reasoning skills to low-resource languages by selectively applying privileged distillation based on reasoning-pivot signals.
The paper introduces a deterministic method to compile raw screen activity into structured, auditable memory frames for computer-use agents.
The researchers developed a method to predict future wrist camera observations to improve how robots perform fine-grained physical tasks.
The researchers developed a method using Concept Activation Vectors to identify if Transformer based speaking assessment systems rely on irrelevant speaker attributes rather than proficiency.
The authors introduce GST-Bench to evaluate and improve how vision-language models maintain consistent spatial understanding across long, continuous video streams.
The paper presents a mechanism for managing AI agent deployment and compute resources through a stakeholder voting model that prioritizes broad support over aggregate wealth.
WorldClaw uses an agentic pipeline to transform text prompts into spatially consistent and editable 3D environments.
The paper introduces UQ-Loc, a method that adds uncertainty estimates to LiDAR scene coordinate regression to improve localization robustness.
AgentOPSD introduces a recursive self-distillation method to provide granular credit assignment for multi-turn agentic tasks by analyzing turn-level evidence.
ChronoVision introduces a visual-focused training framework to help multimodal large language models track and reason about continuous changes in images.
The paper provides a systems blueprint for constructing generative economic simulators that use heterogeneous agents to model market interactions and institutional dynamics.
OPD-V improves multimodal model performance and reduces latency by using visual-based self-distillation to balance how the model uses image and text data.
The authors introduce a reference-free framework that uses LLM judges to automatically assess the quality and consistency of task-oriented conversational agent benchmarks.
SkillTFM enables tabular foundation models to adapt to new tasks and data distributions by dynamically retrieving and evolving skills from a pre-verified skill bank.
The paper demonstrates that current video-language models fail to accurately track event counts and sequences when event frequency and load increase.
The researchers developed a method to increase the clinical realism of synthetic datasets for AI agents while maintaining operational utility through constrained optimization.