Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers introduced SWE-Prime, a method that selects a small, high-quality subset of training trajectories to improve software engineering agent performance.
WikiSkill improves AI agent performance by consolidating execution traces into a structured, persistent wiki that informs future skill development.
The PILOT harness allows AI agents to improve their performance in real time by employing a supervisor that provides live feedback and distills successful strategies during task execution.
The paper introduces UrbanGround, a sandbox environment using real-world 3D mapping data to evaluate how well MLLM agents navigate complex urban settings.
The ACE framework establishes a formal structure for evaluating and improving the data generated to train autonomous AI agents.
The paper introduces Agent Mesh to address unique reliability challenges in agentic software development by defining new primitives to manage non-idempotent tool delegations.
The paper introduces a dual-domain architectural pattern that separates an AI agent's persona from its execution logic to improve governance and auditability in regulated environments.
The authors introduce SymTrace, a framework that improves the reliability of debugging complex multi-agent systems by using controlled intervention anchors to replicate and repair execution failures.
TAU-Agent is an agentic framework designed to improve traffic anomaly detection by using retrieval-augmented generation to integrate video descriptions and object trajectories.
The authors introduce a framework for video-editing agents to generate and verify executable edit plans using a self-improving training loop.
The authors introduce a framework called Cordis that uses a context paradigm to enable reliable dynamic composition of software components.
The paper introduces a framework where a coding agent generates deterministic code to manage world state, which then guides a video model to maintain visual consistency in simulated environments.
RubSE improves AI code generation for web pages by using structured visual rubrics to guide iterative, self-evolving refinements.
The paper introduces AnTrap, a benchmark that tests how Android GUI agents handle dynamic environmental anomalies by injecting perturbations into 236 tasks.
The paper introduces FrontierChallenge, a benchmark for evaluating how well AI agents complete end-to-end scientific workflows.
VoiceMem is a dual-brain architecture designed to provide accurate, low-latency memory retrieval for speech-based conversational agents.
The researchers developed an agentic, iterative framework called VISA to generate high-quality training data for multimodal models by using feedback-driven loops instead of static one-pass pipelines.
The paper introduces a planning framework for AI coding agents that aligns their development processes with human practices to improve task performance.
JIT-Agent improves agent performance by dynamically generating and evolving task-specific control structures just in time to meet individual task demands.
The paper introduces trace integrity metrics to detect silent failures where LLM data agents produce correct answers through invalid logical steps.