Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
WikiSkill improves AI agent performance by consolidating execution traces into a structured, persistent wiki that informs future skill development.
The paper introduces UrbanGround, a sandbox environment using real-world 3D mapping data to evaluate how well MLLM agents navigate complex urban settings.
The ACE framework establishes a formal structure for evaluating and improving the data generated to train autonomous AI agents.
The paper introduces PAWBench, a new evaluation framework that measures how well video generation models predict the physical outcomes of various scenarios.
The authors introduce SymTrace, a framework that improves the reliability of debugging complex multi-agent systems by using controlled intervention anchors to replicate and repair execution failures.
PlanSightRAG leverages a vision-first multimodal framework to automate compliance checking for complex civil engineering drawings.
The authors introduce a framework for video-editing agents to generate and verify executable edit plans using a self-improving training loop.
The research demonstrates that providing prior evaluation scores as metadata to LLM-as-a-judge systems causes systematic anchoring bias that distorts final judgment accuracy.
The researchers introduced XRepoTest to evaluate how effectively large language models generate unit tests within complex, multi-file code repositories across five programming languages.
RotDroid uses a vision-language model to detect GUI rotation bugs by comparing visual states between portrait and landscape orientations.
Researchers evaluated automated fact checking systems across four datasets to reveal how domain differences and retrieval performance impact overall accuracy.
The paper introduces scrydb to enable combined lexical and semantic search capabilities within a single SQLite database file.
RubSE improves AI code generation for web pages by using structured visual rubrics to guide iterative, self-evolving refinements.
The paper introduces AnTrap, a benchmark that tests how Android GUI agents handle dynamic environmental anomalies by injecting perturbations into 236 tasks.
The paper introduces FrontierChallenge, a benchmark for evaluating how well AI agents complete end-to-end scientific workflows.
The paper introduces a planning framework for AI coding agents that aligns their development processes with human practices to improve task performance.
The paper introduces trace integrity metrics to detect silent failures where LLM data agents produce correct answers through invalid logical steps.
The paper introduces EarthVerse, a benchmark designed to evaluate how accurately scientific agents perform end to end investigations involving Earth systems and natural hazards.
The paper introduces ZID, a new evaluation metric for generative models that identifies and ranks failures in image generation where traditional metrics like FID fail.
FedV-KGQA enables multi-hop reasoning over knowledge graphs distributed across different organizations by fusing local entity embeddings without sharing private raw data.