Research Feed Page 2
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The authors introduce a framework called Cordis that uses a context paradigm to enable reliable dynamic composition of software components.
RotDroid uses a vision-language model to detect GUI rotation bugs by comparing visual states between portrait and landscape orientations.
Researchers evaluated automated fact checking systems across four datasets to reveal how domain differences and retrieval performance impact overall accuracy.
The paper introduces scrydb to enable combined lexical and semantic search capabilities within a single SQLite database file.
The paper introduces a framework where a coding agent generates deterministic code to manage world state, which then guides a video model to maintain visual consistency in simulated environments.
StreamPI adds historical context to vision-language-action models to improve robotic task performance without increasing the model parameter count.
RubSE improves AI code generation for web pages by using structured visual rubrics to guide iterative, self-evolving refinements.
The paper introduces AnTrap, a benchmark that tests how Android GUI agents handle dynamic environmental anomalies by injecting perturbations into 236 tasks.
The paper introduces FrontierChallenge, a benchmark for evaluating how well AI agents complete end-to-end scientific workflows.
VoiceMem is a dual-brain architecture designed to provide accurate, low-latency memory retrieval for speech-based conversational agents.
The 4DGS-WAM model enables future video prediction by explicitly decomposing scenes into dynamic objects and a static background using 4D Gaussian Splatting.
The researchers developed an agentic, iterative framework called VISA to generate high-quality training data for multimodal models by using feedback-driven loops instead of static one-pass pipelines.
Prefix Sliding enables large language models to perform reasoning tasks three times faster without requiring additional training.
The paper introduces a planning framework for AI coding agents that aligns their development processes with human practices to improve task performance.
The authors introduce VBVR-Pro, a suite designed to improve visual reasoning capabilities in models by using verifiable generative tasks.
JIT-Agent improves agent performance by dynamically generating and evolving task-specific control structures just in time to meet individual task demands.
The paper introduces trace integrity metrics to detect silent failures where LLM data agents produce correct answers through invalid logical steps.
ProgRouter optimizes multi-agent workflows by dynamically selecting models based on progress and cost to maximize task completion rates within defined energy budgets.
SkillForge introduces a system that distills and verifies reusable skills for agents, significantly improving performance on complex tasks.
The paper introduces EarthVerse, a benchmark designed to evaluate how accurately scientific agents perform end to end investigations involving Earth systems and natural hazards.