Research Feed Page 13
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
EditBridge uses a diffusion bridge framework to enable 4K image editing by reducing attention complexity to linear scaling.
The researchers developed Ouro models that use recurrent stack iterations to improve performance on complex, compositional tool-calling tasks.
The paper introduces a structured operating model for AI assistants that uses persistent rule files to prevent repetitive errors and improve task performance.
The paper introduces a routing system that manages LLM uncertainty by deciding whether to trust a model's internal knowledge or trigger a web search to verify facts with statistical error guarantees.
The paper demonstrates that using LLM progress scores to shape rewards for reinforcement learning agents does not alter the underlying optimal policy of the agent.
The paper introduces RelArena-alpha, a unified open-source framework designed to standardize evaluation, tuning, and benchmarking for relational learning models.
The researchers introduced StartupBench to evaluate AI agents using end-to-end workflows derived from actual market-validated startup products.
The paper introduces SA-MRPO, a method that dynamically reweights reward objectives during reinforcement learning to prioritize underperforming metrics instead of treating all goals as equally weighted.
The paper introduces aDSL, a domain-specific language and agent-based system designed to reliably convert natural language instructions into functional 3D programs and geometry.
TAMP-Nav improves embodied navigation by combining efficient 3D spatial grounding with selective reasoning and a multi-level reward training approach.
Agent Lightning v1.0 provides a declarative framework to manage the complex training loops required for agents that interact with external environments.
The paper introduces Agentic ESOpt, a method for fine-tuning LLM agents that replaces traditional backpropagation with population-based parameter perturbations to reduce memory requirements during training.
The paper presents a two stage grading pipeline that uses high end models for rubric extraction and small models for cost effective, consistent scoring of examination questions.
The paper defines a framework to categorize cognitive risks in agentic AI and proposes mitigation strategies to maintain human control over autonomous systems.
GenRouter optimizes agentic image generation by dynamically routing prompts through a library of primitive operations to reduce latency and execution costs.
This paper reveals that popular multi-hop retrieval benchmarks ignore commercial licensing and deployment costs, misleading engineers who select models for production use.
The authors introduce HarnessEval-W, an agentic evaluation pipeline that decomposes world model testing into verifiable reasoning sequences to overcome the limitations of fixed, non-verifiable metrics.
MOSS-VL introduces an architecture and training curriculum that allows vision-language models to process incoming video frames and generate responses simultaneously.
The paper introduces a runtime system called StateM that uses harness scaling to improve agent task performance without modifying model weights.
UI-Mate-27B improves desktop automation performance through in-context demonstrations and specialized training on task-driven benchmarks.