Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
TTPO improves the reasoning accuracy of language models during test time by using label-free policy optimization that bypasses the need for manual ground-truth data.
The paper introduces a method that uses vision-language models to perform reasoning that guides robot manipulation policies, improving performance on long-horizon tasks.
Prefix Sliding enables large language models to perform reasoning tasks three times faster without requiring additional training.
SkillForge introduces a system that distills and verifies reusable skills for agents, significantly improving performance on complex tasks.
LeFlow optimizes action planning by using a generative model to predict future trajectories, significantly reducing computation time compared to traditional iterative methods.
SPO++ is a refined policy optimization framework that increases online learning efficiency for language agents by aligning data tracking with event timing.
WarpSAC is a scalable reinforcement learning framework that adapts its architecture based on available compute resources to accelerate training and improve deployment success.
The paper introduces IAPO, a method that improves agent training by redistributing reward credit based on how agent actions influence one another within multi-turn service workflows.
The paper introduces a framework called RobustTests that improves AI code generation by synthesizing diverse, failure-inducing test cases to guide reinforcement learning.
The paper introduces CAFE, a system that improves search agent performance and reduces hallucinations by alternatingly updating the agent and its critic through co-evolving feedback.
The SMITH framework uses reinforcement learning to train a single language model to both synthesize reusable Python tools and apply them effectively to solve complex tasks.
The researchers audited 200 world model papers to determine how their capabilities measure up against the structural requirements of professional physical simulators.
The paper introduces BPCO, a method that stabilizes critic-based reinforcement learning, improving performance across various model sizes and tasks.
Agent-G2 optimizes agent training by using Gaussian-based guidance to sample expert trajectory lengths, achieving higher success rates at a fraction of the cost of traditional probing methods.
The researchers introduced a method called OraRL that integrates ground truth annotations as oracle rollouts to improve video model performance and reduce inference latency.
The paper introduces a GUI-grounded simulation agent that uses pixel-level perception and reinforcement learning to generate authentic, multi-turn e-commerce shopping trajectories.
EXIMO leverages a vision-language model to decompose complex robotic tasks into smaller steps, improving the efficiency of training vision-language-action policies.
SPADE improves agent performance by using an automated system that co-evolves training environments and reasoning agents through a continuous reinforcement learning loop.
Lego-RL is a framework that aligns native coding execution harnesses with policy-gradient training to improve agent performance and stability.
SkillGate optimizes agent performance by partitioning training signals to stop task outcomes from interfering with how agents select procedural skills.