Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers introduced SWE-Prime, a method that selects a small, high-quality subset of training trajectories to improve software engineering agent performance.
The authors introduce GameCleaner and the Game2World engine to remove distracting interface elements from game footage to improve the training of world models.
The paper introduces a framework called RobustTests that improves AI code generation by synthesizing diverse, failure-inducing test cases to guide reinforcement learning.
The paper introduces a method called OPDVR that aligns reinforcement learning signal with task success during model distillation to improve reasoning performance.
RecurSE enables LLM-based judges to improve their evaluation performance by creating a bounded, self-correcting feedback loop that eliminates the need for external gold standard rewards.
The paper introduces BPCO, a method that stabilizes critic-based reinforcement learning, improving performance across various model sizes and tasks.
The authors present an end-to-end framework to automatically generate instruction-tuning and benchmark datasets from complex industrial technical documents.
The paper introduces a method called Harness Evolution that improves agent adaptability and performance for live-streaming environments by decoupling execution settings from the base model.
The paper identifies that representational drift in the decoder and output layers significantly impacts out of distribution performance when fine-tuning the MedSAM medical image segmentation foundation model.
The paper introduces Quantization-Aware Healing, a practical recipe for recovering compressed 4-bit large language models, and uses it to produce the open-weight model Hypernova-60B.
The paper presents a method that enables robot policies to self-improve through iterative deployment without the need to modify the original policy weights.
The paper introduces E2-TTT, a new method for Test-Time Training that uses chunk-wise updates to achieve higher performance while maintaining computational efficiency.
The paper introduces a routing mechanism that applies safety interventions only when harmful inputs are detected, preserving model utility for benign prompts.
The paper introduces ThermoDPO, a new training method that stabilizes generative model output by preventing reward-driven distortion of the underlying data distribution.
The paper introduces a method to extrapolate optimal learning rates for Mixture of Experts models using small-scale proxy runs to avoid expensive full-scale sweeps.
The researchers created a 20.3B-token corpus called MidTool-Mix to improve agentic tool-use capabilities in models during the mid-training phase rather than relying solely on post-training.
The researchers developed a staged training method called IAR to improve how models store and answer questions about specific document sets without needing retrieval systems.
Researchers developed a method using attribute guided genre expansion to train language models on diverse creative formats beyond basic narrative generation.
RecipeNet is a hierarchical transformer model designed to process heterogeneous recipe data with variable schemas and sequential procedural steps.
The paper introduces a multi-objective optimization framework to automatically select the best merge parameters for combining specialized neural network models.