Research Feed Page 44
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces a data science world model called DSWorld that uses a mixture of rule-based execution, compilation, and an LLM-based simulator to predict the effects of operations and avoid costly trial-and-error workflows in autonomous agents.
The paper introduces a novel looped Transformer architecture called Loopie that maximizes pre-training compute efficiency to achieve strong reasoning benchmark performance.
The paper investigates the performance differences between multi-agent systems and single-agent systems powered by large language models to address why multi-agent advantages vary inconsistently across settings.
The paper introduces PRISA, a proactive infrastructure LiDAR framework that uses point cloud data and self-supervised training to assess urban intersection safety in real time.
The paper demonstrates that selectively applying the Muon optimizer to hidden weight matrices significantly boosts performance in agentic reinforcement learning tasks characterized by sparse rewards.
The paper introduces Physics-EnhAnced Reinforcement Learning (PEARL), a new paradigm that addresses sample inefficiency and high dimensionality challenges in complex dynamical systems to enable real-time optimal control.
PagedWeight manages GPU memory for Mixture-of-Experts models by dynamically quantizing weights at runtime to balance model precision against KV cache requirements.
SeerGuard is a safety framework for mobile graphical user interface agents that uses an instruction-level screening module and a safety-augmented world model to predict and intercept risks before actions are executed.
This paper investigates how pretraining choices shape reinforcement learning returns and what reinforcement learning actually does to a model policy using chess games and puzzles.
The paper introduces Audio-Visual Flamingo, an open model designed to improve joint perception, temporal alignment, and multi-event reasoning over long videos.
The paper introduces On Policy Delta Distillation, a new method that improves how reasoning capabilities are transferred from a teacher model to a student model.
This paper introduces a memory-efficient training technique called Hierarchical Global Attention to process significantly longer token sequences on constrained GPU hardware.
The HDR model uses a tree-structured hierarchy to balance logical consistency in multi-step visual reasoning with efficient streaming performance.
SUFLECA improves zero-shot CAD-to-image alignment by scaling up geometry-aware feature learning, leading to significantly better accuracy on benchmarks like ScanNet25k.
AeroAct adapts a pretrained video diffusion transformer to connect visual language task specification, predictive action generation, and closed-loop quadrotor execution without generating future video at deployment time.
The researchers demonstrate that current AI coding agents frequently fail to detect malicious package installations when following project setup documentation.
The researchers developed the MM-IssueLoc benchmark to evaluate how visual evidence like screenshots affects AI-driven repository-level issue localization.
The paper demonstrates that decomposing complex queries into smaller attribute-based sub-problems improves an LLM's consistency and alignment with real-world data.
The researchers demonstrate that the bias in specific statistical sampling algorithms can be constrained relative to the system dimension when variables share sparse interactions.
NeuronSoup replaces backpropagation with an asynchronous evolutionary algorithm that uses discrete event simulation to train models for efficient, variable-depth processing.