Research Feed Page 6
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces a protocol to verify if trading strategies generated by large language models align with their actual performance, finding that most claimed advantages fail to materialize.
The authors created the ICU-REACT dataset and a corresponding family of fine-tuned models to improve LLM performance in identifying and reasoning over patient data for critical care.
The paper introduces StrategyBench to evaluate if language models can effectively derive and apply explicit task-level strategies from few-shot examples.
The researchers audited 200 world model papers to determine how their capabilities measure up against the structural requirements of professional physical simulators.
The paper demonstrates that using learned sigmoid gates for key-value cache eviction leads to better performance than existing methods like H2O and KeyDiff.
The paper introduces local distillation, a method that improves the prediction accuracy of simple, interpretable linear models by selectively leveraging predictions from complex black-box models.
SkillAlchemy introduces a systematic approach to converting open world information into reliable, reusable procedural skills for software agents.
The researchers developed CyberFactory, a framework that leverages existing vulnerability data to train an AI model, OpenAegis, to improve security analysis performance.
The paper introduces a framework called DynaRule that enables large language models to dynamically retrieve and apply reusable procedural rules at scale.
The paper introduces an iterative orchestration framework that improves generative search recall by dynamically managing evidence allocation and curbing information dilution.
The paper introduces BPCO, a method that stabilizes critic-based reinforcement learning, improving performance across various model sizes and tasks.
The paper introduces a structural audit procedure that reliably detects and localizes causal leakage in complex sequence models by monitoring intermediate output differences during forward passes.
AutoResearch is an autonomous system that uses multi-model cross-review to improve the reliability of research idea generation and experimental validation.
The paper demonstrates that full-solution interaction between agents can cause proposals to converge too quickly, erasing useful diversity and reducing performance on specific optimization tasks.
Researchers analyzed 115 countries to demonstrate that individualistic national culture significantly enhances how reliably governments follow their own constitutions.
Researchers introduced a system called EvoMap that distills successful, verifier-confirmed AI task trajectories into reusable Genes to improve performance and reduce token consumption across various model families.
The authors introduce FinixDoc, an agentic parsing system powered by a specialized vision-language model, to improve accuracy and structural consistency in real-world financial document processing.
ClawProBench provides a framework for evaluating AI agents by tracing runtime behavior across consistent workspace holdouts rather than relying on final success metrics alone.
The paper evaluates how breaking down VAT determination tasks into different numbers of orchestrated agents affects overall system accuracy compared to using a single agent.
GameXpert-Bench evaluates how well coding agents navigate the full game development lifecycle, from initial generation to defect repair and optimization.