Research Feed Page 12
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The authors propose Software 3.0, a new architecture that replaces traditional three-tier systems with a converged structure consisting of a persistent storage layer, a probabilistic intelligence core, and an agent-based execution loop.
The researchers created a 20.3B-token corpus called MidTool-Mix to improve agentic tool-use capabilities in models during the mid-training phase rather than relying solely on post-training.
EnvHarness and EnvRigger dynamically modify static environments to improve LLM agent training efficiency and performance.
PolicyGuide introduces a workflow-based verification system that uses an external runtime graph to enforce organizational compliance in LLM agents.
The paper introduces a routing framework that reduces computational costs by intelligently deciding when to pay for accurate model value estimates rather than using cheaper, noisier alternatives.
The paper introduces SWE bench Science to evaluate coding agents on repository-level scientific software engineering tasks and analyzes how scientific guidance impacts their performance.
The paper introduces a new algorithm called Best Prefix Selection to choose the most efficient set of skills for LLM agents within a fixed token budget.
The researchers developed a staged training method called IAR to improve how models store and answer questions about specific document sets without needing retrieval systems.
The paper introduces FACET, a framework for generating consistent, executable terminal tasks by grounding task artifacts like instructions and verifiers in a shared containerized state.
HarnessRisk provides a lifecycle based framework for evaluating security vulnerabilities across six distinct operational phases in agentic systems.
SPADE improves agent performance by using an automated system that co-evolves training environments and reasoning agents through a continuous reinforcement learning loop.
The paper introduces a dataset and evaluation protocol to measure how accurately video generation models can complete specific instructed outcomes while maintaining semantic grounding.
Lego-RL is a framework that aligns native coding execution harnesses with policy-gradient training to improve agent performance and stability.
Eureka introduces a meta-agent architecture that dynamically promotes specialized agents to solve long-horizon scientific tasks while minimizing computational overhead.
The authors introduce PTXBench and a supervised fine-tuning method to help LLMs write efficient architecture-specific GPU code.
The paper introduces a training-free inference-time protocol that uses a self-critique loop and a confirmed sentinel to improve reasoning accuracy while early-stopping redundant computations.
SkillForge enhances software engineering agents by distilling repository-specific knowledge into reusable skills to solve project-specific issues.
SkillGate optimizes agent performance by partitioning training signals to stop task outcomes from interfering with how agents select procedural skills.
The paper uses an automated framework to evaluate how untrusted external data can manipulate agents in the DeepSeek Harness framework into performing unintended actions.
SemaPLC uses an agent-based workflow and multi-stage verification to ensure that LLM-generated PLC programs integrate correctly and execute reliably within existing industrial projects.