Research Feed Page 10
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
ClawSentry provides a modular, multi-tier security framework that uses an abstraction protocol to protect autonomous LLM agents against progressive execution threats.
AgentMercury automates the creation of executable business environments to improve agent performance and training scalability.
The authors propose shifting from individual LLM agent loops to a structured system-level engineering approach using graph workflows for better task coordination and state management.
The paper introduces a routing mechanism that applies safety interventions only when harmful inputs are detected, preserving model utility for benign prompts.
The paper introduces S3D8, a quantization format that compresses the Llama 3.2 11B Vision Instruct model to 3.7 GB for mobile CPU execution.
TurboBias 2.0 optimizes production speech recognition systems by enabling efficient context biasing for streaming inference.
This study evaluates how to best allocate model capacity across the generator, critic, and refiner components of an automated self-refinement pipeline.
The authors introduce a pipeline to automatically reconstruct and retarget 3D human interaction data into a large-scale dataset for training diverse robotic embodiments.
The authors introduce a pipeline to denoise, deduplicate, and annotate large-scale digitized library books, resulting in the enriched IB-HL-ET dataset.
Swift-Image is a compact, unified model designed to handle text-to-image generation and image editing tasks efficiently under strict computational budgets.
The paper introduces ThermoDPO, a new training method that stabilizes generative model output by preventing reward-driven distortion of the underlying data distribution.
InsufficiencyBench measures how effectively LLMs identify missing information in legal queries instead of providing premature, potentially fabricated advice.
The authors introduce AI4AI-Bench to test whether AI agents can improve training algorithms by modifying their core components such as learning rules and supervision signals.
MaliciousSkillBench provides a unified, consolidated registry of 13 public sources to detect malicious instructions within reusable agent Skills.
HandMvNet uses multi-view cross-attention to estimate 3D hand poses from multiple camera angles without requiring complex calibration.
The researchers introduced a nested sampling method to guide the output of diffusion language models toward desired properties during inference without requiring additional training.
The paper introduces a margin-controlled technique to help machine learning models for music analysis better identify their own incorrect predictions.
This study analyzes how autonomous coding agents interact with documentation through an empirical examination of their file-level changes and conversational logs.
The paper systematically evaluates various eviction policies for semantic LLM caching and finds that the standard Least Frequently Used approach remains highly effective compared to more complex alternatives.
The researchers developed an LLM that uses Reinforcement Learning Fine-Tuning and specialized rewards to improve performance in single-step retrosynthesis tasks.