Research Feed Page 35
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers developed a new metric and a decoding strategy to more accurately identify and mitigate benchmark data leakage in large language models.
The paper introduces SABRE, a scalable, automated pipeline that generates challenging stress tests to expose weaknesses in how vision-language models reconcile visual evidence with existing world knowledge.
WorldTrace addresses long-horizon visual memory in video models by using a fixed-size cache that prevents positional embedding degradation.
The paper introduces TEPA, a memory management framework that revokes stale active memories and tracks support and conflict counts to prevent memory pollution when the world changes.
Fisher-R1 is a specialized LLM agent trained to perform reliable hypothesis testing by using a new benchmark and outcome-grounded reinforcement learning.
The paper classifies 21 open-source AI security tools against the MIT AI Risk Mitigation and Response Taxonomy to identify gaps in existing risk coverage.
The CoBa framework maximizes LLM inference accuracy by intelligently routing compute resources between candidate generation, verification, and final selection.
The paper introduces a unified evaluation framework called A2E that uses standardized protocols and centralized telemetry to measure agent performance across diverse benchmarks.
This study analyzes the production characteristics and resource usage of AI-generated C++ code compared to human-written code across a large industrial monorepo.
The paper introduces SkillProx, a method to evolve LLM agent skills through verified forward updates and utility-aware consolidation while preventing redundant or regressive skill accumulation.
The researchers introduce Gated Hindsight Distillation to help GUI agents learn from future screenshots when current screen data is insufficient for decision making.
CubicQuant introduces a flexible, GPU-friendly weight quantization format that uses monotonic cubic functions to better represent model weight distributions compared to standard uniform methods.
The paper introduces ReASearch, a framework that uses a single tool-using LLM agent to autonomously handle search and optimization tasks for prompts, code, and machine learning pipelines.
Blast Radius reduces LLM token consumption by identifying and archiving redundant or concluded context in agentic coding environments.
CoinRAG reduces RAG latency and computational redundancy by precomputing and reusing specific information nugget representations within the model KV cache.
LLMRouter provides a unified framework and automated data pipeline to build, evaluate, and deploy routers that select the most cost-effective LLM for a given task.
The ω-0 model enables humanoid robots to perform simultaneous locomotion and object manipulation by learning unified whole-body action coordination.
The paper presents an MR-safe robot designed for percutaneous needle interventions that utilizes fluid-based control to operate within high-field magnetic environments.
The paper introduces a synthetic data generation framework that allows AI to accurately assess patient poses for X-rays by training on generated depth images and radiographs.
The paper introduces a statistical method to represent an athlete's historical data as a reusable latent memory table to improve performance tracking and prediction.