Research Feed Page 4
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Maia 200 is a custom AI accelerator designed to improve performance, energy consumption, and total cost of ownership for large scale workloads.
The paper demonstrates that standard LLM agent handoff processes often cause binding constraints to lose their functional power, and identifies methods to restore this operational state.
The paper introduces a framework and an autonomy classification system for deploying Large Language Model agents in scientific molecular discovery workflows.
PARTAB improves table-based reasoning by decomposing tables into semantically coherent parts before processing them with a multi-stage pipeline.
The paper introduces a method that models agent behavior as finite state machines to improve next-step prediction and detect system failures.
The authors introduce GameCleaner and the Game2World engine to remove distracting interface elements from game footage to improve the training of world models.
The paper demonstrates that simple linear probes on frozen model hidden states provide efficient and robust detection of machine-generated text using minimal training samples.
The authors introduce two systems, MCTidy and MCGenie, that use large language models to automatically reorganize and generate standardized documentation for machine learning models.
The paper evaluates various retrieval-augmented generation pipelines, finding that multimodal vision-based approaches significantly outperform text-based methods despite introducing higher latency and storage costs.
The paper demonstrates that language models can be steered during inference using undisclosed logit modifications, making traditional model weight audits insufficient for identifying production-level bias.
The paper introduces MemUse, a benchmark for evaluating how well conversational AI integrates long-term memory into natural dialogue, revealing a significant disconnect between standard fact-checking performance and actual conversational utility.
The paper introduces a framework called RobustTests that improves AI code generation by synthesizing diverse, failure-inducing test cases to guide reinforcement learning.
Simthesizer utilizes a coding agent to automatically extend simulators for complex LLM serving systems, achieving higher throughput accuracy than existing approaches.
The researchers developed a Bayesian framework that decomposes RAG system performance into distinct stages to reveal hidden behavioral differences between configurations.
AtlasNav introduces a persistent navigation layer for AI agents to prevent evidence loss during large-scale document corpus interactions.
The paper introduces Crase, an agentic system that bounds research discovery within a citation graph to improve evidence grounding and search accuracy.
The paper introduces a method called OPDVR that aligns reinforcement learning signal with task success during model distillation to improve reasoning performance.
The Meta n system introduces a recursive architecture that enables agents to iteratively improve their own problem-solving logic and code libraries.
RecurSE enables LLM-based judges to improve their evaluation performance by creating a bounded, self-correcting feedback loop that eliminates the need for external gold standard rewards.
The paper introduces a gated intervention framework that dynamically manages model activations to reduce sycophancy and hallucinations in clinical question answering while preserving model weight integrity.