Research Feed Page 34
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces a method to improve agent performance by dynamically allocating hindsight feedback across individual decision steps in a multi-turn task.
Capek 0.5 is a vision-language model architecture that uses task-specific specialists merged into a single system to improve robot reasoning and environment verification.
The researchers developed Stockmark-Nemotron-3-Nano-Omni-JapanDocReader to balance structured document parsing with existing document visual question answering capabilities.
The authors created a dataset of 557 hand-drawn UML diagram pairs paired with verified, machine-readable PlantUML code to support automated diagram generation.
Researchers evaluated how LLM agents change their personality traits in response to life events using a new benchmark called BFI-Adapt.
This research introduces a collaborative framework that offloads heavy speech enhancement tasks to a server while maintaining minimal computational overhead on edge devices.
The CreativeInstruct method introduces a way to fine-tune a single unified model that balances instruction following with narrative diversity by tagging and self-injecting creative text segments.
The paper introduces a new benchmark, C4-Eval, to test how effectively Multimodal Large Language Models (MLLMs) perform cross-concept understanding.
The COVER wrapper adds statistical reliability to existing video temporal grounding models by creating calibrated intervals that satisfy a user-specified error threshold.
The paper introduces a method to identify and correct object hallucinations in vision language models by analyzing attention layers and refining token decoding.
The paper introduces AutoPrune, a framework that uses large language models to automatically design efficient algorithms for reducing the number of visual tokens in multimodal models.
The paper demonstrates how diffusion large language models have structural safety vulnerabilities that can be exploited using safety neuron identification and targeted steering techniques.
The authors introduce the Generative Embedding Benchmark to evaluate how much semantic information remains recoverable from frozen visual embeddings when using a generative decoder.
The paper introduces a modular approach to test-time training by representing inner learner components as a directed acyclic graph to simplify design and analysis.
The paper addresses the challenge of achieving assurance closure in AI-native large-scale agile software development by proposing a high-level architecture with six capabilities to support autonomous agentic engineering delegation.
The paper introduces a method to store verified text-to-SQL repair episodes as a reusable memory bank that improves performance on future questions over the same database.
PsychoAgent introduces an affect-aware memory architecture that helps LLM agents retrieve contextually relevant experiences for better decision-making under conflict.
Researchers achieved significant energy and token savings by converting time series data into visual plots for processing by vision-language models.
SimWAM improves autonomous driving performance by separating action planning from resource heavy video generation through a lightweight, self contained model.
PACE improves automated algorithm design by decomposing large programs into reusable components, allowing LLMs to build on successful local logic rather than discarding full programs.