Research Feed Page 31
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers developed a framework called UserIDA that allows developers to precisely control the local conversational intent of LLM-based user simulators.
The paper introduces a multi-layered verification framework that uses an ensemble of LLM judges to validate robot action plans for safety, security, and ethical alignment before execution.
Researchers developed a diagnostic framework called Decoding-Level Taboo that forces language models off their primary prediction paths to test how robust their internal reasoning is when constrained.
The researchers developed a sub 200 USD stereo vision system for capturing large scale egocentric video with synchronized inertial sensor data.
Evidence-RL introduces a training method that forces vision-language models to base their answers on specific image regions rather than relying on language shortcuts.
The paper introduces Avalon-ToM-Bench, a new benchmark designed to measure how well Large Language Models understand human mental states using the mechanics of the game The Resistance: Avalon.
The Model Discovery Agent uses LLMs to iteratively propose and test mechanistic models, enabling data-efficient scientific discovery in fields like physics and chemistry.
BDH-CQ is a reasoning system that enables models to learn new visual tasks from demonstrations using recurrent latent memory instead of updating parameters or using explicit key-value caches.
The paper introduces TIDE, a method to fix model distillation failures caused by degenerate token agreement and teacher-student mismatch.
DistMoE enables visual instruction tuning across distributed systems using a mixture of experts approach that eliminates the need to rehearse or share private datasets.
The researchers developed a self-distillation technique for multimodal large language models that sharpens visual perception by identifying and training on internal counterfactual blind spots.
RoMeRL improves agent performance and storage efficiency by replacing high-dimensional trajectory indexing with a compact, factorized utility state system.
The paper introduces a regression-free, layout-aware matching system that eliminates coordinate hallucinations in GUI agents to improve element selection accuracy.
The paper introduces a verifier-free consensus selection method that improves the geometric accuracy of parametric CAD programs generated by language models.
The paper introduces a cognitive architecture called CEAA that bridges the gap between high-level reasoning models and low-level game engine control systems.
The authors propose treating autonomous research agents like software fuzzers by using intermediate feedback signals to guide experimentation and discovery.
The paper introduces a new scoring metric called Consilience to improve how LLMs select the best output among multiple reasoning attempts, specifically addressing issues where models default to incorrect but confident answers.
The Mendel Gödel Machine improves coding agents by using comparative evidence across tasks and agent versions instead of relying on single failure trajectories.
The paper introduces a framework for nesting smaller sub-models within a larger architecture to reduce training compute and improve speculative decoding performance.
The paper demonstrates that internal residual-stream activations in LLMs contain security signals that are often lost when the model generates final text-based outputs.