Research Feed Page 22
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Researchers evaluated whether large language models can automatically generate executable code to confirm software vulnerabilities within the Autoware autonomous driving stack.
The paper introduces ParliamentRAG, a system that improves retrieval accuracy in parliamentary transcripts by weighting speaker authority based on query relevance and professional background.
The paper introduces a method for composing reusable AI policies in multi-agent environments that maintains safety and flexibility without requiring per-task retraining.
Researchers evaluated how instruction tuning influences the verbalized confidence and lexical diversity of rationales generated by three popular large language models.
The researchers investigated whether using specific rotational transforms that respect RoPE structure improves accuracy during 4-bit model quantization.
StreamTTT introduces a dual-branch architecture that combines real-time attention with a recurrent state to improve video memory recall without sacrificing performance.
The paper introduces a method to track task progress in vision-language-action models by fitting linear probes on internal embeddings to identify completion status and detect out-of-distribution inputs.
UniSwap is a framework designed for low-latency, streaming-ready audio-visual identity swapping that preserves source motion and content while replacing appearance and voice.
PlayWorld introduces an agent-based evaluation framework to test how video world models perform under long-horizon objectives.
AutoDesign improves long horizon agentic design by using an iterative meta harness that learns from failures to optimize system components across tasks.
Researchers developed Synthetic Persona Pretraining to embed desired assistant behaviors into language models starting from the very first token of training.
Researchers developed Mimir v1, a foundation model trained on 161 permissible datasets to ensure compliance without sacrificing performance in English, Math, and Danish tasks.
SkillEvo improves service agent performance by using multi-turn interaction feedback and governance to optimize skills while preventing knowledge base degradation.
SAEVerbalizer automates the interpretation of LLM internal features by directly converting sparse autoencoder decoder directions into natural language explanations.
LiveAnimate is a diffusion-based framework that enables stable long-form human animation for real-time applications by optimizing architecture and inference processes.
Alaya-EVOKE uses an external, camera-indexed state bank to enable long-horizon interaction and persistent memory in world models without expanding the computational footprint.
DreamX-Phi 1.0 is a video world model that generates physically coherent future frames from robot action sequences using a diffusion based transformer architecture.
The paper introduces DARTree, a method that uses causal correction and candidate trees to accelerate autoregressive language model inference using diffusion-based drafters.
A research study demonstrates an AI agent successfully executing a large-scale architectural refactoring of a 717,725-line TypeScript application without human code review or an existing test oracle.
RippleMem introduces an associative memory architecture that connects dialogue history into an event-centric graph to solve the evidence access and completion problem for long-term LLM agent memory.