Research Feed Page 23
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces QuoteBench to show how matched execution scores alone are insufficient to distinguish command generation errors from failures introduced after generation by interfaces.
The paper introduces Vero, a system that tests whether AI agents can build and verify multi-module software repositories with mathematical correctness guarantees.
MARC v1 replaces monolithic LLM prompting with a deterministic multi-agent framework to improve clinical reasoning and enable step by step error tracking.
LycheeMemory V2 introduces semantic segment level consolidation to efficiently preserve long-term conversational memory for LLM agents while cutting construction costs.
The paper introduces StateBridge, a training-free alignment method that lets off-the-shelf LLM agents communicate directly via continuous hidden representations instead of discrete text tokens.
This paper investigates how to set per-token prices and default reasoning-token allocations for an LLM reasoning service where users can accept defaults, customize allocations, or exit.
The paper introduces Reduced Matrix Multiplication, an input-adaptive method to reduce high-dimensional matrix multiplications during transformer inference without modifying model weights.
The paper introduces Hybrid Gated Attention, a technique that improves training stability and performance in language models by modifying how attention mechanisms handle gating, matrix factorization, and head interactions.
The paper introduces Intern-S2-Preview, a foundation model designed for multimodal scientific understanding, reasoning, and long-horizon agentic task execution.
The paper introduces the Spatial Memory Agent, a runtime framework that equips frozen vision-language models with experience-grounded procedure memory to improve spatial reasoning without updating model parameters.
Confucius4-TTS enables zero-shot cross-lingual text-to-speech without requiring transcripts of the reference audio.
The paper introduces a framework that decomposes compound answer options into atomic statements to help language models correctly evaluate explicit logical operators like And, Or, and Neither/Nor.
The paper introduces MiDashengLM-Gen, an end-to-end framework that couples a pre-trained language model with per-token conditional flow matching to generate variable-length mixed audio scenes blending speech, music, and sound effects.
The paper introduces GazeAnywhere, a flexible model that estimates human gaze targets in the wild using natural language prompting without relying on brittle multi-stage pipelines.
The paper introduces a physics-grounded reflection simulation and diffusion-based video dereflection pipeline to remove unwanted glass reflections from videos.
The FQTree method uses fine-grained quantization to reduce resource usage while maintaining high accuracy for boosted decision tree models deployed on FPGAs.
The authors developed an automated framework using Retrieval-Augmented Generation and Large Language Models to construct Dynamic Master Logic models as Knowledge Graphs, overcoming the scalability limits of manual expert interpretation.
The paper introduces SCOUT, a method combining structured chain-of-thought prompting and multi-objective reinforcement learning to fix spatial reasoning bottlenecks in vision-language models.
The paper introduces HAMP-LIC, a Hessian-aware mixed-precision quantization method that shrinks learned image compression models while preserving image quality and eliminating cross-platform decoding mismatches.
The paper introduces a method called Motion-as-Prompt to improve multimodal large language models by adding motion-guided visual markers to video frames.