Research Feed Page 29
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
MedPixel combines visual reasoning and image segmentation into a single architecture to bridge the gap between clinical text and pixel-level data.
InSight-doc uses an agentic system that zooms into document regions to reduce computational overhead and hallucination in multimodal models.
Researchers built GitSkills, a dataset containing over 3.7 million agent instructions scraped from public GitHub repositories.
MultiModal Code-Switching improves multimodal model alignment by replacing text tokens with visual object embeddings during pretraining.
Ex-Omni-2D generates multimodal dialogue responses that natively combine text, personalized speech, and reference-conditioned video to overcome the limitations of visually disembodied models.
The authors introduce EntLORE, a graph-grounded benchmark designed to evaluate how well systems perform complex organizational reasoning beyond simple fact retrieval.
The paper introduces a technique called ASMI that measures model uncertainty by observing how responses change when random paths in the transformer's attention mechanism are disrupted.
The paper introduces a framework that allows GUI visual grounding models to continuously improve after deployment by learning from their own exploration failures through reflection-guided self-distillation.
The paper introduces JEPA-WAM, a framework that integrates spatially structured world modeling with vision-language-action policies to improve performance and robustness against distribution shifts.
The researchers developed a reference-free post-training method to optimize machine translation models using only source-side text, bypassing the need for high-quality parallel data.
DistilVDR creates compact single-vector document retrieval systems by distilling knowledge from large vision-language models into significantly smaller student encoders.
The paper identifies that agentic README files grow indefinitely due to catastrophic remembering, where the rationale for instructions is lost, and proposes a comment syntax to safely manage this metadata.
The paper presents a three-stage taxonomy to classify and structure how multiple agents and their environments can iteratively adapt to one another beyond static, single-entity learning models.
The paper introduces a new framework called ComBodied Agents that focuses on supporting human wellbeing and long term goals rather than just executing isolated tasks.
VeriFin uses a neurosymbolic framework to ground LLM financial claims in source document data and verify them using formal constraint solving.
The paper introduces a runtime architecture that enables diffusion language models to interact with tools asynchronously during the reasoning process.
SkillZip optimizes agent instructions by identifying and removing redundant rules and workflows without requiring external task-based testing.
The paper introduces DSAgentBench, a new benchmark designed to evaluate how effectively AI agents automate end-to-end data science workflows in realistic computer environments.
The paper introduces a new measurement protocol to evaluate if AI agents execute the same tool-use action sequences across different languages.
SPIEval is a new human-curated benchmark designed to evaluate how effectively large language models handle complex tasks using personal information scattered across mobile applications.