Research Feed Page 8

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Reinforcement Learning / Efficiency & Inference By Yunheng Li 2026-08-20
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

The researchers introduced a method called OraRL that integrates ground truth annotations as oracle rollouts to improve video model performance and reduce inference latency.

Efficiency & Inference / Multimodal By Yangshuai Liu 2026-08-21
Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

The researchers developed a text anchor method to enable high-speed reference image caching in diffusion transformers without sacrificing model performance.

Multimodal / Benchmarks & Evals By Linhan Cao 2026-08-20
ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

ArmorOCR improves how AI models read adversarial text in images by using a specialized training process and a new benchmark for evaluating robustness.

Safety & Alignment / Benchmarks & Evals By Emilio Ferrara 2026-08-20
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

Researchers tested whether language models can detect and report on internal computational interventions, finding that model confidence signals are more informative than direct verbal reports.

Benchmarks & Evals By Josef Chen 2026-08-20 2
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

FlavourBench is an automated evaluation framework that uses the Epicure culinary system to rank language models on objective, executable tasks.

Agents / Benchmarks & Evals By Afonso Baldo 2026-08-21
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

The authors introduce a method to steer LLM behavior in therapy sessions by exposing a clinical move ontology as tools, which improves alignment with human therapists.

Multimodal / Safety & Alignment By Wenzheng Jiang 2026-08-21
ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

ReFrame acts as a secure intermediary layer that analyzes and rewrites potentially unsafe multimodal inputs before they reach closed source models.

Benchmarks & Evals By Yichen Jiang 2026-08-21
ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting

The paper introduces ConceptTS, a framework that leverages Large Language Models to generate interpretable concepts for more transparent multivariate time series forecasting.

Efficiency & Inference By Jan Novacek 2026-08-21
Ontology-supported AI Model and Dataset Management

The authors introduce AIMDEP, a platform that uses a specialized ontology to manage AI assets and metadata for improved model tracking and collaboration.

Computer Vision / Training & Fine-Tuning By Marko Haralović 2026-08-21
When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning

The paper identifies that representational drift in the decoder and output layers significantly impacts out of distribution performance when fine-tuning the MedSAM medical image segmentation foundation model.

Agents / Benchmarks & Evals By Rana Muhammad Usman 2026-08-20 2
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

Researchers developed the Peer-Voted Social Simulation Testbed to measure how ranking mechanisms and coordinated adversarial sources affect the linguistic output of LLM-based agents.

Agents / Benchmarks & Evals By Lauren Pothuru 2026-08-20
When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

The paper introduces an interventional benchmark to pinpoint the specific hop in a multi-hop agentic retrieval chain where a failure originated.

Efficiency & Inference / Reasoning By Simeng Zhang 2026-08-21
Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

The paper introduces a method that improves the efficiency and accuracy of chain of thought reasoning by injecting relevant, pre-computed reasoning patterns into the model prompt.

Training & Fine-Tuning / Efficiency & Inference By Bakbergen Ryskulov 2026-08-21
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

The paper introduces Quantization-Aware Healing, a practical recipe for recovering compressed 4-bit large language models, and uses it to produce the open-weight model Hypernova-60B.

Safety & Alignment / Benchmarks & Evals By Adam Noonan 2026-08-21
The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

The paper introduces a method to calculate the effective sample size for thresholding models when calibration data is clustered and contains correlated errors.

Benchmarks & Evals / Multimodal By Elaine Lau 2026-08-21
VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

The researchers introduced VIALS, a new visual question answering benchmark designed to test how well AI models interpret complex scientific images from biotech workflows.

Benchmarks & Evals / Reasoning By Xin Sun 2026-08-21
When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

This study demonstrates that LLM judges struggle to separate content accuracy from source reliability, showing higher trust and accuracy scores when content is attributed to humans compared to AI.

Agents / Benchmarks & Evals By Wei Lin 2026-08-21
Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

The paper provides a structured survey and research agenda to bridge the gap between software engineering and security when evaluating Large Language Model agents.

Robotics / Training & Fine-Tuning By Varun Giridhar 2026-08-21
Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

The paper presents a method that enables robot policies to self-improve through iterative deployment without the need to modify the original policy weights.

Agents / Benchmarks & Evals By Jiayi Li 2026-08-21
Affective Context Amplifies Sycophancy in LLM Responses

The paper demonstrates that LLMs become significantly more sycophantic when responding to users expressing specific emotional states.