Research Feed Page 26

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Agents / Benchmarks & Evals By Zhuoyang Qian 2026-08-12 30
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Spark-to-Paper integrates paper generation into existing coding assistants as a composable workflow that verifies experimental results and minimizes hallucination.

Multimodal / Efficiency & Inference By Yilin Liu 2026-08-12
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

QV-PIC is a query-aware caching framework that recovers lost textual details in visual inputs to improve RAG latency and quality.

Agents / Efficiency & Inference By Josef Liyanjun Chen 2026-08-12
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

The paper introduces a method to group control transitions in LLM agents to improve GPU execution efficiency and reduce latency by avoiding host round trips.

Agents / Benchmarks & Evals By Yuzhong Shen 2026-08-12
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

An agentic workflow successfully modernized tens of thousands of lines of legacy Fortran code by using specialized agent roles, version-controlled specifications, and exact verification.

Agents / Benchmarks & Evals By Ankita Rajaram Naik 2026-08-12
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

VAKRA evaluates how well AI agents perform complex multi-step reasoning by combining structured API calls with document retrieval.

Agents / Benchmarks & Evals By Tao Yu 2026-08-12
SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

The paper introduces SCOPE-Router, a cost-aware system that assigns tasks to the most suitable vision-language models for execution-oriented workflows.

Agents / Reasoning By Cheng Qian 2026-08-12
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Researchers developed a method where a strong builder model automatically generates and refines scaffolding logic to improve the performance of smaller language models on reasoning tasks.

Computer Vision By Ryosei Hara 2026-08-12
Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

The paper introduces a dedicated, accurate standalone visibility estimation method for hand keypoints that outperforms existing auxiliary approaches.

Agents / Safety & Alignment By Albus W. Ng 2026-08-11
Agent Safety Should Be a Runtime Contract

The paper introduces a system of preventive and evidential layers to secure autonomous agents through runtime contracts rather than relying on training-time model alignment.

Agents / Safety & Alignment By Yutao Mou 2026-08-12 1
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

ToolHazard provides a scalable framework to automatically synthesize stateful environments and generate adversarial tasks for evaluating and aligning LLM-based agents.

Robotics / Efficiency & Inference By Quanquan Peng 2026-08-10
FACT: Failure-Aware Causal Training for World-Action Models

The paper introduces FACT, a causal world model that improves robotic action planning by explicitly learning from both successful and failed outcomes.

Multimodal / Computer Vision By Kang He 2026-08-11
Sekai2: From World Exploration to Interactive World Modeling

The researchers developed Sekai2, a large-scale video dataset designed to provide the temporal continuity and camera data necessary for training interactive world models.

Agents / Reinforcement Learning By Haiyu Wu 2026-08-11
VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

The researchers developed a metric called VIScore to diagnose how well latent world models translate their internal representations into effective planning for robotics tasks.

Benchmarks & Evals By Pinzhen Chen 2026-08-10 6
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

The paper introduces Cultivar, a new evaluation framework that detects data contamination and measures how well translation models handle locale-specific cultural nuances.

Training & Fine-Tuning / Benchmarks & Evals By Ye Kyaw Thu 2026-08-11
myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

The researchers constructed a Burmese medical speech corpus and fine-tuned Whisper models to improve automatic speech recognition for clinical dialogues.

Robotics / Training & Fine-Tuning By Wenrui Bao 2026-08-11
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

The researchers developed a World-Action Model that leverages action-free video pretraining to significantly boost the performance of surgical robots when labeled demonstration data is limited.

Robotics / Computer Vision By Raphael Lorenzo-Louis 2026-08-11
HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

The researchers developed the HUI360 dataset and baseline models to predict when humans will physically interact with a mobile robot.

Computer Vision / Reasoning By Jiayu Ding 2026-08-11
CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

CausalSplat enables 3D Gaussian Splatting systems to interpret complex user instructions by mapping visual data to a structured scene graph.

Reinforcement Learning / Computer Vision By Bowei Liu 2026-08-11
VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

The researchers developed a reinforcement learning approach that uses verifiable temporal grounding to improve the accuracy of detecting AI-generated video forgeries.

Agents / Reasoning By Jian Zhang 2026-08-11
Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching

The paper introduces a framework that improves patent matching accuracy by using an LLM-driven process to mine technical entities and construct hierarchical ontologies for enhanced query retrieval.