Research Feed Page 32

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Benchmarks & Evals / Efficiency & Inference By Jutao Xiao 2026-08-10
From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

The researchers introduced a system called DEC that improves table parsing by decomposing, enhancing, and correcting errors through visual consistency checks.

Agents / Benchmarks & Evals By Abraham Gonzalez 2026-08-10
ArchAgent v2: A Case Study with the Data Prefetching Championship

ArchAgent v2 uses an evolutionary search process to automate the design of complex data prefetchers that exceed existing championship-level performance.

Agents / Safety & Alignment By Wanying Qu 2026-08-10
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

The SHE framework allows LLM agent safety systems to automatically evolve over time by analyzing failure trajectories to refine safety boundaries.

Reasoning / Benchmarks & Evals By Wenyao Cui 2026-08-09 1
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

SymDiag improves LLM reasoning reliability by compiling natural language chains of thought into symbolic logic for automated diagnosis and repair.

Robotics / Reinforcement Learning By Dongchi Huang 2026-08-10
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

RynnValue is an open-source foundation model that improves robotic task success rates by predicting the remaining distance to a goal using visual observations.

Efficiency & Inference / Benchmarks & Evals By Junghwan Lim 2026-08-10 20
Motif 3: Technical Report

Motif 3 is a 314 billion parameter model that utilizes a specialized architecture and multi-stage post-training to improve intelligence and task generalization.

Efficiency & Inference By Can Xiao 2026-08-08 2
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV uses predictive prefetching to move necessary KV cache data from high capacity memory to HBM, allowing for larger decode batches and longer context without hitting the memory wall.

Agents / Benchmarks & Evals By Anton Razzhigaev 2026-08-08 5
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Ouroboros introduces a self-developing agent architecture that treats code evolution as a task to adapt tools and core logic autonomously while maintaining strict safety standards.

Benchmarks & Evals / Multimodal By Diandian Zhang 2026-08-10
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

The paper introduces Sci-VBench, a benchmark designed to test if generative video models can accurately simulate scientific mechanisms and causal dynamics.

Agents / Training & Fine-Tuning By Changhao Xiang 2026-08-09
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

OpenVisTool introduces a training method that teaches models to use external visual tools only when necessary, improving performance over fixed image encoding.

Agents / Benchmarks & Evals By Yijun Pan 2026-08-09
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

The paper introduces a controlled marketplace environment to evaluate the long term business performance and decision making capabilities of LLM agents.

Agents / Benchmarks & Evals By Lisheng Huang 2026-08-10 3
Evo-Bench: Can Language Models Improve Agent Harness?

The paper introduces Evo-Bench to evaluate how effectively large language models can autonomously evolve and improve the code harnesses that control agent behavior.

Safety & Alignment / Efficiency & Inference By Alexander Panfilov 2026-08-10
Stealing Reasoning Traces from Proprietary LLM APIs

Researchers discovered an architectural flaw in how major LLM providers handle encrypted reasoning traces, allowing them to decrypt and expose proprietary data.

Safety & Alignment / Benchmarks & Evals By Yingtao Ren 2026-08-07
When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse

The researchers developed D-SCAN, a method to detect malicious documents injected into RAG pipelines by identifying specific anomalies in how the model distributes its attention.

Agents / Benchmarks & Evals By Yuling Shi 2026-08-10
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

The paper introduces SWE-Bench ProMax, a new, expert-curated benchmark designed to evaluate AI coding agents on complex, large-scale, multilingual refactoring tasks.

Agents / Reinforcement Learning By Mind Lab 2026-08-10
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Macaron-V1 introduces a framework for deploying persistent agent models that update themselves through specialized adapters and recursive self-improvement loops.

Benchmarks & Evals / Reasoning By Rodrigo Ferreira Rodrigues 2026-08-07
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

The authors introduce GeoBenchLLM, a unified benchmark consisting of 421,041 questions across twelve datasets to evaluate LLM performance on geographic tasks.

Robotics / Benchmarks & Evals By Haodong Yan 2026-08-07
Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

The authors introduce PSG-JEPA, a model that adds physical grounding to latent world models to improve performance in real-world robotic manipulation tasks.

Benchmarks & Evals / Computer Vision By Yihui Li 2026-08-07
ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling

The paper introduces ArchEGraph, a dataset that maps building geometry and climate conditions to thermal energy loads to improve building energy modeling through graph-based machine learning.

Agents / Benchmarks & Evals By Bo Tang 2026-08-07
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

The paper introduces the PHASE-Tree framework to manage character personality updates across long narrative sequences in role-playing agents.