Research Feed Page 2

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Agents By Yifan Shi 2026-08-26 6
A Programming Paradigm for Spatiotemporal Composability

The authors introduce a framework called Cordis that uses a context paradigm to enable reliable dynamic composition of software components.

Multimodal / Benchmarks & Evals By Mengdi Qin 2026-08-26
RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps

RotDroid uses a vision-language model to detect GUI rotation bugs by comparing visual states between portrait and landscape orientations.

Benchmarks & Evals By Aida Usmanova 2026-08-26
How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

Researchers evaluated automated fact checking systems across four datasets to reveal how domain differences and retrieval performance impact overall accuracy.

Efficiency & Inference / Benchmarks & Evals By Timo Breuer 2026-08-25
SQLite is Enough. Lexical, Semantic, and Hybrid Search with scrydb

The paper introduces scrydb to enable combined lexical and semantic search capabilities within a single SQLite database file.

Agents / Multimodal By Yiwen Chen 2026-08-26
Code World Model: Coding Agent as World Brain

The paper introduces a framework where a coding agent generates deterministic code to manage world state, which then guides a video model to maintain visual consistency in simulated environments.

Robotics / Efficiency & Inference By Zhe Liu 2026-08-26
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI adds historical context to vision-language-action models to improve robotic task performance without increasing the model parameter count.

Agents / Benchmarks & Evals By Tianyi Xiong 2026-08-25 11
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

RubSE improves AI code generation for web pages by using structured visual rubrics to guide iterative, self-evolving refinements.

Agents / Benchmarks & Evals By Guo Gan 2026-08-25 12
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

The paper introduces AnTrap, a benchmark that tests how Android GUI agents handle dynamic environmental anomalies by injecting perturbations into 236 tasks.

Agents / Benchmarks & Evals By Liangcai Su 2026-08-25 91
FrontierChallenge: Evaluating Scientific Workflow Completion

The paper introduces FrontierChallenge, a benchmark for evaluating how well AI agents complete end-to-end scientific workflows.

Agents / Efficiency & Inference By Zhifei Xie 2026-08-26
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

VoiceMem is a dual-brain architecture designed to provide accurate, low-latency memory retrieval for speech-based conversational agents.

Computer Vision / Robotics By Yueen Ma 2026-08-26
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

The 4DGS-WAM model enables future video prediction by explicitly decomposing scenes into dynamic objects and a static background using 4D Gaussian Splatting.

Agents / Multimodal By Min Zeng 2026-08-26
VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

The researchers developed an agentic, iterative framework called VISA to generate high-quality training data for multimodal models by using feedback-driven loops instead of static one-pass pipelines.

Efficiency & Inference / Reinforcement Learning By Niklas Muennighoff 2026-08-26
Prefix Sliding for efficient test-time scaling

Prefix Sliding enables large language models to perform reasoning tasks three times faster without requiring additional training.

Agents / Benchmarks & Evals By Jiarui Yan 2026-08-26
TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

The paper introduces a planning framework for AI coding agents that aligns their development processes with human practices to improve task performance.

Reasoning / Multimodal By Junxiang Xu 2026-08-26
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

The authors introduce VBVR-Pro, a suite designed to improve visual reasoning capabilities in models by using verifiable generative tasks.

Agents / Efficiency & Inference By Guibin Zhang 2026-08-26 25
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

JIT-Agent improves agent performance by dynamically generating and evolving task-specific control structures just in time to meet individual task demands.

Agents / Benchmarks & Evals By Srimonti Dutta 2026-08-26
Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

The paper introduces trace integrity metrics to detect silent failures where LLM data agents produce correct answers through invalid logical steps.

Agents / Efficiency & Inference By Somgyuan Li 2026-08-26
ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs

ProgRouter optimizes multi-agent workflows by dynamically selecting models based on progress and cost to maximize task completion rates within defined energy budgets.

Reinforcement Learning / Agents By Shidong Yang 2026-08-25
SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents

SkillForge introduces a system that distills and verifies reusable skills for agents, significantly improving performance on complex tasks.

Agents / Benchmarks & Evals By Zhiqing Cui 2026-08-24
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

The paper introduces EarthVerse, a benchmark designed to evaluate how accurately scientific agents perform end to end investigations involving Earth systems and natural hazards.