Research Feed Page 7

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Agents / Reinforcement Learning By Zixuan Wang 2026-08-24
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Agent-G2 optimizes agent training by using Gaussian-based guidance to sample expert trajectory lengths, achieving higher success rates at a fraction of the cost of traditional probing methods.

Training & Fine-Tuning / Benchmarks & Evals By Parsa Bakhtiari 2026-08-24
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

The authors present an end-to-end framework to automatically generate instruction-tuning and benchmark datasets from complex industrial technical documents.

Efficiency & Inference By Bin Dou 2026-08-24
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

The researchers introduce a strategy called ALGN to improve user representation learning by reducing redundant behavioral data and optimizing model capacity.

Benchmarks & Evals / Safety & Alignment By Minghao Luo 2026-08-24 2
One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Researchers developed a benchmark called FORGE to measure how easily LLMs can be tricked into recommending fake products through search-augmented content.

Benchmarks & Evals / Agents By Zhenghua Bao 2026-08-24 3
Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

This research demonstrates that advanced multi-hop retrieval systems significantly increase the performance degradation caused by upstream automatic speech recognition errors compared to simpler retrieval methods.

Agents / Safety & Alignment By Hanling Tian 2026-08-24
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

Researchers developed an injection attack method called InjecMEM that can override an agent's memory by manipulating the content stored in its retrieval systems.

Agents / Benchmarks & Evals By Yi Zhu 2026-08-24 26
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

MobilePA-Bench is a stateful, tool-centric benchmark environment designed to evaluate how well mobile planner agents handle complex, multi-step tasks.

Multimodal / Computer Vision By Nan Duan 2026-08-24
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.

Agents / Benchmarks & Evals By Seth Karten 2026-08-24
Prime Agent: A Self-Improving RLM Harness

Prime Agent is a framework that enables language models to recursively invoke subagents and manage persistent state to improve performance on complex autonomous tasks.

Multimodal / Computer Vision By Songchun Zhang 2026-08-24 19
EchoWM: Open and Enterable Omnimodal World Models

EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.

Efficiency & Inference By Zhongpan Tang 2026-08-24
ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

ProxyFormer reduces memory overhead by compressing long input sequences into proxy states to allow for significantly larger context processing.

Multimodal / Efficiency & Inference By Zhifei Chen 2026-08-24
ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld enables interactive video generation with consistent long-term spatial memory by using an efficient chunk-based caching strategy.

Agents / Safety & Alignment By Yibo Peng 2026-08-21 19
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

SecOPD improves AI agent security against adaptive prompt injection by using on-policy distillation to provide fine-grained training signals that distinguish between trusted instructions and malicious data.

Agents / Benchmarks & Evals By Apodex Team 2026-08-24 104
Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 provides a general purpose agentic system that scales intelligence for complex professional tasks across finance and science using a robust execution framework.

Agents / Benchmarks & Evals By Sungho Park 2026-08-24
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

AutoSaddler optimizes agent harnesses by analyzing execution traces to automatically refine performance across complex benchmarks.

Benchmarks & Evals / Efficiency & Inference By Jingjie Ning 2026-08-24 1
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

This study audits how expanding retrieval corpora causes inconsistency in agent responses even when the model and prompt remain unchanged.

Robotics / Multimodal By Siyuan Ma 2026-08-20
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

DECOWAM is a new model architecture that optimizes how legged robots coordinate whole body actions with visual environment predictions.

Agents / Training & Fine-Tuning By TaoLive AIGC LLM Team 2026-08-22 1
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

The paper introduces a method called Harness Evolution that improves agent adaptability and performance for live-streaming environments by decoupling execution settings from the base model.

Benchmarks & Evals By Daniel Arulpragasam 2026-08-21
SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management

SENTRY replaces subjective change management questionnaires with a machine learning pipeline that uses gradient boosted trees and retrieval augmented generation to predict risk.

Benchmarks & Evals By Anagha Gokul 2026-08-21
Truthful Calibration Measures for Sequential Prediction

The paper demonstrates that perfectly truthful calibration is mathematically impossible for sequential predictors and provides new methods for approximate truthfulness.