Research Feed Page 7
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Agent-G2 optimizes agent training by using Gaussian-based guidance to sample expert trajectory lengths, achieving higher success rates at a fraction of the cost of traditional probing methods.
The authors present an end-to-end framework to automatically generate instruction-tuning and benchmark datasets from complex industrial technical documents.
The researchers introduce a strategy called ALGN to improve user representation learning by reducing redundant behavioral data and optimizing model capacity.
Researchers developed a benchmark called FORGE to measure how easily LLMs can be tricked into recommending fake products through search-augmented content.
This research demonstrates that advanced multi-hop retrieval systems significantly increase the performance degradation caused by upstream automatic speech recognition errors compared to simpler retrieval methods.
Researchers developed an injection attack method called InjecMEM that can override an agent's memory by manipulating the content stored in its retrieval systems.
MobilePA-Bench is a stateful, tool-centric benchmark environment designed to evaluate how well mobile planner agents handle complex, multi-step tasks.
The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.
Prime Agent is a framework that enables language models to recursively invoke subagents and manage persistent state to improve performance on complex autonomous tasks.
EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.
ProxyFormer reduces memory overhead by compressing long input sequences into proxy states to allow for significantly larger context processing.
ReWorld enables interactive video generation with consistent long-term spatial memory by using an efficient chunk-based caching strategy.
SecOPD improves AI agent security against adaptive prompt injection by using on-policy distillation to provide fine-grained training signals that distinguish between trusted instructions and malicious data.
Apodex 1.1 provides a general purpose agentic system that scales intelligence for complex professional tasks across finance and science using a robust execution framework.
AutoSaddler optimizes agent harnesses by analyzing execution traces to automatically refine performance across complex benchmarks.
This study audits how expanding retrieval corpora causes inconsistency in agent responses even when the model and prompt remain unchanged.
DECOWAM is a new model architecture that optimizes how legged robots coordinate whole body actions with visual environment predictions.
The paper introduces a method called Harness Evolution that improves agent adaptability and performance for live-streaming environments by decoupling execution settings from the base model.
SENTRY replaces subjective change management questionnaires with a machine learning pipeline that uses gradient boosted trees and retrieval augmented generation to predict risk.
The paper demonstrates that perfectly truthful calibration is mathematically impossible for sequential predictors and provides new methods for approximate truthfulness.