Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers developed a vision-language-action model designed to improve how multiple robotic arms collaborate on complex tasks by using techniques that enforce role-agnostic instruction following.
The paper introduces a method that uses vision-language models to perform reasoning that guides robot manipulation policies, improving performance on long-horizon tasks.
StreamPI adds historical context to vision-language-action models to improve robotic task performance without increasing the model parameter count.
The 4DGS-WAM model enables future video prediction by explicitly decomposing scenes into dynamic objects and a static background using 4D Gaussian Splatting.
The LAWA architecture optimizes robot action planning by using latent intentions to reduce inference latency while maintaining high success rates across robotics benchmarks.
WarpSAC is a scalable reinforcement learning framework that adapts its architecture based on available compute resources to accelerate training and improve deployment success.
OptiSight combines semantic object identification with geometric control to enable efficient robot navigation while minimizing reliance on high-frequency language model inference.
DECOWAM is a new model architecture that optimizes how legged robots coordinate whole body actions with visual environment predictions.
The paper presents a method that enables robot policies to self-improve through iterative deployment without the need to modify the original policy weights.
The authors introduce a method to compress token data in vision-language-action models by identifying and prioritizing information that has the least impact on physical robot movements.
The authors introduce a pipeline to automatically reconstruct and retarget 3D human interaction data into a large-scale dataset for training diverse robotic embodiments.
EXIMO leverages a vision-language model to decompose complex robotic tasks into smaller steps, improving the efficiency of training vision-language-action policies.
RoMAN-Flow introduces post-training optimization and distillation techniques to eliminate the sequential sampling latency inherent in autoregressive normalizing flows for robotic control.
TAMP-Nav improves embodied navigation by combining efficient 3D spatial grounding with selective reasoning and a multi-level reward training approach.
The paper introduces ReflexVLA, a vision-language-action model architecture that uses future prediction and optimized inference to improve performance in time-sensitive robotics tasks.
The paper introduces a framework to ensure safe robotic operation in complex urban environments by defining a dynamic safety envelope rather than using static constraints.
The paper explores the development of autonomous AI systems capable of performing scientific research by integrating neural learning, robotics, and formal reasoning.
The paper introduces a feedback loop between the tracking controller and trajectory planner that adjusts spatial constraints to prevent sub-optimal performance caused by model mismatches.
The paper introduces a toolkit that assesses robotic task execution by analyzing continuous progress curves rather than relying on binary success rates.
The researchers developed a sensing framework that enables surgical robots to estimate cable tension and contact location in real time using a parallelized computation model.