Research Feed Page 20
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The researchers introduced the Edit2TikZ benchmark and a curriculum learning process to help models accurately edit scientific figures using TikZ code.
The authors created a browser-native digital test range using WebAssembly to benchmark ocean-glider planning algorithms through standardized kinematic mission simulation.
The paper presents a framework to select the most efficient intervention method for multimodal models by diagnosing specific task characteristics instead of testing every approach.
SNM-VFI uses a hybrid approach combining structural flow estimation with generative diffusion models to synthesize smoother, high-quality intermediate video frames.
The paper introduces a method called Deliberate Practice to optimally distribute a fixed training budget across various robot skills to improve overall performance.
The paper demonstrates that how you select measurement tokens for sparse autoencoder evaluation significantly alters results and proposes a shared reporting protocol to address this bias.
NestDex improves dexterous robot manipulation by combining pre-trained hand skill policies with operator teleoperation using an action-compressing variational autoencoder.
AlayaWorld improves long horizon video generation by replacing traditional depth warping with a streaming 3D point cache for better geometric consistency.
The paper explores using Small Language Models on NVIDIA Jetson Orin NX hardware to handle cognitive tasks like service routing and memory management for virtual agents.
The paper introduces ARMDIL, a system that uses an MLLM router to dynamically assign images to specialized vision backbones to improve classification accuracy.
The paper introduces a Seeker module that learns to focus robot vision on relevant spatial regions, significantly increasing success rates in complex environments.
TraVEL improves driving video retrieval by training embedding models to prioritize ego-vehicle movement patterns over static visual shortcuts.
The paper demonstrates that latent world models often fail at long horizon planning because their training objectives do not align with the latent information already present in their internal representations.
The researchers developed a method to replace hand-designed improvement rules with a system that learns to generate its own contextual guidance for model training.
The authors introduce Lipschitz-regularized object detection to fix instability caused by functional mismatches when chaining image restoration with detection models.
The paper introduces CW-BASS v2 to fix pseudo-labeling errors that occur when high-performance foundation models become overly confident and biased during training.
PixSDS introduces a method to repair color artifacts and high frequency noise in latent based image generation by forcing pixel space consistency during the optimization process.
UniTraffic-Agent improves traffic video reasoning by using a shared event interpretation approach to handle multi-question queries and sparse visual data.
The authors introduce SPARED, a reasoning-based detector that uses adversarial image editing to train models to identify synthetic content without relying on provenance shortcuts.
The paper investigates whether prompts containing linguistic features associated with women negatively impact the quality of responses generated by large language models.