Research Feed
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
Magpie is a real-time renderer that uses foundation models to generate game visuals by processing white-box frames as a continuous denoising condition.
The paper introduces PAWBench, a new evaluation framework that measures how well video generation models predict the physical outcomes of various scenarios.
TAU-Agent is an agentic framework designed to improve traffic anomaly detection by using retrieval-augmented generation to integrate video descriptions and object trajectories.
PlanSightRAG leverages a vision-first multimodal framework to automate compliance checking for complex civil engineering drawings.
The researchers developed a vision-language-action model designed to improve how multiple robotic arms collaborate on complex tasks by using techniques that enforce role-agnostic instruction following.
RotDroid uses a vision-language model to detect GUI rotation bugs by comparing visual states between portrait and landscape orientations.
The paper introduces a framework where a coding agent generates deterministic code to manage world state, which then guides a video model to maintain visual consistency in simulated environments.
The researchers developed an agentic, iterative framework called VISA to generate high-quality training data for multimodal models by using feedback-driven loops instead of static one-pass pipelines.
The authors introduce VBVR-Pro, a suite designed to improve visual reasoning capabilities in models by using verifiable generative tasks.
The paper introduces a Mixture of Task Experts architecture that uses task-specific modules within a video-language decoder to improve performance across diverse video understanding tasks.
The paper introduces a framework and an autonomy classification system for deploying Large Language Model agents in scientific molecular discovery workflows.
The paper evaluates various retrieval-augmented generation pipelines, finding that multimodal vision-based approaches significantly outperform text-based methods despite introducing higher latency and storage costs.
The paper introduces WeMM-Embedding, a series of multimodal models based on Qwen3.5 that achieve state-of-the-art performance on retrieval benchmarks and demonstrate consistent gains in production applications.
The paper introduces IntentQA and the X-CaVIR framework to help models infer latent human intentions in video content through cognitive context reasoning.
The authors introduce FinixDoc, an agentic parsing system powered by a specialized vision-language model, to improve accuracy and structural consistency in real-world financial document processing.
The researchers introduced JoyAI-Echo-1.5, an audio-visual generation system that maintains narrative and visual consistency over long durations.
EchoWM creates an enterable virtual environment that generates synchronized video, audio, and speech based on user navigation inputs.
ReWorld enables interactive video generation with consistent long-term spatial memory by using an efficient chunk-based caching strategy.
DECOWAM is a new model architecture that optimizes how legged robots coordinate whole body actions with visual environment predictions.
The researchers developed a text anchor method to enable high-speed reference image caching in diffusion transformers without sacrificing model performance.