AI research from July 2026 Page 6
Browse source-linked, plain-English summaries of AI and machine-learning papers published in July 2026.
Research results
VideoChat3 is an open-source video language model that improves generalization and computational efficiency for video understanding tasks through a specialized 3D visual architecture and multi-stage instruction tuning.
The paper introduces Wan-Streamer v0.3 to enable a general-purpose pretraining objective for native-streaming generation by reframing video as a world plus an event stream.
The BadWAM framework exposes vulnerabilities in world-action models by using black-box optimization to force robots into task-failing actions via small visual perturbations.
SearchOS-V1 is a framework that improves information-seeking tasks by externalizing search state and coordinating agents through shared persistent artifacts.
MeanFlowNFT introduces a method to apply reinforcement learning to MeanFlow generators by creating an induced instantaneous velocity predictor to define reward optimization.
The researchers developed RoboTTT to enable robots to handle long sequences of actions by storing historical context through integrated test-time training layers.
HoloGeo is a framework that improves image geo-localization accuracy by training models to reason beyond superficial visual landmarks using evidence-driven reinforcement learning.
The paper introduces a lightweight interface that allows frozen foundation models to dynamically decide when to handle tasks independently or delegate to a stronger fallback model.
The paper introduces MIDI-RAE-JEPA, a model that learns hierarchical music representations by treating piano rolls as images and enforcing geometric constraints on latent space.
The paper introduces generative compilation, a method that provides on-the-fly feedback to large language models during code generation to reduce syntax errors and improve correctness.
The paper introduces AgentCompass, a unified evaluation infrastructure that decouples agent evaluation into independent components to address fragmentation and inconsistent baselines in autonomous agent testing.
The paper introduces VSI-Super-Wild, a benchmark designed to evaluate how multimodal large language models construct and maintain 3D world representations from unconstrained, long-horizon video streams.
Hallo4D introduces a multi-modal hallucination detection and correction framework to fix spatial and temporal inconsistencies in 3D and 4D content generation.
The authors introduce OT-ICA, a new signal separation algorithm that uses optimal transport distances to resolve common failures in traditional independent component analysis.
The paper introduces MOJO, a dual-pathway model that leverages both supervised and self-supervised learning to decode neural activity more effectively than traditional purely supervised methods.
The paper introduces a method to learn difficult robot manipulation tasks by automatically collecting and refining data from reversed easy tasks to lower teleoperation costs.
The paper introduces UESF-Bench to unify the tasks of searching for a target in an unexplored environment and subsequently following that target.
The paper introduces a framework for embodied agents that utilizes visual navigation, interactive intent disambiguation, and reinforcement learning to perform physical manipulation tasks in open world settings.
The paper introduces a hardware and software benchmarking platform to standardize evaluation of industrial dexterous manipulation tasks and proposes a multimodal diffusion-based policy for improved performance.
This paper demonstrates that probes trained on genomic foundation model representations from Evo 2 can effectively detect antimicrobial resistance and bacterial virulence in metagenomic data, achieving high accuracy without retraining for short reads.