AI research from July 2026 Page 6

Browse source-linked, plain-English summaries of AI and machine-learning papers published in July 2026.

Active filters Date: July 2026
Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Multimodal / Efficiency & Inference By Xinhao Li 2026-07-16
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

VideoChat3 is an open-source video language model that improves generalization and computational efficiency for video understanding tasks through a specialized 3D visual architecture and multi-stage instruction tuning.

Multimodal / Agents By Lianghua Huang 2026-07-16
Video = World + Event Stream

The paper introduces Wan-Streamer v0.3 to enable a general-purpose pretraining objective for native-streaming generation by reframing video as a world plus an event stream.

Robotics / Safety & Alignment By Qi Li 2026-07-16
BadWAM: When World-Action Models Dream Right but Act Wrong

The BadWAM framework exposes vulnerabilities in world-action models by using black-box optimization to force robots into task-failing actions via small visual perturbations.

Agents / Benchmarks & Evals By Yuyao Zhang 2026-07-16
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

SearchOS-V1 is a framework that improves information-seeking tasks by externalizing search state and coordinating agents through shared persistent artifacts.

Reinforcement Learning / Efficiency & Inference By Yushi Huang 2026-07-16
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlowNFT introduces a method to apply reinforcement learning to MeanFlow generators by creating an induced instantaneous velocity predictor to define reward optimization.

Robotics / Efficiency & Inference By Yunfan Jiang 2026-07-16
RoboTTT: Context Scaling for Robot Policies

The researchers developed RoboTTT to enable robots to handle long sequences of actions by storing historical context through integrated test-time training layers.

Multimodal / Benchmarks & Evals By Pengcheng Zhou 2026-07-16
HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning

HoloGeo is a framework that improves image geo-localization accuracy by training models to reason beyond superficial visual landmarks using evidence-driven reinforcement learning.

Agents / Efficiency & Inference By Amirhosein Ghasemabadi 2026-07-15
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making

The paper introduces a lightweight interface that allows frozen foundation models to dynamically decide when to handle tasks independently or delegate to a stronger fallback model.

Benchmarks & Evals / Computer Vision By Scott H. Hawley 2026-07-16
MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

The paper introduces MIDI-RAE-JEPA, a model that learns hierarchical music representations by treating piano rolls as images and enforcing geometric constraints on latent space.

Efficiency & Inference / Reasoning By Niels Mündler-Sasahara 2026-07-15
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

The paper introduces generative compilation, a method that provides on-the-fly feedback to large language models during code generation to reduce syntax errors and improve correctness.

Benchmarks & Evals By Zichen Ding 2026-07-15
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

The paper introduces AgentCompass, a unified evaluation infrastructure that decouples agent evaluation into independent components to address fragmentation and inconsistent baselines in autonomous agent testing.

Multimodal / Benchmarks & Evals By Tianjun Gu 2026-07-15
Towards Spatial Supersensing in the Wild

The paper introduces VSI-Super-Wild, a benchmark designed to evaluate how multimodal large language models construct and maintain 3D world representations from unconstrained, long-horizon video streams.

Computer Vision / Multimodal By Hongbo Wang 2026-07-15
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Hallo4D introduces a multi-modal hallucination detection and correction framework to fix spatial and temporal inconsistencies in 3D and 4D content generation.

Efficiency & Inference By Ashutosh Jha 2026-07-15
Linear Independent Component Analysis via Optimal Transport

The authors introduce OT-ICA, a new signal separation algorithm that uses optimal transport distances to resolve common failures in traditional independent component analysis.

Training & Fine-Tuning By Ximeng Mao 2026-07-15
Leveraging unlabelled data for generalizable neural population decoding

The paper introduces MOJO, a dual-pathway model that leverages both supervised and self-supervised learning to decode neural activity more effectively than traditional purely supervised methods.

Robotics / Reinforcement Learning By Qiyuan Qiao 2026-07-15
Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks

The paper introduces a method to learn difficult robot manipulation tasks by automatically collecting and refining data from reversed easy tasks to lower teleoperation costs.

Agents / Benchmarks & Evals By Kun Yu 2026-07-15
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

The paper introduces UESF-Bench to unify the tasks of searching for a target in an unexplored environment and subsequently following that target.

Robotics / Reinforcement Learning By Boyu Mi 2026-07-15
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

The paper introduces a framework for embodied agents that utilizes visual navigation, interactive intent disambiguation, and reinforcement learning to perform physical manipulation tasks in open world settings.

Robotics / Benchmarks & Evals By Honglu He 2026-07-15
Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation

The paper introduces a hardware and software benchmarking platform to standardize evaluation of industrial dexterous manipulation tasks and proposes a multimodal diffusion-based policy for improved performance.

Safety & Alignment / Efficiency & Inference By Jeremy Guntoro 2026-07-15
Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

This paper demonstrates that probes trained on genomic foundation model representations from Evo 2 can effectively detect antimicrobial resistance and bacterial virulence in metagenomic data, achieving high accuracy without retraining for short reads.