AI research from July 2026 Page 7

Browse source-linked, plain-English summaries of AI and machine-learning papers published in July 2026.

Active filters Date: July 2026
Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Computer Vision / Benchmarks & Evals By Mustafa Chasmai 2026-07-15
MetaPerch: Learning from metadata for bioacoustics foundation models

The paper introduces MetaPerch, a bioacoustics foundation model that jointly trains on primary species identification and auxiliary metadata prediction tasks to improve generalization against domain shifts.

Training & Fine-Tuning / Efficiency & Inference By Qingyu Zhang 2026-07-14
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

ShortOPD uses a dynamic distillation strategy to fix structural collapse in pruned LLMs by adjusting training rollouts based on model output quality.

Agents / Efficiency & Inference By Ruhan Wang 2026-07-14
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The paper introduces the Harness Handbook, a behavior-centric documentation system that helps agents and developers navigate and modify large, complex agent codebases.

Agents / Efficiency & Inference By Yunxin Li 2026-07-15
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

KnowAct-GUIClaw is a framework for autonomous GUI manipulation that utilizes a memory-driven, two-tier architecture to manage complex cross-platform tasks.

Multimodal / Reinforcement Learning By Shiyin Lu 2026-07-15
OvisOCR2 Technical Report

OvisOCR2 is a model designed to parse visually rich documents into structured Markdown in a single pass.

Benchmarks & Evals By Serkan Ballı 2026-07-15
The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol

This paper identifies a problem where AI-generated evaluation data for language models can silently contain hidden errors, leading to misleading results, and proposes a mandatory manual check to prevent these issues.

Robotics / Efficiency & Inference By GigaWorld Team 2026-07-15
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld-Policy-0.5 improves real-time robotic control by decoupling action generation from future video simulation using a specialized Mixture-of-Transformers architecture.

Agents / Safety & Alignment By Yiheng Huang 2026-07-15
ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy

ProfMalPlus uses a multi-agent reasoning framework to detect malicious NPM packages by combining static code analysis with dynamic verification.

Multimodal / Benchmarks & Evals By Anders Sjöberg 2026-07-15
Multimodal Empirical Bayes Variational Autoencoders for Joint Longitudinal and Time-to-Event Modeling

The paper introduces a hybrid variational autoencoder framework that integrates longitudinal tumor growth measurements with time-to-event outcomes using genomic data to improve predictive accuracy.

Reinforcement Learning / Robotics By Slava Andrejev 2026-07-15
Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum

The paper uses the Lyapunov characteristic exponent as a reward signal to teach reinforcement learning agents how to stabilize an inverted pendulum with vertical motion.

Efficiency & Inference / Benchmarks & Evals By Daniel Grillmeyer 2026-07-15
Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study

The researchers introduced Cluster-based Sequential Feature Selection (CSFS), a wrapper method that reduces the computational cost of feature selection in renewable energy prediction pipelines.

Multimodal / Computer Vision By Zhihao Xie 2026-07-15
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

VideoRAE replaces pixel-focused autoencoders with a system that maps video foundation model features into more efficient and semantically aware latent representations.

Agents / Safety & Alignment By Sanket Badhe 2026-07-15
Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation

The researchers developed a security framework to protect the lifecycle of reusable LLM agent skills from creation through execution.

Agents / Reinforcement Learning By Leitian Tao 2026-07-15
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

The researchers developed a method to assign rewards to intermediate steps in agent interactions to improve performance on long-horizon tasks.

Agents / Benchmarks & Evals By Wenxiao Wang 2026-07-15
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

The paper introduces RELAI-VCL, a regression-aware optimizer that prevents agent performance from dropping when learning new tasks.

Safety & Alignment / Benchmarks & Evals By Mohammad Allahbakhsh 2026-07-15
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

This paper proposes a new method for penetration testing AI-enabled systems that focuses on violating operational objectives through AI-governed behavior rather than just compromising resources.

Reinforcement Learning / Efficiency & Inference By Mustafa Emre Gürsoy 2026-07-15
Lighthouse RL: Sample-Efficient Circuit Optimization via Strategic Reset Points

The paper introduces Lighthouse RL, a reinforcement learning method that uses strategic reset points to improve sample efficiency in analog circuit sizing.

Training & Fine-Tuning By Katie Everett 2026-07-15
Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

The paper identifies how architecture choices like normalization placement and width expansion prevent gradient rank collapse in deep Transformer models at initialization.

Agents / Benchmarks & Evals By Maliha Noushin Raida 2026-07-15
Early Adoption of Agentic Coding Tools by GitHub Projects

The paper examines how developers integrate agentic coding tools into their workflows by analyzing pull request patterns across 2,361 GitHub repositories.

Computer Vision / Efficiency & Inference By Zhan Chen 2026-07-15
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study

The authors implement a domain routing system that dispatches page images to specialized OCR models, enabling accurate text extraction across varied historical Manchu writing styles.