Research Feed Page 46

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Efficiency & Inference / Reasoning By Niels Mündler-Sasahara 2026-07-15
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

The paper introduces generative compilation, a method that provides on-the-fly feedback to large language models during code generation to reduce syntax errors and improve correctness.

Benchmarks & Evals By Zichen Ding 2026-07-15
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

The paper introduces AgentCompass, a unified evaluation infrastructure that decouples agent evaluation into independent components to address fragmentation and inconsistent baselines in autonomous agent testing.

Multimodal / Benchmarks & Evals By Tianjun Gu 2026-07-15
Towards Spatial Supersensing in the Wild

The paper introduces VSI-Super-Wild, a benchmark designed to evaluate how multimodal large language models construct and maintain 3D world representations from unconstrained, long-horizon video streams.

Computer Vision / Multimodal By Hongbo Wang 2026-07-15
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Hallo4D introduces a multi-modal hallucination detection and correction framework to fix spatial and temporal inconsistencies in 3D and 4D content generation.

Efficiency & Inference By Ashutosh Jha 2026-07-15
Linear Independent Component Analysis via Optimal Transport

The authors introduce OT-ICA, a new signal separation algorithm that uses optimal transport distances to resolve common failures in traditional independent component analysis.

Training & Fine-Tuning By Ximeng Mao 2026-07-15
Leveraging unlabelled data for generalizable neural population decoding

The paper introduces MOJO, a dual-pathway model that leverages both supervised and self-supervised learning to decode neural activity more effectively than traditional purely supervised methods.

Robotics / Reinforcement Learning By Qiyuan Qiao 2026-07-15
Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks

The paper introduces a method to learn difficult robot manipulation tasks by automatically collecting and refining data from reversed easy tasks to lower teleoperation costs.

Agents / Benchmarks & Evals By Kun Yu 2026-07-15
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

The paper introduces UESF-Bench to unify the tasks of searching for a target in an unexplored environment and subsequently following that target.

Robotics / Reinforcement Learning By Boyu Mi 2026-07-15
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

The paper introduces a framework for embodied agents that utilizes visual navigation, interactive intent disambiguation, and reinforcement learning to perform physical manipulation tasks in open world settings.

Robotics / Benchmarks & Evals By Honglu He 2026-07-15
Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation

The paper introduces a hardware and software benchmarking platform to standardize evaluation of industrial dexterous manipulation tasks and proposes a multimodal diffusion-based policy for improved performance.

Safety & Alignment / Efficiency & Inference By Jeremy Guntoro 2026-07-15
Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

This paper demonstrates that probes trained on genomic foundation model representations from Evo 2 can effectively detect antimicrobial resistance and bacterial virulence in metagenomic data, achieving high accuracy without retraining for short reads.

Computer Vision / Benchmarks & Evals By Mustafa Chasmai 2026-07-15
MetaPerch: Learning from metadata for bioacoustics foundation models

The paper introduces MetaPerch, a bioacoustics foundation model that jointly trains on primary species identification and auxiliary metadata prediction tasks to improve generalization against domain shifts.

Training & Fine-Tuning / Efficiency & Inference By Qingyu Zhang 2026-07-14
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

ShortOPD uses a dynamic distillation strategy to fix structural collapse in pruned LLMs by adjusting training rollouts based on model output quality.

Agents / Efficiency & Inference By Ruhan Wang 2026-07-14
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The paper introduces the Harness Handbook, a behavior-centric documentation system that helps agents and developers navigate and modify large, complex agent codebases.

Agents / Efficiency & Inference By Yunxin Li 2026-07-15
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

KnowAct-GUIClaw is a framework for autonomous GUI manipulation that utilizes a memory-driven, two-tier architecture to manage complex cross-platform tasks.

Multimodal / Reinforcement Learning By Shiyin Lu 2026-07-15
OvisOCR2 Technical Report

OvisOCR2 is a model designed to parse visually rich documents into structured Markdown in a single pass.

Benchmarks & Evals By Serkan Ballı 2026-07-15
The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol

This paper identifies a problem where AI-generated evaluation data for language models can silently contain hidden errors, leading to misleading results, and proposes a mandatory manual check to prevent these issues.

Robotics / Efficiency & Inference By GigaWorld Team 2026-07-15
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld-Policy-0.5 improves real-time robotic control by decoupling action generation from future video simulation using a specialized Mixture-of-Transformers architecture.

Agents / Safety & Alignment By Yiheng Huang 2026-07-15
ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy

ProfMalPlus uses a multi-agent reasoning framework to detect malicious NPM packages by combining static code analysis with dynamic verification.

Multimodal / Benchmarks & Evals By Anders Sjöberg 2026-07-15
Multimodal Empirical Bayes Variational Autoencoders for Joint Longitudinal and Time-to-Event Modeling

The paper introduces a hybrid variational autoencoder framework that integrates longitudinal tumor growth measurements with time-to-event outcomes using genomic data to improve predictive accuracy.