Research Feed Page 46
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper introduces generative compilation, a method that provides on-the-fly feedback to large language models during code generation to reduce syntax errors and improve correctness.
The paper introduces AgentCompass, a unified evaluation infrastructure that decouples agent evaluation into independent components to address fragmentation and inconsistent baselines in autonomous agent testing.
The paper introduces VSI-Super-Wild, a benchmark designed to evaluate how multimodal large language models construct and maintain 3D world representations from unconstrained, long-horizon video streams.
Hallo4D introduces a multi-modal hallucination detection and correction framework to fix spatial and temporal inconsistencies in 3D and 4D content generation.
The authors introduce OT-ICA, a new signal separation algorithm that uses optimal transport distances to resolve common failures in traditional independent component analysis.
The paper introduces MOJO, a dual-pathway model that leverages both supervised and self-supervised learning to decode neural activity more effectively than traditional purely supervised methods.
The paper introduces a method to learn difficult robot manipulation tasks by automatically collecting and refining data from reversed easy tasks to lower teleoperation costs.
The paper introduces UESF-Bench to unify the tasks of searching for a target in an unexplored environment and subsequently following that target.
The paper introduces a framework for embodied agents that utilizes visual navigation, interactive intent disambiguation, and reinforcement learning to perform physical manipulation tasks in open world settings.
The paper introduces a hardware and software benchmarking platform to standardize evaluation of industrial dexterous manipulation tasks and proposes a multimodal diffusion-based policy for improved performance.
This paper demonstrates that probes trained on genomic foundation model representations from Evo 2 can effectively detect antimicrobial resistance and bacterial virulence in metagenomic data, achieving high accuracy without retraining for short reads.
The paper introduces MetaPerch, a bioacoustics foundation model that jointly trains on primary species identification and auxiliary metadata prediction tasks to improve generalization against domain shifts.
ShortOPD uses a dynamic distillation strategy to fix structural collapse in pruned LLMs by adjusting training rollouts based on model output quality.
The paper introduces the Harness Handbook, a behavior-centric documentation system that helps agents and developers navigate and modify large, complex agent codebases.
KnowAct-GUIClaw is a framework for autonomous GUI manipulation that utilizes a memory-driven, two-tier architecture to manage complex cross-platform tasks.
OvisOCR2 is a model designed to parse visually rich documents into structured Markdown in a single pass.
This paper identifies a problem where AI-generated evaluation data for language models can silently contain hidden errors, leading to misleading results, and proposes a mandatory manual check to prevent these issues.
GigaWorld-Policy-0.5 improves real-time robotic control by decoupling action generation from future video simulation using a specialized Mixture-of-Transformers architecture.
ProfMalPlus uses a multi-agent reasoning framework to detect malicious NPM packages by combining static code analysis with dynamic verification.
The paper introduces a hybrid variational autoencoder framework that integrates longitudinal tumor growth measurements with time-to-event outcomes using genomic data to improve predictive accuracy.