# Legible Papers Simplified summaries and code implementations of trending machine learning and scientific research papers. ## Research Catalog - [Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI](https://legiblepapers.com/papers/decolonizing-linguistic-policies-in-automated-speech-recognition-a-framework-for-cross-culturally-competent-speech-ai): The paper introduces a framework and auditing protocol to fix automatic speech recognition systems that systematically fail low-resource, Indigenous, and non-standard language varieties. - [Learning visual representations for compositional analysis of artworks and photographs](https://legiblepapers.com/papers/learning-visual-representations-for-compositional-analysis-of-artworks-and-photographs): The paper introduces a human-inspired pipeline using Object-Centric Learning and graph networks to model image composition more effectively than frozen foundation models. - [Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset](https://legiblepapers.com/papers/audio-to-score-transcription-using-pre-trained-features-data-augmentation-and-the-new-sheetsage-a2s-dataset): Researchers developed an audio to score transcription system using a Transformer architecture trained on the new SheetSage A2S dataset to bridge the gap between classical and popular music transcription. - [ErgoSurf: Ergodic Control for the Coverage of Unknown Surfaces](https://legiblepapers.com/papers/ergosurf-ergodic-control-for-the-coverage-of-unknown-surfaces): The researchers developed a method that enables robots to systematically cover unknown surfaces by combining real-time geometric reconstruction with ergodic trajectory generation. - [Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation](https://legiblepapers.com/papers/robot-learning-from-human-demonstrations-handwritten-alphabet-trajectories-and-human-likeness-evaluation): The paper presents a method for robots to replicate human handwriting by learning trajectories from human demonstrations using Gaussian models. - [LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models](https://legiblepapers.com/papers/lawm-3d-learning-3d-aware-latent-actions-from-human-videos-for-generalizable-robot-world-models): LAWM-3D enables robots to learn 3D-aware actions by training world models on human videos using a new geometric alignment method. - [PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say](https://legiblepapers.com/papers/privacypeek-auditing-what-llm-based-agents-acquire-not-just-what-they-say): The PrivacyPeek framework benchmarks how LLM-based agents acquire sensitive information beyond their necessary operational scope during task completion. - [Does FLAIR super-resolution erase or hallucinate small white-matter lesions?](https://legiblepapers.com/papers/does-flair-super-resolution-erase-or-hallucinate-small-white-matter-lesions): The paper evaluates whether super-resolution techniques applied to MRI scans risk erasing real small white-matter lesions or hallucinating false ones. - [Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction](https://legiblepapers.com/papers/hypothesis-testing-with-conditional-queries-learnability-and-the-value-of-interaction): The paper provides a formal framework for determining when distribution classes are learnable and quantifies the exact performance cost of replacing interactive queries with static ones. - [Is Self-Pretraining really useful to improve diagnosis in medical Time Series?](https://legiblepapers.com/papers/is-self-pretraining-really-useful-to-improve-diagnosis-in-medical-time-series): The researchers applied self-pretraining to transformer models to boost accuracy in medical time series classification without requiring external data. - [CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?](https://legiblepapers.com/papers/cogvis-must-open-vocabulary-change-detection-perceive-the-scene-anew-for-every-query): CogVis introduces a modular perception and memory framework that decouples image analysis from query processing to improve speed and accuracy in open-vocabulary change detection. - [CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection](https://legiblepapers.com/papers/cfgpnet-cross-attention-based-fused-gradient-programmed-network-framework-for-multispectral-object-detection): The CFGPNet framework improves multispectral object detection by optimizing feature interaction and gradient flow while reducing computational overhead. - [Muon on the Stiefel Manifold Admits an Exact Closed-Form Update](https://legiblepapers.com/papers/muon-on-the-stiefel-manifold-admits-an-exact-closed-form-update): The researchers developed Skewon, an algorithm that provides an exact closed-form solution for the Stiefel Muon optimization problem. - [Visual Grounding in Zero-Shot Vision-Language Control](https://legiblepapers.com/papers/visual-grounding-in-zero-shot-vision-language-control): The paper investigates whether vision language models serving as robot controllers truly rely on visual inputs or merely leverage non visual shortcuts like simulator rewards. - [BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models](https://legiblepapers.com/papers/bendtwin-robust-dense-to-sparse-physical-reconstruction-with-bending-aware-differentiable-spring-mass-models): BendTwin adds bending stiffness to spring mass models to improve the stability and accuracy of physical reconstructions from sparse video data. - [Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models](https://legiblepapers.com/papers/robust-wam-bridging-generative-pretraining-and-semantic-foresight-in-world-action-models): Robust-WAM introduces a method to align video-generation model latent spaces with semantic features, enabling robots to handle visual changes more reliably. - [Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments](https://legiblepapers.com/papers/prior-sg-task-and-prior-driven-region-segmentation-for-scene-graphs-in-arbitrarily-structured-environments): The paper introduces Prior-SG, a framework that casts scene graph generation as a probabilistic alignment problem using task-conditioned priors and visual-geometric feature fusion to handle arbitrarily structured environments. - [HOPE: Hand-Object Pressure Estimation from Monocular Videos](https://legiblepapers.com/papers/hope-hand-object-pressure-estimation-from-monocular-videos): The researchers developed a transformer architecture that estimates physical pressure during hand-object interactions using only monocular video input. - [Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators](https://legiblepapers.com/papers/design-and-evaluation-of-a-touchscreen-based-teleoperation-interface-for-robotic-manipulators): The paper introduces a touchscreen interface and hybrid control architecture that improves precision and reduces cognitive load during robotic surface interaction tasks. - [RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction](https://legiblepapers.com/papers/rxnclf-contrastive-transformation-aware-reaction-foundation-model-for-improved-reactivity-prediction): RxnCLF is a new reaction foundation model that uses condensed graph structures and contrastive learning to better capture chemical transformation information for reactivity prediction. - [Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data](https://legiblepapers.com/papers/surv-iptb-an-attention-based-model-for-estimating-individual-probability-of-treatment-benefit-with-survival-data): The authors introduce Surv-IPTB, an attention-based model that improves the estimation of individual treatment benefits in survival analysis by converting the problem into a pairwise classification task. - [Scalable estimation of VARMA models](https://legiblepapers.com/papers/scalable-estimation-of-varma-models): This paper introduces a scalable framework for estimating VARMA models by decoupling computation from series length, allowing for efficient processing of high-dimensional time series data. - [OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations](https://legiblepapers.com/papers/otlesmix-wasserstein-barycenter-and-optimal-transport-map-for-synthetic-lesion-generation-with-diverse-shapes-and-locations): The researchers developed a method to generate synthetic training data for lesion segmentation by interpolating shapes and intensities using Wasserstein barycenters. - [Threshold-Based Early Stopping of Accumulations in Neural Networks with Binary Activation](https://legiblepapers.com/papers/threshold-based-early-stopping-of-accumulations-in-neural-networks-with-binary-activation): The researchers developed a method to stop neural network accumulations early by predicting the final sign of binary activations from partial sums. - [Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era](https://legiblepapers.com/papers/reversible-unlearnable-examples-towards-the-copyright-protection-in-deep-learning-era): The paper introduces a method to simultaneously watermark images and apply unlearnable perturbations that prevent unauthorized model training while allowing authorized users to reverse the protection. - [TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN](https://legiblepapers.com/papers/tlnm-externally-validated-tooth-detection-numbering-and-segmentation-from-smartphone-photographs-using-mask-r-cnn): Researchers developed a customized Mask R-CNN model to accurately detect, label, and segment teeth from uncontrolled smartphone dental photographs. - [SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models](https://legiblepapers.com/papers/saga-score-weighted-adaptive-generation-alignment-for-low-resource-nordic-language-models): The researchers developed SAGA, an automated pipeline that uses linguistic scoring to align language models for low-resource Nordic languages without needing human preference labels. - [Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction](https://legiblepapers.com/papers/confidence-matters-leveraging-multi-view-geometric-priors-for-gs-based-reconstruction): The researchers integrate multi-view geometric priors and confidence-based weighting into 3D Gaussian Splatting to fix suboptimal geometry in complex or shiny scenes. - [JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment](https://legiblepapers.com/papers/joyai-ra-0-5-scaling-robot-manipulation-learning-via-dual-action-alignment): JoyAI-RA 0.5 enables scalable robot manipulation by aligning diverse data sources like human videos and simulation into a shared format for consistent learning. - [Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning](https://legiblepapers.com/papers/beyond-simply-environment-scaling-designing-effective-environment-distributions-for-multimodal-agent-learning): The researchers developed a selection and curriculum framework that optimizes training environment diversity and difficulty to significantly improve multimodal agent performance. - [Handling Missing Data in Probabilistic Regression Trees](https://legiblepapers.com/papers/handling-missing-data-in-probabilistic-regression-trees): The authors developed a method for Probabilistic Regression Trees to process missing predictor values natively during tree construction instead of relying on external data imputation. - [QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction](https://legiblepapers.com/papers/quantimedai-quantum-enhanced-time-series-model-guided-by-agentic-ai-for-cardiac-arrest-mortality-prediction): Researchers developed an agentic AI framework using quantum circuits to predict cardiac arrest mortality from longitudinal patient data with significantly fewer parameters than traditional models. - [Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria](https://legiblepapers.com/papers/investigating-artificial-intelligence-digital-sovereignty-in-mobile-shopping-apps-a-case-study-of-nigeria): The paper examines how artificial intelligence in Nigerian mobile applications affects digital sovereignty, evaluated through platform transparency and socio-economic context. - [Mind the Gaps: Mixture-of-Minds for Human Simulation](https://legiblepapers.com/papers/mind-the-gaps-mixture-of-minds-for-human-simulation): The paper introduces Anacreon, a system that uses specialized adapter modules to prevent large language models from collapsing diverse individual personalities into generic averages. - [Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts](https://legiblepapers.com/papers/poli-bias-understanding-and-measuring-large-language-model-biases-in-international-political-conflicts): Researchers developed a systematic framework to audit and measure how large language models exhibit political bias when analyzing international conflicts. - [Reducing belief in conspiracy theories as they unfold using large language models](https://legiblepapers.com/papers/reducing-belief-in-conspiracy-theories-as-they-unfold-using-large-language-models): The researchers evaluated if multi-turn LLM conversations can effectively debunk conspiracy theories as they emerge during crisis events. - [Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model](https://legiblepapers.com/papers/toward-deployable-bangla-sign-language-recognition-with-expert-validated-data-and-a-lightweight-attention-based-model): Researchers developed a highly efficient, expert-validated model for recognizing Bangla sign language that runs locally on commodity mobile hardware. - [Challenges in Evaluating Explanation Methods for Static and Evolving Data](https://legiblepapers.com/papers/challenges-in-evaluating-explanation-methods-for-static-and-evolving-data): This research addresses the lack of standardized evaluation methods for XAI by proposing new metrics and techniques to handle both static datasets and evolving data streams. - [Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture](https://legiblepapers.com/papers/bar-jepa-extracting-values-from-bar-chart-with-joint-embedding-predictive-architecture): Bar-JEPA uses a custom joint-embedding architecture to computationally extract numerical data from bar charts despite visual variability and a lack of real-world training data. - [Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering](https://legiblepapers.com/papers/tracing-the-heart-an-evidence-linked-pipeline-for-heart-failure-feature-engineering): The paper introduces a structured data pipeline that transforms fragmented EHR records into audited clinical features, improving heart failure prediction accuracy. - [Small Foundation Models of Human Cognition and Behaviour](https://legiblepapers.com/papers/small-foundation-models-of-human-cognition-and-behaviour): The paper tests whether small language models fine-tuned on behavioral data use structural reasoning or statistical shortcuts to predict human task performance. - [From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks](https://legiblepapers.com/papers/from-passive-mirrors-to-active-agents-holonic-digital-twins-for-physical-ai-over-networks): The paper introduces a framework called HDT-Nets that pairs physical AI agents with hierarchical digital twins to enable reliable, causal reasoning across distributed networks. - [When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents](https://legiblepapers.com/papers/when-privileged-guidance-misaligns-state-matched-routing-and-contextualized-self-distillation-for-multi-turn-agents): The paper introduces a routing method for multi-turn AI agents that selectively applies reference guidance only when the agent's current state aligns with known valid task paths. - [GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions](https://legiblepapers.com/papers/geniworld-a-generalizable-interactive-world-model-for-robotic-manipulation-via-visual-actions): GeniWorld improves robotic manipulation in unseen environments by using an interactive world model that converts numerical robot actions into dense visual sequences for better control. - [Uncertainty-Aware World Model for Aerial Image-Goal Navigation](https://legiblepapers.com/papers/uncertainty-aware-world-model-for-aerial-image-goal-navigation): The researchers developed a new navigation model that improves drone path selection by accounting for future uncertainties in large-scale outdoor environments. - Code: https://github.com/DurYi/UA-NWM - [EvReflection: Event-Driven Micro-Dynamics for Reflection Removal](https://legiblepapers.com/papers/evreflection-event-driven-micro-dynamics-for-reflection-removal): The paper introduces EvReflection, a method that uses asynchronous event streams to remove reflection artifacts from images captured through transparent media. - [EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation](https://legiblepapers.com/papers/emoworld-a-decoupled-affective-field-for-controllable-emotional-video-generation): EmoWorld is a framework that separates atmospheric, semantic, and temporal elements to enable independent emotional control in frozen video diffusion transformers. - [Depth-Guided Video Object Counting in Crowded Scenes](https://legiblepapers.com/papers/depth-guided-video-object-counting-in-crowded-scenes): The paper introduces a method that incorporates depth cues into video object counting to improve detection accuracy in crowded and occluded environments. - [PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation](https://legiblepapers.com/papers/prism-distribution-gated-flow-matching-for-controllable-unpaired-image-translation): PRISM enables precise, spatially aware control in image-to-image translation by gating feature updates based on distribution discrepancies between source and target domains. - [iARCS: Iterative Agentic RL for Controllable 3D Scene Generation](https://legiblepapers.com/papers/iarcs-iterative-agentic-rl-for-controllable-3d-scene-generation): The iARCS method uses an LLM agent to iteratively refine 3D scene generation through reinforcement learning, ensuring better adherence to physical constraints and task requirements. - [Timestep-Conditioned Transformers for Global Weather Forecasting](https://legiblepapers.com/papers/timestep-conditioned-transformers-for-global-weather-forecasting): The paper introduces a transformer model that allows for variable forecasting timesteps to better balance atmospheric dynamics with long-term predictive accuracy. - [MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction](https://legiblepapers.com/papers/metabollm-a-metabolomics-specialized-large-language-model-for-biochemical-knowledge-integration-and-predictive-metabolite-graph-construction): Researchers developed MetaboLLM to integrate biochemical knowledge and convert it into predictive metabolite graphs for clinical diagnostics. - [RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction](https://legiblepapers.com/papers/rrc-unlocking-generative-reward-models-in-llm-reinforcement-learning-via-ranking-based-reward-construction): The paper introduces Ranking-based Reward Construction to bridge the gap between generative reward models and reinforcement learning algorithms. - [OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction](https://legiblepapers.com/papers/oneemo-a-unified-multimodal-reasoning-model-for-emotion-perception-understanding-and-interaction): OneEmo is a 4.5B parameter multimodal model that improves emotion perception and understanding by using a novel reinforcement learning framework and a human-in-the-loop reasoning dataset. - Code: https://github.com/waHAHJIAHAO/OneEmo - [Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers](https://legiblepapers.com/papers/beyond-sequence-order-syntax-informed-positional-embeddings-for-transformers): The researchers introduce a method to inject syntactic information into Transformer positional embeddings to improve compositional generalization without modifying the underlying attention mechanisms. - [Wan-Animate-2: Pushing the Application Boundaries of Character Animation](https://legiblepapers.com/papers/wan-animate-2-pushing-the-application-boundaries-of-character-animation): Wan-Animate-2 introduces a new architecture to solve inefficiencies in character animation by decoupling reference streams and enabling more efficient training. - [Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction](https://legiblepapers.com/papers/beyond-marginal-validity-finite-sample-guarantees-for-localized-conformal-prediction): The paper introduces Randomly Localized Conformal Prediction to ensure reliable uncertainty quantification in specific regions of the data space rather than just on average. - [Learning When to Trust via Selective Context Preference Optimization](https://legiblepapers.com/papers/learning-when-to-trust-via-selective-context-preference-optimization): The paper introduces the SCOPE framework to prevent language models from abandoning correct answers when they encounter deceptive or irrelevant context. - [KVAE: Family of Tokenizers for Multimodal Generative Models](https://legiblepapers.com/papers/kvae-family-of-tokenizers-for-multimodal-generative-models): The paper introduces the KVAE family of tokenizers to improve latent spaces for text-conditioned audio, image, and video generation by addressing the limitations of existing reconstruction-focused methods. - Code: https://github.com/kandinskylab/kvae - [InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval](https://legiblepapers.com/papers/insightemb-learning-action-intent-embeddings-for-agentic-insight-retrieval): The paper introduces InsightEmb, a method that improves agentic tasks by matching an agent's current state to helpful heuristic insights rather than relying on standard semantic similarity. - [EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal](https://legiblepapers.com/papers/effectlearner-world-aware-object-effect-reasoning-for-real-world-video-object-removal): EffectLearner is a world-aware video object removal system that eliminates both target objects and their induced effects by using a vision-language model to reason about motion and scene interactions. - [Invisible Shortcuts: Why Vision Encoders Know Your Camera](https://legiblepapers.com/papers/invisible-shortcuts-why-vision-encoders-know-your-camera): Researchers discovered that deep vision models unintentionally learn invisible camera metadata as shortcuts, which degrades their performance when image distribution shifts. - [CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks](https://legiblepapers.com/papers/calibforge-adversarial-solver-calibration-for-scaling-learnable-terminal-tasks): CalibForge uses automated adversarial feedback from software solvers to ensure that training tasks for LLM agents are neither too simple nor impossible to solve. - Code: https://github.com/AweAI-Team/CalibForge - [DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation](https://legiblepapers.com/papers/dypes-vla-learning-shared-dynamics-priors-and-embodiment-specific-control-for-cross-embodiment-manipulation): DyPES-VLA unifies robot control by using shared dynamics priors learned from video to enable action generation across different robot types without manual alignment. - [DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models](https://legiblepapers.com/papers/dash-divergence-adaptive-supervision-horizons-for-on-policy-self-distillation-of-reasoning-models): The researchers developed a method called DASH that dynamically adjusts how reasoning models learn from their own outputs to produce more accurate results. - [RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer](https://legiblepapers.com/papers/rp-opsd-reasoning-pivot-guided-on-policy-self-distillation-for-multilingual-reasoning-transfer): The paper introduces RP-OPSD, a method that improves how language models transfer English reasoning skills to low-resource languages by selectively applying privileged distillation based on reasoning-pivot signals. - [Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay](https://legiblepapers.com/papers/activity-frames-deterministic-screen-activity-compilation-for-agent-memory-and-replay): The paper introduces a deterministic method to compile raw screen activity into structured, auditable memory frames for computer-use agents. - Code: https://github.com/nossa-y/activity-frames - [World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation](https://legiblepapers.com/papers/world-to-wrist-task-conditioned-future-wrist-modeling-for-fine-grained-robot-manipulation): The researchers developed a method to predict future wrist camera observations to improve how robots perform fine-grained physical tasks. - Code: https://github.com/yyyyu120/W2-VLA - [Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors](https://legiblepapers.com/papers/bias-analysis-of-l2-speaking-assessment-systems-using-concept-activation-vectors): The researchers developed a method using Concept Activation Vectors to identify if Transformer based speaking assessment systems rely on irrelevant speaker attributes rather than proficiency. - [GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?](https://legiblepapers.com/papers/gst-bench-can-vlms-develop-global-spatial-awareness-from-video): The authors introduce GST-Bench to evaluate and improve how vision-language models maintain consistent spatial understanding across long, continuous video streams. - [Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents](https://legiblepapers.com/papers/resourced-authority-a-mechanism-design-model-for-participatory-governance-of-deployed-ai-agents): The paper presents a mechanism for managing AI agent deployment and compute resources through a stakeholder voting model that prioritizes broad support over aggregate wealth. - [WorldClaw: Agentic 3D Open-World Generation at Scale](https://legiblepapers.com/papers/worldclaw-agentic-3d-open-world-generation-at-scale): WorldClaw uses an agentic pipeline to transform text prompts into spatially consistent and editable 3D environments. - [UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression](https://legiblepapers.com/papers/uq-loc-uncertainty-aware-lidar-scene-coordinate-regression): The paper introduces UQ-Loc, a method that adds uncertainty estimates to LiDAR scene coordinate regression to improve localization robustness. - [AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning](https://legiblepapers.com/papers/agentopsd-recursive-self-distillation-for-agentic-reinforcement-learning): AgentOPSD introduces a recursive self-distillation method to provide granular credit assignment for multi-turn agentic tasks by analyzing turn-level evidence. - Code: https://github.com/ZethWang/AgentOPSD - [ChronoVision: Temporal Reasoning via Latent State Reconstruction](https://legiblepapers.com/papers/chronovision-temporal-reasoning-via-latent-state-reconstruction): ChronoVision introduces a visual-focused training framework to help multimodal large language models track and reason about continuous changes in images. - [From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models](https://legiblepapers.com/papers/from-economic-agents-to-agentic-economies-a-systems-blueprint-for-economic-world-models): The paper provides a systems blueprint for constructing generative economic simulators that use heterogeneous agents to model market interactions and institutional dynamics. - Code: https://github.com/FreedomIntelligence/Awesome-Economic-World-Models - [OPD-V: Visual On-Policy Self-Distillation with Modality Balance](https://legiblepapers.com/papers/opd-v-visual-on-policy-self-distillation-with-modality-balance): OPD-V improves multimodal model performance and reduces latency by using visual-based self-distillation to balance how the model uses image and text data. - [Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents](https://legiblepapers.com/papers/benchmarking-the-benchmarks-evaluating-benchmarks-for-conversational-agents): The authors introduce a reference-free framework that uses LLM judges to automatically assess the quality and consistency of task-oriented conversational agent benchmarks. - [SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models](https://legiblepapers.com/papers/skilltfm-gated-skill-evolution-for-training-free-adaptation-of-tabular-foundation-models): SkillTFM enables tabular foundation models to adapt to new tasks and data distributions by dynamically retrieving and evolving skills from a pre-verified skill bank. - [The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping](https://legiblepapers.com/papers/the-low-frequency-trap-video-language-models-fail-at-simple-event-bookkeeping): The paper demonstrates that current video-language models fail to accurately track event counts and sequences when event frequency and load increase. - [Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints](https://legiblepapers.com/papers/improving-the-realism-of-synthetic-clinical-benchmarks-under-utility-constraints): The researchers developed a method to increase the clinical realism of synthetic datasets for AI agents while maintaining operational utility through constrained optimization. - [Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data](https://legiblepapers.com/papers/tytan-interactive-neurosymbolic-construction-of-analytic-semantic-schemas-from-relational-data): Tytan uses neurosymbolic AI to automatically build semantic schemas from raw relational databases by combining LLM-driven inference with deterministic verification. - [Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI](https://legiblepapers.com/papers/schema-guided-hierarchical-information-extraction-and-semantic-evaluation-using-generative-ai): The paper presents a framework using generative AI and JSON schemas to automate the extraction and semantic evaluation of complex hierarchical data from health technology assessment documents. - [VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing](https://legiblepapers.com/papers/videoargus-agentic-rubric-grounded-unified-evaluation-for-video-generation-and-editing): VideoArgus introduces an agentic evaluation framework that uses dynamic, instance specific rubrics and specialized tools to provide accurate, diagnostic feedback on video generation and editing tasks. - [StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding](https://legiblepapers.com/papers/streamarena-toward-continuous-interactive-and-long-horizon-agentic-streaming-video-understanding): The paper introduces a benchmark and a two-tier architecture to enable agents to handle long-horizon streaming video tasks with improved latency and reasoning. - [LILAC: An Idempotent Neural Speech Codec](https://legiblepapers.com/papers/lilac-an-idempotent-neural-speech-codec): LILAC is a neural speech codec designed to be idempotent, ensuring that repeated cycles of decoding and re-encoding audio do not degrade signal quality or diverge in token streams. - [Learning Globally Reusable Skills for Coding Agents](https://legiblepapers.com/papers/learning-globally-reusable-skills-for-coding-agents): The paper introduces a framework to evolve agent skills as an interconnected global system rather than isolated updates to improve performance and generalizability. - [BaKron: Efficient Quantization with Kronecker-Factored Hessians](https://legiblepapers.com/papers/bakron-efficient-quantization-with-kronecker-factored-hessians): The paper introduces BaKron, a new quantization method that improves efficiency for two-sided Kronecker-factored Hessian approximations in neural networks. - [Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions](https://legiblepapers.com/papers/support-operation-factorization-compositional-readout-of-frozen-vision-encoders-under-controlled-interventions): The paper introduces a method called SO-OPF to decompose vision encoder representations into distinct support and operation factors to evaluate how well models generalize to new visual configurations. - [A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance](https://legiblepapers.com/papers/a-six-dimensional-taxonomy-of-post-training-adaptation-techniques-with-applications-in-ai-governance): The paper organizes fragmented post-training modification techniques into a structured six-dimensional taxonomy to simplify model management. - [Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training](https://legiblepapers.com/papers/sample-adaptive-latent-rewards-for-uncertainty-guided-diffusion-post-training): The researchers introduced a method to calculate predictive uncertainty during diffusion model post-training to prevent unreliable feedback and reward hacking. - [On-Policy Self-Distillation without Any Supervision](https://legiblepapers.com/papers/on-policy-self-distillation-without-any-supervision): The researchers developed a method called u-OPSD that allows models to improve their reasoning by self-distilling their own successful outputs without needing external labels or teacher models. - [Continual Learning in Transition](https://legiblepapers.com/papers/continual-learning-in-transition): The paper presents a three-axis framework to shift continual learning from static model updates to dynamic capability evolution across lifecycles, components, and update strategies. - [SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding](https://legiblepapers.com/papers/smartmage-dynamic-modality-orchestration-for-3d-scene-understanding): SmartMage optimizes 3D scene understanding by dynamically routing input modalities based on query requirements to improve reasoning performance and efficiency. - [MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration](https://legiblepapers.com/papers/microevo-knowledge-guided-llm-sampling-for-efficient-microarchitecture-design-space-exploration): MicroEvo uses large language models to intelligently explore the complex design space of computer microarchitectures, resulting in better energy efficiency and performance trade-offs than traditional optimization methods. - [Robust Context-Aware Detection of Malicious Instructions in Text](https://legiblepapers.com/papers/robust-context-aware-detection-of-malicious-instructions-in-text): The paper introduces a lightweight classifier called Context-Aware Detection that effectively flags malicious instructions in LLM prompts while maintaining high utility. - [The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images](https://legiblepapers.com/papers/the-illusion-of-visual-tool-use-a-causal-audit-of-thinking-with-images): The paper investigates why multimodal large language models see marginal or negative accuracy gains when using active visual operations like crop-and-zoom, using a causal graph and a three-level intervention protocol to analyze returned visual evidence. - [On-Policy Delta Distillation for Multilingual Math Reasoning](https://legiblepapers.com/papers/on-policy-delta-distillation-for-multilingual-math-reasoning): The researchers developed On-Policy Delta Distillation to boost mathematical reasoning capabilities in non-English languages during model training. - [Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents](https://legiblepapers.com/papers/benchmarking-and-enhancing-llms-for-rule-intensive-review-of-national-standard-documents): The paper introduces GB/T-Reviewer, a multi-agent framework that improves the accuracy of identifying rule violations in complex national standard documents. - [VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations](https://legiblepapers.com/papers/vidp-variable-impedance-diffusion-policy-for-compliant-robot-manipulation-from-diverse-demonstrations): The authors introduce a diffusion-based framework that allows robots to learn and execute task-specific compliance from diverse demonstrations without requiring specialized force sensing hardware.