Research Feed Page 38
Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.
Research results
The paper tests whether small language models fine-tuned on behavioral data use structural reasoning or statistical shortcuts to predict human task performance.
The paper introduces a routing method for multi-turn AI agents that selectively applies reference guidance only when the agent's current state aligns with known valid task paths.
GeniWorld improves robotic manipulation in unseen environments by using an interactive world model that converts numerical robot actions into dense visual sequences for better control.
The researchers developed a new navigation model that improves drone path selection by accounting for future uncertainties in large-scale outdoor environments.
The paper introduces EvReflection, a method that uses asynchronous event streams to remove reflection artifacts from images captured through transparent media.
EmoWorld is a framework that separates atmospheric, semantic, and temporal elements to enable independent emotional control in frozen video diffusion transformers.
The paper introduces a depth-aware detector to improve object counting accuracy in crowded video scenes where traditional RGB-only methods struggle.
PRISM enables precise, spatially aware control in image-to-image translation by gating feature updates based on distribution discrepancies between source and target domains.
The iARCS method uses an LLM agent to iteratively refine 3D scene generation through reinforcement learning, ensuring better adherence to physical constraints and task requirements.
The paper introduces a transformer model that allows for variable forecasting timesteps to better balance atmospheric dynamics with long-term predictive accuracy.
Researchers developed MetaboLLM to integrate biochemical knowledge and convert it into predictive metabolite graphs for clinical diagnostics.
The paper introduces Ranking-based Reward Construction to bridge the gap between generative reward models and reinforcement learning algorithms.
OneEmo is a 4.5B parameter multimodal model that improves emotion perception and understanding by using a novel reinforcement learning framework and a human-in-the-loop reasoning dataset.
The researchers introduce a method to inject syntactic information into Transformer positional embeddings to improve compositional generalization without modifying the underlying attention mechanisms.
Wan-Animate-2 introduces a new architecture to solve inefficiencies in character animation by decoupling reference streams and enabling more efficient training.
The paper introduces Randomly Localized Conformal Prediction to ensure reliable uncertainty quantification in specific regions of the data space rather than just on average.
The paper introduces the SCOPE framework to prevent language models from abandoning correct answers when they encounter deceptive or irrelevant context.
The paper introduces the KVAE family of tokenizers to improve latent spaces for text-conditioned audio, image, and video generation by addressing the limitations of existing reconstruction-focused methods.
The paper introduces InsightEmb, a method that improves agentic tasks by matching an agent's current state to helpful heuristic insights rather than relying on standard semantic similarity.
EffectLearner is a world-aware video object removal system that eliminates both target objects and their induced effects by using a vision-language model to reason about motion and scene interactions.