Research Feed Page 38

Search source-linked summaries of recent AI and machine-learning papers by topic, by date, or by whether they include code or a diagram.

Filter papers All papers

Browse by date

Resource filters

Sort options

Research results

Training & Fine-Tuning / Benchmarks & Evals By Nick Oh 2026-08-05 2
Small Foundation Models of Human Cognition and Behaviour

The paper tests whether small language models fine-tuned on behavioral data use structural reasoning or statistical shortcuts to predict human task performance.

Agents / Training & Fine-Tuning By Junzhuo Liu 2026-08-05
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

The paper introduces a routing method for multi-turn AI agents that selectively applies reference guidance only when the agent's current state aligns with known valid task paths.

Robotics / Computer Vision By Chenghao Gu 2026-08-06
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

GeniWorld improves robotic manipulation in unseen environments by using an interactive world model that converts numerical robot actions into dense visual sequences for better control.

Robotics / Benchmarks & Evals By Deyi Zhu 2026-08-06 2
Uncertainty-Aware World Model for Aerial Image-Goal Navigation

The researchers developed a new navigation model that improves drone path selection by accounting for future uncertainties in large-scale outdoor environments.

Computer Vision / Multimodal By Jiaxiao Wang 2026-08-06
EvReflection: Event-Driven Micro-Dynamics for Reflection Removal

The paper introduces EvReflection, a method that uses asynchronous event streams to remove reflection artifacts from images captured through transparent media.

Computer Vision By Bingyuan Wang 2026-08-06
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

EmoWorld is a framework that separates atmospheric, semantic, and temporal elements to enable independent emotional control in frozen video diffusion transformers.

Computer Vision / Benchmarks & Evals By Yuanjing Xu 2026-08-06
Depth-Guided Video Object Counting in Crowded Scenes

The paper introduces a depth-aware detector to improve object counting accuracy in crowded video scenes where traditional RGB-only methods struggle.

Computer Vision By Elad Yoshai 2026-08-06
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation

PRISM enables precise, spatially aware control in image-to-image translation by gating feature updates based on distribution discrepancies between source and target domains.

Reinforcement Learning / Computer Vision By Saugat Adhikari 2026-08-06
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

The iARCS method uses an LLM agent to iteratively refine 3D scene generation through reinforcement learning, ensuring better adherence to physical constraints and task requirements.

Efficiency & Inference / Benchmarks & Evals By Sam Levang 2026-08-06
Timestep-Conditioned Transformers for Global Weather Forecasting

The paper introduces a transformer model that allows for variable forecasting timesteps to better balance atmospheric dynamics with long-term predictive accuracy.

Training & Fine-Tuning / Benchmarks & Evals By Dohyun Ku 2026-08-06
MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

Researchers developed MetaboLLM to integrate biochemical knowledge and convert it into predictive metabolite graphs for clinical diagnostics.

Reinforcement Learning / Training & Fine-Tuning By Chenglong Wang 2026-08-06
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

The paper introduces Ranking-based Reward Construction to bridge the gap between generative reward models and reinforcement learning algorithms.

Multimodal / Reinforcement Learning By Jiahao Huang 2026-08-06
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

OneEmo is a 4.5B parameter multimodal model that improves emotion perception and understanding by using a novel reinforcement learning framework and a human-in-the-loop reasoning dataset.

Training & Fine-Tuning / Benchmarks & Evals By Haris Riaz 2026-08-06
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

The researchers introduce a method to inject syntactic information into Transformer positional embeddings to improve compositional generalization without modifying the underlying attention mechanisms.

Multimodal / Efficiency & Inference By Guangyuan Wang 2026-08-06 2
Wan-Animate-2: Pushing the Application Boundaries of Character Animation

Wan-Animate-2 introduces a new architecture to solve inefficiencies in character animation by decoupling reference streams and enabling more efficient training.

Efficiency & Inference / Benchmarks & Evals By Anton Conrad 2026-08-06
Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction

The paper introduces Randomly Localized Conformal Prediction to ensure reliable uncertainty quantification in specific regions of the data space rather than just on average.

Safety & Alignment / Benchmarks & Evals By Xian Sun 2026-08-06
Learning When to Trust via Selective Context Preference Optimization

The paper introduces the SCOPE framework to prevent language models from abandoning correct answers when they encounter deceptive or irrelevant context.

Multimodal / Training & Fine-Tuning By Andrey Shutkin 2026-08-06 14
KVAE: Family of Tokenizers for Multimodal Generative Models

The paper introduces the KVAE family of tokenizers to improve latent spaces for text-conditioned audio, image, and video generation by addressing the limitations of existing reconstruction-focused methods.

Agents / Reasoning By Tsz Ting Chung 2026-08-06
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

The paper introduces InsightEmb, a method that improves agentic tasks by matching an agent's current state to helpful heuristic insights rather than relying on standard semantic similarity.

Computer Vision / Multimodal By Feier Wu 2026-08-06 16
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

EffectLearner is a world-aware video object removal system that eliminates both target objects and their induced effects by using a vision-language model to reason about motion and scene interactions.