Back to Feed
Multimodal / Computer Vision

Personalizing Human Object Interactions in Video

Original: HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 4 concepts

Key Takeaways

  • HOMIE improves OCR accuracy by 21.8 percent over the SkyReels-V3 baseline.
  • The method uses a multimodal input paradigm that combines video tokens, reference images, and text prompts processed by an MLLM.
  • A global multimodal guidance mechanism injects fused feature representations directly into the video generation process.
  • User study results indicate superior performance in text adherence, video quality, and subject consistency compared to existing methods.

Summary & Methodology Analysis

The HOMIE architecture addresses the difficulty of maintaining subject fidelity and accurate interaction patterns in videos. It processes input via a multimodal large language model (MLLM), which acts as a vision-capable language processor, to integrate video tokens with reference images and textual prompts. The framework employs Global Multimodal Guidance (GMG), a fusion mechanism within self-attention layers, which are components that calculate the relevance between different parts of the input sequence. GMG uses pooling to extract global representations from MLLM features and applies them to video tokens through feature-wise affine transformations, ensuring consistent global guidance during generation.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the primary problem HOMIE solves?

It solves the difficulty of balancing subject fidelity with accurate interaction patterns in human-object centric video personalization.

Q2. How much better is HOMIE at rendering text compared to previous models?

HOMIE achieved a 21.8 percent improvement in OCR accuracy compared to the SkyReels-V3 baseline.

Q3. Does this approach require specific training data?

Yes, it uses a multi-stage training strategy with a curated dataset containing single-subject, multi-subject, and high-resolution video clips.

Q4. How does Modality-Reference Embedding (MRE) function?

MRE is a learnable module that assigns specific embeddings to differentiate input sources and binds reference embeddings to correlate intra-subject images that share the same identity.

Q5. What happens if a user needs to generate long-form video content?

The paper notes that the framework is limited to generating short video clips due to its reliance on the underlying foundation model architecture.

Q6. Which specific architectures and datasets are mentioned?

The paper references various models including Wan2.1-14B, Wan2.2-14B, Qwen3-VL-2B-Thinking, and HunyuanVideo-T2V-13B, and uses datasets like OpenS2V-5M and PhantomData.

Q7. How were the qualitative results validated?

The model outperformed existing methods in a user study involving 40 participants, assessing video quality, text following, and subject consistency.

Q8. Are there specific hardware requirements for running HOMIE?

The paper does not specify the hardware requirements or computational costs for inference.

Q9. Does the paper compare HOMIE against all existing generative video models?

It benchmarks against several methods including SkyReels-V3, VACE-14B, MAGREF, SkyReels-A2, Phantom, HuMo, BindWeave, FFGO, UniVideo, and VINO.

Flag an issue

What is wrong with this summary?

What is wrong?