Back to Feed
Multimodal / Safety & Alignment

Programmable Backdoor Attacks on Vision Models

Original: Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 4 concepts

Key Takeaways

  • The attack moves beyond static one-to-one mapping by allowing arbitrary target captions to be triggered post-training.
  • It uses a novel heuristic poisoning strategy to train models to interpret specific foreign image regions as embedded instructions.
  • The method demonstrates a 92.00% success rate for vanilla triggers on unseen target captions.
  • The approach is resilient against standard backdoor defense mechanisms like Shrinkpad, Flip, and Scale-up.

Summary & Methodology Analysis

The researchers developed a programmable any-to-any backdoor for Vision Language Models (VLMs) by utilizing a two-part strategy. First, they employ a heuristic poisoning method that introduces diverse trigger and caption pairs into the training set. This trains the model to recognize foreign image regions as instructions, effectively exploiting the inherent zero-shot generalization capabilities of the VLM. This process utilizes standard fine-tuning pipelines such as LLaVA alongside LoRA, which is a technique that reduces memory consumption by updating only a small subset of model weights during fine-tuning.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What problem does this paper solve?

It addresses the limitation of existing backdoor attacks that are restricted to static, pre-defined mappings, which would otherwise require costly model retraining to change target behaviors.

Q2. What is the primary contribution of the work?

The authors propose a programmable any-to-any backdoor paradigm that supports dynamic, post-training selection of trigger synthesis and target captioning.

Q3. Is this attack effective in practice?

Yes, it achieves a 92.00% attack success rate on held-out target captions and remains robust against several standard backdoor defenses.

Q4. How does the trigger generation process work?

The method uses a pretrained vision feature extractor to map a selected target caption to a stealthy trigger, such as norm-controlled noise or non-semantic patches, that mimics target-specific features.

Q5. Which models were tested using this method?

The paper includes testing on LLaVA, LLaVA-1.5, LLaVA-1.6, CLIP, ALIGN, GPT-4V, Qwen3-VL, MiniGPT-v2, and MiniGPT-4.

Q6. What datasets were utilized for the evaluation?

The evaluation utilized Flickr8k, Flickr30k, cc-sbu-align, Tiny-imagenet, CIFAR-100, GQA, MME, and stable-diffusion-xl-base-1.0.

Q7. Are there any known vulnerabilities in this attack method?

Yes, the noise-optimized stealthy triggers are noted to be sensitive to defenses based on image preprocessing.

Q8. Does this method require specific architectural changes?

No, it utilizes standard VLM fine-tuning pipelines and standard image-to-caption generation tasks for evaluation.

Q9. What is the computational cost of this attack?

The paper does not specify the exact computational cost or hardware requirements for the attack.

Flag an issue

What is wrong with this summary?

What is wrong?