Multimodal Multi-turn Safety Alignment: From Agentic Interaction to Strategic Alignment
ORIGINAL / Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment
Existing alignment methods primarily target malicious visual QA pairs and fail to address gradual adversarial attacks in multi-turn dialogues. This study introduces the MINT-Safe dataset and TAD-Align framework, which dynamically identify and up-weight unsafe turns via a turn-aware dual-objective reward function, significantly reducing attack success rates while improving safety and helpfulness.
01 ABSTRACT
This study identifies safety risks in multi-turn dialogues of multimodal large language models, which existing alignment methods fail to mitigate. The authors construct a multi-turn visual dialogue dataset MINT-Safe via multi-agent interaction and propose the TAD-Align framework, which uses rollout-based safety score variance to dynamically identify and weight unsafe turns. Experiments on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B show over 10% reduction in attack success rate, with at least 8% improvement in harmlessness and 13% in helpfulness.
02 KEY FINDINGS
- Introduces MINT-Safe dataset: 11,270 multi-image dialogues and 500 refusal VQA pairs for multi-turn visual safety alignment
- TAD-Align framework uses a turn-aware dual-objective reward function to dynamically identify and adaptively up-weight turns with inconsistent safety behavior
- Experiments show over 10% reduction in Attack Success Rate and at least 8% and 13% improvements in harmlessness and helpfulness on multiple models
- Finds that multi-turn interactions enable adversaries to progressively reconstruct harmful intent across dialogues, bypassing single-turn safety constraints
AI GENERATED SUMMARY / DISCOVERED BY ARXIV CS.CL