Vision-Language Models (VLMs) are increasingly deployed in applications that interpret and generate information from visual and textual inputs. While powerful, these models pose emerging privacy risks. In this paper, we introduce the concept of privacy chains: structured narratives that emerge when adversaries aggregate outputs from VLMs across multiple images, often exposing sensitive information even when the individual outputs are seemingly innocuous. Using LangChain, an open-source orchestration framework, we show how identity-linked data extracted via both benign and targeted prompts can be compiled into detailed timelines of private behavior, significantly amplifying privacy threats. To systematically assess this risk, we develop a privacy leakage pipeline within the Visual Question Answering (VQA) framework and evaluate six open-source VLMs across three tailored datasets: Celebrity, Car, and Tattoo. Our analysis reveals substantial and model-dependent privacy leakage, even from general-purpose queries. To mitigate this threat, we propose ChainShield, a white-box adversarial defense that applies targeted, imperceptible perturbations to images. ChainShield reduces privacy-relevant outputs by redirecting VLM responses toward benign alternatives, while preserving image realism. Our experiments show that ChainShield substantially lowers privacy leakage across models and datasets, effectively disrupting the formation of privacy chains.
Funding Agencies|Swedish Research Council (VR); Graduate School in Computer Science (CUGS) at Linkoping University