This Thesis applies formal methods, and in particular the Quantitative Information Flow (QIF) framework, to build scalable models for large datasets and software pipelines that process data. By breaking down complex systems into their components, one can soundly explain privacy vulnerabilities and how tackling them would affect data utility.