Conveners
Disclosure Risk - 1
- Krish Muralidhar (University of Oklahoma)
Disclosure Risk - 2
- Krish Muralidhar (University of Oklahoma)
-
Krish Muralidhar (University of Oklahoma)16/10/2025, 15:40
ε-Differential privacy (DP) is a popular privacy model that has been promoted as the de facto standard in most data intensive areas. However, the selection of the privacy parameter ε (also called budget) in applications of DP remains an open challenge. Even though the meaning and implications of the value of ε are not fully understood, it is clear that large budget values are less...
Go to contribution page -
Jonathan Latner (Institute for Employment Research (IAB))16/10/2025, 15:55
This paper evaluates disclosure risk measures for synthetic data generated by CART-based models, using both a controlled simulated dataset and publicly available data. We find that common disclosure risk measures may fail to detect disclosure risks and, in some cases, misrepresent actual disclosure risks. Additionally, CART-based models, while maintaining high statistical utility, may...
Go to contribution page -
Marieke de Vries (Netherlands (Kingdom of the))16/10/2025, 16:40
The rise in access to public data on the internet, and specifically online social networks (OSNs), is causing new pressures on the statistical disclosure control of microdata. Currently at Statistics Netherlands, a criterium is applied that looks at three properties of variables: rarity, visibility and searchability. Underlying this criterium are, similar to other methods used to assess the...
Go to contribution page -
Dr Sonakshi Garg, Vicenc Torra (Umea University)16/10/2025, 16:55
Government statistical agencies increasingly rely on sensitive tabular data to guide evidence-based policymaking, yet restrictions on data access hinder research and transparency. Synthetic data generated with Generative Adversarial Networks (GANs) offers a promising solution, but conventional GANs often produce unrealistic tables or fail to preserve the statistical relationships that matter...
Go to contribution page -
Gillian Raab16/10/2025, 17:20
Recent years have seen an increased pressure to allow information derived from administrative data to be used to inform policy; see for example the Sturrock Report, 2024. Several organisations have been set up in the UK to develop policies to facilitate this. When data access is given to researchers, who are not part of the organisation that owns the data, there is a concern that there may...
Go to contribution page -
Prof. Mark Elliot (University of Manchester)16/10/2025, 17:35
Introduction: The de-identification of unstructured free-text data is important for sharing large amounts of healthcare information generated by electronic health records, publications and clinical trials. To automate this process, information extraction (IE) and natural language processing (NLP) are essential tools. However, evaluating NLP performance in de-identification requires...
Go to contribution page