Identification, data combination and the risk of disclosure

Komarova, Tatiana ORCID: 0000-0002-6581-5097, Nekipelov, Denis and Yakovlev, Evgeny (2018) Identification, data combination and the risk of disclosure. Quantitative Economics, 9 (1). pp. 395-440. ISSN 1759-7323

Preview

Text - Published Version
Available under License Creative Commons Attribution Non-commercial.
Download (540kB) | Preview

Publisher

Identification Number: 10.3982/QE568

Abstract

It is commonplace that the data needed for econometric inference are not contained in a single source. In this paper we analyze the problem of parametric inference from combined individual-level data when data combination is based on personal and demographic identifiers such as name, age, or address. Our main question is the identification of the econometric model based on the combined data when the data do not contain exact individual identifiers and no parametric assumptions are imposed on the joint distribution of information that is common across the combined dataset. We demonstrate the conditions on the observable marginal distributions of data in individual datasets that can and cannot guarantee identification of the parameters of interest. We also note that the data combination procedure is essential in the semiparametric setting such as ours. Provided that the (non-parametric) data combination procedure can only be defined in finite samples, we introduce a new notion of identification based on the concept of limits of statistical experiments. Our results apply to the setting where the individual data used for inferences are sensitive and their combination may lead to a substantial increase in the data sensitivity or lead to a de-anonymization of the previously anonymized information. We demonstrate that the point identification of an econometric model from combined data is incompatible with restrictions on the risk of individual disclosure. If the data combination procedure guarantees a bound on the risk of individual disclosure, then the information available from the combined dataset allows one to identify the parameter of interest only partially, and the size of the identification region is inversely related to the upper bound guarantee for the disclosure risk. This result is new in the context of data combination as we notice that the quality of links that need to be used in the combined data to assure point identification may be much higher than the average link quality in the entire dataset, and thus point inference requires the use of the most sensitive subset of the data. Our results provide important insights into the ongoing discourse on the empirical analysis of merged administrative records as well as discussions on the disclosive nature of policies implemented by the data-driven companies (such as Internet services companies and medical companies using individual patient records for policy decisions)

Item Type:	Article
Official URL:	http://qeconomics.org/ojs/index.php/qe
Additional Information:	© 2017 The Authors © CC BY-NC 3.0
Divisions:	Economics
Subjects:	H Social Sciences > HB Economic Theory
JEL classification:	C - Mathematical and Quantitative Methods > C1 - Econometric and Statistical Methods: General > C13 - Estimation C - Mathematical and Quantitative Methods > C1 - Econometric and Statistical Methods: General > C14 - Semiparametric and Nonparametric Methods C - Mathematical and Quantitative Methods > C2 - Econometric Methods: Single Equation Models; Single Variables > C25 - Discrete Regression and Qualitative Choice Models C - Mathematical and Quantitative Methods > C3 - Econometric Methods: Multiple; Simultaneous Equation Models; Multiple Variables; Endogenous Regressors > C35 - Discrete Regression and Qualitative Choice Models
Date Deposited:	31 May 2017 13:20
Last Modified:	15 Nov 2025 01:50
URI:	http://eprints.lse.ac.uk/id/eprint/79384

Actions (login required)

View Item

Download Statistics

Downloads

Downloads per month over past year

View more statistics