Reducing the Risk of Re-Identification

Re-identification is a real risk to privacy. Even when data has been de-identified by removing information like name and address, it is often possible to use other details in the data to re-identify someone and figure out who they are. 

The Devil is in the Details

The more detailed the data, the more possible it is to re-identify individuals. Re-identification risk rises when large datasets can be divided into small subgroups – like, for example, when only a few people in a dataset have the same age, race, and location.

Venn diagram of three overlapping circles titled Data Re-Identification: More Details = More Risk. Each circle represents a detail, and in the center where the three circles overlap there is an icon representing a person.

Practical Strategies for Reducing Risk

  • Don’t collect or share human subjects data you don’t need. 
  • Structure your data so it can’t be connected with identifiable public data like voter registration lists.
  • Avoid small groupings. Even groups of 5 or 6 can be easy to re-identify.
  • Get a second opinion from professionals like the Illinois Privacy Team.

Re-identification risk depends on the unique details of the data. 
So, run tests and reach out for help.


All Data Nudges are licensed under CC BY 4.0. You are free to share, adopt, or adapt them and cite the Illinois Research Data Service.

Data Nudge Newsletter Archive
Email: caldron2@illinois.edu