How to apply differential privacy on list of data?

Viewed 113

How to apply differential privacy on a list of data. OpenMined release a differential privacy project called PyDP 2 years ago. On the examples provided, they showed how to compute the PyDP on the data by computing some statistical features such as the mean, Max, Median.

  1. Is there a way to apply a differential privacy to the list of dataset, and get the list of data back, without computing any statistical feature yet ?
e.g. input_list = [1.03,2.23,3.058,4.97]

out_put_differential_privacy_list = dp_function(input_list)
out_put_differential_privacy_list
>> [1.01,2.03,3.8,4.04]
  1. How is the noise added to the data (they use laplacian)? Is the noise added taking into account the whole data set, or is it added considering each single value at a time ?

I couldn't fine the github code for pydp.algorithms.laplacian.

These are the statistical features they showed how to compute.

from pydp.algorithms.laplacian import (
    BoundedSum,
    BoundedMean,
    BoundedStandardDeviation,
    Count,
    Max,
    Min,
    Median,
)

Are they also functions to compute differential privacy percentiles ?

Any other resources will also be welcome.

1 Answers

Here are my two cents on the question,

The idea of differential privacy is to publish aggregated information of sensitive values only if noise is added to the aggregated info. This will in terms make it infeasible to match sensitive values to their owners, and also make the dataset not highly dependent on any particular data from the dataset.

The way that noise is added to the differential privacy is by injecting Laplace noise into each pieces of data at a time which in terms will add noise to the overall dataset, the essential idea of DP would be the following:

A(D,f) = f(D) + noise

  • A = some randomised algorithm
    • This is to ensure that the result each time will be slightly different.
  • f = sensitivity of a function such that this sensitivity will be used to determine to what degree an individual piece of data will affect the output.
  • D = the dataset you want to 'mask', the overall thing, in your case it would be the list of numbers.
  • noise = Laplace noise, i.e. lambda = delta f/epsilon = 1/epsilon
    • The epsilon value here is sort of indicates the privacy loss on adding/removing an entry from the dataset, i.e. making adjusments to the dataset. The smaller the epsilon is, the less privacy loss on the adjustments made on the dataset, which means better protection for privacy.

And as you can see now, the noise are only dependent on the sentitivyt and epsilon value, and has nothing to do with the underlying dataset.

... they showed how to compute the PyDP on the data by computing some statistical features such as the mean, Max, Median.

lets say for example, say we have a bunch of numbers like you have here, we could first find the max number out of the list, which would be 4.97, then we can just try to draw the eta from Lap(4.7/epsilon). I believe the idea is to sort of anchor the data around some certain statistical feature.

Hope this is somewhat useful :)

Related