The simple minded view I have is errors should NEVER be purposely introduced into Census data. Knowingly lying to try to fix a problem is almost never successful anyway, and messing with the accuracy of the Census at any level is playing with fire.
I think it's time to retire the tired trope that private companies know everything about us when the evidence given is some case of a teenage pregnancy back in 2014. If that's all the evidence we can muster, well, then you know . . . My experience is that company tracking is pretty dumb. It's the bounteous ads I get because *I just bought the damn thing they're advertising.* Or you google something once and you get obnoxious ads peppering your screen for it for months, until you swear you will never use that service ever (looking at you, zipcar. com)
As for the real substance of this post, I'd say my bias is toward more protection of personal information from our government than less, but I admit that that's not a slam dunk position of mine. Like so much of life, unlike my hatred of zipcar. com ads, it's complicated.
If you go to any website with chat support, not even logged in, and strike up a conversation, the support specialist (or bot) will generally have a profile of who you are, your other emails and screen names, your employer, etc. Not everyone will have all of that data complete in their profile and it's not that hard to defeat if you're at all tech savvy but it's a mundane, everyday example they won't know if you're pregnant generally but that's because it's not worth knowing in that context compared to the cost of harvesting, curating, and modeling that data.
Here's another example, blood banks aren't governed by HIPAA. My local blood bank, where I'm a donor, gets fussy a few weeks before I can donate again and starts hitting me with ads about every time. Back when I still used facebook, they would put ads on my facebook page, targeted at *me*, with a message like, "The need for B- is at an all time high!". That was in plain-text. So, facebook, in some dark database there's no way of knowing if they care enough to mine, knows some interesting things about me. My blood type (which has covariance with other genetic, racial, and health attributes), information about my likely sex life (I think at the time you weren't able to give blood if you were homosexual but now I think it's been changed to you can give blood as long as you haven't had same sex sexual contact in the last X days), my travel history, etc.
I get the Target example is overused, that doesn't prove the negative, though.
Yeah, the Red Cross tells me when it's been at least eight weeks since my last donation. Would the Red Cross put an ad on Facebook telling me the same thing? Don't know; not on Facebook.
Anyway, this is pretty weak tea for a panopticon world.
And if companies or bad actors are trying to use blood type to supposedly discover things about you you'd rather not have them know, then I suspect they're being snowed. Blood type? Really?
I used 60 seconds to come up with top of mind examples. The missed point is that these things are mundane and banal and everywhere. The largest barrier to companies "knowing" something about you is that the apparent value exploitation of that knowledge could deliver is lower than the cost to organize that data. The examples of value are becoming more accessible and the cost to organize is getting lower, both incrementally.
If you read that context as uninteresting, that's okay.
Besides reducing the real accuracy of the small area counts and estimates and undermining faith in the data, it's more expensive to do this AND has resulted in a long delay for the release of the data. The newest ACS was supposed to release in December and has been rescheduled twice already. I think it will be rescheduled once more. It's an expensive project to make the product less valuable.
I lead the data science team at a company that makes use of Census data and derivative products. We already have data to basically undo the differential privacy errors applied. The net effect is to add additional complexity (and thus monopoly) in the community scale data market and undermine faith in the publicly purchased and published numbers.
I don't understand the question. It's always been published de-identified (it's aggregated and most statistics are only available at block group and up). It's used for "profit" (commercially and by local governments, who often pay for expert help in doing so) every day by thousands of firms and local government entities.
PUMS data is still aggregated. It's just making smaller crosstabs available at smaller geographies for higher density areas. A limited set of researchers have access to raw (non-aggregated) Census data.
Census data would not be my 3rd or 4th or 10th choice to identify individuals. At scale, I'd look at credit, tax, and county assessor's data. There are firms that even specialize in joining your data across sites and creating a likely individual profile for you they sell. Almost everyone reuses screen names and emails and carries their facebook, google, and twitter cookies around to every site they visit. 99% of the time when someone visits any of the apps or websites a modern firm operates, you know a ton of information about them just from the identifying characteristics of their visit you can link to other data. Not at scale, you combine some of that with simple google searches, phone number lookups, forum history, etc. It's not difficult and Census data is not essential or even a big boost at the individual scale.
It is true that private concerns could gather the information you cite.. But gathering that information would also be extremely costly just like it is for the Census which is exactly why they would de engineer expensively gathered census date. The expensive part is done. I have to think that you are overstating the case here. As long as the census preserves unaltered data for critical applications like redistricting I don't see an enormous problem. Especially since the decade interval census snapshot does not make the granular data particularly useful for shorter term analysis.
In many ways, using census data to apportion congressional seats can never really produce equal representation, since the data gets distorted by people who are not eligible to vote, whether it be prison inmates or illegal immigrants. There is no logical reason why citizens who happen to live geographically near such groups should get more political power than citizens who don't, but that's effectively what the current system does.
At first glance, it seems as though the presence of noncitizens in calculating political power favors Democrats (Republicans certainly think so, as indicated by the census meddling under Trump); however, if you dig deeper, it's not necessarily clear cut. Texas has a huge immigrant population, whose presence results in more Republican electoral votes. At the House level, districts are drawn by partisan legislatures, so immigrant groups with a lot of noncitizens can still be drawn into Republican districts, even if most of their geographic citizen neighbors are Democrats. And, if the term "noncitizen" is generalized to include prison inmates (who are ineligible to vote), the prison population effectively boosts the electoral power of voters who happen live near the prison - and since prisons are usually located outside of large cities, these voters are usually Republicans.
Overall, which party actually benefits from people who are not eligible to vote still counting towards the drawing of districts is not clear. I can easily imagine Democrats benefiting in some states, Republicans in others, and the two errors mostly canceling each other out.
> the data gets distorted by people who are not eligible to vote, whether it be prison inmates or illegal immigrants.
No: representation is of the total population, *not* just of voters. See 14th Amendment, section 2: "Representatives shall be apportioned among the several States according to their respective numbers, counting the whole number of persons in each State, excluding Indians not taxed." --- and even before that, the original text of Article I, section 2, included the whole population, even slaves who were obviously ineligible to vote (albeit at a lower rate). When the drafters of those provisions wanted to distinguish between people in general and citizens, specifically, they knew how to do it; the fact that they didn't implies something about who's to be counted.
There’s a concept of “virtual representation” - if non-voters in a given geography can be assumed to have similar interests to voters in that same geography (a big “if” for prison populations, but much more plausible for children in family neighborhoods and not-yet-naturalized immigrants in immigrant neighborhoods, and hard to say for eligible voters who choose not to vote) then you can get the right number of representatives for their interests.
The largest group that is not eligible to vote are people under 18 I would assume. Weirdly while young people (and even young parents under ~40) skew democratic, the youngest states are generally redder, the youngest states are Utah, Texas, Alaska, North Dakota, Nebraska, Idaho, Oklahoma. Not including DC because it's just an urban core and it also doesn't have real house representation. Next is California, but even so it's a pretty right wing group. So even here the influence of districting to include non-voters is awkward given the politics of non-eligible voters.
if there are 500m internet users in the US and Europe (there are more), and they take just one second per day clicking "Allow Cookies", then that means we're wasting 15.85 years of human life every day. Every single day, these privacy nazis are doing the equivalent of locking someone up for a decade and a half. That is unconscionable.
On one side, there’s a bunch of political scientists and economists who think they have the Constitutional right to have the taxpayers do data collection for them, and will now have a slightly harder time publishing papers that “discover” “causal effects”.
On the other side, there’s a bunch of theoretical computer scientists who, in the tradition of their field, have a very paranoid threat model and always focus on the worst possible outcome. It’s pretty rare for theoretical CS to have any real impact and they’re not going to give up on that opportunity.
I’m biased towards the computer scientists here, although the anti-differential-privacy crowd’s arguments seem stronger now than a year or two ago (as the two sides come to understand each other better).
If the non-noisy data is released, someone will do the database reconstruction attack, and it will become available on the Internet. It won’t be hypothetical for long. I don’t actually have a sense of how bad this would be in practice.
Hi All. I run a privacy think tank (FPF.org) - a pragmatic centrist think tank focused on helping support responsible uses of data, working with the Chief Privacy Officers of many data driven companies, researchers, cities and schools....I am a data optimist and enthusiast, not an overly cautious fanatic. I have closely followed this issue. My informed view is 1) the actual privacy risk of releasing large overlapping sets of data is real. The Census will never be trusted if we release data as in the past and it is used to disclose details people consider confidential. 2) For many years, the Census has added noise to data to support deidentification. Many users of these data sets havent been aware, or simply relied on the final data. Now the Census is transparently explaining the techniques used. 3) Changing the way data is released is painful due to the transitions that will need to be made by those relying on the previous techniques, and the tradeoffs required by the protections needed. 4) Differential privacy is currently the most sophisticated way to assess techniques that add some noise to data, in a manner that maintains accuracy at aggregated levels. My interview with one of the "inventors" of differential privacy can be viewed here. https://www.linkedin.com/video/live/urn:li:ugcPost:6783434965402075136/
In that long paragraph, you don't even come close to making an actual argument in favor of your position. What exactly is the "risk" and how does it compare to the massive amounts of data that people are constantly handing over to private companies every second of every day.
Happy to post links to the extensive debates by technical experts on this...will do so as soon as I can today. With regard to the data I provide to Google and the rest of the companies I interact with - they do not make it public, and I dont have to use those services, or can choose less data intensive alternatives. This data is published to the world. How many people are in my household is no big deal to some of us - to far more vulnerable people, the census data being identified can be damaging.
My impression from the Census Bureau's big reconstruction and reidentification attempt was that (1) it only successfuly reidentified a minority of people, and (2) a hypothetical attacker would have no way of knowing how successful their reidentification attempt was, because that attacker would not have access to the original identifying information that the Census Bureau does.
My takeaway from this was that reidentification doesn't seem to be a huge risk yet - am I getting something wrong here?
Also, it's hard for me to imagine what harmful things would befall people if Census data were released, given most of it is pretty mundane (age, race, etc.). What are your thoughts on that?
As always, the question I want an answer to is RISKY COMPARED TO WHAT?
Consider that for most of the latter half of the 20th century, AT&T -- as a treasured public service! -- distributed printed books listing the name and address of the vast majority of people in the country. Everyone who lived in the same city got the names and addresses of all of their neighbors, and most major libraries got copies of the whole shebang. They also published or licensed the publishing of "reverse directories" that linked numbers and addresses back to names.
We somehow managed to survive this era without whatever privacy apocalypse is being worried about here, so I'd really like to know what the scenario is with the census that is so much worse that it justifies eliding the raw data.
Especially in a world where Equifax hacks have already exposed all of the genuinely interesting information about individuals already, right on down to the complete salary histories of 125 million of us.
This is a fascinating topic for me to discuss because while I was in applied math grad school at MIT, my research was tangentially related to the academic discipline of differential privacy, which is generally seen as one of the hot up-and-coming fields in computer science. It is genuinely exciting for many of those researchers to be able to see real-world applications on something as high profile as the Census, and while I'm not personally motivated by privacy concerns, I sympathize with the attempt to try to quantitatively balance privacy and accuracy.
One of the things that I would emphasize in this discussion is that all of the various knobs in these algorithms are adjustable. If they're indeed overvaluing privacy as you claim, they can always turn the amount of noise down to produce generally more accurate counts the next time around. If the adjustment algorithms prioritize getting the wrong counts right, they can adjust the algorithm to get the counts that matter more right. Just as you're questioning the value of the privacy offered, one could also question the value of exact accuracy, and perhaps attempt to offer an actual cost-benefit analysis. I fully expect academics to dissect this a thousand ways and be able to come to some sort of general consensus of the different options and their tradeoffs over time.
I'd also just emphasize how new all of this. DP as an academic discipline really only got kicked off with a 2006 paper establishing how noise injection leads to privacy. There's been a lot of academic work since then, but handling the complexities of applying this to real world data is just getting started.
Isn't part of the problem that it's hard to prove tight bounds on how well an algorithm preserves privacy? So e.g. if we decide to use an algorithm that allows an at most 5% probability of someone's identity being uncovered, it might actually be much lower than 5%, and maybe we could have achieved 5% with much less noise introduced. (Maybe even the old "swapping" technique was already sufficient to achieve eps-differential privacy, but we just couldn't prove that it works!)
I think this answer is true in a general sense -- I'd just shorten it to saying that it's hard to prove things -- but these criticisms don't necessarily hold in this particular case.
First, the definition of differential privacy doesn't reason in terms of a probability of someone's identity being uncovered; it's quite a bit more robust than that. I would explain it as an upper bound on the amount of information anyone (with any amount of time or computing power on their hands) can glean about any individual person from the data released. (Here by "information" I technically mean something like "log odds ratio" for those with more technical background.)
Second, yes, there can often be some slack in the sorts of bounds that can be proven, but computer scientists are also always trying to characterize the other side of the equation with concrete examples. At first they always start with theoretical scenarios, because that's what you can most easily prove things about, but with all of this real world data, I fully expect plenty of people to try to quantify the actual amount of information released by this technique. We know based on the proofs that there's some upper bound, but how close do we get to it in reality? There are lots of interesting (and now, important) questions to be uncovered there, and I fully expect a robust literature to develop to answer that question, if it hasn't already.
Third, while I'm not directly familiar with the swapping technique, I would fully expect that they've already demonstrated that previous techniques had problems. That's like a prerequisite to the field even being interesting in the first place, and speaking from experience hearing various talks about DP, they often have several stories to share about privacy bring broken in unexpected ways.
This is all very reasonable, but John Abowd at Cornell Economics led the Census Bureau's charge on this, and while he may be wrong on this issue, he's neither evil nor stupid. For all the controversy and criticism about this change, I'm shocked that there's no "John Abowd faces his critics" type interview anywhere. Presumably he has SOME responses to these critiques, but I feel like the conversation never directly engages both sides of the argument and people kind of just talk past each other.
Another example of making policy without cost benefit analysis.
My related beef is that medical communication has to be conducted by clunky "Portals" rather than old fashioned email. What is the expected value of the harm of someone hacking my email an learning my PSA at the dame time I do? Does it exceed the cost of establishing and using the "portal?" Much communication with the government suffers the same problem. Why does it need to be more difficult to look up something from my Social Security account than to buy something from Amazon?
Email is uniquely terrible in a way that makes this necessary. If it were e.g. Whatsapp it would be fine. But email is transparent to too damn many intermediaries.
That’s one of those too-convenient sour grapes assumptions, like “don’t worry about biased job interviewers because you don’t want to work for any company with biased people anyway”.
I don’t know how many children share email addresses with their parents, but if they start getting STD tests they probably don’t want their parents knowing, even if they didn’t mind sharing stuff with their parents when they were 13.
Maybe I'm to categorical, but my feeling is that on balance more costs are imposed than benefits created by these privacy rules. Perhaps if one could opt out. Like the easier to open and more secure medicine containers. :)
I definitely believe that there are more costs than benefits here (especially when it comes to things like my doctor saying "we can't e-mail you the x-ray, but I won't object if you take a picture of my computer screen right now") but I just want to be sure that we're not pretending that the benefits are zero.
Here's a partial solution to the issue - as well as a slight correction on Matt’s description of noise at larger geographies versus smaller ones.
The biggest distortions don’t come from differential privacy on its own. The noise added is pretty small, and is normally distributed which is good.
The worst distortions come from modifications made to the correct for problems the normally-distributed noise causes:
(1) Values that look weird – say, a Census Block with negative 3 houses, or a county with 800.354 people – are changed to look normal.
(2) Data for different levels are modified so all levels line up (for example, the population of states might be modified to sum to the population of the country). They have to do this because the noise gets added to summary data at every level, making different levels not cohere with one another. They start modifying at the top and work down, which is why smallest geographies are the worst - they're affected by all the modifications at the levels above them.
A bunch of researchers have suggested the Bureau also release data with just the original normally distributed noise - not the corrections for weird values and for alignments across levels. Last I heard, the Bureau wasn’t considering this, which is super weird - the data with just the normally distributed noise is still differentially private, and would be *way* more useful to researchers.
Two quick points on this: "But as John Roberston noted on Twitter, the Atlanta Fed is relying on the ACS for their data. If they use rounded data, then their median wage growth metric gets totally broken. Not great!"
First, the Atlanta Fed wage tracker uses the CPS (Current Population Survey) not the ACS (American Community Survey). The current discussion about rounding the reported wage data refers to the CPS, not the ACS.
Second, John Robertson tweeted the chart that shows the impact of the rounding prposal on Atlanta Fed wage tracker, but the chart was created by Ben Zipperer at the Economic Policy Institute. Here is the original tweet thread, with the chart and a link to the code that was used to produce it: https://twitter.com/benzipperer/status/1485665258438078467?s=20
I get the intellectual privacy concern, but people need to realize we lost that war without a fight over 20 years ago. All your data is easily acquired by anyone who wants it. There is nothing you can do about it, unless you are willing to live in a cave. And even then, wait til the satellite cameras get better resolution....
No doubt, this is bad for people who really do need protections - such as people trying to flee abusive ex spouses, people in witness protection, etc. But being mad about that is like being mad at an asteroid coming at you from space... be mad all you want, but the asteroid doesn't care about your feelings.
Hi there! European living in the US here. With the exception of American healthcare, I can think of few things more stupid than GDPR. When I’m home, I frequently use an American VPN to browse the internet without all the pop ups. I understand that we wanted to hurt US tech companies, but we should have found a more user friendly way.
I wouldn't hold my breath. They've just about wrangled Microsoft now about bundling, only 10 years after it was 10 years too late. Google and Apple are going to have so much data just from being the mobile providers that it's almost besides the point now.
For big tech firms, it's a mixed bag as it makes operating more complex which basically disadvantages small and medium firms. CCPA is more stringent than GDPR in general but neither is incredibly impactful to Facebook or Google. They are already very good at getting their users to agree to [whatever] using features and network leverage.
Having led projects to comply with both under multiple products. It's generally been sufficient to focus on CCPA first and then the product is mostly GDPR compliant as a consequence (gap analysis reveals very little effort needed to bring it inline). That's conventional wisdom AFAIK, too. Do you have counter-examples? I'm sorry, it's not a discussion interesting enough to me to cite sources in the comments about. Basing it on multiple personal experiences and similar anecdotes from professional sources I trust.
Given even odds, I'd put money on click-wrap licensing enduring in the US.
Also, the environment that allowed that cadre of small-town lawyers to target false advertising has shifted significantly. There's a massive pile of services and experiences that case law suggests should be WCAG 2.0 compliant in order to comply with the ADA but are not. There are some lawsuits filed over it but the quantity and quality of suits and outcomes is, at least for me, unsatisfying.
If big companies are going to have access to this information anyway, then why not make it available to the people who will do public good with this information too?
The resources, time, etc for a company to get this data is literally so low that you might as well consider it a rounding error.
I would argue against doing a census due to how inefficient it is, since you could easily get the same data, at higher quality, through brokers. The census is largely a legal and political matter of how you count, not what you can actually know.
There are things we do in private, things we do in public, and things we do we do with specific groups of people. I think it's bad to invade private behavior, and extremely bad to leak across life-compartments (work vs. friends vs. family vs. internet), but the kind of stuff that's included in the census or the phone book is neither of those. It's just public.
So are you saying it's not true that that info is available? I think it's likely true that it is available, but only available with effort and only available on average. No guarantees that you can get it for any specific person.
Thus I'm not willing to increase the chances to 100 percent that my info is obtainable.
1. Well, if the information is already out there, the damage is already done. Nothing to be done about it, so you *should* stop fighting.
I'm not confident that the information is or is not already out there, but what we have here is a disagreement on an empirical fact. Some one just has to figure it out.
2. I do not think the root comment made this claim.
I have no clue what I just read. I'm totally lost.
This was fascinating Matt. Who knew???
The simple minded view I have is errors should NEVER be purposely introduced into Census data. Knowingly lying to try to fix a problem is almost never successful anyway, and messing with the accuracy of the Census at any level is playing with fire.
Great, informative article.
I think it's time to retire the tired trope that private companies know everything about us when the evidence given is some case of a teenage pregnancy back in 2014. If that's all the evidence we can muster, well, then you know . . . My experience is that company tracking is pretty dumb. It's the bounteous ads I get because *I just bought the damn thing they're advertising.* Or you google something once and you get obnoxious ads peppering your screen for it for months, until you swear you will never use that service ever (looking at you, zipcar. com)
As for the real substance of this post, I'd say my bias is toward more protection of personal information from our government than less, but I admit that that's not a slam dunk position of mine. Like so much of life, unlike my hatred of zipcar. com ads, it's complicated.
If you go to any website with chat support, not even logged in, and strike up a conversation, the support specialist (or bot) will generally have a profile of who you are, your other emails and screen names, your employer, etc. Not everyone will have all of that data complete in their profile and it's not that hard to defeat if you're at all tech savvy but it's a mundane, everyday example they won't know if you're pregnant generally but that's because it's not worth knowing in that context compared to the cost of harvesting, curating, and modeling that data.
Here's another example, blood banks aren't governed by HIPAA. My local blood bank, where I'm a donor, gets fussy a few weeks before I can donate again and starts hitting me with ads about every time. Back when I still used facebook, they would put ads on my facebook page, targeted at *me*, with a message like, "The need for B- is at an all time high!". That was in plain-text. So, facebook, in some dark database there's no way of knowing if they care enough to mine, knows some interesting things about me. My blood type (which has covariance with other genetic, racial, and health attributes), information about my likely sex life (I think at the time you weren't able to give blood if you were homosexual but now I think it's been changed to you can give blood as long as you haven't had same sex sexual contact in the last X days), my travel history, etc.
I get the Target example is overused, that doesn't prove the negative, though.
Yeah, the Red Cross tells me when it's been at least eight weeks since my last donation. Would the Red Cross put an ad on Facebook telling me the same thing? Don't know; not on Facebook.
Anyway, this is pretty weak tea for a panopticon world.
And if companies or bad actors are trying to use blood type to supposedly discover things about you you'd rather not have them know, then I suspect they're being snowed. Blood type? Really?
I used 60 seconds to come up with top of mind examples. The missed point is that these things are mundane and banal and everywhere. The largest barrier to companies "knowing" something about you is that the apparent value exploitation of that knowledge could deliver is lower than the cost to organize that data. The examples of value are becoming more accessible and the cost to organize is getting lower, both incrementally.
If you read that context as uninteresting, that's okay.
Besides reducing the real accuracy of the small area counts and estimates and undermining faith in the data, it's more expensive to do this AND has resulted in a long delay for the release of the data. The newest ACS was supposed to release in December and has been rescheduled twice already. I think it will be rescheduled once more. It's an expensive project to make the product less valuable.
I lead the data science team at a company that makes use of Census data and derivative products. We already have data to basically undo the differential privacy errors applied. The net effect is to add additional complexity (and thus monopoly) in the community scale data market and undermine faith in the publicly purchased and published numbers.
I don't understand the question. It's always been published de-identified (it's aggregated and most statistics are only available at block group and up). It's used for "profit" (commercially and by local governments, who often pay for expert help in doing so) every day by thousands of firms and local government entities.
PUMS data is still aggregated. It's just making smaller crosstabs available at smaller geographies for higher density areas. A limited set of researchers have access to raw (non-aggregated) Census data.
Census data would not be my 3rd or 4th or 10th choice to identify individuals. At scale, I'd look at credit, tax, and county assessor's data. There are firms that even specialize in joining your data across sites and creating a likely individual profile for you they sell. Almost everyone reuses screen names and emails and carries their facebook, google, and twitter cookies around to every site they visit. 99% of the time when someone visits any of the apps or websites a modern firm operates, you know a ton of information about them just from the identifying characteristics of their visit you can link to other data. Not at scale, you combine some of that with simple google searches, phone number lookups, forum history, etc. It's not difficult and Census data is not essential or even a big boost at the individual scale.
It is true that private concerns could gather the information you cite.. But gathering that information would also be extremely costly just like it is for the Census which is exactly why they would de engineer expensively gathered census date. The expensive part is done. I have to think that you are overstating the case here. As long as the census preserves unaltered data for critical applications like redistricting I don't see an enormous problem. Especially since the decade interval census snapshot does not make the granular data particularly useful for shorter term analysis.
In many ways, using census data to apportion congressional seats can never really produce equal representation, since the data gets distorted by people who are not eligible to vote, whether it be prison inmates or illegal immigrants. There is no logical reason why citizens who happen to live geographically near such groups should get more political power than citizens who don't, but that's effectively what the current system does.
At first glance, it seems as though the presence of noncitizens in calculating political power favors Democrats (Republicans certainly think so, as indicated by the census meddling under Trump); however, if you dig deeper, it's not necessarily clear cut. Texas has a huge immigrant population, whose presence results in more Republican electoral votes. At the House level, districts are drawn by partisan legislatures, so immigrant groups with a lot of noncitizens can still be drawn into Republican districts, even if most of their geographic citizen neighbors are Democrats. And, if the term "noncitizen" is generalized to include prison inmates (who are ineligible to vote), the prison population effectively boosts the electoral power of voters who happen live near the prison - and since prisons are usually located outside of large cities, these voters are usually Republicans.
Overall, which party actually benefits from people who are not eligible to vote still counting towards the drawing of districts is not clear. I can easily imagine Democrats benefiting in some states, Republicans in others, and the two errors mostly canceling each other out.
> the data gets distorted by people who are not eligible to vote, whether it be prison inmates or illegal immigrants.
No: representation is of the total population, *not* just of voters. See 14th Amendment, section 2: "Representatives shall be apportioned among the several States according to their respective numbers, counting the whole number of persons in each State, excluding Indians not taxed." --- and even before that, the original text of Article I, section 2, included the whole population, even slaves who were obviously ineligible to vote (albeit at a lower rate). When the drafters of those provisions wanted to distinguish between people in general and citizens, specifically, they knew how to do it; the fact that they didn't implies something about who's to be counted.
There’s a concept of “virtual representation” - if non-voters in a given geography can be assumed to have similar interests to voters in that same geography (a big “if” for prison populations, but much more plausible for children in family neighborhoods and not-yet-naturalized immigrants in immigrant neighborhoods, and hard to say for eligible voters who choose not to vote) then you can get the right number of representatives for their interests.
The largest group that is not eligible to vote are people under 18 I would assume. Weirdly while young people (and even young parents under ~40) skew democratic, the youngest states are generally redder, the youngest states are Utah, Texas, Alaska, North Dakota, Nebraska, Idaho, Oklahoma. Not including DC because it's just an urban core and it also doesn't have real house representation. Next is California, but even so it's a pretty right wing group. So even here the influence of districting to include non-voters is awkward given the politics of non-eligible voters.
think about how much collective time has been wasted by all of clicking "Allow Cookies"
like, who exactly are these privacy nazis? Who is it that cares this much?
if there are 500m internet users in the US and Europe (there are more), and they take just one second per day clicking "Allow Cookies", then that means we're wasting 15.85 years of human life every day. Every single day, these privacy nazis are doing the equivalent of locking someone up for a decade and a half. That is unconscionable.
I share your hate for these boxes. I use a Chrome extension that takes care of most of these: https://chrome.google.com/webstore/detail/i-dont-care-about-cookies/fihnjjcciajhdojfnbdddfaoknhalnja?hl=en
Germans. I think they are technically privacy anti-Nazis and anti-Stasi.
On one side, there’s a bunch of political scientists and economists who think they have the Constitutional right to have the taxpayers do data collection for them, and will now have a slightly harder time publishing papers that “discover” “causal effects”.
On the other side, there’s a bunch of theoretical computer scientists who, in the tradition of their field, have a very paranoid threat model and always focus on the worst possible outcome. It’s pretty rare for theoretical CS to have any real impact and they’re not going to give up on that opportunity.
I’m biased towards the computer scientists here, although the anti-differential-privacy crowd’s arguments seem stronger now than a year or two ago (as the two sides come to understand each other better).
If the non-noisy data is released, someone will do the database reconstruction attack, and it will become available on the Internet. It won’t be hypothetical for long. I don’t actually have a sense of how bad this would be in practice.
Hi All. I run a privacy think tank (FPF.org) - a pragmatic centrist think tank focused on helping support responsible uses of data, working with the Chief Privacy Officers of many data driven companies, researchers, cities and schools....I am a data optimist and enthusiast, not an overly cautious fanatic. I have closely followed this issue. My informed view is 1) the actual privacy risk of releasing large overlapping sets of data is real. The Census will never be trusted if we release data as in the past and it is used to disclose details people consider confidential. 2) For many years, the Census has added noise to data to support deidentification. Many users of these data sets havent been aware, or simply relied on the final data. Now the Census is transparently explaining the techniques used. 3) Changing the way data is released is painful due to the transitions that will need to be made by those relying on the previous techniques, and the tradeoffs required by the protections needed. 4) Differential privacy is currently the most sophisticated way to assess techniques that add some noise to data, in a manner that maintains accuracy at aggregated levels. My interview with one of the "inventors" of differential privacy can be viewed here. https://www.linkedin.com/video/live/urn:li:ugcPost:6783434965402075136/
yeah so like what's the marginal benefit that's being achieved here?
In that long paragraph, you don't even come close to making an actual argument in favor of your position. What exactly is the "risk" and how does it compare to the massive amounts of data that people are constantly handing over to private companies every second of every day.
Happy to post links to the extensive debates by technical experts on this...will do so as soon as I can today. With regard to the data I provide to Google and the rest of the companies I interact with - they do not make it public, and I dont have to use those services, or can choose less data intensive alternatives. This data is published to the world. How many people are in my household is no big deal to some of us - to far more vulnerable people, the census data being identified can be damaging.
I'm not reading extensive technical debates. If the argument can't be explained clearly and shortly then its probably not very good.
My impression from the Census Bureau's big reconstruction and reidentification attempt was that (1) it only successfuly reidentified a minority of people, and (2) a hypothetical attacker would have no way of knowing how successful their reidentification attempt was, because that attacker would not have access to the original identifying information that the Census Bureau does.
My takeaway from this was that reidentification doesn't seem to be a huge risk yet - am I getting something wrong here?
Also, it's hard for me to imagine what harmful things would befall people if Census data were released, given most of it is pretty mundane (age, race, etc.). What are your thoughts on that?
As always, the question I want an answer to is RISKY COMPARED TO WHAT?
Consider that for most of the latter half of the 20th century, AT&T -- as a treasured public service! -- distributed printed books listing the name and address of the vast majority of people in the country. Everyone who lived in the same city got the names and addresses of all of their neighbors, and most major libraries got copies of the whole shebang. They also published or licensed the publishing of "reverse directories" that linked numbers and addresses back to names.
We somehow managed to survive this era without whatever privacy apocalypse is being worried about here, so I'd really like to know what the scenario is with the census that is so much worse that it justifies eliding the raw data.
Especially in a world where Equifax hacks have already exposed all of the genuinely interesting information about individuals already, right on down to the complete salary histories of 125 million of us.
This is a fascinating topic for me to discuss because while I was in applied math grad school at MIT, my research was tangentially related to the academic discipline of differential privacy, which is generally seen as one of the hot up-and-coming fields in computer science. It is genuinely exciting for many of those researchers to be able to see real-world applications on something as high profile as the Census, and while I'm not personally motivated by privacy concerns, I sympathize with the attempt to try to quantitatively balance privacy and accuracy.
One of the things that I would emphasize in this discussion is that all of the various knobs in these algorithms are adjustable. If they're indeed overvaluing privacy as you claim, they can always turn the amount of noise down to produce generally more accurate counts the next time around. If the adjustment algorithms prioritize getting the wrong counts right, they can adjust the algorithm to get the counts that matter more right. Just as you're questioning the value of the privacy offered, one could also question the value of exact accuracy, and perhaps attempt to offer an actual cost-benefit analysis. I fully expect academics to dissect this a thousand ways and be able to come to some sort of general consensus of the different options and their tradeoffs over time.
I'd also just emphasize how new all of this. DP as an academic discipline really only got kicked off with a 2006 paper establishing how noise injection leads to privacy. There's been a lot of academic work since then, but handling the complexities of applying this to real world data is just getting started.
Isn't part of the problem that it's hard to prove tight bounds on how well an algorithm preserves privacy? So e.g. if we decide to use an algorithm that allows an at most 5% probability of someone's identity being uncovered, it might actually be much lower than 5%, and maybe we could have achieved 5% with much less noise introduced. (Maybe even the old "swapping" technique was already sufficient to achieve eps-differential privacy, but we just couldn't prove that it works!)
I think this answer is true in a general sense -- I'd just shorten it to saying that it's hard to prove things -- but these criticisms don't necessarily hold in this particular case.
First, the definition of differential privacy doesn't reason in terms of a probability of someone's identity being uncovered; it's quite a bit more robust than that. I would explain it as an upper bound on the amount of information anyone (with any amount of time or computing power on their hands) can glean about any individual person from the data released. (Here by "information" I technically mean something like "log odds ratio" for those with more technical background.)
Second, yes, there can often be some slack in the sorts of bounds that can be proven, but computer scientists are also always trying to characterize the other side of the equation with concrete examples. At first they always start with theoretical scenarios, because that's what you can most easily prove things about, but with all of this real world data, I fully expect plenty of people to try to quantify the actual amount of information released by this technique. We know based on the proofs that there's some upper bound, but how close do we get to it in reality? There are lots of interesting (and now, important) questions to be uncovered there, and I fully expect a robust literature to develop to answer that question, if it hasn't already.
Third, while I'm not directly familiar with the swapping technique, I would fully expect that they've already demonstrated that previous techniques had problems. That's like a prerequisite to the field even being interesting in the first place, and speaking from experience hearing various talks about DP, they often have several stories to share about privacy bring broken in unexpected ways.
This is all very reasonable, but John Abowd at Cornell Economics led the Census Bureau's charge on this, and while he may be wrong on this issue, he's neither evil nor stupid. For all the controversy and criticism about this change, I'm shocked that there's no "John Abowd faces his critics" type interview anywhere. Presumably he has SOME responses to these critiques, but I feel like the conversation never directly engages both sides of the argument and people kind of just talk past each other.
How do we make this issue go viral email our congressman. This is important but hard to get on the front page.
Another example of making policy without cost benefit analysis.
My related beef is that medical communication has to be conducted by clunky "Portals" rather than old fashioned email. What is the expected value of the harm of someone hacking my email an learning my PSA at the dame time I do? Does it exceed the cost of establishing and using the "portal?" Much communication with the government suffers the same problem. Why does it need to be more difficult to look up something from my Social Security account than to buy something from Amazon?
Email is uniquely terrible in a way that makes this necessary. If it were e.g. Whatsapp it would be fine. But email is transparent to too damn many intermediaries.
Whatsapp would be OK with me, but it's pretty hilarious to imagine MY provider with a whatsapp account. He's just about as likely to be in Tic Toc! :)
I think the medical stuff is more about your spouse or parent who might share your email account getting your private medical records.
Anyone who shares an email account probably does not value that kind of medical privacy very highly. :)
That’s one of those too-convenient sour grapes assumptions, like “don’t worry about biased job interviewers because you don’t want to work for any company with biased people anyway”.
I don’t know how many children share email addresses with their parents, but if they start getting STD tests they probably don’t want their parents knowing, even if they didn’t mind sharing stuff with their parents when they were 13.
Maybe I'm to categorical, but my feeling is that on balance more costs are imposed than benefits created by these privacy rules. Perhaps if one could opt out. Like the easier to open and more secure medicine containers. :)
I definitely believe that there are more costs than benefits here (especially when it comes to things like my doctor saying "we can't e-mail you the x-ray, but I won't object if you take a picture of my computer screen right now") but I just want to be sure that we're not pretending that the benefits are zero.
And I'm making sure we don't think the costs are zero. We are non-zero guys. :)
Here's a partial solution to the issue - as well as a slight correction on Matt’s description of noise at larger geographies versus smaller ones.
The biggest distortions don’t come from differential privacy on its own. The noise added is pretty small, and is normally distributed which is good.
The worst distortions come from modifications made to the correct for problems the normally-distributed noise causes:
(1) Values that look weird – say, a Census Block with negative 3 houses, or a county with 800.354 people – are changed to look normal.
(2) Data for different levels are modified so all levels line up (for example, the population of states might be modified to sum to the population of the country). They have to do this because the noise gets added to summary data at every level, making different levels not cohere with one another. They start modifying at the top and work down, which is why smallest geographies are the worst - they're affected by all the modifications at the levels above them.
A bunch of researchers have suggested the Bureau also release data with just the original normally distributed noise - not the corrections for weird values and for alignments across levels. Last I heard, the Bureau wasn’t considering this, which is super weird - the data with just the normally distributed noise is still differentially private, and would be *way* more useful to researchers.
Two quick points on this: "But as John Roberston noted on Twitter, the Atlanta Fed is relying on the ACS for their data. If they use rounded data, then their median wage growth metric gets totally broken. Not great!"
First, the Atlanta Fed wage tracker uses the CPS (Current Population Survey) not the ACS (American Community Survey). The current discussion about rounding the reported wage data refers to the CPS, not the ACS.
Second, John Robertson tweeted the chart that shows the impact of the rounding prposal on Atlanta Fed wage tracker, but the chart was created by Ben Zipperer at the Economic Policy Institute. Here is the original tweet thread, with the chart and a link to the code that was used to produce it: https://twitter.com/benzipperer/status/1485665258438078467?s=20
Apologies to John Robertson! John Robertson and Ben Zipperer produced almost identical charts independently.
I get the intellectual privacy concern, but people need to realize we lost that war without a fight over 20 years ago. All your data is easily acquired by anyone who wants it. There is nothing you can do about it, unless you are willing to live in a cave. And even then, wait til the satellite cameras get better resolution....
No doubt, this is bad for people who really do need protections - such as people trying to flee abusive ex spouses, people in witness protection, etc. But being mad about that is like being mad at an asteroid coming at you from space... be mad all you want, but the asteroid doesn't care about your feelings.
Hi there! European living in the US here. With the exception of American healthcare, I can think of few things more stupid than GDPR. When I’m home, I frequently use an American VPN to browse the internet without all the pop ups. I understand that we wanted to hurt US tech companies, but we should have found a more user friendly way.
I wouldn't hold my breath. They've just about wrangled Microsoft now about bundling, only 10 years after it was 10 years too late. Google and Apple are going to have so much data just from being the mobile providers that it's almost besides the point now.
For big tech firms, it's a mixed bag as it makes operating more complex which basically disadvantages small and medium firms. CCPA is more stringent than GDPR in general but neither is incredibly impactful to Facebook or Google. They are already very good at getting their users to agree to [whatever] using features and network leverage.
tl;dr Google and Facebook are not much harmed
On what basis do you view the CCPA as "more stringent" than the GDPR?
Having led projects to comply with both under multiple products. It's generally been sufficient to focus on CCPA first and then the product is mostly GDPR compliant as a consequence (gap analysis reveals very little effort needed to bring it inline). That's conventional wisdom AFAIK, too. Do you have counter-examples? I'm sorry, it's not a discussion interesting enough to me to cite sources in the comments about. Basing it on multiple personal experiences and similar anecdotes from professional sources I trust.
Given even odds, I'd put money on click-wrap licensing enduring in the US.
Also, the environment that allowed that cadre of small-town lawyers to target false advertising has shifted significantly. There's a massive pile of services and experiences that case law suggests should be WCAG 2.0 compliant in order to comply with the ADA but are not. There are some lawsuits filed over it but the quantity and quality of suits and outcomes is, at least for me, unsatisfying.
If big companies are going to have access to this information anyway, then why not make it available to the people who will do public good with this information too?
The resources, time, etc for a company to get this data is literally so low that you might as well consider it a rounding error.
I would argue against doing a census due to how inefficient it is, since you could easily get the same data, at higher quality, through brokers. The census is largely a legal and political matter of how you count, not what you can actually know.
Well, it's not free for you, me, or a company. We pay taxes.
(I'm sure we will all be paying less taxes now that we're getting a worse product, right?}
Nope. More. It was more expensive to do this.
I'm not sure where the "nope" part comes in. I didn't make any claims about the magnitude of the price.
> I'm sure we will all be paying less taxes now that we're getting a worse product, right?
You asked a (likely) sarcastic question. I answered "Nope" and elaborated. I'm not sure how that flow was hard to understand.
Oh, I didn't realize you were responding to a parenthetical joke aside rather than disagreeing with my comment.
You specifically said it's available to a company for free.
But it's not free. They all paid taxes for it.
Association of my real-world identity with my internet opinions != disclosure that an individual with my real-world properties exists.
There are things we do in private, things we do in public, and things we do we do with specific groups of people. I think it's bad to invade private behavior, and extremely bad to leak across life-compartments (work vs. friends vs. family vs. internet), but the kind of stuff that's included in the census or the phone book is neither of those. It's just public.
So are you saying it's not true that that info is available? I think it's likely true that it is available, but only available with effort and only available on average. No guarantees that you can get it for any specific person.
Thus I'm not willing to increase the chances to 100 percent that my info is obtainable.
Good luck with that.
Okay, but you seem to be under the impression that you were contradicting the root comment. Being able to lower information exposure doesn't do so.
1. Well, if the information is already out there, the damage is already done. Nothing to be done about it, so you *should* stop fighting.
I'm not confident that the information is or is not already out there, but what we have here is a disagreement on an empirical fact. Some one just has to figure it out.
2. I do not think the root comment made this claim.
SSNs are already easy to guess. Probably don't even need to post it.
https://www.science.org/content/article/social-security-numbers-are-easy-guess
I’m now officially grateful my mom lost my card when I was a kid and she had to get me a new SSN.