Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You're a step too far with your assumptions. I'm a Googler. I've noped away when Pebble sold to Fitbit. But seeing from the inside how Google treats user data, I'd be 100% ok with the data going to Google instead.


I just can't imagine how you would think that others will just ignore Google's history of violating it's privacy promises.

>In today's Big Tech antitrust hearing in front of a Congressional subcommittee, representative Val Demings questioned Google CEO Sundar Pichai about the company's merger with DoubleClick.

Specifically, the way Google combined data from the advertising company -- bought in 2007 -- with Google's own data. Founder Sergey Brin had told Congress it would not combine the personal information, but the company quietly did so in 2016 anyway.

https://www.engadget.com/google-antitrust-hearing-doubleclic...

>Google’s privacy policies as of March 1, 2012 established that no combination between DoubleClick’s advertising data and Google’s personally-identifiable information would take place without the prior consent of its users, but an updated version of those policies subtly allowed the integration of both databases regardless of prior consent of its users

https://www.promarket.org/2020/08/21/why-we-should-be-carefu...


I think the parent is implying Google is better than Fitbit, rather than Google doing everything correctly. If the grandparent’s comment is correct, then I have no trouble believing Google will be an improvement.


With Google, the more likely problem is that you get algorithmically banned with no recourse, and your Fitbit just stops working.


> But seeing from the inside how Google treats user data, I'd be 100% ok with the data going to Google instead.

how does this make the following OK?

> No Pebble user ever consented to their health data being sold to Google, but all of this is legal.


I extrapolated this into the hypothetical "if asked, users would not consent to this sale". I would consent to it though.

Entirely possible I misread the spirit behind the letter of what GP wrote.


Google is a large company. You know how your team treats user data, but it would be difficult to have visibility into what all teams are doing. It only takes one or two to ruin user trust for years.


Sort of? Unusually for a company of this size there's a ton of standardization and cross-company consistency. In this case, I believe the whole company (aside from very recent acquisitions still being migrated) used the same logs access framework.

I think this is part of why Google has earned itself a reputation for killing things: if you require everything to use the same set of internal systems, and can't just leave old projects limping along on a fork of a deprecated system, the cost of keeping old things around is higher.

(I worked at Google until June)


I cannot speak for my employer, but I can speak in general terms that I agree that companies of this size have processes to try to avoid this kind of thing from happening. This is true at Google scale but also true once your company passes a certain size+age, although of course it can continue to mature.

The thing is that these processes don't actually work, because in the end someone is going to do something dumb anyways and you really can't stop them. The same way your processes aren't going to stop an outage, someone is going to correlate data they shouldn't and somehow lack the tact to go "hey maybe I should not be doing this". Often various pressures will make it seem like something they should do. And when they do it they will cost the company an incalculable amount of user trust, to say nothing of the irreversibility of a privacy incident. Someone will do a writeup of it after it is discovered by the press but at that point the damage is done.

(FWIW, I'm not even including the stuff that goes on with executive approval, where they make decisions on how likely they think they are going to be discovered and sued for large amounts of money. Of course this would bypass any policy…but it's usually kept under wraps for obvious reasons.)


Paradoxically, trust-ruining incidents is what makes Google trustworthy. There's a strong postmortem culture. "We messed up, what systems and processes are now needed to never mess up like this again?" I hear that's also reinforced by how legal achieves settlements. And usually it's one team messing up, but all of Google being in scope of the remediation.

This has the downside that the number of organisations that need to approve any launch ends up being borderline overwhelming. This is a big part of why it takes so long, if at all, to get Google products outside the US. Another implication is needing the expensive security org and culture.


I have yet to see a postmortem that includes things like "there are insufficient checks on how a team grows metrics".


Are we talking about the same Google who empowered blanket keyword search warrants?

You could argue they had no choice, but I could argue that they had the choice to allow search history to be end to end encrypted and/or anonymous. That is if their business model did not rely on selling plaintext PII.

https://www.nbcnews.com/news/us-news/police-google-reverse-k...


Yes, because Google’s advertising based revenue model provides the right incentive scheme to discourage use of personal data to increase revenue and profits?


Google signed an agreement with the EU that said we wouldn’t. Bad things would happen if they did.

As a Googler that has access to some Fitbit stuff, I have to sign a form every 30 days that indicates I understand this.


That EU agreement is only for 10 years, starting from 2019. So it’s got 7 years before expiring?


But will Google respect this agreement? If the expected fines are lower the the value extracted from the data or if they can drag the eventual lawsuits for a long time, I'm pretty sure they will just exploit that information anyway.


A whole lot of the world uses google but isn't under EU law (I'm in the USA and EU laws don't mean anything here) so how would that protect the rest of us?


Can you expand a bit on the quality of Google's internal practices with user data? eg what would an engineer on Google Photos need to do to be able to access my photos, if possible at all


In general, SWE/SRE do not have access to user data. SWEs just have no prod access at all, and SREs generally only have the ability to run signed binaries with code that has been code reviewed and submitted, though there are obviously break glass features with auditing. Though I am not super familiar with them.

ML is where things tend to be a little bit more grey overall since being able to look at data is very useful for development, so some things are scrubbed for PII, but then accessible in some form. But for things like GMail or Photos, I would assume nobody (including ML engineers) can read your data as these are basically impossible to sanitize.

Some products have systems train ML models without engineers seeing the data, e.g. spam filtering, even when the underlying data is considered sensitive.


Basically all the data is available with no oversight if you ask permission and have some allusion to a relevant $JOB reason to need it.

The fairly recent case of people's private conversations being shipped out to basically unvetted contractors for labeling and analysis (and subsequently leaked) should serve as sufficient evidence that "shit happens," and if private conversations without having even initiated an interaction with your Google devices are being tossed around and leaked, forgive me if I don't believe that when producing the tagging, timeline, and album features in Google Photos there wasn't some underpaid, unwatched contractor snooping through my photos without my permission.


I believe I can share an anecdote: a while ago I have uploaded a bunch of photos of a work event to the corporate account. This being corporate, it runs pre-release of everything. The gallery managed to hit some bug in the jillions of lines of Javascript, which I never cared to understand.

I reported the bug. Knowing how security works technically, I added to the bug the words "I'm happy with whoever works on this to take a look at the gallery, here's a world-readable sharing link". A couple rounds of bug comments later, I have been asked to sign a legally binding consent form allowing an engineer to look at the gallery. Then somehow they decided I need to sign a different form to satisfy whatever other legal spirit needed appeasing. Only then someone finally looked at the bundle of photos. They figured out whatever was triggering the bug. They generated a gallery reproducing the bug with generic sample images. Whoever worked on the bug and adding a regression test worked off that synthetic gallery instead.


My concern—and I think what everyone's concern should be—is not that Google employees have access to my data. Or even that it's used to train ML models. It's that my data is sold to advertisers via a shady behind-the-scenes marketplace, and later used to profile me in order to show me content that manipulates me into spending money.

And that, moreover, I get none of the profits from these transactions, and have no control over whom it's sold to and under what terms.


Ok, I'll bite. How does Google, in your opinion, treat user data?


Data is divided into many different logs, with engineers having access only if it's needed for their role. Access can be at the per-log level or the per-column level: working on Ads JS I had access to raw frontend logs, but I didn't have access to the identities of the users sending those queries (if I'd needed that for a project I could have applied and had my need reviewed). Every few months you need your access renewed, so if you've switched roles you stop having access. What queries you run are logged, and I believe there was some sort of random auditing of those logs to make sure people were running queries that made sense, though I never interacted with that part.

(I left Google in June.)


So basically very primitive security compared to say how banks are handling client data in jurisdictions like Switzerland (one of my daily tasks). The fact that dev can have access to PROD logs just by himself unsupervised (even if for obviously good reasons from his/company point of view) is a big fat red flag.

That all mentioned here clearly means google doesn't have data privacy at the core of their priorities and this won't change unless forced by fines/regulation, just like banks. Slightly disappointed when reading this, but I guess I shouldn't have expected more.


> The fact that dev can have access to PROD logs just by himself unsupervised (even if for obviously good reasons from his/company point of view) is a big fat red flag

What's wrong with granting engineers access to PII-less logs[1] by default? How does that compromise privacy?

I'm willing to bet your bank differed on the following ways from Google in absolute numbers and per-engineer:

* handled significantly less requests per second - at least 2 orders of magnitude - therefore lower log volumes

* shipped less changes to prod per unit time, so fewer problems to investigate

1. No IP client address, raw session id or username


Likely better than 99.999999999% of the companies you interact with.


It doesn’t matter. When we talk about “having access to user data,” we’re talking about a humans being able to look at, I.e, the user logs of a former lover and doing something bad with it.

But that’s not the real danger of Google having this data. The danger is in ML. Google’s entire business model is predicated upon using information about you to change your behavior, and to sell on the market predictions of your future behavior.


> It doesn’t matter. When we talk about “having access to user data,” we’re talking about a humans being able to look at, i.e, the user logs of a former lover and doing something bad with it.

Google is one of the few companies known to have caught, fired, and officially publicly named an employee who did something like this:

https://techcrunch.com/2010/09/14/google-engineer-spying-fir...

Most tech companies about which people don't routinely raise this type of concern have far weaker security controls against (and detection systems for) this threat model than Google does.

That said:

> But that’s not the real danger of Google having this data. The danger is in ML. Google’s entire business model is predicated upon using information about you to change your behavior, and to sell on the market predictions of your future behavior.

I'm of two minds of this. I'm not thrilled about how much data Google combines and unnecessarily insists on collecting in order to allow things like the Google Assistant and Google Maps to provide full functionality. At the same time, many of Google's assistance and search services are better than their competitors exactly for this reason. I primarily wish they were more transparent in this area with fewer dark patterns and more user control, with forcing users to pick between excessive data sharing and inadequate access to Google services.

Disclosure: I have worked for Google in the past, but not since early 2015. I certainly am not speaking for them here.


> Google is one of the few companies known to have caught, fired, and officially publicly named an employee who did something like this

Not disagreeing with your broader point, but that specific employee was not caught by Google but was reported externally by parents of the minors after abusing access for months. I suspect the incident you linked predates - and was the impetus for - many controls that were subsequently added

Speaking speculatively: I've never worked for Google


Obvious typo corrections following the end of the edit window:

> I'm of two minds of this.

I'm of two minds on this.

> with forcing users to pick between

without forcing users to pick between


Google Maps and Google Assistant could work with that data and Google could also choose not to be surveillance capitalist robber barons.


Vacuum everything constantly and then horde it like Smaug, if Smaug was an unscrupulous salesman in a polyester suit?

But somehow this is ok because within Google the common employee doesn’t have access, but is rather incentivized to make the vacuum or selling more efficient?


Are you under the impression that the rest of FAANG doesn't do this? If so, I've got some big news for you...

When people talk about 'security' at Amazon/Microsoft/Google/Apple, they're describing audit performance and ISO compliance. From a business perspective, that's really the only thing that matters. Everything else is played pretty fast-and-loose. Both Apple and Microsoft have been caught backdooring their infrastructure for foreign governments, if anyone out there still has faith in these companies then I hope they stop taking life advice from VC heads.


> Are you under the impression that the rest of FAANG doesn't do this?

All data collection practices are not equal.

>Google Tracks 39 Types of Private Data, the Highest Among Big Tech Companies

Google takes the cake when it comes to tracking most of your data. This should not surprise, given that their entire business model relies on data.

Twitter and Facebook both save more information than they need to. However, with Facebook, most of the data they store is information users enter.

Apple is in a league above Amazon in protecting user privacy. It is the most privacy-conscious firm out there. Apple only stores the information that is necessary to maintain users’ accounts.

https://stockapps.com/blog/google-tracks-39-types-of-private...


You are giving Apple too much credit here. Apple can plaintext any iMessage if they wish with their signing keys for the servers that exchange user encryption keys. Why do you think the CCP demanded control of those servers for chinese residents? Just because Apple has a profit model other than adtech does not mean they are making choices that prioritize immutable privacy.

Apple could have folded, like even (old) Google did, and refused to comply with the CCP but shareholders would have been too upset seeing the price go down.

Successful publicly traded companies will almost always do the most profitable thing legally possible, no matter how evil.


Businesses must follow the laws of the jurisdictions they operate in. Technology cannot outrun the legal system for long.


I agree with your second paragraph, but I don't understand the point of the first? Are you saying that Google can be forgiven, since its competitors do the same?


They're saying everyone does it and it's NOT OK for all of them.


Disagreeing with the practices of one company doesn’t indicate support of companies with potentially similar practices.


Since you have the inside scoop does this mean that as a result your health data is hooked into your regular gmail/docs account and the info sold to whomever wants to buy it as well as used for advertising? Say I'm 40lbs overweight and 55 so I'm quite possibly have diabetes and would be a mark for diabetes medicine?


An easy one: Google doesn't sell your data.

For using it internally: health data is widely considered toxic. As in "I'm not touching that thing with a barge pole" toxic. I would personally be pretty surprised if Google ever started monetising it.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: