I once did an application of Benford's Law to USDT transactions between crypto exchanges, which seemed to indicate some exchanges had mostly "organic" transactions and a handful of exchanges seemed to have heavy transaction volume of seemingly-random but not really random amounts, indicating some level of wash trading on those exchanges.
To be fair, the reason the CA laws are much more expansive on all uses of data is because companies have tried a number of arrangements to get around the definition of "sale". This was Sephora's defense back in the first CCPA case, that their data sharing relationship in exchange for targeted marketing services was not "selling" data:
I'll second this observation, as well as add that apart from AI slop most people around here associate the data center push with the sudden proliferation of Flock cameras at every major intersection and along every highway. Provo defeated a major data center project that was going into an empty industrial park, arguably the kind of place that would fit that sort of development. The actual cost-benefit calculation for most people is heavily weighted towards the negative and this should not continue to surprise people. The perceived downside with no upside is just going to get worse if the government gatekeeps the most useful models.
In 2017 LLMs weren't powerful enough to generate working code on their own, but my goal was to at least create a chatbot that could help you rubber-duck-debug your way to a solution. Unfortunately the tech wasn't quite strong enough for that, and not enough engineers even knew what rubber-duck-debugging was. RIP Duckly.
Trying to train an LLM on two 1080ti's on the StackOverflow corpus in my living room was a vibe though. Good times.
Duckly deserved to actually work. There’s a small irony here: the closest study I found to this, robots specifically built to simulate attentive listening, found they performed no better than an actual inanimate rubber duck for adult engineers. The mechanical signal of listening doesn’t seem to be the active ingredient. Makes me wonder if Duckly would have needed real disagreement to close a gap a duck can’t, not just better natural language.
You're probably on to something with the value of disagreement. I think it's one reason why chatting with current models doesn't create the same stimulation as rubber-ducking used to bring. The models are typically too quick to agree and amplify what you think rather than truly break it down and push back.
And thanks for saying it should have worked, I agree. My chagrin has increased over the years as I have realized the magnitude of my ill-timing.
Has anyone seen a good set of prompts for that disagreement? For the "skeptical eyebrow-raise" or "confused/doubtful head tilt" aspect of rubber ducks?
Agentic uses adversarial expert, steel-man opponent, risk-mitigation and failure-mode analysis. But what about almost brainstorming, but with thought-provoking nudge questions? Or on the other hand, arm-waving fight-club style discussion? Or... It's a big design space. I used to go to lots of research talks at MIT, in assorted departments. The post-talk Q&A question cultures varied a lot. Like encompassing both "leaves the speaker in tears", and "nudge so subtle, you won't quickly get it if you've not already spotted the fatal flaw in the work".
So aside from dialing down the "transformative insight!" silliness, there seems a rich multi-agent multi-persona space to explore.
I think agreement has value here too. An LLM that's starting to get a bit sycophantic will rephrase your ideas in a few different ways, and seeing the different presentations is helpful for reconsideration.
I wonder how much is actually needed to create an automated rubber duck. How well would ELIZA work? (https://en.wikipedia.org/wiki/ELIZA) (might need some adjustment to not talk like a therapist, but you get the idea)
2017 is a bit early to refer to them as LLMs. I'm not sure when exactly we started to refer to LMs as 'large', but I don't think it was before GPT2 (2019). That said, from the NLP work I've done, it was much more interesting working on small specialized models.
I built a half-baked CRM that has a lot of custom fields and visuals for statistics that are relevant to my potential customers. I'm selling primarily to registered data brokers, so being able to pull up their self-published compliance stats (gleaned from their own privacy pages or public filings) and contextualize them in terms of the rest of the industry ("your deletion request volume has been in the 95th percentile year over year") has been extremely helpful when starting conversations. I also gamified it a bit by giving myself targets for cold outreach and gathering hard numbers on my cadence for outbound calls and emails per lead.
I also built this site for educating potential customers and other privacy professionals about the increasing tempo of CCPA enforcement actions driving compliance: https://ccpa.world/enforcement
I could have probably coded this from scratch quicker considering that it took me two weeks to remove all of the hallucinated imaginary enforcement actions against real companies and also the citations to non-existent California law that the models kept injecting into my enforcement summaries.
More importantly, many companies will follow California rules even outside California. My car was built to California emissions spec at a time when very few states had stricter rules.
(The one major exception seems to be the "sell my data" opt-out and such privacy rules, that industry is sleazy enough that they'll go through extra trouble to keep screwing over non-CA residents.)
Well, CT and VT passed their own version of the California DROP system last week and there are 5 other states in play for the current 2026 legislative sessions. I think it will be a slow patchwork for more states to take similar action, but it is coming.
I will note that many "data brokers" will just honor non-California residents' requests as if they were California residents and subject to the CCPA, simply because they would rather remove a potentially litigious consumer from their databases. Given the relatively low potential revenue for a single consumer's data it just doesn't make sense to hold on to information for the kind of person who currently goes out of their way to make that kind of request.
At the same time, many data brokers do go out of their way to deny as many privacy requests as possible. Given that the CPPA/CalPrivacy is starting audits very soon I don't see this as a winning strategy for them in the long run.
Watching "The Price is Right" made California a mythical place for me as a child in the Midwest. All the cars being given away, they were sure to mention, followed "California emissions standards!"
The FTC settlement with GM allows GM to sell precise location as long as it's anonymized by attaching it to anonymous identifiers rather than personal info. It also allows non-precise location (e.g. zipcode/census-block) attached to identifying information.
Apparently no one at the FTC is smart enough to realize if Bob and anonid both move through the same sequence of approximate locations that the anonid is Bob. Or maybe they aren't that ignorant and just wanted to look like they were doing their job while protecting the surveillance status quo.
Selling anonymized precise location of a car that spends ~half the day at a residential location sure will make it impossible to de-anonymize that data.
The FTC under this administration that just doesn't care about people and only care about helping corporations.
The real impact of this case is that it's the first time we've seen a serious data minimization case in the US. California's investigation showed that they will prosecute if legitimately collected data is repurposed and resold after the fact.
My wife used to think that I had terrible sleep apnea because I'd repeatedly quit breathing for a minute or two at a time and then gasp for air, but it turned out I was just dreaming about freediving for lobsters.
reply