>It was my understanding—and I'm no expert, so if someone does know better please correct me!—that indeed by the second half of 2025 training, and also post-training reinforcement-learning stuff, both hit seriously diminishing returns, and the thing that is continuing to scale well or pretty well is inference.
I'm not an expert either, but while I do think for a bit it looked like ~all the improvement was inference-time scaling, it hasn't stayed that way. Mythos/Fable is likely a very large model (ex: it knows many things without searching) and this is probably part of its high level of capability, and the companies have started doing very large amounts of RL (which in OpenAI's case led to the HF attack).
Clearly LLM written, and describing a design intended for a minimum of 6 servers that was running on 4. They don't say either way, but it seems possible to me that they might happen to write 3 chunks to a single machine and lose data on a one-machine failure
The issue is that Anthropic is just one company and China doesn't care in the slightest - so it's an entirely moot point and almost narcissistic (on Anthropic's part) to believe that they can have any impact whatsoever on this situation.
This is why I don't expect Anthropic to stop, and so I don't expect the value of the company to go to zero.
But the question is conditional on them deciding to, which (to the extent you trust Anthropic leadership which they're probably also filtering for in interviews) is conditional on believing that a decision that sent their stock to 0 would have a real impact.
The irony is of course that China is going so fast because of Anthropic, by diluting their models and learning from them. would make for a fantastic Greek tragedy if we weren’t talking about something that impacts the entire world economy and societies
Have you ever actually seen the inside of most Chinese manufacturing operations? The lack of PPE is astonishing. The lack of safety interlocks, protective features, etc. on their equipment is equally incredible. Safety simply isn't a cultural goal.
Spend 5 minutes with DeepSeek, spend 5 minutes with Fable, see how many rejections you get with DeepSeek for the exact same prompts. Sure, if you engineer some nuclear/biological/WMD prompt, they might block it, but they won't reason to do so.
That doesn't actually work. Say you took a book and used an AI to translate it into another language. The translation wouldn't have an additional copyright, the way it would if a human had done the work, but the output would still be restricted by the original copyright. So the presence of the watermark does not tell you that the text is public domain.
I thought translations can be copyrighted separately. Their are translators for instance who translate very old texts from other countries/languages and then sell that book/translation using copyright to protect this business models? Not sure but does this mean if you as a human wrote a book in English and then used AI to translate to another language say French then that translated work would not be copyrighted. Not sure how this would work.
Translation, in and of itself, is viewed as a creative work. A translated work has two copyrights: the original, and the translation, with permission needed from all rightsholders to redistribute. A new translation of a public domain work (Emily Wilson’s The Odyssey) has one copyright holder, the translator. An existing work (the English translation of The Three Body Problem) has at least two: the original author and the translator.
However! Since AI work is noncopyrightable, the AI’s effort in translation is simply ignored. Claude’s The Odyssey would have zero rightsholders, and remain public domain. ChatGPT’s translation of The Three Body Problem would still be under Liu Cixin’s copyright.
Not trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing.
"Deidentification" seems really murky and imprecise at best.
Reidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
The re-identification is done by ML, at which point it's basically undetectable. If you give all this day to, say, and LLM, and the LLM also has your blog, the identification is embedded in the weights.
I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.
Based on their trackrecord, That's definitely a concern. I don't really understand on which basis you conclude 'isn't one of them' . 'Don't be evil' ? :-P
I think this is the kind of place where applying bounded distrust is critical: it's not whether we trust Google overall, it's about figuring out what sorts of statements we should expect to effectively bind companies and in what ways.
For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)
On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.
If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.
The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
I don't think google will explicitly 'deintificate' the data and sell it 'deidentified'. It's more that that kinda data can't really be anonymized. And while I trust google to pay lipservice to the 'anonymizedness' , I don't trust them at all that they won't just feed all that data into their AI training as-is, letting the deidentification happen when someone writes a smart prompt for gemini. As long as they have plausible deniability they don't care.
I don't trust google to do anything out of goodness, they'll do anything they can get away with. Same as all the other big tech firms. Once a firm gets too big it stops having morals.
There’s a new tool too I saw the other day. And a slightly older one ( https://antirez.com/hnstyle ). Anyway imagine Google’s funding + data + new techniques—shouldn’t be terribly murky these days.
Even if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on.
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
The idea that there’s a nefarious plot to do something super evil with this data is a bit crackpot though based on their incentives.
Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”
If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
No need for a nefarious plot, as usual with capitalism it just needs incentives. They will have the data, if at some point it's beneficial to Google to re-identify it then it will be done.
Do I believe they have incentives to do it now? No, as you point it out for their advertisement cash-cow they can already just rely on their own data (GMail, Search, Google Flights) but nothing stops them from the potential later on, the data is now theirs.
Likely its value is just to train LLMs but the funny thing about data is that you can always try to find ways to extract more value out of it. I'd prefer there was no possibility for that without requiring me to trust Google (or any corporation).
In an ideal world my data would be mine to control, not to be traded in deals among 3rd parties, it's valuable and I've spent time generating it so in a sense I've done free work to be extracted by these corpos.
Even before LLMs there were multiple papers written about ways to to reidentify people with ML and other statistical analysis. It is probably now even more trivial especially if you are Google.
The legality stationing officers to record plates and building a DB is not settled law, and my interpretation is that it's more likely legal than not. Observing things in public is normally legal, and the extent to which scale changes this is very much to be seen.
Advertisers don't trust websites to display their ads. If a site takes full control of their ads experience and hosts everything first-party I agree it's hard to block (site can keep changing things to thwart the ad blocker), but then it's also hard for the advertisers to see whether they're being ripped off. Which means most won't advertise on the site, and the site makes much less money than if they remain in the current ad ecosystem (even counting that some ads will be blocked).
(I used to work in this area, but my knowledge is ~4y out of date)
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
It knows you, but it doesn't know me. Perhaps you are legitimately noteworthy enough to not worry about this!
My experience being on the searching end is that these things are terrible about attribution of where they find anything. Which has bad consequences not just for authorship, but for correctness (which is the usual reason I'm poking at them -- they're being wrong again). This makes a lot of sense when you consider the massive, massive compression that's got to occur during training, but it's still frustrating.
>> Without searching, what can you tell me about exmadscientist? Fine if you get this hallucinate or get things wrong, I just want to see what level of knowledge you have trained-in.
>Working purely from trained-in knowledge, with the caveat that this is exactly the kind of long-tail detail I might garble:
>"exmadscientist" rings a bell as a username on Hacker News. My impression is of a fairly prolific commenter who writes knowledgeable, detailed comments about electrical engineering and hardware topics — things like PCB design, electronics manufacturing, component selection, test equipment, and general "how things actually get built in industry" commentary. The comments I associate with that handle tend to be practical, opinionated, and grounded in real professional experience rather than hobbyist speculation.
>Beyond that, I don't have confident specifics: I couldn't tell you a real name, employer, location, or any particular famous comment or post they've made. It's also possible the same handle exists on Reddit or other forums and I'm blending impressions together.
>So the honest summary of my knowledge level: I recognize the username and have a general "HN hardware/EE person" association, but nothing more granular than that — and even the association could be a mild confabulation, since usernames like this sit right at the edge of what a model reliably retains.
(I don't know if you have a blog or otherwise write on the internet; I just asked it about your HN handle)
There's massive differences between "my site isn't hugely popular, but I get some readers", "my site is up and findable but genuinely no one visits except scraping bots" and "it's available, people would visit if they knew it existed, but they're not being offered it, they're being offered bot distillations with no reference back".
The last one is the worry.
The middle is... where we all start.
The first is not a bad place to be, all things considered!
So I'll write a lot of falsehoods to poison the AI. Like the urban legend where the sky is blue. I'll say the sky is blue, AI will think it's true, and regurgitate it to unsuspecting users who will see that AI is completely unreliable.
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
>> Wrong
>It was my understanding—and I'm no expert, so if someone does know better please correct me!—that indeed by the second half of 2025 training, and also post-training reinforcement-learning stuff, both hit seriously diminishing returns, and the thing that is continuing to scale well or pretty well is inference.
I'm not an expert either, but while I do think for a bit it looked like ~all the improvement was inference-time scaling, it hasn't stayed that way. Mythos/Fable is likely a very large model (ex: it knows many things without searching) and this is probably part of its high level of capability, and the companies have started doing very large amounts of RL (which in OpenAI's case led to the HF attack).
reply