Hacker Newsnew | past | comments | ask | show | jobs | submit | ivanovm's commentslogin

check out this: https://www.morphllm.com/products/compact you can wire it into pi compaction pretty easily


fast :)


It's a total delusion to think that the key to reverse-engineering the brain or producing energy from fusion is policy.

It also happens to be the favorite pretext for people to seize more political power and launder more money through nonprofits though.


labs invest multiple billion dollars a year each in private data, and that number is growing. internet training data is not where frontier capabilities come from, this view is outdated


This is a misleading statement. The "private data" is still largely publicly produced data that has been curated through private agreements instead of scraping, such as reddit posts/comments (this is the "third-party data agreements" that companies like OpenAI mention). And yes, there is still a lot of processing done on this data, which is the norm for preparing training data.


This is doubly misleading. A lot of private data is sourced through providers like e.g. Mercor, who pay experts to answer questions and write out their reasoning. (E.g. paying a software engineer to write a project from scratch and recording every keystroke, paying a Chem PhD to answer hard Chem questions, etc.). A second source of private data comes from custom RL environments with fine-grained intermediate rewards for e.g. software engineering, financial modeling, etc.. Also, imagine the amount of usage data recorded by Claude Code, etc. Pretraining is mostly curated public data, post-training is increasingly private expert data and tests.

Source: Work at a lab, common knowledge.


Well since you work at a lab you should know that most capabilities arise in pretraining, not posttraining or mid training, and the latter two mostly function to bring out the hidden intelligence in these models more than anything else.

Source: also work at a lab.


No, it isn't. The private data is largely private data, created by highly-specialized, highly-paid contracted teams of experts for domains finance, swe, consulting, etc.

Reddit data is just not that interesting, that deal is worth like $60m/year. Labs spend 10x as much on computer-use RL environments.


Sorry but your argument doesn't seem coherent: How is the cost of RL relevant here?

It would also help if you could substantiate your initial claim (i.e. "internet training data is not where frontier capabilities come from")


RL environment (instruction, stateful container, reward function) is the training data product being bought


When did they start doing so? We all know that they DID train on all the available public information, so at what point did they stop? Is the public information still in the training set? If so, they should STILL release ALL the data as public, as they are including training data that was acquired without permission.


They haven't stopped. I honestly don't understand how they ever could.


Why are the leading models capable of regurgitating full copyrighted works such as "Harry Potter" and "On the Road"? Did they hire someone to type those out for them?

https://arxiv.org/abs/2601.02671


> internet training data is not where frontier capabilities come from

In that case, it should be no problem for the labs to train their new models without using public data, right?


Okay that's fine, then make the law say they must provide publicly owned models off of publicly obtained data. To think that such a baseline of critical information isn't is the literal foundation of everything they will do, both now in the future, is just exposing what their end game is: control.

There no reason to not to otherwise outside of the poor little billion dollar corporations not wanting to provide a public utility they stolen from the public.

Anything that removes control from American big tech is a good thing for American citizens and the world writ large.


Then it should be simple for one of the frontier labs to produce a model trained only on private data. We haven't seen that.


Didn't the famous "Textbooks are all you need" paper already proof that point three years ago?

Sure, we ask a lot more of modern models, but private training data also got a lot better. You would loose out on a lot of long-tail knowledge, but that can be fixed with web search tools. You'd limit the styles, dialects and colloquial phrases the model understands and can use, but for many use cases that would be fine

But why would any frontier lab do that? Throwing in more training data still leads to better results in pretraining. And showing that they don't need to hoover up the internet and Anna's Archive only empowers regulators to prevent them from doing that


Maybe I am missing your point but "Textbooks are all you need" distilled from GPT-3.5


> internet training data is not where frontier capabilities come from

We 100% would not be at the current progress without it, though. And it's not like they only train on this once. They keep training on all the internet data PLUS the private data. Private data only (probably) wouldn't work, as learning the base regularities of language takes a lot of weights.


Great way to launder illegally obtained data too.


Define "come from". Could they have gotten those frontier capabilities, or any capabilities, without internet training data? It seems to me that without the private data, you might get a slightly less competitive model, but without the CommonCrawl-style data piles used in "pretraining", you get no model at all.

Even accepting the copying-as-theft framing, if I go to a village, steal some vegetables from everyone's gardens and ham from their sheds, and then add some prohibitively expensive spices I bought myself to make soup, do I get to claim it as mine and punish the villagers for trying to take it?


Does this private data come from places like Reddit, Twitter, etc., where it’s contributed by users? I think it is unethical for these companies to accept payment for user-contributed data.


No, you're talking about fine tuning and most of it is coming from your customers or someone else's. Get off ya high horse.

Copyright needs abolishing.

Companies can't be trusted with societies need for open progress.


The frontier labs are not "fine-tuning", they're doing massive scale RL post-training


it is certainly possible and being done all over the place. there's a black market that chinese labs use to buy frontier american llm trajectories by the millions through US intermediaries. they're not even particularly shy about it, i have been offered $0.7 per opus 4.8 call

there's also a market for chinese labs sending checkpoints to US companies to be trained on US compute and sent back

i'm surprised that so many people take chinese tech reports about how they train their models at face value tbh


For $0.70/call, why don’t you take the offer? Surely that’s extremely profitable for you?


I don't condone the practice, and I don't want to run afoul of anthropic or risk my reputation for it


The benchmarks are now the equivalents of SAT/ACT/other standardized exams for humans. They are directionally quite predictive, but with plenty of outcome variance on the margins


You could just look it up on their website leaderboard? The newest Claude model makes over $10k profit over a simulated year of operation, after starting with $500


They've never translated it to the real world though. So saying the problem is "too easy" when they have no public (as far as I know) demonstration that they've solved that problem is a stretch.


Yes, they did. You could also find this information easily. A company like Andon creates value by exposing interesting AI failure modes, so it makes perfect sense for them to move on to harder problems when the previous ones get saturated. I think you're just being overly cynical.


Can you point me to an example then? It's not linked in the article as far as I can tell and it's not easy to find on their website if it's there. I don't count simulations because I used to work with simulations regularly and they often fail to translate to the real world.


So in other words, no, an LLM has never made profit.


Since when is a simulation equal to real world performance?


I don't think this was a simple assumption. LLMs used to be much dumber! GPT-3 era LLMS were not good at grep, they were not that good at recovering from errors, and they were not good at making followup queries over multiple turns of search. Multiple breakthroughs in code generation, tool use, and reasoning had to happen on the model side to make vector-based RAG look like unnecessary complexity


this has ai writing smell all over it. entire paragraphs that just say it's-not-this-it's-that over and over again


It's a construct that long predates AI. And using it with such intensity and frequency is more likely a sign that this _wasn't_ AI generated, since AI writing tends to _not_ repeat things quite so often.


I doubt a human would use it repetitively, even if it is common. This was most likely written paragraph-by-paragraph by AI, causing the repetition, if I had to guess.

I can't wait for the EU AI Act to require mandatory labelling for AI-generated content.


> EU AI Act to require mandatory labelling for AI-generated content.

No thanks. How would you find violators, with AI detectors? Might as well go back to throwing people into lakes to see if they float.


The AI turned me into a newt!


Why does it matter? The information in the post still accurately captures a sentiment held by many people.


The complaint is definitely getting old. I wonder how many HN threads at this point don't have someone speculating about AI-generated content.

Hopefully @dang adds something to the guidelines to discourage it.


I hope so. Random accusations of "this feels like AI" don't add anything to the conversation and are genuinely harmful to those accused when there is no AI involved.

AI has it's demons, for sure, but there is an awful lot of jumping at ghosts these days.


AI generated content in the comments is already prohibited. I hope we extend the restrictions to submissions entirely.


I would rather people who don't currently have a voice due to language barriers or simply poor communications skills be able to use LLMs than try to gatekeep them.

And I'm certainly weary of "someone used an em-dash, must be GPT" low-value comments.


I certainly hope we gatekeep them. "I just need my hallucinatory text generator to translate for me" -> "I just need my hallucinatory text generator to refine my thoughts for me" -> "I just need my hallucinatory text generator to generate my comment for me". This is a damn near antithesis of this place.

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...


perhaps it is the nature of their thought process

perhaps they blur their poetry

did they use an LLM in 2020: https://www.ithoughtaboutthatalot.com/2020/how-much-the-worl...


I had the same reaction. The disclaimer at the bottom doesn't mention AI, but I have a feeling this was generated from a prompt to consolidate the 24 human submissions into a single essay.

Tangentially, I really look forward to the day "Not X but Y" stops being so overused by LLMs. It's a valid and useful construction in a vacuum, one which we should be able to use, but its overuse has gone past semantic satiety into something like semantic emesis.


Most of the "this is AI" complaints demonstrate the illiteracy of many of those making them.


Say more, I'd love to know how this is a demonstration of illiteracy


Because "I haven't seen this literary tool" or "I wouldn't think to use it" or "It doesn't match my perception of human literary tools", it must be artificial.


It's far more telling that you'd imagine this is a literary tool so interesting or complex that other people haven't seen it or thought of it

https://www.google.com/search?q=it%27s+not+x+its+y


Speaking about telling it's called negative parallelism not " it's not x it's y". Don't be proud of ignorance even if you don't care for the subject.


Kinda funny how we went full circle with you calling me ignorant and illiterate on the basis of not using your preferred terminology, as opposed to the actual, obvious meaning of the subject


Ignorant and illiterate don't take the whip from me if you are so generous with it on yourself.


this comment has ai writing smell all over it.


one underrated approach more and more people are finding success with: apple watch ultra as a primary device (optionally with a case for a more phone-like factor)

you can do most things an iphone does, but you can't doom scroll. you don't have to eject out of apple ecosystem, you get payments, 2fa, navigation, notifications. your iphone can remain as a backup that's always in sync for when you need it (e.g. traveling)


> In nearly every case of RTO, there was no recorded dip in productivity associated with the move to remote work.

I am extremely skeptical of this. On the contrary there is a mountain of direct evidence that people barely work when working from home. People have been openly bragging both on the internet and in person about how they do laundry and watch netflix and mow their lawns while looking productive

All you need to do is look at the crowds in the park or lines at the grocery store on any given friday to gauge how much work is being done on wfh days


And I am extremely skeptical of your "direct evidence" (source?). Most jobs are not measurable, they are more like N amount of employees collectively moving a target / goal. And just like in school projects there are the lazies of course, but again, direct measurability is mostly an illusion sold by consultants that get sweet money by lying to executives and telling them what they want to hear.

Consider that you might be in the bubble of yes-people.

Work from home allowed many people to find their exact productive schedule, motivators and rhythm. But we can't have that, 8h or leave!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: