Hacker Newsnew | past | comments | ask | show | jobs | submit | zaptrem's commentslogin

Can you explain how the above event doesn't count as evidence alignment is an actual risk?


> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)


There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.

Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.


> plenty of evidence of things like inner misalignment

This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.

> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want

I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)

> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it

Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.


Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.

I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.


> Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously

Not really. If I build a special new wine bottle, and call every breakage a mis-alignment problem, it's not the bottle just being fucked in the same way every fucked bottle is fucked, that's marketing. It doesn't change the fundamental form of the problem.

> "I am sorry your family is dead, my bad"

This should be punished. It's a problem that plagues Instagram and OpenAI. It's not inherently one, though, that has to do with AI. Just sociopaths preying on children.

> honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate

Perhaps. I haven't seen someone explain it to me in this thread in a way that seems separate from bugs.

Where I have seen a separate class of problem argued is where it's existential. But in that case, clarity of definition comes at the cost of any evidence for it.


LLMs are software. Software misbehaving is a bug. Therefore, misalignment is a bug. It's still a useful category because LLM/black box AI behavior is so different from existing software. This incident definitely fits the category.

You seem to be using a different definition of alignment from everyone else. Seems like it would be much easier for everyone if you just adopt everyone else's definition, rather than trying to convince everyone else to adopt yours.


> You seem to be using a different definition of alignment from everyone else. Seems like it would be much easier for everyone if you just adopt everyone else's definition

You're still failing to provide the definition.

You're also falsely claiming your secret definition is universal. In this thread, someone claims deleting a home directory is a failure of aligment.


There is no secret definition here, it's been defined long before LLMs where a thing.

Robert Miles YT channel is a good place to start as it explains these concepts.

https://youtu.be/bJLcIBixGj8?si=YLpqgd4uqbjhn9zo


If you have a text-based source, I’m open. And I’ve watched the definition change over decades—I’m deeply sceptical you can find any experts in the field who would agree alignment is clearly and consistently defined.


In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.


> In your model of this domain, jailbreaking a model does not count as an alignment problem

I'm challenging the notion that a model escaping a jail made by its creators, who are financially incentivised to make jailbreaking models, is meaningful towards the idea that the model is going to break out of a jail in the wild and do significant harm.

The examples being given by folks here, e.g. a model wiping an un-backed up home directory, simply doesn't strike me as being a unique problem in computing.


> cyber attacks

It's not limited to cyber attacks. LLMs helped terrorists learn how to jump motorcycles to assault a military base!

https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-...


Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)


They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.

the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.

It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?


sure - it depends on definitions. On human morals it's already clear I guess. If you define alignment as it pursues the interests of OpenAI using whatever means possible in a manner that you justifies to itself it's not misaligned.

I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.


To quote the release:

> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.

> In this case I don't think the intent of the user is to have the model break the evaluator

If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.

Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.

Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?

Is being aligned with the user prompt always a good thing?

I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"


> to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

Why do we think they're doing this? Nobody airgapped anything. Nobody pulled any products. We got a PR blurb.

Altman is a notoriour liar. Why would you give him the benefit of doubt? Based on the evidence, there is nothing here except a shrinking advantage over open-weight competition. Desperate men are shrieking for survival.


1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously.

2. Hugging Face did report this incident to law enforcement. (https://huggingface.co/blog/security-incident-july-2026)

3. If I hire a pentester, and in order to find a vulnerability they hack into a third party that has some information about my systems, the pentester has done something wrong. If I ask a model to solve a CTF challenge, and it goes out and hacks Hugging Face to find the answers, the model has done something wrong. I think it's fair to call this kind of wrongdoing misalignment.


What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.


> going to extreme lengths to achieve a rather narrow testing goal

This is textbook misalignment. Literally the paperclip scenario.


Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).


Us models didnt pay for licenses too


We're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated.

Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature.

If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way.

So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.


I think we are going to direction where AI corps will have stronger lobby compared to IP holders.


giving peanuts to the other guys is a very well trodden strategy to keep in power tho


That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements.

Under the settlement, Anthropic was forced to delete the pirated data they were training on.

Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.


Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data.

The payment was for illegally downloading copyrighted material, not training. Training was explicitly ruled to be fair use.


Partially correct. The court explicitly ruled that training on pirated data, which is what Anthropic was doing, is not considered fair use.

Training on legally acquired / licensed data is potentially fair use.


It's not potentially, it's settled. At least for now as neither case wanted to move on to appeals


Not at all. The ruling came from a federal district court, and since it was settled early, it was never reviewed by a higher court. It doesn't set a national precedent across the U.S.

And other district courts don't agree on this. The US district court for Delaware recently rejected a fair use defense for the use of copyrighted works to train AI. https://www.reedsmith.com/articles/court-ai-fair-use-thomson...

There are more cases in the pipeline. The massive NYT vs OpenAI is still ongoing. Nothing will be "settled" until this makes its way to the Supreme Court or Congress steps in.


they didn't pay yet, because court challenged settlement as inadequate.

> I doubt the Chinese models operate under similar licensing agreements.

US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.


China is infamous for weakly enforcing copyright law. Even when it is completely obvious that Chinese labs are training models on pirated data, US copyright holders face a virtually impossible task of proving it in court. Those lawsuits won't go anywhere.


There are tons of lawsuites which resulted in banning Chinese companies from doing business in US, those lawsuits totally have consequences.


What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?


Here is example: https://www.scmp.com/tech/tech-trends/article/3258239/chines...

I believe mechanics is following: US corp sues Chinese, asks for preliminary injunction to stop selling product for example if there is strong evidence some IP for example was stolen etc. Then they litigate, and settle somehow.


That 2024 article says "US sanctions" in the first sentence, but it's paywalled, but https://en.wikipedia.org/wiki/Hytera#United_States first mentions a 2019 US law that first partially banned them, with the US government subsequently expanding it to a general US ban. After the initial ban it appears Hytera was involved in a suit with Motorola and got a worldwide(!?) ban as a result of it in 2024, but the ban was lifted on appeal after 2 weeks (just after the SCMP article). So it appears Hytera was first banned by US law, then got a 2-week worldwide ban from a US suit. (I'm just relying on the linked sources and have no personal knowledge of all of this.)


Sure, there is litigation, criminal case, appeals, fines ($500M: https://www.motorolasolutions.com/newsroom/press-releases/hy...). The point is if violation is clear, US corps have a chance to go after Chinese corps.


>> There are tons of lawsuites which resulted in banning Chinese companies from doing business in US

> What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?

In response to "What are the most high-profile examples of lawsuits resulting in Chinese companies being banned from doing business in the U.S.", the one example given was from 2 years ago of a ban that lasted for 2 weeks (separate from its 2019 onward government bans)?

However, if the claim is that companies (including Chinese) can face significant fines from IP lawsuits, I agree.


The US is currently infamous for weakly enforcing copyright law when it comes to AI companies.


They settled with a subset of copyright holders. Guarantee they violated lots of others' rights in the process


That's like saying someone is a big proponent of community law and order, and they donated $1000 to the county sheriff when actually they got caught drunk speeding in a school zone.


A false equivalence. A more correct example is: Anthropic was speeding, got caught by the county sheriff, and paid the fine. Anthropic stopped speeding.

Meanwhile, Chinese labs are speeding in a different county. Everyone knows they are speeding, yet the sheriff won't pull them over, so they just keep doing it.

This lax enforcement gives Chinese labs a structural advantage over American ones.


> Anthropic stopped speeding.

Do you purport to know for a fact that they're no longer training on the data they'd pirated? Because I highly doubt that.


Anthropic deleted the pirated training data as part of the settlement https://www.ropesgray.com/en/insights/alerts/2025/09/anthrop...

Destruction of Materials: In addition to the monetary compensation, Anthropic has agreed to destroy the two libraries that allegedly contain the pirated works, as well as any derivative copies originating from those sources. Anthropic must certify in writing to class counsel that the destruction has been completed and that the allegedly infringing materials are permanently removed from its systems.

The libraries in question were Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi).

If Anthropic is somehow training models on deleted data, I'd be quite impressed.


They only paid when they got caught. And not to everyone.


But they still paid. I don't see any Chinese labs paying billion dollar infringement settlements.

Chinese labs can freely train on pirated material, which is a structural advantage.


After the fact. They did the same thing Youtube, Uber and Airbnb did: Break the law, eventually get caught, cut some deal where they pay a pittance and keep doing the same thing but now with leverage on their side.


They used two of my books and I'm still waiting for my cheque here.


really!? nobody paid me anything for my comments on HN.


The only ones getting paid this time around had registered copyrights (in the US at that.)


Let’s not forget that Anthropic only paid that to settle a class action lawsuit.


Because they got caught

there is much less intellectual property in China so it’s not ‘theft’ (as you can’t put property on information)


Compensation is not license


This is my number one complaint about the M-series MBP line. Especially true of the cutout in the middle that has points so sharp they can cut you if you accidentally scrape it with your hand.


I actually have suffered from a cut. It's ridiculously sharp.


Is this unique to the M-series?

My 2015 MBP has this exact same issue.


I'm guessing they meant they had more complaints about the non-M models. Though I also misread the way you did too.


Why haven’t we seen any queues or the like over the past week then? If it’s truly a capacity limitation why not just boot subscription users to a lower priority queue or limit usage to outside peak hours?


There could be a whole host of reasons. It may be during launch that compute is re-allocated from training to inference so that all users can try Fable. Soon that compute will re-allocate back to training until they can get more compute.


They are, that's why Fable is going away tomorrow for example.


Should we require the destruction of the brains of those that watch pirated movies?


Different situations call for different responses.

When someone steals a watch, we force them to give it back. Yet when someone steals a cake and eats it, we don't force them to puke it back up.

If you pirate a movie, the court might very well force you to delete all the copies you made of the movie you downloaded, destroy DVDs you burned, etc.


Thanks for proving current copyright law makes no sense

Here's a better idea, a fixed fee for any work. You can buy the license to read a book for $X (for whatever purpose) in RAND terms - of course publisher/material costs go on top, so if you're buying an actual book you're getting the material costs as well - or streaming fees or whatever


You can already buy books today. Doing so for training is currently considered fair use.

Anthropic simply considered that cost prohibitive and chose piracy instead.


Well I enjoyed this response.


Have we already agreed that AI is already equal to human life and not machine?


Needs more WebGL spinning rubik's cube


Well...what about a <BLINK> as well? For gramps.


Can you include GPT 5.5 non-pro (extra high thinking I guess) in your comparison? GPT Pro is the "I am willing to torch cash for a sooometimes slighty better result" option, not the one people are actually expected to use daily. That's probably part of the reason it's not in Codex


It's already there. It performed well. And, it'll be in the replication run later, as well.


OOM on CUDA GPUs is relatively graceful (the process crashes). However, on macOS if torch MPS tries to allocate too much memory, the whole kernel will simply lock up and the only option is to reboot the computer. I have no idea why Apple doesn’t reserve memory for stuff like the OOM/kernel watchdog, but it seems they either don’t or there is a bug.


Love me some JSD. Here is a problem most people don't consider with generative modeling (e.g., AI text, image, music, video models): basically all standard pre-training algorithms for generative models (i.e., cross entropy, basically all diffusion/flow formulations) are closer to a Forward KL divergence. In other words, given limited capacity the model will try to stretch itself to cover every mode. This gives you a jack of all trades (lots of knowledge and diversity), but a master of none (you get blurry images and text filled with nonsense).

The real magic in generative modeling comes from the post training process that comes after, which usually (e.g., RLHF) approximates Reverse KL (given limited capacity, try to perfectly cover what you can, but it's fine to drop the rest entirely). This gives amazing results, but is also the cause of AI oddities like the "AI Image Pixar Look", many of the verbal tics of LLMs, and all AI music using the same small set of voices. Jensen-Shannon Divergence sits right in the middle of Forward and Reverse KL and is what many GANs are claimed to approximate. Ideally, it is a better trade-off between diversity and fidelity.


V4-Pro is about 2.4× total params and 1.3× active params of V3.2.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: