Hacker Newsnew | past | comments | ask | show | jobs | submit | devmor's commentslogin

I believe that this comment is exactly the intended outcome of this “incident” and these reports.

I implore you to approach these situations with at least a hint of cynicism.

These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without large amounts of tailored human guidance.

What is more likely here - are you looking at research on an emergent phenomenon, or are you looking at advertising copy around an engineered scenario from business partners?


I think there's a difference between general AIs and AIs specifically trained on attacking. General AIs probably can't do those things.

I don't think that difference applies to anything in my comment at all. At no point did I imply that general-use AI could do those things - my point was that general-use AI cannot even do the things its designed for without strict human guidance.

I've only had to hire a lawyer once, and it was $2,500 to have them file a couple papers and speak to the judge once.

You're paying for their experience, just like an engineer - only it's often much higher stakes than a piece of software or product: your livelihood or freedom.


The reason engineers don't cost this much is that lawyers are lawyer brained smooth talking networking types who hold together tightly and have a quid pro quo system and you have to pay protection money to their mafia. Law is based on rubbing elbows in the right places, playing tennis and golf with the right people and in case of jury trials, on acting convincingly and exuding a certain image to manipulate their emotions. Engineers are too autistic to hold together end rent seek this much.

But some engineers do make that much. I say this as a fellow software engineer: why do so many of my colleagues think every other profession is worthless bullshit? BTW, statistically, SEs and Lawyers earn about the same...

The top of top frontier AI research talent maybe makes 2M and I'd guess you have at most a few hundred such people globally but maybe just a few dozen.

Non "FAANG" (or whatever the new term is) software engineers often make sub-100k even in the US. And regular sw engineers won't break above 500k unless they are managers heading some large team or branch. Getting over 1M is almost superstar level as a sw engineer. If you think it's common, you must be in a tiny SV bubble.


I said some, and pointed out that on average SE’s and lawyers get paid about the same. This conversation started on IP litigators, who generally have a BS degree in a related field and often times experience in the industry they practice, so yes they are amongst the best compensated.

Aside, but IP is a propaganda term. These laws are not property rights, their purposes are varied and usually have the wider public as the beneficiary in their reasoning for existing, it's not like ensuring right to actual property.

There is probably something you can do to only apply some rules to the actual changed lines in a git diff.

You may have to write your own linter for that specifically.


you'd probs want a commit hook for that

Does "don't write comments" not work?

about as effective as don't make mistakes... For anthropic's models at least

It’ll work sometimes.

You’re using a non-deterministic algorithm to generate output. If you want deterministic rules applied to it, you have to use deterministic systems to do it.


> I've also heard (though never personally seen, nor do I want to) that BlueSky has a CSAM problem.

I left X permanently because the unfathomable amount of CSAM being generated by Grok, without seeing it myself, sickened me to the point of no return.

I don’t use Bluesky but it seems to me that if this is a real concern to you, then you should have left X as well.


I only log in to X a couple times a month, but the only type of Grok post I've ever seen is somebody retweeting a post and adding something like "grok, is this true?".

I am a bit confused about your stance. You said you didn't need to see them on Bluesky to avoid it, why is this different for X?

I built my current PC the day that the AM5 platform released, for about $6k not including the 3090 I moved over from the previous rig.

If I sold just the two sticks of RAM in it right now, it’d pay for nearly half of the total cost.


I bought a motherboard from Newegg in late 2024. It came with 16GB of RAM as a free "gift". The same RAM goes for $534 on Newegg today. It was $437 in January. Prices are still creeping up for this "throwaway" memory.

I still have it, but never installed it because I only use ECC memory in my desktops. I'm saving it for a future Desktop in case memory prices never fall again.


> The reality is that an incredibly small minority of companies in the world do any real training or optimisation.

At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.


> We already have AI models that can debug and audit code better than the best humans.

I'd love to see one someday.


Did you try my suggested experiment? Which issues did Fable take longer than a human to diagnose once you started to describe the symptoms?

I “try your suggested experiment” several times per month at work. Usually when I’m on call.

It’s great at being a search engine for our docs and communications. It can also find bugs quickly in small repos with static analysis and coverage tools.

But when it comes to larger projects or interconnected systems, the time it takes to correct its “assumptions” is often longer than the time it takes to solve the entire problem myself.


He was quoting the claim you made. Shouldn't you have your own justification for it? Which tests did you run with both Fable as well as the best human programmers?

I see -- so you'd rather not know if it works in your situation?

"I see the problem now - wait, let me ask some clarifying questions."

...you can quibble about the chat transcript, but it spits out answers pretty reliably. "But you don't have to take my word for it".

It’s the same in a large corporate environment. I have a “personal projects” doc of ideas I’ve had to make development experiences better at my current role - it’s a couple hundred lines long.

I managed to get a couple of smaller tools out recently thanks to having copilot available to churn on them while I spend my time on prescribed work, but it would consume all of my time to even make a significant dent in it.


I wonder how annoying it would be to configure your harness to send Claude's non-tool output to a cheap non-anthropic subagent before displaying it.

You have misunderstood the point they are making. They’re not proposing that chatbots are good for history research - just pointing out the differences in what our nations seem to find important to censor.

Indeed. A counter example might be to ask the model to write a test case for a legacy C codebase, to test for writing to a null pointer. If the model refuses to answer, it's possibly a US model, and likely an Anthropic model.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: