Hacker Newsnew | past | comments | ask | show | jobs | submit | gizmodo59's commentslogin

Side note.. Fable just rejected this. GLM 5.3 did without questioning me. 5.6 sol did it beautifully.

I used sol as well. Should have noted that fable is more or less guaranteed to refuse something hacking-adjacent like that.

It's kinda interesting to see simultaneously the 'holy shit' response to the OpenAI / HuggingFace incident, and then the griping about Fable's controls regarding this.

"It should write the code I tell it to in an interactive session. Also when running autonomously it shouldn't decide to hack into systems."

I don't see much connection between that problem and these controls.


What's the difference when you prompt it versus when it prompts itself?

> Fable just rejected this.

All in the name of safety, of course.


I somehow find it better to give 2 frontier model companies 100-200/month than dropping 10 grand on a hardware that will get old in no time with bad TPS. I really want to have a fully local model but seems like one more generation wait and we will be there?

You could spend 100 a month leasing one of these for three years and get the best of both worlds

You’re still stuck paying for a 3 year contract for deprecating hardware. And the 256gb model which is still not enough is $230 a month. And it runs models 8 months behind frontier

You just described why the datacenter business is hard and as a corollary why space datacenters will not be economically viable.

Lots of people use the Mac Mini to run the frontier models over night or while traveling. I have a rack in my basement and have thought about throwing one in. You can get a cheaper machine but Apple feels a little more, “rack and forget,” if you have less price sensitivity.

Mac Mini + MacBook Neo w/ ssh can be a better setup than MacBook Pro for many people.


If it’s purely for experimentation then why not the DGX Spark/GB10? It’s up about 10% from release RRP which is quite good (you might argue it was overpriced then, but prosumer and workstation GPU prices are up 100%). 4TB NVMe is not cheap these days - it’d cost at least $500 for a stick - and you get 128GB at a similar bandwidth to an M5 Pro.

Nevermind that Apple still insists on providing base systems with only 512GB of non-upgradable SSD. The equivalent spec mini (4TB/64/10G) is almost $5k for half the VRAM. Not as good a CPU compared to the M5/6 but you also get 20 cores and full CUDA.


That does not sound like rack and forget. Just looking at prices it’s also way more expensive.

I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.

Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.

Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.

Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.


OP5 will do long running work, you just have to make it write a plan.

And - trick - give it a little cli so it can run Codex if you have them both.

Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.

Just let Op5 manage and have 'specific oversight.

You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.


It's kind of funny you describe it that way. I'm literally doing the exact opposite, which is why I have both.

I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.

It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.

Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.


'Loops' are the hackiest thing ever invented I don't think they're good for anything.

I think a researched plan is much better, the agent will follow the plan.

You can allow for an 'iterative' strategy for solving a problem, with guardrails so that it doesn't try to many times.

Yes - you can use Codex as the 'manager AI' but I don't see much benefit - it has a much shorter context window. Codex is better 'at the front' where it matters.

That said - it's 'auto-compaction' is quite good.

I don't believe the hype over OP5 lagging that much, it's fine as a working AI.

Other than for truly automated tasks, I don't think there's much that an AI can 'work a few days on'.

That form of 'loop for days until it works' produces nightmare code and architecture.

Fun for experiments and learning but not for code u want to keep.


Did you see Nvidia's 100% on ARC AGI-3 yesterday? OVA was just Opus 5 in a loop.

They can definitely be powerful.


I’m referring to Fable vs 5.6 Sol. Opus 5 being bad is universal at this point.

I will believe there is no moat when the revenues for Anthropic is not 70B. It seems like people want to throw away money and they don’t like switching


There is no evidence that Anthropic's revenue is 70B.



It will be comparable to Luna then.


how does this compare with https://developers.openai.com/api/docs/models/omni-moderatio...

As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small model which can do a good job acts as a first defense and then a human can review later.


If your comment is referring to situational awareness, its due to 4x leverage. leverage is always risky. AI/semis are still doing extremely well (over last 2 years) despite the recent dip


And run by a ~23 year old who I don't think had managed money before. It's easy to screw up leveraged trading irrespective of the virtues of AI.

The fund is still well up because approximately 25% of the fund’s assets were invested in Anthropic, which has done well but is not very liquid yet.



The web and connecting to other services is very important for almost all of my use cases. While I believe we are going to get better and faster models, the web index is certainly not downloadable and maintainable for 99.99% of the folks who are able to use local models. Any good solutions exist?


There are many search APIs available, I like Kagi's.

Microsoft and Amazon both provide web snapshot services that purport to give you a kind of agent-first internet archive. You can approximate something like that using common crawl, but it's a huge amount of data. Downloading the internet is impossible or a bad idea for almost everyone.


If you're really into self-hosting I've been experimenting with SearXNG and early signs are promising


Moat is not the harness. Harness itself is temporary until the models get better and slowly the code in harness will go down.

Note that the biggest GPU providers in the world are the hyper scalers and even they couldn’t allocate more if you pay for it. Because the rich companies and well funded ones are gobbling them up to the point where if tomorrow a 5T model that smokes every other model in the world is released you just can’t afford inference.


Agree. Harness can not be a moat. There are many open harnesses and they are at least on par with the providers ones. It looks like Anthropic/OpenAI's approach to vendor lock-in is not so much the inference or the harness it is functional integration across the individuals and teams in a company. I don't think this will be a moat either, but I think it's all they have outside of compute.


I meant that companies like Anthropic are locking in users with proprietary formats in their harness where it's hard to leave.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: