It's kinda interesting to see simultaneously the 'holy shit' response to the OpenAI / HuggingFace incident, and then the griping about Fable's controls regarding this.
I somehow find it better to give 2 frontier model companies 100-200/month than dropping 10 grand on a hardware that will get old in no time with bad TPS. I really want to have a fully local model but seems like one more generation wait and we will be there?
You’re still stuck paying for a 3 year contract for deprecating hardware. And the 256gb model which is still not enough is $230 a month. And it runs models 8 months behind frontier
Lots of people use the Mac Mini to run the frontier models over night or while traveling. I have a rack in my basement and have thought about throwing one in. You can get a cheaper machine but Apple feels a little more, “rack and forget,” if you have less price sensitivity.
Mac Mini + MacBook Neo w/ ssh can be a better setup than MacBook Pro for many people.
If it’s purely for experimentation then why not the DGX Spark/GB10? It’s up about 10% from release RRP which is quite good (you might argue it was overpriced then, but prosumer and workstation GPU prices are up 100%). 4TB NVMe is not cheap these days - it’d cost at least $500 for a stick - and you get 128GB at a similar bandwidth to an M5 Pro.
Nevermind that Apple still insists on providing base systems with only 512GB of non-upgradable SSD. The equivalent spec mini (4TB/64/10G) is almost $5k for half the VRAM. Not as good a CPU compared to the M5/6 but you also get 20 cores and full CUDA.
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
It's kind of funny you describe it that way. I'm literally doing the exact opposite, which is why I have both.
I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.
It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.
Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.
'Loops' are the hackiest thing ever invented I don't think they're good for anything.
I think a researched plan is much better, the agent will follow the plan.
You can allow for an 'iterative' strategy for solving a problem, with guardrails so that it doesn't try to many times.
Yes - you can use Codex as the 'manager AI' but I don't see much benefit - it has a much shorter context window. Codex is better 'at the front' where it matters.
That said - it's 'auto-compaction' is quite good.
I don't believe the hype over OP5 lagging that much, it's fine as a working AI.
Other than for truly automated tasks, I don't think there's much that an AI can 'work a few days on'.
That form of 'loop for days until it works' produces nightmare code and architecture.
Fun for experiments and learning but not for code u want to keep.
I will believe there is no moat when the revenues for Anthropic is not 70B. It seems like people want to throw away money and they don’t like switching
As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small model which can do a good job acts as a first defense and then a human can review later.
If your comment is referring to situational awareness, its due to 4x leverage. leverage is always risky. AI/semis are still doing extremely well (over last 2 years) despite the recent dip
The web and connecting to other services is very important for almost all of my use cases. While I believe we are going to get better and faster models, the web index is certainly not downloadable and maintainable for 99.99% of the folks who are able to use local models. Any good solutions exist?
There are many search APIs available, I like Kagi's.
Microsoft and Amazon both provide web snapshot services that purport to give you a kind of agent-first internet archive. You can approximate something like that using common crawl, but it's a huge amount of data. Downloading the internet is impossible or a bad idea for almost everyone.
Moat is not the harness. Harness itself is temporary until the models get better and slowly the code in harness will go down.
Note that the biggest GPU providers in the world are the hyper scalers and even they couldn’t allocate more if you pay for it. Because the rich companies and well funded ones are gobbling them up to the point where if tomorrow a 5T model that smokes every other model in the world is released you just can’t afford inference.
Agree. Harness can not be a moat. There are many open harnesses and they are at least on par with the providers ones. It looks like Anthropic/OpenAI's approach to vendor lock-in is not so much the inference or the harness it is functional integration across the individuals and teams in a company. I don't think this will be a moat either, but I think it's all they have outside of compute.
reply