But that assumes this is a net improvement on linguistic efficiency rather than an artifact. Given that they tried to RL away from this style in 5.1 I'm not terribly bullish of Claudlish becoming something people try and learn. It being dense is less the issue than it being vacuous (as another commenter mentioned here). It's just very unclear and ambiguous writing. I think it has no place anywhere that needs language to be put to productive use.
I agree with you here, but to make it clear what I meant, I'll reiterate what I said in a sibling comment: I wasn't referring to the current state of Opus 5 (or even Fable 5.1) output. I was referring to possible future information density (vocabulary and sentence structure) that LLMs may evolve to use.
My guess is that there's probably a missing abstraction layer in the 3d editors somewhere. I am pretty bullish on AIs ability to solve this problem, even just based on existing image to 3d models like TRELLIS and HY3D. They work pretty darn well and take a lot of the burden out of not just modeling, but texturing and uv unwrapping.
That being said, I have a correlate here for music production and you could definitely make a lot of the same arguments about it, and they'd be true -- really high learning curve to make something halfway decent even assuming you have the taste part sorted. Some of it is the ability to just creatively think in another language once you become fluent in it.
I think AI can give you leverage when it does stuff you could actually do it manually yourself because you can manually evaluate it and correct it competently. The hard part is how would you evaluate work that is beyond your current competence level?
If that isn't easy, if it's incompressible relative to your current skill level, then that's the very human part that will always be a bottleneck. And that's fascinating to me because if there were better AI tools for music production when I picked it up a few decades ago, I don't know that I would have developed the same fundamentals.
What will fundamentals look like in a post-agentic world? If a lot of the manual gruntwork disappears or is more heavily abstracted, I believe that will matter a lot more.
First of all, it's so cool to see you on HN and creating this. I'm not sure how much the HN community is aware of this, but JUCE is the standard for cross platform music plugin development and needs to work efficiently in hard realtime settings. Many of my favorite plugins are JUCE based (such as the Valhalla stuff) and I'm a huge fan. Tracktion is a great DAW as well that I got a lot of mileage of in my younger days (ended up with Renoise because I'm a tracker guy after all long term).
I mention all this to say someone like you picking up the desire to build an agentic platform really piques my interest. Right now I am using Opencode for most of the stuff I am trying to do at $WORK and it does a good job on the whole at having sufficient functionality. But the release pace is blistering and it does feel bloated - both in terms of functionality as well as system prompts.
Moreover, I observed all of the same issues you mentioned and certainly wanted more of a tree like experience as well as a more UI forward experience. In order to get the functionality I wanted (good worktree support, sandboxing, etc) I eventually just had to let go of using opencode's UI and embrace the TUI because it was the only thing I could embed into a workflow that let me set up all of that in a sane manner. But problems still remain with the "doom scroll" experience when to your point clearly a tree based experience would be better.
I was ready to just settle with my cobbled together opencode flow and maybe migrate to Pi later on and just accept i'd have to roll my own GUI harness for my nontechnical team members. But seeing what you've put together so far (and knowing it's you who wrote it so I'm probably going to just see a step function level better quality in architecture/efficiency) is going to make me reassess in a good way. Some things that I'll be considering:
- Worktree support
- Sandbox support
- Skills/subagent handling
- Hashline based editing (feel like this is a huge part of why
people get better results from pi/omp over opencode/codex/claude code)
- Ability to customize tool calls + have rich embeds
- Web UI support (if i'm building this out for team members, native GUI can get messy and web client is ideal)
- Long horizon efficiency (IE i regularly get to 200k-400k context length sessions; while the model handles it fine, opencode gui will get laggy while the tui keeps chugging along)
For a lot of this stuff, it's less critical that all of this works perfectly out of the box and more critical that the architecture makes it easy to build (ie as with Pi ecosystem). What I'm after long term is something a bit like https://github.com/ColeMurray/background-agents in capability but without the overly tight coupling and design decisions that product has made.
The way I want to get there is to find the right base (whether that's Pi, OpenCode, or your project Juggler) and build the background agent harness layer. Previously it was just Pi and OpenCode and neither was really perfect (GUI story was probably the weakest for both) but it's great that I have another option to diligence that might actually be a better fit for what I'm trying to do.
Excited to see how this develops and kick the tires on it myself. The tree paradigm feels like the killer feature to me; not sure of anything else besides pi/omp that has it.
I've got things like worktree/sandboxing/skills on my TODO list.
I'd heard of hashline based editing - I will dig into that, and it's probably easy to add, though TBH I've not had any hassle with the editing tools so far.
If you get stuck into customising tool calls + their UIs, would love to hear how you get on, as that's one of the big goals for this. I've implemented all the built-in tools as plugins so hopefully it'll cover everything you need.
In terms of long-horizon stuff, yes, I also often hit 3-400k and haven't had any issues, but let me know if you spot anything untoward
I understand the concern and it's fair but I am very curious about what happens when the two notions of "free" (free as in beer, and free as in freedom) start to diverge because the former gets easier to do.
The latter as always been more durable. Linux doesn't have the mindshare it does because it's "free" as in beer - it's because it's "free" as in freedom.
The price of freedom, of brewing your own beer, is sometimes higher than buying it from the store. But for many folks, the control over the supply chain is what makes it worth it. In LLM-land, it might take a little bit of time for folks to catch up -- or maybe a lot of that is already in motion as companies get paranoid (and rightfully so) to frontier labs getting a little grabby about data. If you need a ZDR environment, "free" as in freedom has a very high premium that you will pay and rightfully so.
I kinda like the music comparison because it's so rich with sociological analogues. You have the gear snobs who can't stop buying more premium gear but never sit down and write anything. The kids who pick up a copy of Garageband and put together hits with their natural talent, sense of taste and and interesting story to tell. Soundtrack and videogame composers who have hybrid instincts of session musicians and architects. Avant garde musicians who you're never fully sure of whether they're playing a joke you're in on (or on you) and that's kinda the point. Music critics who have never played or written music a day in their life and yet end up becoming arbiters (or more accurately, delegates) of taste. Fandom stans and ringleaders who absorb it as an identity and run extraordinarily well organized cults with an iron fist.
You can probably find correlates here with coding and AI any which way you look. Coding is so rich that you can use it to do artistic, creative pursuits because it really is an interactive and world building medium if you want it to be. And it can also be a practical, reliable machine that helps you get useful business objectives done. And anywhere in between!
Perhaps the author is indexing on the former because there's an intrinsic value to that, and intrinsic values seem to be quite drowned out by the noise of extrinsic values in this media supercycle.
But I don't think it'll be that way forever. Whenever things get too noisy, people have a way of seeking peace and quiet.
You also have people who just like bangers and love throwing some tracks together, like to create sets or playlists, think they might be a DJ one day, to move the crowd, no time/inclination in making music per se but care about the created experience and will learn the 'plumbing' if they have to....
Coding is so rich that you can use it for artistic pursuits, but I argue that the output is the real art, and the code itself is not art and itself is not expressing anything.
I don't disagree with your conclusions (enterprises will pay top dollar for service guarantees, integration, and someone they can sue) but by that same logic there is no clear winner with Anthropic/OpenAI. Claude has a habit of going down on me when I need it most and seems to be struggling to even keep 3 nines of availability. They're actively hostile to integration and seem more convinced they should be suing others than behaving in a way that doesn't get them sued.
That's not to say I don't believe that there won't be a closed source correlate. I just don't know if OAI and Ant are all that exists.
"Anyone who depends on code review to find bugs is living in a fool's paradise. As everyone should know by now, it is not in general possible to find bugs by examining the code."
I think this is pretty clumsily stated. The way I would best summarize what it should be is "The person best suited to address bugs in the code is the author/owner." That's usually some blend of detecting and fixing them. Code review can certainly surface bug-prone patterns being introduced (or extended), and even catch them directly.
But as many of the peer commenters state here, the depth of code review that finds buggy behavior and risky structuring is pretty involved and an expensive use of time. At most places I've reviewed code prior, this was always amortized as cost of doing business. Maybe partly because there wasn't a "theoretically" faster alternative, and maybe because we were trying to avoid 2 steps forward 1 step back.
Now, I think we are struggling with "the clinical trial problem" where there is more pressure to ship using AI but it doesn't actually abridge the QAing and review part. The problem with trying to abridge that part is it just creates backpressure like a spring and then catastrophically explodes 10x worse later on than if it was dealt with directly. It can be very easy to build a circle of pitfalls that were completely unnecessary. I think that automated code review techniques are going to have to evolve to keep up with the new paths that code can be generated through.
And to anyone that has to go through new, arms race enhanced forms of slop cannons -- my heart goes out to you, because the old ones were never particularly pleasant to deal with.
I've seen folks make it work with a 3090 on 4 bit quant using turboquant for KV cache. That's key because 3090s remain the most cost effective gpu metal for enthusiasts (albeit 24g) and the jump to 5090 (32g) is quite expensive and not always worth the LLM specific performance; sadly, good 32g metal is somewhat lacking in the price point at or above the 3090.
reply