It’s not perfect, it has shortcomings, it sometimes produces bogus outputs. All of that is fine for a tool, it’s not fine when it pretends it’s a conscious being, because errors start to feel like lies and it becomes a bit too personal.
People want it to be Data from Star Trek, when it really should be the ship's computer. I want to tell it to run a simulation accurately, create some solved tool, etc.
I don't think there's any correction that can return LLMs to a purely tool-space. Too many AI boyfriend/girlfriends.
You can want your e-scooter to be a jet ski, but you'll wind up having issues when you try to use it as one. LLMs are _really good_ at pretending to be something they aren't, but not always so good at being that thing, so you should be careful what you ask for.
The difference is what happens when you come to depend on it. Do you want your airplane to be flown by a pilot, or someone who's so good at talking like a pilot that they can fool almost everyone?
If they are only good at pretending to TALK like a pilot, then of course that is bad. However, if it is so good at pretending to BE a pilot that it can actually fly the plain really well, then that is good enough.
He may be able to talk the talk, and walk the walk, but today's guest pilot has a crippling heroin addiction. Let's hope that auto-landing system works.
No because then the non-deterministic grifter Data tells you to kill yourself.
Data is an individual. LLMs are time-shared hallucinatory algorithms that will ultimately be used to sell ad space.
It is Stupid to create "the machine that melts people's brains", but these assholes are trying to sell you "the machine that melts people's brains" and love to market it as being "anything the user wants".
I think the problem is that we’ve anthropomorphised the LLMs in their training and through system prompts to the point that that seems reasonable.
We don’t thank our other tools like grep, or the compiler.
If you train your LLM to produce words that seem like human responses or conversations, they will. If you train them to seem like god-like superintelligences to people who think that’s what they’re creating, they will.
I think a lot of that is people who were always very insecure because they’re mediocre engineers. Previously asking questions or not understanding something was a bit painful but now you can hear “you’re absolutely right!” and make progress all day every day.
Until you need to interact with actual humans and that’s why you try to minimise it, hiding behind ai generated content.
I’m not a developer, but when I watch devs whose work process is prompt, copy-paste, try to run, paste error into code, try to run, etc. I can’t help but think they’re unskilled. There’s no brain engagement, no understanding of the bigger picture, just being a worse slower agent.
Sure, but even worse, is the state-of-the-art: they use a cli agent and the "copy-paste, try to run, paste error, try to run" loop is called "Agentic engineering." Now we have to believe they're 10/10?
I see no difference there except speed at the cost of whatever little understanding may have been gained by the manual inspection between steps.
Any competent developer knows there are limits to prompting.
Personally I've found LLM's suck at multi-threaded applications. (Because I've been tempted by the ~agentic loop~ and been burned. Then I hand code the core logic and all is well).
Woe to the developer who tries to prompt their way through this.
I hit a wall with creating a browser based video editor. Up until then, over the last year, I have been taking my hands off the wheel more and more and really have just become more and more productive. This video editor experience has forced me to get more into the code again.
They suck at multi-threading, they suck at maintaining a coherent goal through multiple files, they suck at basic math (claude could NOT figure out how to get the hypotenuse of a triangle given angle and side lengths...).
Sometimes I mess around with prompts just to see what I get back, more often than not, it's mostly a waste of time. God forbid you use it to refactor anything.
But seriously, lots of apps are just glorified NextJS apps which have tons of training data. Something like rust would likely churn out nonsense that compiles eventually but isn't optimal.
Agreed with the caveat that I think if you know what you're doing and are very cognizant, LLM generated Rust is amazing. I feel like it's hard compilation requirements gives a guardrails for a LLM and if it compiles, you're pretty safe against memory issues.
Maybe it's because my org gave us permission to fail and make mistakes, but I transitioned to AI driven development pretty quickly and am having fun learning how to best drive the AI to produce good quality code, with harnesses and patterns to prevent mistakes and a process to learn and rollback. It's a different kind of engineering now. I find myself thinking more product level and system level than if else branches. I think I've been able to upskill in ways that I wouldn't have been able to if I were still writing code by hand, mainly because I have the bandwidth to do so now.
I only have 10 years of experience, so I'm not trying to say that your lived experience is invalid, but personally I figure if this is the way the industry is heading then I may as well try to learn how to thrive within the new environment.
A lot of the joy I used to have doing this job is gone. I wrote some recursive code to turn nested API parameters into Elastic search queries, it was the last really cool thing I did pre-Ai. Sometimes I look at that code with a fond nostalgia, I'll never write code like that again unless I go out of my way to do it for fun. I don't even know what being a software dev will look like in a year from now.
This is my biggest problem with the current state of things. The loss in quality of life out weights the gains in productivity. And they are concentrated on highly conscientiousness individuals, those who already got silently taxed with most of the 80% of the work to get the last 20% of the results. So it is now compounding, if you care your workload will approximate 99,99..% of the total workload divided by the number of high conscientiousness individuals in your team/organization.
The current hype cycle might even represent a net global gain like some people argue. But it represents yet more externality driven exploitation. Remains to be seen what the new equilibrium will be given that now the population being exploited is not only close to the core of the world system but also already overburdened.
I do think it represents a mere acceleration of the previous trend that resulted since around 2005 in an explosion of average (not median) pay for software developers due to similar dynamics. If the system settles on a new level of pay that manages to convince enough people of enough skill to go along, it might just hum along and not implode. As for the rest of us, welcome to the growing permanent underclass and brace for the impact of the climate wars. We will be the fodder that will insulate the chosen ones under their air conditioned domes.
Someone very close to me is in their second year as a dev. I hope things don’t turn out as bad as it looks like it might. We are all going to need some luck here.
I'm in the same boat as you... 30 years of experience writing software, and now I'm just supposed to click "approve" on PRs without reading the PR. I'm just supposed to press a button. I'm looking for a new job, but it's going to be difficult to find one in tech that isn't a complete AI-psychosis shitshow.
Same. I'm being forced into a position where I'm supposed to do everything with AI agents, not write any manual code, and only act as a reviewer. It sounds like you know the drill. It may push me into early retirement.
The last straw (of many last straws) was over the weekend a co-worker sent me a chat "hey can you click approve on this PR real quick?". This only makes me click "Apply" (jobs) instead of "Approve".
I've been clicking apply myself, but only have had a few interviews. It's bad out there. I've been working since the late 90's and this is the worst job market I've seen in decades.
I mean, we're just the neo-luddites in this case. Or maybe the rust belt when jobs moved elsewhere.
The real question isn't if things are going to change around our jobs, its are we going to be able to move fast enough to avoid starving in the streets?
Yeah, the economy wants a dystopia, so we're just gonna have to let that happen industry by industry, calling each group that is exploited for maximal corporate gain a neo-luddite. Inevitability and whatnot. Hell, we're gonna die one day anyway, that's also inevitable, so we might as well all just drink the Kool-Aid too.
I think I’m pretty skilled, or at least I was, but much of my work now looks like this. The fact is in many cases Claude can diagnose and fix the issue quicker than I can even read it. It would be crazy not to take advantage of this. Of course it does get hard to resist the temptation to just become ever lazier over time.
This is my exact process nowadays when debugging some weird Linux wifi driver issue or something that came up after upgrade.
But that's because I don't have any real familiarity with the systems involved and I don't expect that gaining such familiarity will benefit me. If I am working on a system or product I'm responsible for at my job, it should be a different situation.
This is a sign of poor harness configuration and/or org level constraints (lack of MCP support, etc.) With properly configured and prompted agents error copy-pasta should be the exception not the norm.
Over the past year maybe 1.5x to 2.0x for me. As in: I can work on two projects at the same time with reduced amount of context switch. But that's it for me.
Maybe I don't have the brains for 1000x terminal agent coding, but 2 parallel projects seems like my saturation point.
I do believe there's a tiny subset of people who are more productive with AI, in the same way that Erdős was more productive with amphetamine.
But most who keep on going about their 10x productivity gain are indistinguishable from that one obnoxious guy at a party who won't shut up about his Ayahuasca retreat last spring. And they think they're Erdős.
this is a great metaphor precisely because (outside those with paradoxical stimulant response) everyone thinks they're more productive on amphetamines and there's some naive evidence to that effect. the number of things you've actually done is significantly higher. the change in value created, as measured by the decrease in distance between where you are and the the goal you're shooting for, is unpredictable at best and slowly degrades as reliance increases.
AI is like having a junior dev with an adderall addiction and an encyclopedic knowledge of coding syntax at your beck and call.
The ones who get more productive are either the very incompetent who get pulled up to the AI-floor level, or the very competent who know when and how to use it and for what. The midwits are too proud to use it and instead placate themselves that their precious skill is more special and immune to mechanization than it really is.
Many many people have jobs where their contribution is granting access to deliberately undocumented things, like knowing where the config files are and some such. They hate the idea of AI. For my non IT friends its great for diagnosing wifi issues. It's also great for competent network engineers. It's not great for those who gain a salary due to having memorized some actions or settings that they don't even understand much. Note that this group has also already resisted traditional script automation, just like US dock workers who resist automation.
I would bet real cash money that many many people, including your coworkers, think that anyone who unironically uses the term "midwit" is completely and totally obnoxious.
Recently saw a podcast about the rise of "AI midwits". The guy's point was that it's good to become a midwit, because it's a stage between knowing nothing and knowing a lot, but an AI midwit just has the aesthetics of a midwit and none of the actual knowledge and is stuck there.
But I don't remember how he defined a midwit, because podcasts are a terrible way to convey knowledge. It might've been something like someone who acts like they know everything.
People who uses the word "obnoxious" are also obnoxious. There are many unpleasant words, you have to use some of them and people don't like any of them.
Would you refer to someone as a midwit in a typical office environment and expect your professional reputation to remain intact? If you were to do this on an ongoing basis for whatever reason, how long before you're asked to have a little chat with the folks in HR? Would you then attempt to explain it away as one of those unpleasant words that you just had to use? I want to be there for that conversation.
Corporate offices are low trust environments where people do not really openly speak directly in any negative way (for basically the reason you've illustrated), so I wouldn't use them as any kind of standard for normal social interaction.
> I wouldn't use them as any kind of standard for normal social interaction
Workplaces are where many of us spend a large percentage of our waking lives. So, it is "normal," unfortunately. This is why HR language and norms have broken containment into the non-work world. (Whether that's a good thing is another question.)
Different context, different politeness rules. I bet you're fun at parties. An online forum is a relaxed space, like a town pub. HR isn't breathing down our necks here. Personal attacks are a different sort of thing. Using a word to describe in a comment some abstract situations and kinds of characters you see at the office is not a personal attack.
There is not one single discourse norm to rule them all and calling out every little thing as too offensive is just annoying. Yeah, yeah there are no midwits, everyone is a unique little genius flower, yeah.
People apparently can get a very different picture of the same site. Anti-AI people think this site's commenters are mostly pro-AI and vice versa. I'm generally pro-using-AI-in-proper-ways, and mostly have the impression that people on here tend to complain and grumble about AI mostly and are generally skeptical and dismissive of everything which is right in the HN tradition (the famous Dropbox comment etc).
Karpathy, Carmack, Terence Tao, Simon Willison etc. are all smart people and manage to use AI effectively and productively because it doesn't hurt their ego.
It's also very easy to dismiss everyone falling into the "AI trap" as being mediocre in the first place, though. But one just can't know this without having seen their work pre-AI.
I think the "meat proxy" people are mediocre regardless of whether or not the were brilliant pre-AI. They've reduced themselves to a copy-paste go-between for claude and slack (or github or jira or whatever) and are mediocre now. If only they could turn back the clock...
Some skills are easy re-pick-uppable, like bicycling. Mathematics is probably more challenging. Not sure about programming, especially "borint" business-like programming.
There's a reason games like Factorio or Exapunks are so popular among engineers!
Especially if you get a little burned out and can't bring yourself to contribute to a side project, but still want to do programming-ish things that get your brain moving
The fundamental skill can evolve. It used to be OOP mania as promoted by Uncle Bob, now it's Casey Muratori DOD. At least that's been my perception. I'm sure there was sometimes before Uncle Bob too, but I wasn't alive. I did however learn OOP in school, class collaboration cards or whatever they're called, and then learn why it all needs to be thrown out the window after finishing school.
Casey also pointed out three kinds of programmers: those who just want to get something done and will use an LLM because it's faster, those who actually like programming, and those who actually like LLMs. Last group always gets left out of these discussions.
(Just as importantly to note - Bob-style OOP was just fine on microprocessors of the 80s with no caches. The state of the industry has changed.)
I think people tend to forget that not only do we ourselves have different skillsets and can be amazing at one thing but horrible but another, but this also applies to other people in the world! Far all we know, there is an amazing developer out there who without LLMs, might have been the single best developer in the country, but even this person might not be able to figure out how to effectively work with LLMs. And vice-versa too.
Maybe this «spearheading AI person» just sucks at AI related stuff, as clearly that approach is bananas, but they could still be a OK developer.
> Far all we know, there is an amazing developer out there who without LLMs, might have been the single best developer in the country, but even this person might not be able to figure out how to effectively work with LLMs. And vice-versa too.
If working with LLMs effectively means accepting subpar results or be a reverse centaur, then I’d be glad not to be able to work with them.
I’ve never seen a good example where AI is a net positive to any development workflow. No one argues against compilers, build tools, IDEs, task runners, deploy and orchestration tools. Because they are great levers that lets you create more with less effort.
Many developers have narrow job descriptions, and the creative, productive uses of AI aren’t really obvious.
Where it shines is glueing systems together or building one-off automations that would take days, or weeks, to figure out. It’s for things you don’t have time to figure out or didn’t think were possible.
I'll argue against half that stuff because you probably don't need it. Do you need kubernetes or does "scp service.exe server: && ssh server systemctl restart service" work for you? GitHub Actions is the worst thing I've ever had to work with and I'd rather have a shell script. What is a task runner, is it an overcomplicated way to SSH?
> If working with LLMs effectively means accepting subpar results
Why would it mean that? That's one way of using them, sure. Personally, my code is better as I have more time to think about the software design than before, and I'm less avoidant of refactoring in my personal projects.
> I’ve never seen a good example where AI is a net positive to any development workflow
Alright, does that mean you also believe it's impossible then that anyone out there is using AI in a "net positive" way for their development workflow? Or just that you've never seen it, but you're open to it existing?
> Or just that you've never seen it, but you're open to it existing?
This one. Only a sith deals in absolute.
I don’t mind experiments to try to find methodologies for those tools. And I believe there are instances where they’ve been successfully used. The issue I have is the kind of generic statements that they are good enough to replace currently established methodologies. Like using AI is a panacea.
> Personally, my code is better as I have more time to think about the software design than before, and I'm less avoidant of refactoring in my personal projects
That’s a bit what I’m talking about. Have you investigated how it has helped you? And if there are other, more economical way to get the same result? Your statement seems more ritualistic than logical.
> Have you investigated how it has helped you? And if there are other, more economical way to get the same result?
No I haven't, but I'm happy to just freeform walk you through my thinking on it: I typically write (wrote?) software for two purposes: consulting/freelancing for others so building what others want, or for simplifying and making my own life easier and more enjoyable. "Stupid" stuff like Home Assistant for example, isn't really life-or-death, or Jellyfin for that matter, both things my family relies on now, but our daily life just gets easier all throughout the day when everything works in sync with what we're doing.
It used to be I had to make a decision what to spend time on, either I work on my professional stuff so we have enough money to survive (maybe more) and I get new challenges and all that, or I spend time improving and maintaining my home infrastructure, or whatever software I feel like I'd need to be better at doing my professional development.
I no longer am making that choice, I'm spending less time in front of the computer, yet the output and quality of my work remains the same, and the code and design when I look at it, even stuff I shipped 6 months ago, I'm still happy with how the code is, which for me I guess is the way I validate if what I produce is good enough.
Nowadays, my entire home-lab is configured with Nix and almost everything except my workstation and some random stuff, runs NixOS. Everything is hosted on a local Forgejo instance, which also has it's own (custom "written" of course) agent acting on issues and PRs, and I have my harness basically maintain my entire home lab at this point. Now I just open issues, have a conversation until everything is 100% clear, end up with a PR to review and merge if it looks good, and I can do this while juggling other things.
I agree with you that there are tons of people who are selling LLMs as a panacea to lots of things, and there is so much over-hype in the industry and ecosystem, I also feel like every "new thing" kind of comes with this type of almost scamming, which sucks, and makes it hard to discern from real positive opinions vs just regurgitated opinions someone read somewhere. I'm not sure what the answer to that is, except perhaps as what you say, only a sith deals in absolutes.
What are the exact "currently established methodologies" you're talking about that cannot be replaced by LLMs + a harness today, just as some examples? You're probably right that those exists, but I'm curious to hear what you think would be the most difficult to replace today.
> What are the exact "currently established methodologies" you're talking about that cannot be replaced by LLMs + a harness today, just as some examples? You're probably right that those exists, but I'm curious to hear what you think would be the most difficult to replace today.
I was explaining [0] under another post that programming is mostly translation works. You take a specs and you formalize it using code, like going from sketch to a proper engineering drawing. Software design is more creative, where you take a problem and then comes up with a solution (creating the specs). Software Engineering is ensuring that those two are done well enough while consuming the least resources.
So a program is always a formal system. It's also static. It will be executed by a computer which will actually have a tangible effect in the real world. That effect is what's valuable. The program is the seed which let us control that effect. Aka it's the map that let us plan the journey, but it's not the territory that we will have to travel in.
The issue I keep pointing in most of my comment is thinking that the map is the territory. That the novel are the words and not the story so we need more words. Or that the code is more important than the user' workflows, se we are adding more buggy code, while not ensuring that the workflows are undisturbed.
> Nowadays, my entire home-lab is configured with Nix and almost everything except my workstation and some random stuff, runs NixOS. Everything is hosted on a local Forgejo instance, which also has it's own (custom "written" of course) agent acting on issues and PRs, and I have my harness basically maintain my entire home lab at this point.
It's also highlighted here where you focus more on the process than the output here. The goal is to have a working homelab. NixOS managing it is only the process (accidental complexity). If it's where truly about the goal and not NixOS and using AI, by this point, adding new nodes (software, devices,...) should be as easy as selecting it and adding it to the current system, like a strategy game.
You can see that philosophy in OpenBSD, where the focus is to have a working OS, not to work on developing an OS. A lot of software are done and it's mostly just bug fixing every once in a while. You can also see the same attitude in industrial engineering where you develop a solution and then use it for years. You don't spend all your time tweaking it and thus disturbing the production flow.
So yes, when I see a LLM methodology, it's mostly about the work itself, not the output of the work. There is no definition of done or even the idea of having one. It's work for the purpose of working.
I think many people did argue against these things when they were new. Compilers, for instance, were seen as a waste of the computer's resources and produced less than optimal code.
That argument still stands, and is still right. Compilers do allow for some programmers to remain uninformed about the actual behaviour of the code they write and this has allowed for a number of actual harms in the outcome of code. (see: Fujitsu computers floating point errors and the British Postal System's persecution of her own Postmasters as one example among many cases) Also, the most performant code either has to be written such that a compiler doesn't incorrectly unroll it's loops or otherwise mangle the intent, or it has to be fine tuned after the fact to correct such mangling. Even Linus does this for the kernel in some cases.
So, drawing the parallel, 'these are the new garbage' but 'old garbage became acceptable so we should accept the new garbage' as an argument in support of being a meat proxy for Markovian stochastic lossy compression-decompression chatbots is isn't very persuasive, especially among this community with a greater concentration of systems level programmers than the general population.
That being said, i think both chatbots and compilers have a specific level of utility; neither of them should have unrestricted access to production filesystems or networks. That way be dragons.
They were right. Only after quite a lot of evolution did they start being wrong. And they're still not that good at SIMD. ffmpeg still uses assembly-code kernels.
Eh. The useful LLM usage I see is in internal tools (bugs don’t matter because the output or UI is the only thing that counts and they are throw away) or personal projects that otherwise wouldn’t exist.
Both of these can be quite invisible. But the benefit is there.
All paradigm shifts and regime changes happen because of a coalition of people that have nothing to gain and everything to lose by the current system.
But I think we should decouple mediocrity from laziness. I haven't seen any team invest in the mentorship required to develop juniors in years, for example.
The incentives for developing juniors have become misaligned as the expected stay in a company dwindled from decades to years to maybe year.
It would make sense to develop juniors if most of their comp was a four year vest but that doesn't happen until later. And in your first year or two you're usually a net negative... This is even more true with AI.
To be fair it’s also worth noting that it’s much easier to find a buggy edge case with existing code than it is to write bug free code that doesn’t have any edge cases at all. It’s so much easier to read some concrete logic and find holes in it, than it is to start from nothing and end up with perfection. It’s true both for humans and agents, but agents are better are validating correctness.
There’s been infinitely times where I was stuck at some problem and every single solution was complex, messy, over-engineered and somehow wrong, until I went on a walk or moved to a different issue and suddenly it would hit me that I was looking at it all wrong, and there’s a simple solution there but my tunnel vision didn’t let me see it.
LLM rob you of that, everything is instant, there’s no time to reflect, there’s no time to realise it’s a dead end, or that it won’t work with the next problem on your todo list.
I would say that the mechanism for this is actually due to the nature of working with LLMs: they free your attention.
Say you're working on a project, and it does an ok job like what the GP says. Each step looks fine, but put together it looks off. You give it some instructions, and then you let it work.
But what are you doing in the freed up time? You context switch into another project and give that some instructions.
Now you have two or three or ten projects that superficially look fine, but you yourself have context switched so much you don't have the overview anymore of why exactly each project needs fixing.
Which is also why it's so unbelievably useful for the ones for whom coding is the side career. My main job is being a psychiatrist but i've always been coding in my free time and now my productivity gains are unfathomably high as they allow me to do things in days that would've taken me years of commute time.
I've been successfully using local LLMs for the past few months on equipment ranging from AMD 395+ strix halo, to dual nvidia gpus with 96GB ampere, to 24GB dual 3060, etc. They all produce text gen at almost readable rates. You can read the thinking traces, the editting of code in the opencode harness, etc.
Along that same use, I've never touched any cloud service for code gen, and looked at docs or arena LLM. From my POV, it seems like a lot of people arn't even paying attention to the LLMs output when they generate their projects. So I don't find it hard to believe there's seemingly intelligent people being swindled by LLMs, big or small, because they do make bizarre assumptions and go into code edits that don't make sense even when they appeared to be running.
Then there's people that say they're in the ballpark of 500k context before they even get to work on something where these local models struggle when they get up to 100k both from generating and keeping scope on what they're doing. This is managed by opencode plugin dynamic-context-pruning that has the LLM rewrite sections of the context into summaries, while keeping the working context. It works pretty wel where a session then gets into ~500k where the working increments from 30k-80k of useful life.
So, I think we all underestimate how easily people can be swindled into poor development behaviors and _one_ reason, is the large cloud models just dump way more changes along a scope and can easily poison a whole chain of a feature or app implementation.
So the OP above who got 3 months into an app they think is hitting a dead end is recognizable as having one idea in their head of what's being built and the LLM building something subtly different because they're inevitably only reviewing the things that confirm the specs in their head and don't look for the edge cases that break it.
That's a choice (unless your management thinks otherwise). I don't need my agents running non stop. I tell them to start running once my walk is completed and I have a clear vision of what I want them to do.
It seems to me this is the most reasonable approach. Using LLMs to remove the cognitive load from implementation details, scaffolding, and taking it from there to production grade code. In short, like an IDE on steroids.
I'm curious and hear a lot about agentic programming, yet hear little about the cognitive load from it. Having more code generated doesn't make you more productive, in fact, I would argue the contrary is the case. There's more code to understand, more code to discard, an extra decision (which result do I pick? How do I split the workload through my agents), which is not at all how we operate (I. E context switching, limited by our cognitive energy, working task by task).
I really don't buy the whole agentic thing. The only reasonable use case I can see for it, is to scaffold multiple projects at the same time.
Anthropic Needs to have the most intelligent and scary agents, so if OpenAI does something bad they need to prove their models can do even worse. Without that the whole valuation collapses. There’ll be more “our model outhacks others” for the next hype cycle.
I find the era of "apocalyptic alarmism as marketing" super strange and maybe not good either. How long until one of these firms intentionally "forgets" to airgap the model during a cybersecurity test, just to get a good headline about how powerful theirs is? Perverse incentives all around.
Is it "anti-LLM" or is it anti "here's some code I don't understand that a machine generated for me kthxbye"?
Code is not just code, it's also liability and trust.
You sound like a webdev, e.g. the "let's just install 1000s of unvetted external dependencies" is nowhere else nearly as extreme as in web development land.
I don't think that's accurate. Until recently the code we rely on wasn't blindly trusted. People wrote and QA'd it. We may not have reviewed it personally, but that's one of the functions we outsource to software maintainers or "manufacturers".
I'm not "blindly trusting" code on my computing devices. I'm trusting the vendors / maintainers to do their job.
Until very recently the norm has been that the vast majority of code had human eyes and hands on it.
Edit:
The owners of those human eyes and hands had some type of accountability (either reputationally, in the case of free/open-source software, or occupationally, in the case of proprietary software).
The LLM has no accountability as to the output it generates.
The companies who make the LLMs also seem to have very little accountability, too. We've assumed a "blame the victim" stance when people use LLM-generated output in some inappropriate ways (legal briefs with "hallucinated" citations, articles "written" by LLMs). Whether that's the right location for accountability to be placed isn't for me to say, but that seems to be how it is.
I'm not sure that we're applying accountability to LLM-generated code in the same way we are for, say, the LLM-generated legal brief.
A compiler is not an LLM, and I do not want to equivocate, but there are aspects of similarity. We do not assume people read the bytes of machine code to ensure it’s correct — there could be mistakes. We also write and run automated tests to ensure the code outputted from a compiler and an LLM behaves correctly. At some point, we won’t have to literally read every byte of code that comes out because we have a reasonable assurance that it’s correct.
That we don't have to heavily scrutinize the code generated by compilers is a result of the huge amount of human toil that went into the compilers and test suites. Compilers are deterministic mechanisms so tests can be constructed.
LLMs, at least as they're currently constructed, aren't deterministic (the whole "temperature" thing). I don't see how to build a mechanistic test for something that has non-deterministic output. It feels a little bit like solving the halting problem.
I have no doubt we'll move away from human code review. The idea of large amounts of software edifice being built upon foundations that no human has reviewed or, perhaps even understands, is horrifying to me, though.
Compilers also usually give you the same output for the same input. And fwiw I do spend quite a lot of time reading compiler output to check that it's not doing something stupid or unexpected (usually as part of optimization work).
Also this sort of 'technological whataboutism' really isn't helpful, compilers are entirely different from LLMs. I agree that it doesn't make much sense to read or review LLM output in detail, but I also don't plan to use LLM output for anything important or mission critical. That would be irresponsible.
In addition, when we usually say that triggering undefined behavior on C can start a game of Tetris or format your hard disk we're usually joking (or at least exaggerating), i.e. the most common failure conditions of compilers are really limited in scope, and very likely caught by whatever testing mechanism you use for the software.
No such limits for LLMs where losing all your files is about par for the course for everyone who uses them regularly.
JIT compilers certainly do not give the same machine code for the same input.
The actual machine code depends on several parameters, and it is very hard to replicate them, hence why many devs get benchmarks with JITs wrong.
Additionally, compiler optimisation passes with machine learning is starting to be a thing, yet another way how the machine code differs for the same input across compiler executions.
It has nothing to do with 'elite teams', it has to do with decades long and very careful maintenance.
It's too early to say whether LLM generated projects will ever reach that sort of maturity, most examples I've seen so far are basically "fire and forget". But lets talk again in one or two decades, maybe there will be counterexamples of successful open source projects which will be just as well llm-maintained as human-mainained.
But I suspect that to reach that sort of maturity, the resulting human effort will be mostly the same (e.g. not much of a productity win - except maybe on the 'edges', e.g. maintaining the test suite, documentation, helping to analyze bugs..., e.g. these are examples where LLMs are genuinely useful and where plagiarism hardly matters).
> But the vast amount of software written, react components and rest endpoints, are very ripe to be entirely written by agents.
In that I agree, nobody should be forced to write React code manually, that's almost a human rights violation ;)
REST endpoints (and the code talking to those endpoints) should be code generated anyway though, no need for LLMs, and instead of human language prompting, a precise IDL should be the spec and basis for a mechnical code generation process. That problem was solved decades ago with much more pedestrian technology.
E.g. it basically comes down to "it's fine to use LLMs for software that shouldn't have been written in the first place", and funny enough that's where LLMs are really good at: creating software that has been written a million times before with only minor variations, and doing this type of work manually (cranking out one cookie cutter React webpage or REST API after another) is essentially what's called 'bullshit jobs' (which bring food on the table though, but that's another topic).
PS:
> and where software developers carefully will detail stear the work.
...I think the further a project evolves, the less this "detailed stearing" will be any more productive than doing the same without LLMs. The older a project, the more the work shifts from implementation to decision making, and in most cases the result of that decision is just a very tiny code change. I already see cases in my daily work when I use 'agentic workflows' where a tiny change takes longer and involves more 'collatoral updates' then just fixing that one frigging line of code by hand like in the olden days, and for LLM-generated code bases I really do prefer to not mix LLM and manual work, I think that's the worst of all options.
The company that employed the developer ultimately holds the accountability in the marketplace. The employed developer maintains (or loses) their job because of their accountability to their code (or, at least, they should). There's an economic incentive for all parties involved.
In the free/open-source world the incentives aren't economic, but they're still there.
At the start of this you said: "For 99.999% of people, it is literally kthxbye on all code they execute on all their devices."
I think that's inaccurate. The vast majority of code running on "all their devices" is code made by employees of companies being held accountable through traditional industry methods, or free.open source projects where reputational integrity was at stake. Those developers have been held accountable, for some value of accountable.
Maybe there's less value in human accountability than I think there is. Only time will tell. That's a different conversation.
The code running "for 99.999% of people" is not "literally kthxbye" LLM-generated code without someone behind it holding accountability. Maybe it will be in the future, but it's not now.
A bit of good old “it’s really good idea, I love it, great effort … BUT”
Mixed with “if kids can’t get semi dangerous drugs from the back of my van then they will be forced to buy from even shadier, more dangerous dark web van”
I'm sure Anthropic would love for everyone to believe that soon they will be the sole provider of "skills" for the entire planet and every single business transaction is done as a part of the subscription they offer - people sell AI generated content and other people's agents buy that - they're the entire world economy. I'm sure such narrative would help them reach infinity market valuation for the upcoming IPO.
I have never seen a plan written by LLM that wasn’t vague and light on details, every time I need to ask for more details and every time I ask to implement it trips over some dead end in the plan, discards the plan and continues as if there was no plan to begin with. Planning feels often just narrating the request and a wish list then what a human would do - methodically build enough understanding so that you’re confident of the direction you choose.
LLM plans are overhyped and overrated.
It’s not perfect, it has shortcomings, it sometimes produces bogus outputs. All of that is fine for a tool, it’s not fine when it pretends it’s a conscious being, because errors start to feel like lies and it becomes a bit too personal.
reply