Depends on what’s being written, and who the audience is. Anything of any length would be hard to simulate in a way that would fool an author - writing has a certain flow to it. A cadence. The editing and restructuring, deleting of words, typos you don’t catch until some random reread, rephrasing of sections because you want to use the original phrasing later in the piece.
Could you simulate something be typed? Trivially. Could you simulate something be drafted? Honestly, even if you wanted to put in all that time and effort, I’m not even sure LLMs are sophisticated enough to send the logical drafts, loops and edits that would pass a writers sniff test
I think you could simulate something that passes a sniff test. A writer would probably spot implausibilities in the simulation if they paid attention, but then we're back to square one, because you can spot that something was written by an LLM after you've wasted your time reading it and realize that you've been led around in circles with superficial information and no coherent train of thought, but by then your time has already been wasted.
To add to this, one can’t ignore the relationship between signal and receiver. I’d imagine most people on HN have enough pre-LLM reading experience to have a decent sense of what was written by an LLM versus a human.
And as LLMs get better at producing human-like text, that same pre-LLM reading experience, which helps people tell the two apart, will become less and less common.
I think it’s safe to say that this will not be consensus. Personally, I am getting increasingly (irrationally) angry at AI generated content. AI generated art quite literally makes me nautious. I mean an actual, physical reaction where I feel queasy.
I know I’m not the only one who feels this way, and notice more and more people reporting the same. Several of my non-technical
AI generated content is bland and soulless. There’s only so much bland and soulless most people can take in their life before they start to get fed up.
When everything feels the same, nothing is interesting anymore.
Everyday AI writing was not a thing with GPT 3.5. It happened more around GPT 4o. And now some people are entirely comfortable with using AI writing and not even trying to hide it (while, I would agree, it's obviously still fairly garbage and easily identifiable, which helps with triggering strong averse reactions).
However the models are getting better at everything, including writing for the past years. Why would that stop now? It's reasonable to assume that the makers also know about bad writing, dislike it, and thus the models will get trained to get even better at it.
Eventually how will you be able to tell? You won't. You can't. And that goes for the rest of us. And I suspect everything will just feel somewhat nicer.
To claim there was amazing progress in the past therefore there will be amazing progress in the future is an inductive fallacy.
And as someone who gets dozens if not hundreds of AI generated emails I have to go through every day, it is _incredibly_ easy to spot which ones are AI generated and which are human written.
By its nature AI generated content is statistically consistent, the narrative equivalent of monotone speech. I don’t know anyone that can’t spot it a mile away at this point, and the more people are confronted with slop, the more attuned they become to it.
> AI generated art quite literally makes me nautious. I mean an actual, physical reaction where I feel queasy.
Have you tested yourself on this, e.g. https://www.astralcodexten.com/p/how-did-you-do-on-the-ai-ar... ? The works shown here are IMO very impressive overall, especially in the impressionist style. And the test was originally given in October 2024, i.e. an eternity ago in model-advancement years.
I do think Ed in intentionally ignorant of the capabilities of LLMs. But I also don't know that I would classify LLMs as 'wildly useful' for coding. Most productivity gains seem to be hallucinated, and while it's too early to make any claims on long term outcomes, there are plenty of studies indicating they might be even more negative.
There are definitely use cases for LLMs in coding. And at times, they can be wildly useful. But I feel like the industry atm wildly overestimates their broader/long term utility.
Anecdotally, I have not seen an explosion in quality/bespoke software since LLMs. In fact I've noticed the opposite to quite the extreme. Not only are new products worse in quality, but the quality of existing products is falling off a cliff.
> Anecdotally, I have not seen an explosion in quality/bespoke software since LLMs. In fact I've noticed the opposite to quite the extreme. Not only are new products worse in quality, but the quality of existing products is falling off a cliff.
This is the big one. It's clear that AI can generate huge volumes of code by KLOC. It is not clear that spending a lot of money tokenmaxxing will eventually result in increased real revenue for software businesses, and eventually even an MBA has to look at a "money in vs money out" chart.
Have been thinking about this a lot recently. AI could be an absolute game changer for a small start-up rushing a product to market – you could quickly build an MVP that would take years and tens of hires before.
But how much ROI is there for large businesses with established products and huge development teams burning through tokens making subtle tweaks that can’t be directly tied to revenue?
Not much ROI, if any. My employer's been making some studies and come up with very modest productivity gains - so of course they want us all to use it, but I'm not sure they're taking the true costs into account. Especially not once token pricing actually reflects reality, and we all get brain rot from using the things instead of thinking. If this thing doesn't collapse before we can run a solid coding assistant model on a developer's machine, maybe it's got some legs.
That doesn't seem very likely.
The legacy of LLMs will live on in various models doing various specialist things (they seem like a really good progression on speech synthesis for example) but the current edifice will come crashing down and if we're very very very lucky they won't take the global economy with it.
I've been thinking about this too. The quality maintenance of large systems isn't something you can just completely automate away with AI. Even if the code is written with AI, you still have to read through and verify it.
Even though that is still faster than regularly writing code, I end up losing that nuanced knowledge that I get from going through documentation and writing it out by hand - actually doing the work. I just don't see it actually replacing developers unless managers are willing to produce MORE code with the SAME level of quality.
> I do think Ed in intentionally ignorant of the capabilities of LLMs.
I think it's more complicated than that too. He's pretty well versed in the stated capabilities of LLMs.
The fact that he isn't a deeply involved technical developer who knows the ins and outs and nuances of using LLM tools is the point, because the stated capabilities of LLMs are that they are trivial to use, extremely powerful, and getting so much better every month that you personally can replace developers without even trying as a completely non-technical person with basic writing skills.
Given the hype and extreme claims being made, the fact that he remains ignorant and gets practically no use out of LLMs immediately disproves those statements. The counterargument boiling down to "you're using it wrong" is actually just a further indictment of Sam Altman and his like, because it shouldn't be possible to use LLMs wrong!
The rest, well, the hype needs to die before anyone can make sane estimates of what LLM tech can do for us in various fields. Right now it's all a complete mess.
I don’t get me wrong, I’m on Ed’s side and get where he’s coming from. I just think his arguments are normally taken to the extreme, making them less defensible, when he could make the same arguments from a more moderate stance and ultimately be more convincing.
His arguments, albeit valid, can often sound like reductos ad absurdums the way he presents them.
One of the worst things about LLM writing is how it makes big promises of what it can prove in some piece of writing, and then never really follows up on that, or has specifics that go all the way towards the original, grandiose statement.
And frankly, Zitron is guilty of that pattern of writing too, or of relying on some unstated "baseline" knowledge which is clear from his other writing but not in the specific piece.
So, basically yeah, agreeing about the ad absurdum thing.
(I will note, the tone, the swearing, etc. really doesn't matter nearly as much as these problems, and everyone instead obsessing about the swearing and personality is really boring)
Personally I only find his swearing and delivery annoying when he is not delivering his point well (ie reducto ad absurdum). I’d be welling to bet a good amount of the complaints about his swearing are really just from a poor delivery, and people don’t know why, so they latch onto his swearing.
His early stuff was just as degenerate and vulgar, but was much less of an issue for me.
Part of the problem is the like - sociomedia factor. Ed’s figured out how to break through the noise. I’m not surprised that Ed Zitron is the kind of counterpoint you get in the Musk-Trump attention economy / CEO’s that sound more like prophets than business executives world.
One person's ignorance of something can never be evidence that it doesn't exist. It's far too easy to be willfully ignorant; no one can force you to abandon ignorance if you don't want to.
On the other hand, the hype of "Sam Altman and his like" being plainly exaggerated doesn't mean there's nothing at all behind it. It's plain to see there's something important about LLM capabilities. I don't even use them myself, as emotionally I find them entirely repugnant, and I can still see that.
We need to wait to get the whole story about LLMs, but we don't need to wait to confidently reject both extremes of opinion about them.
> Anecdotally, I have not seen an explosion in quality/bespoke software since LLMs. In fact I've noticed the opposite to quite the extreme. Not only are new products worse in quality, but the quality of existing products is falling off a cliff.
Very little new or ground-breaking (I struggle to think of things AI has produced that aren't themselves just more AI), but various previously-stable sites and services breaking.
The studies you are talking about are probably outdated, it's difficult to deny the actual productivity boost of coding agents.
I'm not talking about the quantity of code produced, but about actual user needs that are now resolved that would not have been before.
The main productivity gain will not come from existing software engineer, but from people that couldn't code at all before but are now able to do things by themselves. We are still very early.
It still takes the same mindset and skills to use AI productively and effectively as regular programming. The productivity boost only applies if you know what you’re doing and can actually steer the agent carefully.
Vibecoding hits a glass ceiling very quickly and this will not be solved incrementally. Besides, if the agent could work autonomously to that degree then it would no longer need any prompting at all and we’re living in a very different world. On the other hand that would make the debt actually meaningless, so I guess that is 'a' solution.
As a developer who has not been able to get any boost in productivity from coding agents, I find it incredibly easy to deny.
I’m a solopreneur, if I could lighten my load I would. However I have yet to save time using coding agents, with the exception of “I made this change to my model file, update all model to match the new format.” Which is cool, but maybe 0.01% of my job, and took a 1 hour task down to 10 minutes.
We previously had a backlog of 2 years of features planned. Now we have no work and are just working on tech debt and planning to get more involved in the product side because they need help getting the developers work to do.
To continue paraphrasing Steve Jobs, focus is the most important thing. When the cost to produce new features/implementations goes down, focus is even harder (and even more important).
It took me probably 5 years of writing Clojure before it clicked. Once you get used to structural editing and repl driven development, it’s really hard to go back to syntactic languages.
It’s kind of like in treesitter style editing, where you can “swap these two arguments,” “select this function,” “wrap this in a try block” with a single keyboard command… but way more standardized and granular. Plus with the ability to execute anything you highlight
All that and then you realize you can store code as data (since it’s just a data structure) and run data as code.
I think most programmers don’t realize how arbitrary the difference is between code and data until they get used to using LISP.
Spot on. For me, it clicked with Common Lisp, 15 years after I graduated from university. Now, Clojure is my daily driver. And it’s extremely difficult to explain to people. I’ve gotten to the point where I don’t even try. You’re right about all the things you mentioned. Once you discover structural editing, everything else seems primitive, on the level of cavemen playing with rocks. But it’s not just one feature that makes Lisp better. It’s all of it which interrelates and creates a powerful synergy (I hate that word, but in this case it’s appropriate) that just isn’t matched by anything else. There are other languages that have a similar vibe, notably Forth and Prolog, but they are often misunderstood, too. Honestly, that’s my real test of whether someone is a senior programmer: do they understand and at least have an appreciation for these languages, even if they don’t program in them everyday.
Not in many tasks. I use deepseek as a fallback in https://phrasing.app and it’s always very apparent when it happen (due to mistakes/clear performance drop off)
I think it really depends on what you’re doing. I use mistral for many tasks in https://phrasing.app and they blow models many times their size out of the water.
None of my tasks use reasoning though (reasoning actually kills the performance) so perhaps that’s why. Still, I just had to rewrite my pipeline, and mistral was both faster, cheaper, and substantially better than any alternative
Actually, deepseek v4 was 1/3 promotional price for the first month or so. This was pretty clearly communicated. The promotions window just ended is all.
If you run out of 50% coupons to your local pizza joint, did they double their prices? Does every company double or triple their prices after Black Friday?
There’s a pretty significant difference between saying someone tripled their prices, and a temporary promotion ended. It’s even more so the case if someone is using it as an example for raising prices as a trend.
I’m 100% in the camp that prices are going up and quality is going down; companies are retiring models and requiring you to use more expensive ones. This has happened to me and there are dozens of examples that one can point to.
But a promotion ending is a strawman argument and does the point a disservice.
> If you run out of 50% coupons to your local pizza joint, did they double their prices?
Yes. Did they double their msrp? no. They did double their effective price relative to me which is all that matters unless you're doing economic math or something.
The original comment was used as proof of a trend that vendors are raising prices. Would running out of coupons indicate a trend in rising pizza prices?
It depends whether it's me personally who's running out of coupons or the entire supply of coupons is being reduced. If my ability to get the product for the same price is diminished then the price is being effectively raised.
In this case I'd agree that pricing is effectively raised as 10$ > 10$ - 50%, there's no need to complicate it. However this is not even the right metric for this problem, a better one would be total spent / work produced. If all customers spend more money for the same amount of work (adjusted to progress) then clearly the price is increasing. This would be true in this example as well.
Could you simulate something be typed? Trivially. Could you simulate something be drafted? Honestly, even if you wanted to put in all that time and effort, I’m not even sure LLMs are sophisticated enough to send the logical drafts, loops and edits that would pass a writers sniff test