Hacker Newsnew | past | comments | ask | show | jobs | submit | boron1006's commentslogin

My Eng lead has no coding experience, 25 years of management experience, yet has driven 3 separate projects into technical bankruptcy to date.

He just accepts anything that Claude says as truth. He vibecoded over 60,000 lines of code in 3 weeks, but couldn’t get it to do what he want and made a project overrun for 3 extra months. When the pissed off stakeholders called a meeting to ask what was going on he didn’t show up and sent his junior engineer to answer questions and take the blame. Now thats leadership.


> When the pissed off stakeholders called a meeting to ask what was going on he didn’t show up and sent his junior engineer to answer questions and take the blame.

He doesn’t have 25 years of management experience for nothing, that’s a crafty vet move!


Thanks for sharing this; I'm hoping your horrific experience will help some of my friends, who are afraid LLM's will take everyone's jobs.

My boss is a designer and does the same. He has come into a mature product's codebase, the only project the company has, he called me about 6 months ago and started boasting with a screenshare showing me how he had 5 terminals open and smashing through ideas straight into production.

Now the system is failing, he thinks I am going to go in and fix all of his problems. Told him straight what I told him 6 months ago, he owns it, so get it fixed.

Not sure what happens to myself or the company by the end of the year.

Needs to be more stories about these people injecting AI Hopium and destroying their product, there isn't enough of them.


> you break it you buy it.

Works well for lower stack systems or juniors merging stuff into a development environment not so well for production.


> My Eng lead has no coding experience, 25 years of management experience, yet has driven 3 separate projects into technical bankruptcy to date.

This is C-suite material.


The most serious vulnerability that any LLM has exposed to date is in the human brain, where a whole lot of people cannot help but anthropomorphize it and they seem to be willing to totally forego critical thinking. I think this vulnerability has existed for a long time but the mass exploitation is going to destroy quite a few organizations.

  > totally forego critical thinking
if there is anything i've learned working in this industry is that everyone hates thinking; they want a rote formula or pattern they can repeat for every project and every functionality... and ai mania is perfect cat-nip for this type imo

Future president in the making.

I don't think we should expect such unreasonable standards from future presidents.

bro the president suggested we nuke a hurricane and raped someone and is dismantling our natural resources, the bar is substantially worse than a couple months technical debt.

Yeah, I think N_Lens was saying the manager's behavior in that anecdote was too high a bar.

Fellow Apple product user. Join me in jeering at people who drink instant coffee.

I grew up on instant coffee, i don't like it much anymore EXCEPT for tossing a bit and some ice in a protein shake 100% recommend.

I use Apple and drink instant coffee, though!

I don’t like the term taste, but the problem that I have is that LLMs don’t seem to work “good enough”. They seem to be able to solve the immediate problem, but stacking this on the scale of 3-4 devs over 6 months or so doesn’t seem to produce anything.

One thing that I’m particularly frustrated with is the writing quality of LLMs. Like this is the thing that they should be able to do, but I would say almost everything they write has almost no signal.

Over a mid sized AI generated codebase that means I’m reading like 500 words to figure out what a module is even doing.


I'll risk sounding antagonistic and ask you this: if LLMs are not "good enough", why are they still around?

Obviously "good enough" was a poor pick of words of words on my behalf---and I'll gladly own it---but they should be "good enough" for something if a significant amount of resources keep getting allocated to them. Yes on a very personal level you look at LLM-generated code and think to yourself "wow this is garbage" but what about the people, as pointed out many times in this thread, that simply do not care? Does that not count as "good enough"?


> Yes on a very personal level you look at LLM-generated code and think to yourself "wow this is garbage" but what about the people, as pointed out many times in this thread, that simply do not care? Does that not count as "good enough"?

Aside from the "amateur" qualifier, the Booch quote below provides insight to those questions:

  The amateur software engineer is always in search of magic, 
  some sensational method or tool whose application promises 
  to render software development trivial. It is the mark of 
  the professional software engineer to know that no such 
  panacea exist.[0]
"Code" is shorthand for "encode", which can be defined as "encoding a solution to a problem", which implies knowing what problem to solve, and thus ultimately implying understanding of same.

So "good enough" depends on the stakeholders' definition of that valuation.

IOW, if a person tasked with delivering a solution can instead deliver garbage and still get paid for it, then it is "good enough" for some definition of same. If the people approving a person's remuneration are not satisfied with this type of work product, then the answer as to whether it is "good enough" or not will be quickly answered.

0 - https://en.wikiquote.org/wiki/Grady_Booch#Quotes


The use of “amateur” is indeed a very poor choice. An amateur is someone who seeks to master a skill or perform a craft for the love of the art, not for money. It’s the professional who is looking for the shortcut, not the amateur.

The amateur is the one who hones a tool’s edge until it’s razor sharp, and cuts away all unnecessary detail until they’ve found the essence of what they’re trying to achieve. The work of an amateur is often highly impractical and uneconomical, but sublimely beautiful.


I think you are torturing this comparison by using different meanings of amateur and professional

A professional does mean "someone who does something for money" but it also often just means "someone who is very good at something"

Similarly, amateur does mean "someone who does something out of passion without wanting/needing to be paid" but also often just means "someone who is just starting out, is new at something, and is probably not good enough to be a (paid) professional doing it"

You're talking about the "passionate" amateur and comparing it to the "getting paid" professional. I think many people talking about amateurs and professionals are mostly talking about "amateurs (newbs)" and "professionals (very experienced)"


You just twisted the meanings of the words.

An amateur is not a master nor a disciple of a master (e.g. future master).

The vast majority of amateurs have no ambitions nor passion and simply give up early on, that's a rose tinted view.

Since they don't get paid to do it, they have to allocate their precious free time toward the skill, which means they don't have those 40 hour weeks where they spend all their time honing their skill to begin with.

Even someone who does a skill half heartedly without passion nor ambition is going to learn faster with those 40 hours at his disposal.

>The amateur is the one who hones a tool’s edge until it’s razor sharp, and cuts away all unnecessary detail until they’ve found the essence of what they’re trying to achieve.

That's not an amateur and you know that. By your logic Japanese master craftsmen doing this are amateurs.


It’s not I who have twisted the meanings of the words, it’s the people who use the word amateur when they mean hack. The word amateur is very old, and for most of its history [1] it literally meant lover and in the case of practitioners of some art or sport, it meant those who participate out of love for the activity, not for money.

[1] https://www.wordorigins.org/big-list-entries/amateur


The market logic isn’t going to be defensible. The market is highly irrational and prone to be influenced by those with the most capital, not the, “invisible hand.” Cars are freaking everywhere not because they are the best solution of transportation but because people with vested interest in the sale of oil and cars made it so by buying politicians, starting political and marketing campaigns, and on for generations. It took decades for the car to become the de-facto mode of transportation in North America let alone the rest of the world. It’s not good for us. It’s highly irrational to prefer cars when freaking wildfires and climate change threaten our very existence. In a similar vein, LLMs as they are, are not around because they are necessary and superior.


The current kind of car is extremely good at doing what it does. It's just that most of the benefits are short term, and for the car user (and the oil salesman), while the costs are borne by everyone else, and some of the worst ones are long term.

(Some similarities to hard drugs here.)


"The market chose it" doesn't mean it was the best solution, just that it fit the incentives


Cars are actually the best solution to transportation once you stop ignoring inconvenient but real requirements like autonomy and isolation.


Cars are not a system.

You need roads, traffic lights, drivers and pedestrians that respect rules, not too many holy cows in the way, not too much CO2 in your quasi-self-sustaining spaceship, lungs and livers that can deal with the soot, a tolerance for high-velocity impacts risks, not too many other cars around, ...


> Cars are actually the best solution to transportation

"the" is the error here.

There's places and people for whom cars are so. There's other cases where public transport covers the people while business-owned vans cover deliveries of bulky items, which is essentially how I've lived the last 8 years, though for "bulky" I should note I've seen someone take an actual kitchen skink on public transport here in Berlin.

I've also found some islands here Berlin with no bridge connection to the mainland and hence no practical use for a car at those locations; that said, my main reference for "car free" domiciles would be Sark: https://en.wikipedia.org/wiki/Sark


They're here, because even "bad" code has uses. We can have quantity over quality. We can generate throw-away programs (where the lack of maintainability doesn't matter). We can solve some problems by brute force (keep trying until tests pass). For people who can't program, even bad code can do more than no code.

LLMs are also useful as a natural language interface. They're pretty useful for search (as RAG and autoresearch, not the glue-on-pizza Google search crap). They're good for filtering and classification against rules that can be specified in natural language instead of special syntax or by training dedicated models.


> They're here, because even "bad" code has uses. We can have quantity over quality. We can generate throw-away programs (where the lack of maintainability doesn't matter). We can solve some problems by brute force (keep trying until tests pass). For people who can't program, even bad code can do more than no code.

If this is an acceptable definition of software engineering productivity, where "quantity over quality" is prized, then hire me.

Because I guarantee I can produce dozens of PRs daily, each having tens of thousands of LoC deltas (pick any language you desire), none of which having research and/or understanding underpinning them, and all easily quantified as being "brute force."


There's a quote I forgot that says: "There is a Quality to Quantity".

I think it was from a Chinese person, as they took manufacturing of cheap things, low quality stuff, but managed to scale it to mass quantities.

I agree with that, it may not be reliable, full featured, well designed, ergonomic, etc. But quantity enables things and thus is a quality in itself. It allows use-cases that benefit from accessible cheap software. Those use-cases were not viable prior, because cost/time was prohibitive.


> There's a quote I forgot that says: "There is a Quality to Quantity".

> I agree with that, it may not be reliable, full featured, well designed, ergonomic, etc. But quantity enables things and thus is a quality in itself. It allows use-cases that benefit from accessible cheap software.

The underlying assumption here is that the quantity delivered is acceptable and does not worsen the customer experience. To continue with the manufacturing domain example; if a company produced widgets quickest and for minimal cost, yet had more lawsuits stemming from their use, is that the kind of enabling an organization should pursue?

> Those use-cases were not viable prior, because cost/time was prohibitive.

Sometimes a cigar is a cigar and sometimes what is thought to be a desirable use-case is prohibitive because, once sufficiently vetted, it is neither cost nor time which disqualifies it.


You're just saying quantity is not always a desirable quality. That's true, and hence it's not every use-case that would favor quantity over other qualities. But there also exists those which do, and trading other qualities for quantity can be a good trade off.

> the quantity delivered is acceptable and does not worsen the customer experience

No, it will explicitly worsen the customer experience in other ways, but it will also give them a benefit: access to something they couldn't afford at all before.


Often attributed to Stalin in reference to the mass production focus of the Soviet war machine. https://quotees.co.uk/war/joseph-stalin-quantity-has-a-quali... When you treat scale as a design criteria, you get a different product


Ah interesting, it does say probably, so seems it's not fully confirmed. In any case, I guess my brain re-interpreted it in the context of manufacturing and making products and remembered it as that.


> Because I guarantee I can produce dozens of PRs daily, each having tens of thousands of LoC deltas (pick any language you desire), none of which having research and/or understanding underpinning them, and all easily quantified as being "brute force."

I can guarantee that you cannot produce such PRs in the same quantity/quality than a modern LLM. The speed at which today's LLMs output code is superhuman.


> I can guarantee that you cannot produce such PRs in the same quantity/quality than a modern LLM.

Who said I would not use a combination of LLMs, scripts, and any other automation technique available?


I'm not happy about the drop in quality either, and I'm not advocating for just making software shittier.

However, look at history of DSLR cameras vs mobile phone cameras. Mobile phones were always worse than large format cameras, and we've got a massive quantity-over-quality explosion in photos taken. It didn't replace professional photography, but let people upload photos of their lunch to social media.

I think in programming we're getting frustrated because we're trying to naively use high-volume low-quality code generators in our previous low-volume high-quality workflows. We're still discovering what's the programming equivalent of bathroom selfies, photos of receipts and QR codes.


It's textbook historical revisionism. It's not like we didn't have mountains of bad code in the pre-LLM era. I'm old enough to remember endless complaint posts about having to wotk with legacy codebases full of horrible spaghetti code. Somebody took a long time to write it by hand.


Have you realized you are attacking a straw-man? Nobody said all “hand written” code is good. I _think_ the implicit argument is: “all vibed code lacks cohesion”.


> I'll risk sounding antagonistic and ask you this: if LLMs are not "good enough", why are they still around?

To quote the parent commenter:

> They seem to be able to solve the immediate problem,

but not long term


Breaking down long term problems into small shorter problems we can solve one after another until the long term problem is solved is the essence of engineering.


No, the "essence" of engineering is having a cohesive vision for a larger plan.

No quality large scale project can exist without it. And any seasoned engineer should understand by now that no SOTA LLM can produce quality engineering at scale.


Lol. That's not true at all. And not how projects are done on startups or big companies.

Execute fast or get nowhere, a series of small wins get you to live long term.


Nah. You can succeed in every single individual thing but fail at the overall project because the parts don't align into anything that makes sense.

If you need examples, look at game dev. There's plenty of games that have good execution but aren't fun to play or they're a confusing mess because there wasn't good overall direction.


Videogames are a weird comparison because they are basically the ultimate culmination of the combination of both engineering and art. The two can crossover in weird ways but still have distinct aspects that do not really compromise the other.

Just because a videogame isn’t fun doesn’t mean there wasnt a cohesive engineering vision. God of War: Ragnarok is a good example of this. Excellent technical execution and (arguably) great direction but a terribly boring game.

And many games that are terribly engineered are also incredibly fun. Dark Souls is a good example - runs like ass, looks very rough even by the standards of the time when the game released, with terrible enemy AI in a combat focused game - one might even call it a confusing mess (and I laugh at the notion that it doesn’t have good direction)… despite all this it’s a classic and arguably the most influential game since its release.

Elden Ring carries a lot of legacy baggage from years of iteration on the same engine as ~~Dark~~ Demons Souls - and it’s widely considered From Software’s magnum opus and a masterpiece. Yet considering the scope of the game it is a great example of how small wins added up can solve large problems when executed with the skill and judgment gained from experience. So one might still consider it a success in terms of “engineering”.

Would From Software have been able to make the same game using Unreal Engine 5 built from scratch with dime a dozen “Unreal experts” brought in off the street? Not a chance. Want to know why western game development studios are suffering right now? Compare American layoff culture and race to the bottom economics to the retention rates of Japanese game studios like Capcom, Nintendo and other heavy hitters.

For many of the most marvellous things humans have ever built, if you take a peak behind the curtain you’ll still probably find some amount of duct tape and popsicles sticks somewhere keeping it all together.

Fun, however, is a matter of taste and you can’t engineer that.


Long term projects that survived on short term wins are still around. But you're going to want to check how big the graveyard is of "project collapsed under tech debt from short term decisions" to see whether it's a good strategy.


But that means they're good enough for the short them. This is why I find it such an interesting question to wield, but also, it is why my wording was poor in the post.


Short term can, IME, be very short. I've seen people generate, say, a bash script with an LLM. It's generated: short term, the problem is "solved": we've generated a bash script.

… but does it work? Someone comes along, reviews it, "this is garbage, and does not do what it says it purports to do". Perhaps it even gave an output: the script computed … something, but it's just GIGO.

But that "check if this works" friction is the same friction that is what people try to avoid by generating it with an LLM in the first place. If you're too lazy to write the script, you're practically by definition too lazy to verify it.


If this is your working environment, it sounds like quite an unusual place.

I literally can't imagine generating a script with an LLM without testing it at all.

Bash is one of those situations where LLMs can do really well. No human on Earth can remember all of the commands and even fewer humans can remember all of the switches for all of the commands.

This kind of remembering, searching, and assembling is exactly what LLMs are good at - as long as you're not writing a gigantic build system with hundreds of moving parts, in which case you should probably be using something more streamlined anyway.


> I literally can't imagine generating a script with an LLM without testing it at all.

Then you're extremely unimaginative as well as unusually fastidious.

Certainly someone - several someones - are generating lots of scripts and not testing them, given the PRs I'm seeing.


You can have another agent write the tests and verify the former agent ? This is pretty basic stuff. Makes me question if people are actually trying to use AI


Now you have two problems lol, in that you don't know if the tests are any good or actually test the thing in question.

Sooner or later you run out of turtles to put on the stack.


That’s a solved problem, you just add another agent to check if the tests are any good, and one more to oversee the test-checker, and one more…


They're "good enough", but they're not good. If you're trying to run a business, programmer skill isn't worth the money you spend, LLMs are cheaper. If you care about the quality of what you produce, they're trash.

Programming is no longer skilled labor. Quality is a hobby.


If your business involves storing any kind of customer data or handling payments, you probably should still hire a programmer who at least reviews the results of your vibe coding. Getting hacked is bad for business.


AI is just fine at security audits these days. It's better at finding bugs than it is at writing code, overall. You need to prompt for it, but that's not skilled labor.

I'm also not sure getting hacked is so bad for business. Most businesses that got hacked seem to still be around just fine.


Because if you criticize the LLMs, you get headbutted into the Dante's Inferno that is the current job market


I'd agree with the OP you are responding, what he means is that this isn't just about taste. It's that the code, while it works, and the text, while it has the ideas in it, is still not to the level of a human author. It's not that it isn't tasteful, in the sense of having the flair of specific author X, but it is verbose, goes around the point to get to it, etc.

Taste seems to imply an art, but I don't think we're at the point of cleverness, of LLM innovating a new beautiful programming paradigm or pattern out of immaculate taste for code. We're just talking more basic, even bad devs have more terseness, more clear intent in their prose and code, etc.

It's "good enough" for many use-cases, but it could be better, even before expecting the most beautifully clear code from the best devs out there, it has not reached the level of the average one in readability and clarity of expression (both code and prose). That annoys the author, he has to intervene manually to trim all the fat, and therefore not "good enough".


It’s a good question and I don’t doubt they’re useful.

But I mean they’re not good enough to do the things I need them to. Not enough that any difference can be chalked up to “taste”.

Put it this way, if you hired a very good handyman to build you a cathedral, you wouldn’t stand looking at the smouldering wreckage of the construction site saying “oh well that’s just a matter of taste”.


The market desperately wants to replace wages and the idea of a machine doing the work you previously needed to do before is so enticing, and the initial results so promising, that the rest doesn't matter.

Look at how many AI lay offs there already were a year or two ago - back when almost everyone would agree they weren't good enough. If good enough didn't matter then does it matter how when AI is better?

And despite what you might be trying to prove, it's obvious I'm replying to an LLM (a green account at that), but it's not for you, I don't expect you to bother reading it.


> if LLMs are not "good enough", why are they still around

They are not good enough yet, the people using these tools every day, and with a high bar for quality and design 'taste' notice the rough edges.

They are still around because LLMs are plenty good enough for casual consumers, but where 'taste' really starts to matter in design or in writing, the rough edges show up much more obviously and users who have been working with these tools feel it immediately.


One part of the answer is marketing it to people looking for an easy way out without having to do the work or thinking at all. There's value in that as shown in everything like things for the home that purport to make your daily life easier. LLMs market themselves in much the same way


Conversely, if they are good enough, why do they degrade frome consuming their own output?

How come humans consuming their own output gets better over time, but llms and any other form of ai consuming their own output only get worse?


Not disagreeing with you but the "resource allocation" which I presume to hint at the fallacy that capitalism is somehow great or optiomal at allocating resources is not really true. Just look at any big company such as meta and see how much inefficiency and bad choices there is.

Stock market as a measuring stick is completely irrational.


Facebook/Meta is a great example indeed when one remembers that they took part in a genocide, and still haven't faced consequences for it.


There's the hope that they'll make all employees redundant and that businesses that adopted AI quickly will be able to exist as a pure money-printing entity without the liabilities of products, customers, or employees.

Put another way, LLMs are just another part of enshittification.


Since they can summarize, you'd really think they would be better at condensing their own output and cutting out the filler after. I wonder if you could use a specialized second pass for it or something?


In my recent experience it seems to be a contextual issue. Llms are constantly being dropped into new situations where they have to rediscover high level and cross cutting information about the system they're working on from contextual clues.

The internal representations of this state and its projection back out to human language wouldn't be as concise as that of a practitioner or team that develop their own verbiage and ontology over time molded to their system.

This verbosity might get better as we figure out better ways for agents to learn long term and use that knowledge to adapt to the users and projects over time.

There might also be some good harness improvements we could consider like forked output streams or multiple long lived filter subagents to ensure that output appropriate for thinking is separate from code output and separate from output given to the user driving the session.


Prompting is a skill.


This is just a problem of articulating what you want. Most people can't write for shit, I'm sorry. Because they maybe took one writing class in college. It's a skill in it's own. LLMs write better than anyone I've ever worked with and the reasons are obvious. If it rambles or is too "purple", that's on the person prompting it, because they can't write for shit to begin with.


Then why is it that ALL content generated by LLMs, no matter who prompted it, is of such low quality?


> LLMs write better than anyone I've ever worked with

If you're talking about prose, this is obviously false. All the LLM prose I see is so mentally fatiguing to read that I usually just give up.

There are no lack of examples of poorly written LLM prose.


Right, because of the person promoting it.


> LLMs write better than anyone I've ever worked with

And yet every time I see AI-written documentation, there's paragraphs worth of throat clearing and space-wasters like the word "genuinely" for things that could have been outlined in a couple of bullet points.


Right, that's on the prompter being ok with that. It does what you tell it to do. I don't understand how you dorks on this site are so obtuse about this shit. Give it examples of what you want, authors, etc. Or don't and just cry. The world is your oyster.


I'd much rather see badly written text that someone invested time into instead of perfectly written and infinitely soulless AI slop


Lol why, this is just silly


I’ve started requiring that unless absolutely necessary, all comments and docstrings be one-liners in our code 10k+ codebase(s). LLM-generated code has a tendency to “over-justify”, which is reasonable when first reviewing the code as a human.

But then, once it’s passed the first human it should be considered human-to-human communication rather than LLM-to-human communication, and the long, winding, and often repeated justifications can be significantly condensed.

Also you can’t let it run loose with unit tests unless you want to be blasted in the face with hundreds of lines of absurd test fixture preparation.


> LLM-generated code has a tendency to “over-justify”

This is the thing that drives my nuts about LLM-generated code. I will see PRs that fix a bug and the entire bug fix is re-explained in 10 different places in a code comment.


I like judgement. Taste is a subset.


For me taste is judgement about things that don't matter. Like you have taste in clothes. But a judge doesn't have bad taste when he makes a wrong decision, he has bad judgement.


I think there are really two types of use cases for LLMs here, and people that get value out of them.

First are those that care very little about the quality of the work, or the process behind it, and just want a 'finished' product no matter what. These are many of the folks vibe-coding everything without even looking at the output, and a depressingly number of those getting hacked or what not.

But hey, if you've got an idea for a complex app or website but no interest in actually building it, an LLM will get you... something vaguely like what you wanted.

Second are small projects and small sections of existing ones, most of which easily fit inside the LLM's context window. If you're making a fairly basic WordPress plugin or React component or one page website, an LLM will probably be able to handle it just fine. Heck, it might not even look all that different from what a human might have created code wise.

One of the big issues we have those is that plenty of people and companies are using these tools despite having requirements that aren't met by them in the slightest. If you're working on a Google/Microsoft/Meta scale product where performance and security and code quality are of the utmost importance, then an LLM isn't going to be a great fit. Unfortunately, that's exactly where this technology is being used, and the results are becoming more and more obvious.


I get this is a cheap response but “you’re doing it wrong”. Early last year they were shite. Then they became good enough to write tests. Then they became good enough to write code. Currently they’re good enough to design APIs for “typical” applications. Are they good enough for vibecoding a production app? No. Not yet. Don’t let that confuse you. Many of us are using them effectively, and learning what works and what doesn’t. We’re also learning how to evaluate a new model: can it do more than the last one? Does it need the same level of detail as the last one? Can we go faster?


This is true but also false.

In my experience (scientific programming) AI is a giant multiplier for people with specialized knowledge.

But it’s also a giant devaluer for that same knowledge as people with no idea what they’re doing can clog the field with plausible bullshit.

It’s now the case that if someone tells me they’ve done something, and I look into it and find out it’s completely AI slop, then I will have spent more time on the project than the person who “made” it. The situation is completely untenable and only serves to drain time and resources from people with better things to do.


we are slowly punishing reading comprehension

this will have educational consequences (that I'm trying to solve). I don't think that we can adjust without rapid education and making extreme specialists of us all.

This requires coordination, certification, licensing, and other tiers of authenticity. False experts can ruin sample gathering, can ruin training. False expertise is exemplified by the current American Administration. Look at Robert F. Kennedy Jr.; he's a false expert. He is responsible for the measles outbreak. He is responsible for ivermectin abuse by humans. False expertise is overtaking real expertise. And the results are continuously disastrous and large-scale.


Or in a way, nothing at all.


On its own, yes, but not in context. I think one of the problems with Claude's stock output is it assumes the reader has a firm grip on the context of the output, which is often untrue. Stock output is exhausting to read because you have to unwind metaphors in an unstated context.


What I’m saying is that it often uses metaphors in ways that are subtly wrong or don’t make any sense.


I think that’s just a fundamental tradeoff though.

Being able to run Jupyter cells independently is a feature until it’s not.

I’d say for most of my one-off work, it’s fine. But for stuff I want to share it’s not.


Yes I understand it's a tradeoff, but I'm saying I would prefer a different tradeoff. It would be more intuitive to me if running a cell always invalidated the cells below; this would make variable reassignment unambiguous, just like in a script, but still with all the visualization goodies of a notebook.


I just dont think LLMs are very good at judging importance or summarizing code.

I tried experimenting with what is ultimately a treesitter based approach - https://github.com/0x007BA7/codebook

And really liked it. Definitely nowhere near production ready but I think theres room for a player to come in and do something similar.


yeah that was another thing i hoped would pour through here - that deterministic systems are much better for evaluating quality (test, linters, cyclomatic complexity, etc) - but that we don't have such a system for code maintainability, at least not one that's widely accepted or adopted


I think anthropic with its enterprise strategy and google with its integration in everything have a bit of a moat.

But I switched from ChatGPT to Claude 3 months ago because my account was down for like 6 hours. I haven’t used it since. It’s too easy to switch away from chatbots on a whim. There is no moat for that.


> I think anthropic with its enterprise strategy and google with its integration in everything have a bit of a moat.

But... Anthropic doesn't have a moat. It's clear at this point that SOTA models are not a moat, and Opus 4.6-level (or GLM 5.2) is sufficient.

Google, though... they own the entire vertical, from the semiconductors to the end-user software. They may have a moat.


The narrative that superintelligence is imminent is partially at fault here.

There are competing definitions of what intelligence even is, and the one that I find most striking is from Francois Chollet which is that intelligence can be boiled down to skill acquisition efficiency. This type of definition makes intelligence more akin to polishing a ball than growing a watermelon.

The superintelligence doomers warn that the watermelon is going to start growing exponentially and crush everyone. But what might actually be happening is that we are not growing a watermelon but rather polishing the ball until its really smooth and shiny. There's a point where you can get it to micron levels of polish but for most tasks (white collar text domains tasks), it's smooth enough! You will be able to go to the ball store and buy a low cost made in china ball for most tasks.

The real challenge is actually branching out domains and modalities to tackle things like blue collar labor. Over time, white collar work automatable or able to be made hyperefficient by LLMs will see LLM commoditization.


Observationally, for people that /aren't/ using models to code but to just do their white-collar job, claude.ai /is/ AI, now. The entire perspective for how to use AI is through claude skills, claude projects, claude cowork, etc. They've massively won the corp buy-in at the moment I believe.


Fully agree, that’s how I see most of my less technical coworkers reason about using AI. “Is there a Claude skill for that?” is a question I hear multiple times a week.

I see a lot of comments (incl 2 sibling comments) are always discussing whether there is a moat on AI. We agree there is no technical moat, there is nothing that Anthropic or any other AI lab could do that wouldn’t be quickly offered by other competitors too. However, _market penetration is the moat_. The deeper Anthropic is in relationships with orgs, the higher the cost of switching. Sure, for an individual it’s a 2-second job, but for a business it’s actual work of changing permissions, provisions, updating vendors etc. Nothing catastrophic but still real work, implying that Anthropic would need to drop the ball significantly to be swapped out, and wouldn’t be just because there’s a competitor who does everything kinda the same for kinda the same price (or slightly less).

Moat can be in execution and not just in technology. McDonalds has no specific burger-making technology that no other restaurant can acquire, what they do have is a well-scaled execution. Sure, there are competitors, and sure there are new comers to the burger space with different recipes (e.g. smash) but that doesn’t mean McDonalds is going under. Market penetration is the moat.


People always underestimate the moat that is institutional agility/ossification. There may be no theoretical moat in software, but every time you switch HR or financial systems you find out the hard way just how many integrations have been built (directly or indirectly) around a specific way of thinking about organizational data imposed by the software containing it, all of which require significant rewrites for the new system's conceptualization. The costs of switching add up fast and the risks of delayed payments/paychecks significant, which is how you end up with companies paying $$$ for extended support on discontinued HR and financial software.

It's the same reason companies will pay for GSuite and O365 subscriptions concurrently because a handful of departments refuse to give up their desktop applications and others have Excel-based workflows that do not work in Google Sheets and do not care enough to learn Python or find someone in the company who can write a better system.


> The entire perspective for how to use AI is through claude skills, claude projects, claude cowork, etc

But as they have repeatedly pointed out, creating software is almost zero-cost now, so software cannot be a moat.

After all, all of the Claude software can be vibe-coded by any competitor; that's the dream that Anthropic has been selling anyway...


doesn't matter. that just means they've incentivized all competitors to enter the market and let's be honest none of their tools are that novel.

https://www.youtube.com/watch?v=2J2Fb1bBufA


I guess I’m thinking a lot of companies seem to be getting Claude code subscriptions. It usually takes some time and effort for an org to switch away from one solution. In the meantime a lot of workflows get more and more tied to Claude in particular.

It’s not much of a moat, but it’s more than a lot of orgs have.


obligatory correction: the semiconductor layer is still owned by TSMC and Samsung. Google sketches chip designs for them to implement - that's the lowest layer they control. I am not denying that this is impressive.


google might have tons of integration. But if it invested too heavily into AI then it will also suffer when increased competition causes returns to fall:

https://www.youtube.com/watch?v=2J2Fb1bBufA


I posted about my project below - https://github.com/0x007BA7/codebook

But maybe you should checkout the tools it’s based on, sem - https://github.com/Ataraxy-Labs/sem and ultimately treesitter. They at least give a more structured approach to dealing with code than simple text.


This is great stuff; I've been prototyping with a few language-specific parsers like the Golang and the polyglot approach looks really helpful for me


Yeah I hope the links are helpful


Codebook - https://github.com/0x007BA7/codebook

It’s a better code reader built on top of sem (treesitter). I’m getting a lot of massive PRs at work now, and this has helped a lot with reading them. It decomposes the changes into entities and sorts based on what has the most dependencies. This tends to put the most important functions first. Plus I can click through the dependencies for each function and mark things as reviewed as I’m reading them. It’s a big improvement over the GitHub review flow for me at least.


this looks incredibly useful! I've also run into similar problems with code review and been building similar tooling with the same idea (reconstructing the changes into entities, and finding focal/important changes), but haven't gone as far as this.


This has always been the fun part of programming to me. I know most people hate it, but I really don’t mind being on-call (ok I hate being woken up) and fixing weird bugs that users run into. All these small edge cases that people run into because reality is odd. Of course I’m in scientific programming so that probably colors my view.

It’s always a little disappointing to me when I think I’ve run into something unique but it ends up being user error or something.


I echo this. The kind of entropy that real users bring has been refreshing to face as a founder.

Being a founder has a lot of SRE like activities. Fortunately I used to actually like troubleshooting and hence love being a founder but I know a lot of people quit this path because of the "suprising amount of details" in reality!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: