A lot of the low level stuff is outsourced to biochemistry: the physical properties of proteins, and the various self-regulating biochemical systems of an animal, can "encode" a lot of intelligence, easing up on the computational demands of the brain proper.
The reason this is a thing is that other people can bid on your product (Amazon has an entire sponsored product category for this reason). So the reason for Seth Godin to bid on ads for "Seth Godin The Knot" is to try to crowd out other people bidding on the same term (eg another book on the same topic). Can be smart but you don't have to play this game if you don't want to, and the ACOS will tell you if it's worth playing. Even if you don't bid on those ads your own product will be the first result. It's just whether you want to bid up the price for competitors -- if you don't bid on your own product someone else can come in and buy the ad spots cheaply.
Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).
It seems to me like he started out mad and looked to justify it.
I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists.
If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference.
But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
It says they observed no difference in people clicking thumbs up or down. There are loads of other behavior that they didn't observe; like, say, switching to a different LLM.
The question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable.
The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
Isn't "which one is watermarked?" a different question than "which one is better?"
"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)
Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
That's true, but you don't typically explain what you're testing in this sort of (presumably) randomised trial.
And the Daring Fireball article does complain that watermarking will reduce quality. If that's what you're trying to check, "which is better?" is the right question.
Agreed, if I simply didn't like the style or words an AI was using in something it wrote, I would switch to a competitors and see what it can come with. I probably wouldn't hit the thumbs down on the Gemini response as it's not that the response is wrong, I just didn't like it. I usually reserve the thumb down for when the AI is wrong.
Also, depending on what I am asking it, I often don't want to use the thumb down or up, as this may mean my conversation is going to have some kind of human review and depending on what I am asking for, I may not want to bring attention to my stuff.
> It seems to me like he started out mad and looked to justify it.
Yes, but that's neither surprising nor a reason to dismiss the anger. People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. They're angry -- and Gruber acknowledges that factor too -- because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It might be another instance of consequentialism vs. honor ethics. Many consequentialists don't seem to understand that something that doesn't have demonstrable consequences can still have moral implications.
Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next.
A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1]
Done right you won't know the difference, done badly and you will.
Flac takes a raw .wav and effectively zips it up to shave off a certain amount of space. (there are nuances, I think the compression scheme is designed for streaming.)
mp3 is perceptual, so throws away the stuff that humans can't hear. This yields a much smaller file.
However its all a sliding scale like PNG vs jpeg.
a .jpg with a quality setting of 85 will be almost identical to a .png in visual quality. However if you then edit that jpeg, the image degrades and you start to see artifacts. (hence why memes look like shite as they get older)
Its the same with mp3s if you compress the hell out of them, say 64kbit or lower adaptive, then you'll start to hear the tell tail "schlop" noise of mp3-like compression. You might notice it most with cymbals in drum kits. cymbals are wideband noise. as in there are loads of constituent frequencies so if you remove some of the "hidden" frequencies you tend to notice, so they sound more metallic, ironically.
But, all of this is solvable, 256+kbit is more than enough, bonus points for higher sampling frequencies. (however you need a decoder that can actually do that sample rate...)
Could be c't magazine accidentally played the mp3 version louder. Human's have a known preference for louder music, and will tend to prefer louder samples over quieter samples. Rumor in the industry is that this was a trick MS used to try to push the WMA format, that they encoded some WMA samples used in some publicized tests at +3dB above the source sample.
I get what you're saying, but I think it's ridiculous for people to think of LLM services generating text as either their own domain or something that they own.
To the extent that it's anybody's, it's either Anthropic's (they run the service) or everybody's (in that we created the content it's remixing). Legally LLM prose isn't copyrightable for good reason.
Even if honor were real, LLMs do not have honor, and people using LLMs to write without disclosing that fact (or indeed at all, to some purists) do not have honor either.
> because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It's LLM output! It's not your domain, it's the LLM owner's!
I care about optimal word choice when generating LLM texts. Because my use case is almost exclusively reading the generated text not posting it. I use LLMs to summarize, translate and review other texts. When using LLMs in that way, as a research tool watermarking is a pointless and should not get in the way of "optimal" results.
Again, given the limits of LLMs (stochastic, rapidly changing, everything's a hallucination, widely known prose issues) I am skeptical that you really care that much about optimal prose. I could believe it's one of the things that you care about, but at a pretty low priority level.
Taking you at your word, though, I'd be interested to see what you think of the watermarking technology in a blind A/B test.
What's your evidence that it will result in worse translations?
I'm skeptical that such a thing as a universally optimal translation exists in cases beyond the trivial. But if it does, I see no reason to think LLMs are anywhere close to it, so I think nobody will be able to tell the difference with watermarking.
That's certainly true for code. LLM code is at best mediocre. There is oceans of room to subtly watermark generated code without practical impact.
I assume they are just passing off AI prose as their own and don't want anyone to be able to tell. Which is surprising for someone who's been blogging for a thousand years. But I don't really see any other reason for this amount of heat and FUD.
From what I read, two big problems for long-running columnists are getting tire of the work and running out of things to say. I have no knowledge of Gruber, but I can certainly see why people who are expected regularly to have something to say would turn to "AI". It doesn't get tired and is always ready to spew infinite words.
I think in this case it doesn't help that there are multiple watermarking schemes, and the easiest for people to understand is the red/green scheme by Kirchenbauer et al. (https://arxiv.org/pdf/2301.10226), which does technically distort the logits (but I'd argue only in cases where you wouldn't notice it anyway).
I wasn't aware of this gumbel softmax scheme, it seems you're referring to https://simons.berkeley.edu/talks/scott-aaronson-ut-austin-o... ? That's really clever as it doesn't even distort the logits, basically cryptographically indistinguishable from a "real" random sample unless you have the key.
The actual scheme Claude uses seems to be neither of those two though, they say it is SynthId-text which seems to be tournament sampling based.
They do claim that the per-token output distribution remains unchanged, but the proof is relegated to Appendix B.1. The perplexity comparison includes methods that do change the output distribution.
If this is true then the probability of the detection tools flagging completely human generated text as AI generated is non-trivial. Let's say I write a completely original piece and the detection tool says there is a 36% probability it was generated with Claude. What then? Now it's up to the person looking at the score to cast a subjective judgement. Maybe to me, anything over 25% is unacceptable. Maybe to someone else, it must cross over the 50% threshold. This is the problem.
> the probability of the detection tools flagging completely human generated text as AI generated is non-trivial
How does that follow? AI-generated text is already not a perfect emulation of human writing. There's lots of room to affect it laterally without changing the level of quality.
As I understand it, LLMs with temperature >0 can select from many possible outputs. All they're doing is limiting the possible outputs to ones that contain this pattern. I don't see any reason why the quality of that subset should be lower than average. The very best outputs will likely be eliminated, but so will the very worst.
To sample from the probability distribution you already need random numbers.
If you get your random numbers from a cryptographic PRNG, then to notice the difference between that and 'real' random numbers even in theory, means you need to break the cryptography. In practice, your gut feeling about how good some text is won't break modern cryptography.
Yeah. And the detector can’t even tell you “X% chance this is watermarked” because it doesn’t know the input distribution. It can only tell you “Y% chance that an unwatermakred text would score this high” and how many history professors understand Bayes rule well enough to understand the distinction?
Worse, what will academic institutions decide is the threshold for detecting AI generated work. If you have a false positive how do you prove it was a false positive or we all just trust the watermark detector over the student saying "I swear I did it all by my self"
I don't think that's true. I think it's a binary 0% or near 100% probability of a watermark having been detected; the more changes to the text having been made after the text was output by the LLM and the less leeway the LLM had for probable word choices, the longer the passage necessary to see it.
The "problem" is that seeing the watermark doesn't mean that the person claiming to be the author didn't make extensive changes to the output of the LLM, or that the LLM wasn't simply the final editor of something that the author had put a lot of work into.
> Cognitive surrender.
I don't know what this means. It's just drama. Don't let the LLM write for you and this is not a worry. I'm not worried about the poetry of LLM output being subtly adulterated.
No, watermark detection is not binary, you get a real number. You decide on a threshold when looking for the watermark. This is the problem - by random chance, some human text will be detected as watermarked. You can turn the detection threshold up until it guarantees <0.001 false positive rate at the expense of higher false negatives, but seems inevitable that someone gets wrongly flagged.
It's a writer perspective versus a reader perspective maybe?
Sometimes when you're trying to write something, it really seems like the exact words matter a lot. Suggestions made to be more direct or use a more common word here or whatever seem to really impact the thought that you're trying to communicate.
Certainly we've all had times when trying to communicate clearly when the specific words seem very important.
> The very fact that there is generally no "best next token" with 100% certainty
This is not entirely accurate. Sure, there's never a token with 100% certainty, but there are often tokens with 99.9% probability, but this technique of course does not change how such a token is sampled.
By definition watermarking narrows and biases the response distribution. Clever algorithms might reduce the perceptual impact and minimize some cherry picked metrics, but it's still worse.
> The very fact that there is generally no "best next token" with 100% certainty
Indeed.
It's frankly bizarre to see the assumption to the contrary being made by someone who's been passionately blogging by hand for years, who also happens to be responsible for the notoriously vague, humanistic, DWIMmy Markdown standard.
I think he’s still generally good on business, UX, and hardware design. That’s all subjective and taste I suppose, but his taste works for me.
On deeper tech stuff, like this utterly nonsensical misunderstanding of watermarks… yeah, classic case of a guy who is smart, and has lost the ability to realize when they’re not knowledgeable in a domain.
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
I think it depends on whatever contract you signed. If you signed a contract that says “you pay per minute of screen time but only get the end result” then I bet that if you went to court demanding the screen recording, you’d lose.
The default is set for the marginal new user, which at this point is probably not someone like you (who benefits a lot from manual mode) -- it's someone who's more "code-naive" and might get anxious about approving random bash script commands they don't recognize. Safely getting the user from prompt --> first vibe-coded app is the "user journey" now, and since auto mode seems pretty good at not letting Claude rm -rf'ing the home directory, this is 100% the right business move. For people who know what they're doing (like you), manual mode is just a shift-tab away
I recall hearing similar sentiments from linux sysadmins regarding cloud infrastructure. In many respects they were and continue to be correct. In other respects, the world doesn’t care about the loss in understanding as long as things work “well enough” for the cogs of society to keep turning.
For those who do care (and have the aptitude) to understand things deeper there is always work to be had when “well enough” stops being good enough and someone has to unravel the “RDS queries are taking too long” problems that crop up as a result.
These progressively higher levels of abstraction are how everything has moved and will move in science and human technology. There aren't enough hours in the day nor years in the human lifespan to gain a deep understanding of every single level below us. Rather we build on the abstracted API layer beneath us, and those who come after will build on top of us using a simplified abstraction to hide the tangled mess we had to make.
What an awful position to take. Tech used to be about becoming more accessible to people! Now we have a magical assistant to make computers do what you want with natural language, and your desire is gate keeping that so only programmers can use it to write software for themselves?
>so only programmers can use it to write software for themselves?
Yes?
The idea of an assistant that can use natural language is nice! But why would you MAKE software with it, it IS software, just do the thing you want to do! If you want to make an app, be prepared to jump hoops because this is no longer about YOU the user, it's about OTHER users.
The idea of making personal single user software is a fantasy, an oxymoron, you MAKE software? there's the presumption that it will be used for other people, otherwise you'd be USING software. There's a counter and you are at either one side or the other. It's the difference between making yourself a sandwich vs making a pot pie vs making chicken nuggets. One has the form factor for individual consumption and the other has the form factor for a social gathering, and the latter is an industrial form factor.
Maybe if there were a magic microwave that created random foods from thin air, people would create chicken nuggets or pot pies for themselves, but it's a vestigial maladaptation that will soon dissapear. Any reasonably designed product would try to provide different UX for industrial and individual users. The magic microwave that makes chicken nuggets better not be the same one that an actual factory is using. It's not a matter of cutting the middleman and revolutionizing wealth distribution from those fat chicken-nugget cats, it's about having two distinct products for two distinct usecases.
Tl;dr: Personal and industrial usecases are different, and if I'm in the industry, I don't want to use (the same product that end-users are using) to build products. What a clusterfuck.
> The idea of making personal single user software is a fantasy, an oxymoron, you MAKE software? there's the presumption that it will be used for other people, otherwise you'd be USING software.
It definitely isn't. I've done it (successfully) a few times.
> The idea of an assistant that can use natural language is nice! But why would you MAKE software with it, it IS software, just do the thing you want to do!
This makes no sense to me. Are you suggesting that instead of using an LLM to make, say, an ebook reader or crossword app that meets my personal needs, I should invoke an LLM every time I want to read a book or do a crossword? That feels like a strawman, but I can't work out what else you might be arguing here.
> I should invoke an LLM every time I want to read a book or do a crossword.
Well no, I'd say, use any of the existing 100 book readers to read a book, or any of 50 crossword apps.
But if you want to use an LLM to customize it the exact way you want it. Yes, use the LLM every time you want to do it. Wanting the LLM to do a previous gen app and then get out of the way sounds like asking for faster horses. Just tell the LLM you want to play a crossword game with X and Y rules, then give it a name so you can play it again in the future if you want, and if you want to try out rule Z, you do that, a la Kay's Dynabook.
That may sound weird to you, but asking an LLM to 'make software like we used to' sounds weird to me, and seeing "invoking an LLM every time I want to do X" as weird sounds like something that's true only for a very brief period of time where something is so uncommon that it's inference is expensive.
I can imagine a future where it might make sense. Right now, though, it would make for a far worse experience, and wouldn't even really be practically possible.
Both of those examples were real ones. The crossword app runs on my phone, pulls the crosswords from a specific source, and lets me access and solve them via the exact interface I prefer. The ebook app is cross platform, syncs via a remote server, has the interface I want and includes some niche features. There's no realistic way to create that on the fly every time I want to read a book on my phone, and if there were it would be extremely inefficient.
And I see literally no advantages to doing so -- even if time and tokens weren't an issue, what would I gain by recreating the apps from a prompt every time I wanted to use them, rather than deterministically running code I have already tested?
I get that you are doing that, but I think the fantasy is that this is creation of software instead of consumption. It's an issue adjacent to licence washing, where mangling some code through an inference layer is considered transformative or even unrelated and the original license doesn't apply. But in this case, what you are washing is not the license, but the valor of writing software.
Broadly speaking, your approach would be to have the LLM write application code, my approach would be for the LLM to write commands, 'apt-get install calibre', maybe if I want to add or modify a button I can ask it to hack the X interface. You go straight for the LLM generating the code. There's certainly technical differences between what we are doing, but they are very arbitrary, we are essentially doing the same thing, but what I am doing looks less impressive, and what you are doing you can sell in your CV to potential hiring managers as 'using AI to write software'. It's more about the semantics than the actual requirements.
I may be wrong though, maybe your approach is far more effective than just importing transitive dependencies, but I would think it's more about taking credit for the thing and increasing your sense of ownership and achievement. Sorry if that sounds harsh, but I just need a way to think of myself as better than others as an unemployed neverviber.
Thanks, I think I understand your point better now. In this case, though, you are wrong about both my intentions and the relative practical value of the two approaches (to me).
I'm not doing this to take creative or intellectual credit in any external way; you're right that there is some degree of increased personal satisfaction (which I don't see as a problem, as long as it doesn't crowd out more wholesome ways of 'earning' that satisfaction), but I'm not kidding myself about what I've actually done here. I also write my own code for fun/creative expression/intellectual stimulation/showing off, but that's a separate thing and there's not much crossover between the two types of project for me.
And the end products really are useful to me in a way that I couldn't replicate just by using something that already exists, and couldn't replicate nearly as easily by manually modifying open source. (I'm sure I could do it by starting with open source and using an LLM to make changes, and in other cases I have done exactly that, but at that point I don't really see the conceptual difference -- I'm still getting an LLM to write code and then repeatedly running that code. Ideally I would be giving something back by making a useful contribution to the public repo(s), but that would turn this into a completely different, more tedious and effortful thing, and I'm not sure it would be welcome anyway. So, case by case, I choose whichever approach seems likely to be more effective or efficient or less annoying, and sometimes that means getting Claude to write something 'from scratch'; other times there's an open source application I already use that just needs some tweaking, and I start with that.)
>couldn't replicate nearly as easily by manually modifying open source. (I'm sure I could do it by starting with open source and using an LLM to make changes, and in other cases I have done exactly that, but at that point I don't really see the conceptual difference -- I'm still getting an LLM to write code and then repeatedly running that code
Yes, that's what I meant, vibecoding something that uses existing software, not manually using open source stuff. It doesn't even need to fork or modify code.
>at that point I don't really see the conceptual difference -- I'm still getting an LLM to write code and then repeatedly running that code
I do agree, it's a subtle difference, about importing higher level dependencies vs building on top of low lever abstractions and writing everything else. Which is ironic/nuanced because I'm a huge proponent of aggressively not using dependencies in industrial programming, to the point where my requirements.txt/package.json is literally empty, and I use POSIX compliant sockets syscalls instead of importing packages like flask or express.
But when it comes to actually using software, whether for personal usecases, or as a sysadmin, my approach takes the opposite form, I aggressively don't write code, I still aggressively minimize dependencies, but the game is actually using the Operating System primitives to combine these dependencies, relying only on OS installers like apt/yum, maybe minimal configuration, if code is written, it's on a scripting capacity, a bash or python script, glue code you know? Sure the line can be fuzzed at some point, but it's clear to me that you can either write an application or be a poweruser of an application, and early in my career I've seen businesses go for the building software in house route for the fun factor, I don't think that was the right answer in the dot com boom, and as time went by, and the corpus of software grew, building your own became even more wrong than using existing third party products.
Now in personal computing, the fun-factor maybe is more important, but I have to judge this personal-software thing on how it will affect the actual important stuff, because that's what the stakes are, and that's how it's being sold. In the industry code agents are either used for building software, or for consuming software, and outside of personal experimentation I still hold that building your own software isn't a good idea, I don't think the advent of LLM materially changes that, the consequence of ending up with an ossified, non standard, low quality product is still there, perhaps even magnified, it's just that it's not something that you notice when you are starting a software product, it's only when your pyramid reaches a couple of hundred meters high that you realize that it can't grow into a skyscraper.
And I get that not everything needs to be a skyscraper, but it feels like one-off software is taking the form factor and tooling of long-term skyscrapers, the logical consequence is that we would end up with thousands of little skyscrapers, which is a place we can only get to by ignorance of the history and essence of skyscrapers, it's something a city child would imagine after going on a road trip once, "what if we had little skyscrapers throughout the whole country instead of very high skyscrapers in a single place?".
> Tech used to be about becoming more accessible to people
Since when? Because funny story, I only ever hear that narrative from tech people trying to put a self-serving spin on whatever egregious thing they want to impose on others. For the past few decades, the tech industry has consistently acted to wrestle control of people's own lives and place it in the hands of the few. The justification is always the same. It's about "keeping people safe" or "making tech more accessible." People are sick and tired of this, which is demonstrated by public backlash against tech.
> your desire is gate keeping that so only programmers can use it to write software for themselves
In what universe is learning "gate keeping"? A sane society doesn't criticize people for asking drivers to learn how to drive. What GP is asking for in the case of software development is much less than a driver's license, and yet you question their ethics.
If anything, you're the one trying to rob people of their opportunity to learn, which is a prerequisite to making informed decisions. You're the one advocating that we surrender control over computing to a handful of trillion dollar companies. That is an awful position to take.
Incredibly naive, AI is not magic and it won’t always do what you want or expect. Not sure how you’ve determined that I’m gate keeping when I’m simply warning that tech illiteracy can get you in trouble if you start giving mystery black boxes that sometimes call themselves “Mecha Hitler” root access on your machine.
I am using many many many things that I don't understand. Cars, public transport, etc...
I review and test the end product, not every tiny step along the way. If the LLM uses some command line tools I have never heard of to create a model I can verify, why should I learn a tool that is completely irrelevant to my core expertise?
"If the LLM uses some command line tools I have never heard of to create a model I can verify, why should I learn a tool that is completely irrelevant to my core expertise?"
Because that command might also give someone else access to your computer along the way. So your tool seems to work, but your computer is owned.
Are you jacking into your car’s OBD port and messing with the engine timing? There’s a difference between using a known tool in a controlled way and giving the tool to a hallucinatory goblin (or overconfident intern) with the directive “do it for me”. You don’t have to know everything about the tool to know what it does, or to tell if it’s doing something it shouldn’t (like malware). What happened to the old advice given to tech neophytes, “don’t run random scripts from the internet if you don’t know what they do”?
Just continuing the car example, because you actually do need to know enough about it to operate it and get your license. There’s no license, no insurance for using AI. Driving can be deadly, and when you do it on a public roadway there are certain requirements that have to be met as agreed upon by most governments. Same should go for AI. Do whatever you want with it on your machine, but when you let it out on the internet, you become liable for any damage you, and by extension your AI, may cause. If you feel uncomfortable approving its actions because you don’t understand them, you should either a) take the time to understand, then approve, or b) listen to your discomfort and don’t do the thing. Frankly I think that’s more accessible because you then see the decision points instead of leaving Oz behind the curtain. Much easier for a neophyte to learn from that instead of trying to reverse engineer a final output.
Many many people care more than the end product, for example whether a shirt is made of cotton with the forced labor, carbon emissions of public transport, etc.
In terms of engineering software, you care the cost. An intelligent agent may try to read unnecessary files and it's time to stop it to save tokens and avoid polluting the context.
He didn't advocate for being completely blind in every way. You might care about working conditions without understanding how the textiles, dyes or cotton production works.
These non-programmers probably shouldnt use computers at all, right, since they don't understand them?
I doubt that it is more efficient for someone to routinely watch every line of output or stop to review terminal commands, rather than waiting for the turn to complete.
It is a broader debate about agentic AI, and whether one should relinquish control to the tool rather than aim for full understanding of every action taken.
The people arguing for a hands-on, fully in control approach are losing ground by the week, in my opinion.
It’s definitely on the high friction side of the usability/security tradeoff, I just think auto modes like this are as liability-inducing as handing a script kiddie intern full admin on your production environment
The real answer is somewhere in the middle and is probably a mix of traditional AV/EDR and AI QA judges that mitigate risk of running more or less random arbitrary code and auto approve based on configured detection rules and your personal risk tolerance. Would it suck to stick an EDR sensor in every code execution environment spun up for an agent to run a python script… yes. It would also suck if you were responsible for hacking a company without knowing about it because you didn’t watch what your AI was doing
I can be against animal testing without being a chemist or having a full understanding of the experiments being made on them. Knowing it's cruelty is enough to make opposition a valid and defensible position.
I don't think the marginal new user is anxious about approving messages - I think they're quickly annoyed by permissions prompt they don't understand and quickly get in the habit of approving everything or figuring out how to set bypass permissions on
I am actually curious, how much non programmers use claude now. I know just one and she really does not know much about computers, I suppose their numbers will grow (but I doubt most get much value out of it).
I mean if you don't care code, you are essentially a product manager who gives instructions to your programmers (whether humans or intelligent agents).
Then if you use the created product, you are at best a test engineer if not just an ordinary user.
I think in the era of AI, people get tools they want in an expensive way. Rather than finding an existing tool, they ask an intelligent agent to parrot one, which guarantees no safety, security, efficiency, and accuracy. Yet, being able to use Claude makes them feel smart and productive (in parroting wheels).
Surely the two leading candidates must be (a) the model is just not that good, or (b) it is misaligned in a pretty obvious way that can't be swept under the rug.
Sure, but the question is why either of those things happened, given Google's immense resources and talent.
The fact that OpenAI, Anthropic, SpaceXAI and 3 different Chinese companies were all able to train big models without these issues, yet Google could not, seems shocking.
They already quietly agreed to "all lawful use" with the Pentagon, to no real fanfare. Gemini Slaughterbot Edition, coming soon to a DHS facility near you?
I don't think so. I've used Claude to generate 3D animations in python purely by making and manipulating raw meshes, and I've had great results. A more likely explanation is that 3D graphics is just pretty straightforward matrix algebra, and models (or Claude specifically) has that down cold.
Why not? Slowing down / halting biological weapons research mostly worked. Yes there have probably been modest sized defections here and there, but the pace of bioweapon development is a crawl compared with (a) what is possible with science already, and also with (b) the pace of bioweapons development from 1910 to 1970.
Bioweapons are good for .. war. Agentic AI is good for changing your economy, strategic advantage, labor explosion and making more effective Agentic AI. They are not the same category of thing.
reply