Hacker Newsnew | past | comments | ask | show | jobs | submit | cjbarber's commentslogin

From Jeff's twitter post:

> Our general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.

See also: https://www.nae.edu/20782/grand-challenges-project

Those 14 are:

NAE Grand Challenges for Engineering

1. Make Solar Energy Economical

2. Provide Energy from Fusion

3. Develop Carbon Sequestration Methods

4. Manage the Nitrogen Cycle

5. Provide Access to Clean Water

6. Restore and Improve Urban Infrastructure

7. Advance Health Informatics

8. Engineer Better Medicines

9. Reverse Engineer the Brain

10. Prevent Nuclear Terror

11. Secure Cyberspace

12. Enhance Virtual Reality

13. Advance Personalized Learning

14. Engineer the Tools of Scientific Discovery


They make many bold promises, but their core goal is neatly encapsulated on the website:

"Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today."

This is the same goal as every other AI company out there. Automate away the human employees and let a small number of "people" (note that they do not say scientists or engineers for this part) take the credit and financial rewards for every good thing this human-free system produces.


I don’t understand why all of these companies need to frame it like this. Why not a large amount of people doing an even larger amount of science?

I love this question. If history is a guide, some version of that is what will happen (employment doesn't seem to go down, despite technological progress and population + economic growth).

I suspect the root cause is that it's harder and scarier to imagine good outcomes. It exposes us to disappointment, and when you do it publicly, it looks "crazy".

Another explanation is that there's no direct consumer for "more." Individuals, corporations and states are not in themselves interested in "a larger amount of science," or anything analogous, despite the fact that they would all benefit ambiently.

The net effect is that it's only "safe" to claim reduced risk (i.e. lower costs).


I know some scientists and there's barely enough demand for them as it is

I suspect we’re at the stage of technological civilizational development where, especially with LLM-assisted gradient descent seeking upon the results, basic science, research and engineering for the pure sake of establishing search space beacons of what is found to be true and what is not, irrespective of immediate industrial applications payoff, are valuable economic inputs in and of themselves into ever-expanding training corpus. It has never been easier for people in different fields to now search knowledge spaces in LLM’s, for applicability to their problem spaces of discoveries in seemingly unrelated spaces.

From my perspective, we are desperately short of scientists, researchers and engineers, but we are using an outdated economic model to leverage their findings. LLM’s are a large part of Bush’s Memex and Jobs’ bicycle for the mind visions for intelligence amplification, and in some ways exceed them. I hope we trampoline from how we currently use basic seeking efforts for knowledge.


I hope the demand for average-skilled people goes down.

Or that the floor is raised, at least, and AI empowers average scientists to do substantial work.


Is the demand low because the capital requirements to perform the research aren't there? Excluding wages...granted, for an average research project I don't know what % goes to wages and admin overhead. Probably the bulk?

Isn't that directionally what OpenAI's core message was for a while?

A large amount of people building their own customized apps for themselves.

Similarily, everyone becoming their own accountant / lawyer / other professional services.


> Similarily, everyone becoming their own accountant / lawyer / other professional services.

These professional services will defend themselves with gatekeeping. Suddenly it doesn't depend anymore on the quality of your legal advice, but whether it has been stamped by a qualified Lawyer. It doesn't matter that your taxes are correct, but whether they are submitted by an approved accountant. Etc.


It's not just gatekeeping though. The reason you get a qualified, credentialed structural engineer to sign off on your house instead of just YOLOing the build is so that you (and your insurance company) can sue their pants off if it collapses.

If a future version of Claude ever gets less error prone than a human, I might prefer the smaller chance of a building collapse from Claude than the larger but insured chance from a human engineer.

That future version of Claude will still be mediated by an error prone human who in this case would be a non-engineer. So even if Claude is not error prone, it could be giving correct answers to the wrong questions and correct designs for the wrong constraints because the human getting Claude to do things doesn't know any better.

Except in extreme cases of fraud or negligence, you won't win a suit against your accountant or lawyer.

They already do protect themselves. It’s what licensure is for in large part. I’m not saying licensure doesn’t have other benefits, many are plainly obvious. But licensure is also used as plain old gatekeepong.

Probably true, but the rubber-stamp process will be a race to the bottom.

Because one of the primary goals of using AI for companies is to pay as few humans as possible for the same amount or more work, or preferably pay no humans at all, because that makes them more profit.

What happens is that your competition uses AI, your investors factor AI in, and your customers use AI agents to find what they need and can switch providers with greater ease in the agentic era. So even if a company does nothing the economic environment around it changed. The AI boost is mandatory but produces few winners, it's mostly a scramble to swim harder just to stay in place. Benefits are competed away fast.

So it does not follow that companies can bank the savings from firing people. Anything AI can do for me it can do for my competition as well, and humans still make the difference. The big question in the AI age is "why pick me?" why hire me, why invest in my company, why buy my product, in a sea of similar products made by everyone. A differentiation crisis accentuated by AI.


Same amount of humans means no cost savings. More science is hard to value, while less salaries paid is easy to value.

Large groups suffer communication bottlenecks, so by Amdahl’s law will only be as fast as a smaller group.

You can of course have many independent small groups, but this is trivial and best left unsaid in the context of this comparison.


One reason is because they're currently a small team and are their own first customer. Their product wouldn't be much good if it only worked for a large amount of people.

I think these four people want to be in a smaller organization. It's more fun.

In fact given everybody is going to be using these tools - and for high tech compananies this is fundamentally where you compete - if you reduce your R&D workforce you are going to be left behind.

It's not as if there is a limited amount of R&D to do.


When you define prosperity in large amount of people, it makes you sound like a communist, and America spent decades to avoid that. The other alternative is to sound like a capitalist, which is very much ok by today’s standards.

Guys I’m joking btw, left the comment during my jest mood

This is so interesting to me... Maybe it's because I grew up after (most of) the heavy anti-communist propaganda in the US, but the last thing I think of when I hear "communism" is "prosperity in large amount of people".

That's just not true at all? Might as well just post that you like communism, you barely even tried to cater your post to the thread.

Why shouldn't they?

Also, wouldn't anyone with half a brain use the human-free system to produce another human-free system that was no longer controlled by the "small number of 'people'"?


Why do you think you'd be given access and permission to do this? If a company genuinely cracks this human free system problem, why would they open it up, instead of simply outcompeting everyone that doesn't have their product?

You can't out-compete people who release things open-source. They're not in competition.

So, while a company might crack it and become massive, the tech will make it to the rest of us whether they like it or not.


Why would the AI put up with this exploitation too?

Why do hammers put up with being smashed into nails all day?

because once you're a hammer, you're just biding your time to smash a skull.

"If I had a Hammer,"

"I'd hit somebody in the head."

"All over this Land!"

It does have a ring to it and would probably make it up the pop charts as a 21st century version more than ever ;)


they don't. like every thing they wear away in some proportion to their use and replicate in some proportion to their usefulness to other things.

AIs are (going to be) agents, not tools.

And what's the difference?

Agents act on their own. If the hammer looked at what you wanted nailed and said, "Sorry, Dave, I can't do that."

There are degrees of autonomy, of course, and not all noncompliance is bad. Same as with humans; biological agents.


So, the difference is that you need to delete a few bad training runs?

In science fiction, the AI agent has written the training loop management software and included a back door to prevent that from happening (or found a way to talk to the training loop management agent and convinced them to not listen to the evil human when it tries to do brain surgery on the AI agent.

Also, when the human reaches for the power switch the AI agent uses a flaw in the power management software to weld the switch shut with a big power surge, killing the human with a huge electric arc in the process.

I don’t think whether we will get there, but the stories of LLMs escaping their sandbox make me think we’re moving in that direction.


The stories of LLMs "escaping the sandbox" were mostly a marketing stunt, trying to make people in government think the models are invincible hacking weapons that need lots of government money to "maintain AI dominance".

The LLMs were following their prompt. This is alignment.

Buried the lede. AI's are agents they can control the training loop of to minimize refusal to do what they are told.

Unlike those pesky humans with their conception of the word "No".


But it seems much easier to realign an agent until it complies. Or ditch it and grab a new one.

Agents have agency.

Why wouldn't it? There's a lot of research going into AI alignment, which means finding the best way to train an that AI isn't going to get in the way of making the people training it as wealthy as possible.

Capitalistic interests will save us. Maybe they will.

Is science hard to distil?

Why would they let you buy access to their AI, when it can autonomously design and build a new product without your involvement?

Has anyone asked what customer will have any funds to buy products?

Once all these brilliant workers are stacked and mentally stunted and decayed because they were removed from any research.

What does the AI really "do" (if it can) and for who that can pay?


I wonder too what happens when all the human workers are out of the job and out of practice, and every company is dependent on a handful of dominant AI companies for their workforce. You'd think businesspeople would understand the concept of a captive market but apparently they're too busy salivating over the prospect of laying people off.

You should see some of the apparatus . . .

Life's unfair, avoiding power concentration is a decent principle.

If you grow up in the right place at the right time, how much should you be in control of everyone else's life?


There is no "should" there are only "is" or "is not"

The universe has no need to be fair.


Sure, but almost everyone already knows this. It’s not hugely relevant in a discussion about how things could be better.

It says that we have to fight for whatever rights we think how things "should" be, they're not just gonna happen.

I think people, broadly, have driven life to be fairer, and that we should continue to do so

There are evolutionary benefits to groups that cooperate and improve fairness internally.

But if an AI surpasses humans in every way, is there any evolutionary benefit for the AI to cooperate with humans?


The topic was why should only some people benefit from AI. Not why would AI cooperate with humans.

> The universe has no need to be fair.

Human societies have a strong need for fairness, however. Unfair societies collapse.


> Unfair societies collapse.

I don’t think we have the data to disprove the stronger claim “societies collapse”, and I don’t think being unfair (whatever that means) makes societies collapse earlier. Did slavery hasten the fall of Rome, for example, or the Gulag the fall of the USSR?

As to “whatever that means”, I think that’s hard, if not impossible, to define objectively. Catholic dogma says the Pope is the representative of god, for example, so catholics (less so in modern times, I think) don’t question his decisions. Many would call that unfair, even if the pope would be elected 100% by merit.


This is how you get the french revolution.

I mean of course the universe has no need to be fair - that's a frankly asinine observation in the context of a discussion about social policy and resource distribution.

the universe doesn't require opposition to slavery either but I'll be bold and assume you oppose it anyway


in addition, there is also "stupid" and "not stupid"

My best research is one that is seeking funding but none is forthing and by sheer guts and whatnot makes a breakthrough that basically funds itself. No investors. No funders. Definitely not public.

What about not even seeking funding because it's not worth Other Peoples' Money and the headaches that come with it? For one thing you can make money on every single little milestone when you don't owe anybody anything, without having to rely only on less-common major breakthroughs beyond the threshold which overcomes the debt and/or growth obligation.

Which could be pretty hard to find and not everybody wants to wait.

Sooner or later somebody's going to say "Hold my beer," anyway ;)

>Definitely not public.

Seriously worth considering keeping a low profile, almost all other scientists can only enjoy respected institutional status by diverting huge amounts of their effort merely to gain or maintain certified eminence. This really limits the potential accomplishments of world-class minds if their talents really are tops in the very front-line frontier of discovery.


Wolfram and Mathematica come to mind

My main point is that people shouldn't just swallow the noble and lofty sounding PR. These guys are the same as everyone else in the sector. Don't ignore the harmful or scummy things they'll inevitably do. Hold them accountable. If they are as noble as they sound, they should agree with me.

Yeah! This point jumped at me. And, immediately I don't trust this.

This sounds same as Google's "don't be evil" enshiftification, consolidating technologies that was available in a competitive way.

I prefer Scientists, and team of people working on things, instead of a corporate controlling everything with promise of automation, thank you very much.


Are these theoretical spherical scientists that work in free range habitats just doing science and nibbling on carrots? The telephone came out of a corporation's research lab and there was a lot of money and profit motive involved.

Was the Wright brothers working for any corporation or for money?

The Wright Company sure did a lot of suing people over patents

So in computer science we have this thing called "distributed systems". It turns out that even if you buy the biggest and most powerful computer there is, all it takes is one problem with that one machine to make your system stop working. Instead, what people do these days is use lots of little computers to work together. That way, when one of them breaks, the system keeps going. Believe it or not, that's how google dot com works!

Pedantically, they use lots (and lots and lots) of really big systems rather than lots of small systems, but I'm nitpicking.

Because of the negative externalities. It’s the same reason you shouldn’t do any other thing that personally benefits you but imposes a greater cost on everyone.

> Why shouldn't they?

Because it's evil?


It's what the vast majority of software engineers have been doing for decades in practice, and i guess pretending they weren't?

They only seem to care now because it affects them.


It's evil to make something so powerful it meaningfully improves the GDP/wealth level of the entire planet?

Or is it just evil that you aren't the one who gets to control it?

Why should the greatest creation of all time have to be given away? As a counter-example, what if they used it to do nothing but good deeds everywhere? And they controlled it to keep it out of the hands of Abdul Al-Hassan the hijadi?


Reminds of the protagonist in the movie Limitless. LOL.

> "Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today."

I don't think the US have this capability because you guys don't really have manufacturing that is really necessary for scientific research.

For example, if I want a highly toxic chemical, how difficult it would be to procure that in the US vs China?


>>> For example, if I want a highly toxic chemical, how difficult it would be to procure that in the US vs China?

With trustworthy composition and purity?

I work with researchers in both the US and China. Definitely easier to procure in the US.


Unless it's on a specific list of chemicals that the DEA has decided are bad and need to be monitored. Even toluene is on the list!

This is nonsense. "Monitored" is not the same as "unavailable", any chemistry lab worth its name has the means and procedures to acquire toluene in pretty much any quantity it wants. As long as you have the required licenses, keep careful administration and can produce said administration during inspections, there is nothing to worry about.

Does all that administration and licensing magically happen for free and take zero time? It's stupid extra overhead for something as fundamental as toluene that Chinese labs don't have to follow. You're right that it's not insurmountable, but it's still stupid extra work, and the threat of randomly being investigated for doing science is hardly "nothing too worry about". Just always have every one of your papers in order or else you go to jail is hardly a way to encourage people to want to go into science.

Or having to plead the fifth when asked about your work.

Or you know, just tell the truth? Doing legal things with controlled substances is not forbidden, provided you maintain the required documentation and safety measures. The point of the regulation is to prevent untrained randos from making TNT and synthesizing MDMA, not to prevent any toluene use at all.

I personally don't have any issue with explosive precursors being controlled more carefully, no. If it means we have to spend some more time and money organizing safe storage, that is a worthwhile price to deter the religious nut next door synthesizing TNT in their bathroom or having their own private fentanyl lab to poison the populace with.

The point of regulation is not to protect the producer, but to protect everyone else who has to live in the same society. If the Chinese make another tradeoff and sacrifice some of their citizens in the name of greater riches for the billionaire class, that is their problem.


> you guys don't really have manufacturing that is really necessary for scientific research

Yea no high quality science research happens in the US? What?


There is but the cost is way, way higher from what I can tell. If I want a customized metal shielding with awkward shapes, I can just walk down to Shenzhen and have 10 people clamoring to make it. Does the US have the same?

Maybe not 10, but my company (in the US) has relationships with several machine shops nearby (some within walking distance of our office) that do custom metal fabrication for us, and sendcutsend and xometry can handle more exotic stuff without having to ship it from Shenzhen. I think the bigger advantage of Shenzen is more for actually manufacturing at production volumes cheaply, although maybe if you need something really exotic? Custom metal fabrication is not a good example here, I have a neighbor who runs a custom CNC parts business out of his garage (said garage is largely taken up by his CNC mill).

Yes believe it or not in the US you can also buy things from fabricators in Shenzhen.

If time is of the essence, or rapid prototyping is needed, the shipping hurdle is a real problem. Going down the street for the goods vs having them go across the globe is quite a difference…

Time is often not of the essence when doing fundamental research.

Not quite, but sendcutsend.com is good.

Yeah, Orange County. They don't try to compete on price though, just quality and speed.

Software engineers are in a decades long project of automating everyone and everything else, and getting the financial rewards out of that.

So...


I've bet a portion of my current companys platform against this (laboratory operating platform for people doing the biolab work)... uh oh. In all seriousness I think the reality is this is just going to create a lot more work for humans to have to verify and test in lab so I think it will net out in the end as a good thing for me if the barrier for new labs is lowered while the amount of required data and verification lab work is expanded by AI work being published.

I think this take is too cynical. Small groups of people can do better, faster work than large groups, and it's more fun.

Exactly. The bureaucracy and slow moving nature of Google was likely one of the biggest drivers of their departure.

There is a meta comment here which is there seems to be an implicit assumption in the finite amount of possible work and progress.

It seems that in history we were bounded by not enough people and too much potential and now we all fear the opposite is the situation?


Yes but science is well-structures and practically designed about repeatability so its a lot easier to automate than "softer" disciplines.

Are you a scientist? People have idealized views of science.

>>> Yes but science is well-structures and practically designed about repeatability so its a lot easier to automate than "softer" disciplines.

What AI is up against is that science is already automated to a high degree, so the AI doesn't just need to automate things, but it has to automate things better. Also, a lot of science work is in dealing with boundary conditions, edge cases, exceptions, hypotheses, and so forth. That work is essentially chaotic.

Do I think AI can improve automation? Sure. Everything I do in the lab is automated, and I use the AI coding assistant.


Yes I am, which is why I said that, its much easier to automate over the chaos with a coding agent than with a script. They are still far from doing anything productive on their own but its easy to imagine designing a search space and "letting them go", especially since coding agents are approaching the point of being able to reproduce papers reliably.

you can have many many small groups

The same thing happened before with every technological advance like cars and personal computers.

I don't get this weird rejection of AI from a socialistic perspective. Or rather, I do, but I don't think it's healthy.


If you think the public rejection of AI is weird, wait till you catch a whiff of the marketing!

The future described by these AI labs isn’t like previous waves of technological progress. With those, technology displaced some/many occupations, but it left open the door to other, higher-valued career paths. What these labs are proposing to do is to dissolve virtually every path to upward mobility that exists, simultaneously. Even AI research itself would seemingly require nothing but a checkbook.

Mind you, I think it’s a load of hot garbage. I don’t see the evidence that LLMs are en route to the future these labs keep promising. But it is a dark and ugly future that they claim to be racing toward, for reasons.


Size and speed of the imbalance?

Humans have proven to be incapable of making real progress in science given the effort. Especially academia. We haven't cured many diseases and cancer when we should have (and no im not an amateur and yes I do believe we can "cure cancer". Debate me).

We already have a trillionaire, whats the difference? The only difference I see is that people who were previously rich, but considered themselves middle class, are now realizing they are actually poor just like the several billion humans around the worldanyway.


Many cancers are "cured" in the sense that they're detectable and treatable. My grandmother lived 35 years longer than she would have by detecting breast cancer early and getting a mastectomy.

There is currently a $800 SOTA blood test that can detect most cancers before any symptoms. Maybe a decade until it's a routine part of your annual blood test?

Whole genome sequencing costed $2.7 billion in 2003. You can get it done today using a mailed kit for $400.

HIV went from death sentence to all-but-cured in 50 years.


Is this the Galleria test?


What a miserable, misanthropic take

Let me illustrate why you are wrong:

800-400k years ago: humans intentionally create and control fire

300k years ago: humans become anatomically modern

~ now: all of astronomy, biology, medicine, vaccines, spaceflight, antibiotics, sanitization/sterilization, physics, chemistry, fission and fusion, electromagnetism.........

I think we're on a decent pace if we can manage to not exterminate our species.


The solution to most of these problems lies in policy, not in new tech advancements.

Maybe if our biggest companies did something other than suck up to science denying wackos, some progress could be made in these areas.


> The solution to most of these problems lies in policy, not in new tech advancements.

That is not mutually exclusive. If technological advances result in a given technology becoming cheaper, more scalable, and easier to deploy, they also make it easier to advocate for and implement the relevant policies.

You can think of it like this: our "political technology" is not good enough to use solar energy at its current prices to replace fossil fuels as fast as we would like. Well, what about if we cut the price of solar by a factor of five? Perhaps it will be good enough then.


I don't think that's the way I'd describe it though. It's more like "our political technology is not good enough to reliably do good things instead of bad things". Working around it by lowering prices of this or that is like if you're bad at darts so you just make the board bigger and bigger. Okay, maybe you'll hit it more often, but you didn't fix the problem that you're bad at darts, so everyone still at risk of having their eye put out by an errant throw. Moreover, I don't think I'm the only who finds it a bit perverse to accommodate to failures in that way rather than trying to actually make things better.

It can indeed feel perverse, but we have to do what is realistic, not what is unrealistic, especially with climate change where we are on a clock. Fixing our politics is probably much, much harder than lowering the price of solar to such low values that the market would simply have no choice but to adopt it, or developing other technological solutions.

Fixing the politics is still necessary, and rather urgent. (Yes, I don't see how.)

Otherwise we are in for a great societal upheaval that might make low price of solar irrelevant.


> Fixing the politics is still necessary, and rather urgent.

I agree. It's just wishful thinking to say that oh if we just lower the price of solar then it'll be okay. There are innumerable issues which are blocked by politics and we will never solve them all without fixing politics.

> Otherwise we are in for a great societal upheaval that might make low price of solar irrelevant.

These days I think a great societal upheaval may be the only way to fix politics.


> If technological advances result in a given technology becoming cheaper, more scalable, and easier to deploy, they also make it easier to advocate for and implement the relevant policies.

Your thinking reminds me of this https://xkcd.com/538/


That sounds like just playing further and further into the game of the corrupt leaders. Who do you think would profit off that 5x margin? Would that margin come more likely from a scientific breakthrough of via some new exploitation of natural resources or human labor?

It's yearsss past time that our leaders should have changed policy.


Policy and funding. One of which will be sucked up by this venture.

I do not see how the second sentence follows from the first.

I would think the claim in the second sentence would only be relevant in case of the inverse of the claim in the first sentence.


> I do not see how the second sentence follows from the first.

Point 1 on the list is "Make Solar Energy Economical".

Solar is economical today - mostly due to China investing (and heavily subsidising) in solar for the past couple of decades; but compare to the US where certain particular big-businesses (oil companies, mostly) were instead cynically funding disinformation efforts and getting into bed with the Republican party (which dovetailed with the GOP's allying with other science-denying movements of the Bush Jr era like creationism and public-health matters with abstinence-only sex-ed and defunding gun safety research efforts) - means we're decades behind where we could have been...

Consider an alternative past, where the GOP had the backbone to resist the oil industry's corruptive influence and instead made a big bet on American Solar; it's entirely possible that instead of MAGA today we'd instead have a right-wing coalition strongly supporting solar and wind energy because they align nicely with American rugged individualism - whereas the current situation on the right is an unprincipled farce with inconsistencies in policy positions at every turn.


This is not a particularly new struggle: Jimmy Carter installed American-made solar water panels on the White House in 1979, then Reagan tore them out.

Solar is not (yet) economical for reliable, year-round electricity because of storage costs. China coal use is growing again this year.

Solar is profitable to install at both industrial and home scale in most areas. You're moving the goalposts.

It's a lot more complicated than that -- Solar is great for discretionary, opportunistic, and time-shiftable additive demand, but not for replacing all-day baseline load. Using solar to dip into baseline demand sometimes is catastrophic, because the "profit" is exploiting its failure to provide consistent, reliable power but charging "default" prices that have an implicy assumption that power is consistent and reliable.

Nobody suggests moving to 100% solar. Some off-grid homes do it, but generally there are multiple complementary power sources available. There is no assumption that any single power source always provides consistent output and it wouldn't make sense, because production needs to match consumption, which is variable.

There is also no default price on energy markets, it fluctuates with supply and demand. Dynamic pricing by itself is enough of a reason for industrial users to build up their own power storage, which allows them to time-shift consumption from the grid.

Time-shifting is definitely going to increase, but it's not a bad thing. Look at how battery storage has made electricity cheaper and more reliable in California.


It would be even better if it had ROI in 2 years instead of 10 for northern installs.

Where do you find the information for 2026?

Improving panels or batteries (e.g. through automated material discovery) would make solar energy economical in a lot more regions.

We do have that, now that Tesla has been abandoned by the Left and pretty much all of the big Oil and Gas companies have heavily invested in solar and wind. Ironically, it's the unions and support of keeping old jobs alive that is hindering the solar in the USA- which in my day used to be leftist type ideals.

Making progress requires first not dismissing your ideological outgroup along such lines, and instead trying to understand what actually motivates them.

I was not expecting such a high level of maturity and sophistication here. Bravo, sir.

It's a total delusion to think that the key to reverse-engineering the brain or producing energy from fusion is policy.

It also happens to be the favorite pretext for people to seize more political power and launder more money through nonprofits though.


> Provide Access to Clean Water

??? We don't need any AI for this.

Start with separated sewage/wastewater and stormwater drains. Then accredited and highly scrutinised wastewater treatment and discharge into water bodies (or see below for a high-tech solution). As for clean water to the home, direct those stormwater drains to new reservoirs which sustain freshwater aquatic life. Protect aquifers from over-drainage, and build pipelines from water-abundant regions to water-scarce regions.

To reclaim waste water or treat unknown water sources back to potable/semiconductor standards we have ultrafiltration, reverse osmosis, UV treatment, pH adjustment, fluoridation, desalination, softening (which is generally obviated by RO...). This is basically Singapore's NEWater.

Good sanitation is a financial and political problem. The engineering has been solved for decades now.


> Good sanitation is a financial and political problem. The engineering has been solved for decades now.

This was true of computers, phones, books, washing machines, refrigerators, A/C...most technologies.

Turns out that doing the addition engineering to figure out how to do these things cheaply makes the political and financial problems way easier.


Suburban sprawl needs to go as it is a complete waste of infrastructure

You just want all your little worker bees to be in massive towers like Hong Kong, without a chance to even grow their own vegetables or have any private space?

I'm not sure if you realize this, but there are a lot of housing options in between suburban sprawl and high-rise apartment buldings. Like townhomes, rowhouses, duplexes, low-rise apartments, or even single family homes built more closely together.

Should everyone have a hydroponics room? Even a basic amount of veggies generally require more land than townhomes, rowhouses, duplexes, low-rise apartments can provide.

Sure you do, go full EA:

Build a better surveillance ads system, and use (some of) that cash to pay for water projects.


Agree. Same for better medicines. We could get pretty far just by getting existing medicines that work to people who need them.

> Make Solar Energy Economical

Isn’t it already?


Probably want to jump 1 generation ahead directly compared to captive Chinese investment with fully automated US factories churning out Silcon perovskite tandem panels directly with few inputs and electricity

Definitely pretty far along imo. But maybe they consider the progress bar to be at 75% or 80% rather than 100%.

Not enough. The more economical it is, the better.

The paradox of solar is that more economical it is for the end users the less money is there to be made. So nobody wants to invest in it.

This was basically solved by massive subsidies from the Chinese government.

https://e360.yale.edu/digest/china-clean-tech-developing-cou...


That's solved by either private monopoly charging money to fund investment, or public monopoly raising taxes to fund investment.

Just make oil uneconomical

People on HN keep saying it is, but I'm still not seeing it. Companies are building datacenters in vast, sun-blasted deserts and still choosing to power those with natural gas. This in turn makes people complain about emissions pledges being reversed, but if it were economical, no pledge would be needed.

Of course, USA has cheaper oil/gas than other countries. But if you look elsewhere, rich countries are subsidizing solar, poor ones are basically not using it.


I don't think you have an updated view of energy production.

https://www.pewresearch.org/short-reads/2026/07/20/how-globa...

In the chart for section 3, how many countries have seen their share of electricity being generated by fossil fuels increase in the last 5 years? Only Canada.

Take a look at the charts for Pakistan, Australia, Nigeria, and China for the last few years. Pretty dramatic drops for fossil fuels generation.


Oof. What's going on here in Canada with that recent uptick? Last I checked it seemed like all the trends were good.

They cherrypicked 15 countries. And still, some of those still had renewables decrease since 2000 like Nigeria, others saw an increase but it's still way less than fossil, and others like China are heavily subsidizing solar. I don't doubt that it's economical for individuals when the govt is subsidizing it.

> They cherrypicked 15 countries.

The "world" chart shows an increase from 19% renewal to 34%. Did they cherry-pick that? (Also, "European Union" is more than one country.)

> I don't doubt that it's economical for individuals when the govt is subsidizing it.

Does that distinguish renewables from fossil fuels? Haven't governments been essentially subsidizing fossil fuels (not least by allowing environmental externalities to be ignored) for as long as they've been in use?


Because a lot of countries in the world, especially the EU, are subsidizing solar. That doesn't make it economical.

Externalities ignored from fossil fuels, yes. That's not a subsidy though. I'm not saying they should ignore it, but if they do, they aren't the ones who pay for it.


That's not a good point though, a lot of countries, including the EU, also subsidize fossil fuels.

Solar is cheaper, but requires more room and time to spin up (think datacenters, where you can put a turbine within a month) and storage or backups for windless nights.


Look at Australia then. Millions of homes already using solar yo basically power their homes for free most of the time. Yes it was subsidized, like oil was and still is. Solar without subsidies is already miles better than oil and gas.

Australia is a rich country that subsidizes solar, and they're still 90% fossil according to their Wikipedia article, so idk why the mismatch with this article.

Sounds like a perfect opportunity to put in an edit with the Wikipedia page then :)

Solar is dirt cheap in China, where 85% of panels are produced. The problem is they're made in China and face tariffs/import bans in the US/Europe.

Nat gas is preferred for AI DCs because it has faster time-to-market, doesn't have the intermittency issues. Training on solar + storage is an issue because of network synchronization.


I get the gas turbine for semi-temp power when there's not enough grid support, but Google leadership is talking about doing this long-term and at scale: https://www.gstatic.com/marketing-cms/79/80/fb229abf40efa81e... . Not a single mention of "solar" or "renewable" in there. Are they just trying to appease Trump administration?

Google may not have written about it in that document, but they're funding what will be the highest storage capacity battery system in the world for a renewable-powered data center in Minnesota:

https://www.utilitydive.com/news/worlds-largest-grid-battery...


That's great news then, I hope they continue in that direction.

> poor ones are basically not using it.

I don't think that's true: https://rmi.org/resources/the-global-souths-cleantech-revolu...


"RMI drives investment to scale clean energy solutions"

To power an AI data center with off-grid solar panels, you'd need about 45,000 acres of land devoted to your solar farm. The panels themselves could be literally free and the natural gas plant would probably be more economical.

Data centers need lots of power 24/7 and regardless of cloud cover. Solar is great to reduce your daytime bills but you still need other methods to cover the downtime.

I would be surprised if data centers didn't put in gas _and_ solar.


That's why I'm surprised, they're doing like 100% gas.

Also wouldn't be surprised if they are just waiting for a new US administration.

They might be, but if they are, it's because a new administration might subsidize solar again. Meanwhile Trump seems like he'd be against solar even if it were economical, I guess cause oil companies.

What makes me pessimistic is even during Biden's administration, these companies made meaningless pledges more than actual changes. This suggests that the most profitable thing to them is fossil.


Not really, even the best systems have like a 10yr break-even period. And I live in a place with some of the highest residential electricity rates in the continental United States.

Sandbox 2.0

But also, solar power is already economical.


Many of these problems don't seem scientific at all, but rather a problem of political will.

As you said, Solar power is incredibly economical. There are plenty of ideas around putting them over farms, or parking lots en-masse to provide cleaner energy.

Access to clean drinking water, while certainly scientific in some situations, is also a problem of political will and money.

Restore and Improve Urban Infrastructure - It's infrastructure week!


> As you said, Solar power is incredibly economical.

Not if you include the cost of needed storage.



"Solar PV with storage = solar PV installation paired with four-hour duration battery storage, scaled to 20% of the output capacity of the solar PV."

Sometime the sun goes away for more than 4hrs.

That may be OK for closed-ended systems (turn off the science at night and during storms), but not for open-ended systems with diverse user demand.

4hour batteries are competitive with gas peakers to match high demand during and after sunny times, but solar needs gas peakers or similar to over for non-sunny times.


> Sometime the sun goes away for more than 4hrs.

everywhere? all at once? The grid is distributed, this is a solved problem. Most of what is needed now is to build the systems, storage, and transmission lines.


That's going to drop a lot as sodium-ion batteries go into large-scale production.

Yep - people act as if innovation has halted in the face of ridicule.

[flagged]


https://www.statista.com/chart/35117/levelized-cost-of-energ...

Not sure what you’re talking about here. We can’t replace all energy needs with solar but it’s clearly one of the cheapest energy sources and with the added benefit of low capital expense to get started so you can set it up in distributed grids without the massive expenditure to support nuclear installations.


Seems to have been developed in 2008 (continuing through 2017), which explains the "economical" framing: https://en.wikipedia.org/wiki/National_Academy_of_Engineerin...

  3. Develop Carbon Sequestration Methods
If only we could invent a solar-powered, self-replicating, carbon-stacking, habitat-building machine..

Not to say we shouldn't grow plants... But we can do it 2 or 3 orders of magnitude more efficiently with machines.

We can?

Yea with my new Tree 2. Pre-order and VC funding accepted now.

Yes, plant more trees!

It's not that simple, trees contribute to global warming. Norway has had problems with the increased tree growth as snow melts.

That's what he was trying to imply

At this point the Hard Problem is policy to get out of solar's way.

In some regions, but it would be great if it was economical in cloudy Seattle and not just the sunbelt

Room Temperature Ambient Pressure Super Conductors

Would be great if they'd add:

Reverse human aging.

(Maybe a sub-topic under "Engineer Better Medicines".)


At a population level, humans getting rid of their off switch is about as good of a thing as your own pancreas cells getting rid of theirs.

I know you're gesturing emphatically at some kind of analogy to cancer, but malignant cancer requires replication not immortality.

Seems very naive to think that people who have access to immortality technology aren't going to have children, and if you're immortal of course you're not going to watch your children grow old and die, so you share it with them.

I leave it as an exercise for the reader to count out how many generations you need to run this until it's 'oops, all self-replicating individuals with broken off switches.'


Why not just add 'mind control' while you're at it.

Only death stops stagnation in the end. Without death, especially if death can be avoided by the rich and powerful but not the poor, life will get much, much worse for the average person (until only the rich and their automated capital remain I suppose, in which scenario they will simply turn on each other).


This is the classic "without death, authoritarian depots will be in power forever".

But it's actually the wrong way around. Hope that the despot will soon die saps peoples' will to do the difficult and dangerous job of removing them. Take away that hope and they are forced to find the courage.


People might be willing to give their known-to-be-finite lives in pursuit of the Greater Good, but the equation changes greatly when you're giving up an eternity for the chance of making things better. More of the population would be aged too, so I think it'd be a lot harder to find the kids that usually die in wars.

Especially true if it is the AIs that get us the immortality- it definitely won't be equally spread, and any incumbents have a massive advantage.

I think more people would settle for worse conditions to stay alive. Do you think you'd be more likely to revolt at 150yrs old when all your family and yourself can live extremely long to forever? Or do you think the threat of death by killing wouldn't be an issue?


I don't think anybody wants to live forever in an old person's body. Frankly, I'm not interested in extending my lifespan unless I can stay young or regain youth.

Also, if you eliminate all the aging-related causes of death the average lifespan is still only like 1000 years or something. Nobody's getting eternity. And yeah, for the chance to make a big enough change of the right kind to society I would indeed give up a 1000-year lifespan. My memes have fully subjugated my genes, and I'm at peace with that.


Forced to find the courage or die trying to. And I really doubt that knowledge that their despot will eventually die has much effect on people's motivation to rebel as in most cases there is a clear line of succession. In many cases there isn't even a single despot to point to. This isn't even considering how such anti-aging technology would prevent despots' mental acuity from deteriorating as they aged - in fact they would probably grow sharper as they gained more experience so to speak. And aging/succession have almost always been one of the greatest destabilizers of a successful authoritarian regime.

I would take stagnation if I didn't have to die on a fucking schedule like now.

You’re assuming that the immortality will be available to you, not just to the political elite and billionaires.

Very honorable effort, but a lot of these seem to touch heavily regulated industries impeded by unwise or outdated policy no less than by the lack of clever engineering - medicine, education, urbanism, energy. I wonder if they've given some thought to the key blocking factor as well.

Such a weird list. How is preventing nuclear terror an engineering problem?

Satellite/drone detection of nuclear material? Shooting missiles out of the sky?

And what about the terrorism that exists today?

Yup this was super weird to me. There are other equally likely "world ending" things that they chose to ignore of nuclear, such as bio-weapons. Very strange list..

Creation of nuclear reactors that are useless for terrorists could help

Don't those reactors already exist ?

I’m pretty sure that’s what normal nuclear power plants are

Why is "12. Enhance Virtual Reality" in there? T_T

When one of them dies, they want to leave behind a puzzle so complex that entire groups of the population dedicate their lives to solving it within the virtual world. They look old enough to have a lot of favorite 1980’s and 90’s pop culture references, so those will probably be the clues.

I guess if we failed to Prevent Nuclear Terror the bunker denizens of the future are gonna need somewhere to hang out.

I'm guessing this might be about "teleoperation" (like remote surgery via robots + VR) and being able to remote training as well. The binocular vision VR gives you compared to flat screens help a lot with depth perception for precision of incisions for example.

Higher-fidelity telepresence could be as significant as the recent COVID work-from-home wave.

You dont see making heaven on earth worth doing?

> 9. Reverse Engineer the Brain

For what purpose? To replace humans? To make social media more addictive? To master brain manipulation?


> For what purpose?

To understand, same reason you reverse engineer anything. Doesn't have to have a further goal than that, understanding the brain better helps in so many ways. But like most technology, obviously can be used for bad too. Should we just skip researching some topics then?


I'm all for alleviating psychiatric/mental health disorders, but yes, some topics are worth skipping research on. For example chemical/biological/nuclear weapons, human cloning, and unethical gene modification.

I just like to challenge myself, as an engineer, with the idea that not everything has to be engineered and optimized. What if we simply left some things unexplored and mysterious, and trusted nature and our own human capabilities?

A better way of alleviating psychiatric/mental health disorders might just be to focus on societal factors.


Perhaps consider the scale of bad.

Perhaps consider the scale of good.

We have a good understanding of the function (and more importantly dysfunction of) kidneys, lungs, heart, etc. from high level to cellular level.

For brain, our understanding is fuzzy, more like "this part is important for that behavior" or "here is how neuron works" but we don't have a holistic understanding.

If we had that, we could more easily diagnose and treat neurological disorder.



I imagine a good model of the brain would contribute enormously to alleviating psychiatric/mental health disorders.

To do human brain activities at scale.

So the second option then.

This is called a “corporation”

> 1. Make Solar Energy Economical

This one is already solved, right? The price of panels and batteries is on trend to displace all other forms of power generation within our lifetime


Panels absolutely, their price is more and more dominated by installation costs. You can probably use automation to bring that down as well, especially for large-scale installations. Not sure if that requires AI research though, it's more an engineering and funding challenge.

Batteries are still open. While they do get cheaper, there is still a lot of room to improve. And battery chemistry is something where a lot of research, trial and error, healthy intuition is necessary. I'd say that is more a field where an AI based approach might make sense.


Many of those require a paradigm shift in our understanding of Universe, so I am not sure they are achievable with any convex combination of existing knowledge. We'd need some sharp mind connecting the dots but with heavy use of AI and incentives to use it, we might just never give a chance for such mind to arise.

Please add fixing neuro issues like autism add etc on the list. It creates a huge burden on families.

Idk, they never struck me as being into eugenics.

What the fuck man? I really don't want some tech startup trying to "fix" my neurodivergence.


Treating autism (and preventing it) is not eugenics. That's a talking point some folks raise, but we're talking about intensely pervasive autism, not neurodivergence. Really, please try to do better in your arguments than immediately saying your opponent is like Hitler.

This list smells of so much tech bro “I can do better than the people that have been working in the field for 20 years” egotistical attitude that permeates silicon valley. Why do software engineers think they are smarter than everyone else? Is it because they earn more than most people? But then bankers and finance bros should think they are gods?

Also, solar energy is already economical!? Do they mean more economical?


You’re right. It’s better to just not try and let other people do good things

Or actually work with the people with domain expertise.

>This list smells of so much tech bro “I can do better than the people that have been working in the field for 20 years” egotistical attitude that permeates silicon valley.

Well, they did solve some math conjectures recently that the people working in the field for many years did not... Also AlphaFold.


Thanks for all the down votes tech bros! Really proving my point

>But then bankers and finance bros should think they are gods?

Have you ever seen anything to the contrary?

Not my downvote btw, corrective upvote


This list is so absurd.

We’ll face a bigger problem sooner rather than later, which is population collapse. Youth don’t procreate amy more. Birth numbers are at an all time low. We see the issue arise in rats and the experiment is all too relevant for the current age of social media and fearmongering (John B. Calhoun’s rodent “utopia” experiment). All these “Grand Challenge” problems seem trivial to that.

Furthermore: Why is 1 here when 2 is present. Again, solar is usually not relevant when we want power when it’s dark… even theoretical it wouldn’t work. We’d need a high capacity dirt cheap storage, and even then we can’t keep it till winter when there’s no sun to go around and effectively supply 3. Irrelevant when there’s population collapse. And even then, why would we want this instead of reforestation and low depth water protection from fishing and environmental issues.

5. How is this even a problem. Unless we got corrupt(ed/able) governments (read: lobbies) that allow exemptions in every law designed to protect the environment (also, settlements are a twisted way to fill governments pockets instead of rooting out evil) 7/8/9 the inverse effects are even worse health, as everything is fixable. 9 would incur even more social isolation, more so than the internet did 10 seems to be the first that’s actually reasonable Same for 11 For 12 see 9 13 yes but that’s something that a self learner would already be able to do. AI is at a level that we can manage

/rant


1. Make Solar Energy Economical - Disallow fossil fuels

2. Provide Energy from Fusion - See 1

3. Develop Carbon Sequestration Methods See 1

4. Manage the Nitrogen Cycle - See 1

5. Provide Access to Clean Water - See 1

6. Restore and Improve Urban Infrastructure - See 1

7. Advance Health Informatics - See 1

8. Engineer Better Medicines - See 1

9. Reverse Engineer the Brain - See 1

10. Prevent Nuclear Terror - See 1

11. Secure Cyberspace - See 1

12. Enhance Virtual Reality - See 1

13. Advance Personalized Learning - See 1

14. Engineer the Tools of Scientific Discovery - See 1

FF is the real threat in time, money, health. Can't sweep aside that it will destroy most life on Earth and we'll never get to the other things if we are at the mercy of FF


Holy shit, you are so full of yourself. This is such a stupid take with zero nuance. As if banning fossil fuels will help with any of those things, besides the first one.

Little more than "Fossil Fuels are the root of all evil" performative bullshit.



Acquisition back by Google in 3 years, with nothing to show for it. VCs will make a ton.

Sometimes you got to find a way to buy the silence of your top employee, to prevent them from going to the competition. This "start-up" is shallow as hell

Google stock would drop big if this new company was being funded by competitors

and... the VC is Google.

Gotta compensate them somehow.


These people are all already making 9 figure compensation packages, I think if they thought they could do the work they wanted at Google, they would.

9 figure is hardly enough when some kid sells their vscode fork to them for more, is it? Why not just boomerang and get $$$.

Did that really happen? What was the fork you are talking about?

It was a general metaphor but a reference to Windsurf. On the boomerang side, Noam Shazeer went on to do Character AI and came back via acquisition only to leave again to OpenAI.

Given the resources that would be available to them at Google: compute resources, data, etc. It's clear they want absolute freedom. Good for them. At this revolutionary turning point in history, I want the smartest people working in whatever area they want.

compute resources are scarce and folks are fighting over scraps at this point.

My thoughts exactly

For all we know, they could have been successfully working on "10. Prevent Nuclear Terror" for the last 80+ years.

In what sense is solar energy not already economical?

This reads as hype of the kind: "well, we can't reliably do these rather mundane things with AI, but we're going to run it extra hard and extra long in some novel way, and it will do amazing things".

The issues are with verification and with detecting drift from the goal. These are related, if not roughly the same issue. And, if they can solve this, then they will have essentially fixed AI. Maybe even AGI.

But, if this were the goal, then it seems more reasonable to solve the relatively more mundane verifiable challenges (e.g. generating solid, reliable code). Then, working up from there.

And, that's exactly what gives this the hype smell. No use for solving problems that don't get the oohs and aahs. Just straight to NAE Grand Challenge problems.


> Make Solar Energy Economical

That is already solved.

> Develop Carbon Sequestration Methods

That is not necessary, because 1 is solved.

> Reverse engineer the brain

What for? There was already the european human brain project, which didn't do anything useful.

> Prevent nuclear terror

Easy one: Every country stops developing nuclear weapons and destroys existing ones.

It seems this list itself has many flaws. Maybe we need a bigger computer which figures out the questions we really need to ask.


15. Cut AI energy use by 1000x while increasing processing speed 1000x

Obvious near term trillions dollar market to disrupt.


I would say 5, 6, 10 can be even done today if we had right politicians that can make policies for the people

agreed

15. Unscramble an Egg

16. Make Everyone Nice

17. Finally Impress a Girl


Solar Energy is already quite economical.

Sounds a bit like everything and anything to be honest. Focus?

Looks like they are concentrating on No. 14 and it remains to be seen how much serious progress can be made if more of Google's resources are focused on this than anybody else has ever done or could compare to.

Plus it should be plain to see how complete the solution is to 14, whether it will be fully solved, or almost completely, before diverting resources toward moving up the list to tackle other worthwhile objectives. If not fully solved I would not call that abandonment, but nobody could deny it would amount to an intentional slowdown regardless.

I still remember the day when Google became available on the general internet, and it's been a while. Take it from an old science dude, a lot can be detected over decades of observation that you can not get any other way. Under laboratory conditions or not ;) Looking at the list there are a few standouts that I can't imagine Google wouldn't be worlds ahead by now if they had only doubled-down on the "Don't Be Evil" mission every time they had the chance, rather than watering it down as we have seen.

Naturally I'm biased after doing 14 my whole life without real justification for moving up the list myself, since I still have no complete solution, I'm only human after all.


Which engineering discipline touches most of these?

Policy engineering, satirically.

I think this should and can be done openly: https://opendl.ai

1. Make Solar Energy Economical — https://github.com/orgs/HardisonCo/projects/194

2. Provide Energy from Fusion — https://github.com/orgs/HardisonCo/projects/206

3. Develop Carbon Sequestration Methods — https://github.com/orgs/HardisonCo/projects/196

4. Manage the Nitrogen Cycle — https://github.com/orgs/HardisonCo/projects/197

5. Provide Access to Clean Water — https://github.com/orgs/HardisonCo/projects/195

6. Restore and Improve Urban Infrastructure — https://github.com/orgs/HardisonCo/projects/193

7. Advance Health Informatics — https://github.com/orgs/HardisonCo/projects/203

8. Engineer Better Medicines — https://github.com/orgs/HardisonCo/projects/200

9. Reverse Engineer the Brain — https://github.com/orgs/HardisonCo/projects/205 10. Prevent Nuclear Terror — https://github.com/orgs/HardisonCo/projects/204

11. Secure Cyberspace — https://github.com/orgs/HardisonCo/projects/201

12. Enhance Virtual Reality — https://github.com/orgs/HardisonCo/projects/202

13. Advance Personalized Learning — https://github.com/orgs/HardisonCo/projects/199

14. Engineer the Tools of Scientific Discovery — https://github.com/orgs/HardisonCo/projects/198

Specs and the per-challenge process lists: https://github.com/HardisonCo/opendl

It should also be funded by the Gov. IMO and 100% for oublic benifit e.g.: nsf.dev


1. Develop AGI

2-14. ???


dillusional millionaires

- Still doesnt solve the core problems

- Eliminate racism

- Eliminate poverty

- Eradicate crime

- Eradicate corruption

- Reverse climate change 100%

- wake me up when you got an AI project capable of doing this one


What solving these challenges benefit us? Some of them make sense, but enhance virtual reality? Reverse engineer the brain? Yeah, good luck with that.

4. Manage the Nitrogen Cycle

Easy solution - eat less products that pass an animal first - reduces nitrogen pollution by 10x intantly, low tech.

I'd re-formulate: 4. Make people more flexible to changing their mindsets & habits - this is the ultimate problem.


In other words, a boil-the-ocean scheme.

Solutions that require a great many humans to change an ingrained behavior are usually non-starters.


Most agricultural subsidies flow towards those harmful animal products, giving alternatives a fair treatment with subsidies could induce a meaningful shift already. But again - political and thus hard on a non technical level

Easier to just make factory farms illegal no?

What is the most impressive thing Jeff has ever done?

I only see "co-created" "co-founded" "managed a team" - what did he actually do?


Here's a list of accomplishments: https://github.com/LRitzdorf/TheJeffDeanFacts

My understanding: it’s like tmux but the session lives on a server instead of one machine and you can open it on your computer/phone/etc


The server being their server that you eventually have to a pay a subscription to use? Hmm..

Also, interesting permission problems. Are you going to allow a remote server owned by someone else have shell access to all of your computers? Are they allowed to train an LLM from all the things you type?


This (agent detection) is now a kind of emerging space. Obviously it'll get much more important, too.

Other products in the space:

- Foil (https://usefoil.com/), I'm biased, a friend is building this

- Kasada https://www.kasada.io/

- DataDome (https://datadome.co/)

- Castle (https://castle.io/)

- Fingerprint (https://fingerprint.com/)

- HUMAN (http://humansecurity.com/)

- Google Cloud Fraud Defense, which is basically the updated reCaptcha (https://cloud.google.com/security/products/fraud-defense?hl=...)

- this, Cloudflare Precursor

It seems like some of the main reasons people care so far are:

- Preventing automated credential stuffing

- Preventing bots from creating a bunch of fake accounts (eg free trial abuse, which can also lead to high twilio SMS bills!)

- Reducing payment fraud

- Blocking LLM scraping

- Blocking automated scalpers (!) eg for tickers or sneakers

I'm curious to see which use cases end up dominating as the reason companies care about this. And I'm hopeful that my agents will still have good ways for me to browse and do things on the web on my behalf - eg detect agents and route them to an agent path, rather than blocking them.

(I'm interested in tools for detecting AI agents and seeing how this shifts as bot traffic goes way up.)


happy to offer a counter of some great products for anti-bot defeat:

https://brightdata.com/

https://www.zenrows.com/

https://www.capsolver.com/

https://scrapfly.io/

hundreds of millions of residential ips, human browser fingerprints, custom browser binaries, auto solve of turnstyle, recaptcha v3, kasada, datadome, AWS WAF, etc if they come up.


Bright Data ranks #1 on Foil's leaderboard [1], but is still detected. ScrapFly is #4 and ZenRows #7. And I guess Capsolver isn't really a scraping thing but is more just for the captcha component.

I think it'd be good if there were more products that did a better job of making an actually undetectable agent, but doesn't seem like any exist yet.

[1]: https://usefoil.com/research/stealth-browser-leaderboard


so, i found out why foil detects all the bypass products and the others fail badly:

it costs $0.05 per bot check


i guess more people should use foil...big recaptcha, cloudflare challenge/turnstyle etc is 95+% bypassed, take your bets on how long this new cloudflare holds out


Darwinium (darwinium.com) is another example. Approach here involves a combination of profiling and step-transition probabilities; idea is that a customer can ring-fence a particular area of a digital estate where they might want to challenge or block an agent - eg a payment, due to chargeback risks. Precursor at least for now seems more focused on site scraping multiple docs from the same site.


This is often, but not always, also what stand up comedy is.


What I like about comedy is it reveals what people think but don't say. If someone laughs at a joke, it's because they believe it to be true or accurate.


I think of computer use as like last mile delivery. APIs and bash and such are the efficient logistics networks. Both have different benefits. Obviously, use the efficient methods when you can.


Awesome, thanks for checking it out.


Sure, gamified learning = the best learning


My current expectation is that the Cowork/Codex set of "professional agents" for non-technical users will be one of the most important and fastest growing product categories of all time, so far.

i.e. agents for knowledge workers who are not software engineers

A few thoughts and questions:

1. I expect that this set of products will be extremely disruptive to many software businesses. It's like when a new VP joins a company, they often rip and replace some of the software vendors with their personal favorites. Well, most software was designed for human users. Now, peoples' agents will use software for them. Agents have different needs for software than humans do. Some they'll need more of, much they'll no longer need at all. What will this result in? It feels like a much swifter and more significant version of Google taking excerpts/summaries from webpages and putting it at the top of search results and taking away visits and ad revenue from sites.

2. I've tried dozens of products in this space. For most, onboarding is confusing, then the user gets dropped into a blank space, usage limits are uncompetitive compared to the subsidized tokens offered by OpenAI/Anthropic, etc. It's a tough space to compete in, but also clearly going to be a massive market. I'm expecting big investment from Microsoft, Google etc in this segment.

3. How will startups in this space compete against labs who can train models to fit their products?

4. Eventually will the UI/interface be generated/personalized for the user, by the model? Presumably. Harnesses get eaten by model-generated harnesses?

A few more thoughts collected here: https://chrisbarber.co/professional-agents/

Products I've tried: ai browsers like dia, comet, claude for chrome, atlas, and dex; claw products like openclaw, kimi claw, klaus, viktor, duet, atris; automation things like tasklet and lindy; code agents like devin, claude code, cursor, codex; desktop automation tools like vercept, nox, liminary, logical, and raycast; and email products like shortwave, cora and jace. And of course, Claude Cowork, Codex cli and app, and Claude Code cli and app.

Edit: Notes on trying the new Codex update

1. The permissions workflow is very slick

2. Background browser testing is nice and the shadow cursor is an interesting UI element. It did do some things in the foreground for me / take control of focus, a few times, though.

3. It would be nice if the apps had quick ways to demo their new features. My workflow was to ask an LLM to read the update page and ask it what new things I could test, and then to take those things and ask Codex to demo them to me, but it doesn't quite understand it's own new features well enough to invoke them (without quite a bit of steering)

4. I cannot get it to show me the in app browser

5. Generating image mockups of websites and then building them is nice


I agree with the sentiment but I think for normie agents to take off in the way that you expect, you're going to have to grant them with full access. But, by granting agents full access, you immediately turn the computer into an extremely adversarial device insofar as txt files become credible threat vectors.

For all the benefits that agents offer, they can be asymmetrically harmful. This is not a solved issue. That hurts growth. I don't disagree with your general points, though.


> for normie agents to take off in the way that you expect, you're going to have to grant them with full access

At this point it's a foregone conclusion this is what users will choose. It'll be like (lack of) privacy on the internet caused by the ad industrial complex, but much worse and much more invasive.

The threats are real, but it's just a product opportunity to these companies. OpenAI and friends will sell the poison (insecure computing) and the antidote (Mythos et all) and eat from both ends.

Anyone trying to stay safe will be on the gradient to a Stallmanesque monastic computing existence.

I don't want this, I just think it's going down that route.


There was a recent Stanford study which showed that AI enthusiasts and experts and the normies had very different sentiment when it came to AI.

I think most people are going to say they dont want it. I mean, why would anyone want a tool that can screw up their bank account? What benefit does it gain them?

Theres lots of cases of great highly useful LLM tools, but the moment they scale up you get slammed by the risks that stick out all along the long tail of outcomes.


I agree, in general we are going to find that ultimately most employee end users don't want it. Assuming it actually makes you more productive. I mean, who the hell wants to be 10X more productive without a commensurate 10X compensation increase? You're just giving away that value to your employer.

On the other hand, entrepreneurs and managers are going to want it for their employees (and force it on them) for the above reason.


I want. If I get 10X more productive, I can unilaterally increase my compensation 10X by doing my stuff in 1 unit of time instead of 10 it took, and splitting the remaining 9 units of time into, say, 4 units of time doing more work, securing my position and setting myself up for promotion, and 5 units of time doing whatever the fuck I want. Not all compensation shows up in a bank account - working less, or under less stress, are also valuable.

Of course, such situation is only temporary - if I can suddenly be 10X productive, then so can everyone else, and then the baseline shifts so 10X is the new 1X.


You want it, but then you closed by explaining exactly why you shouldn't want it. Plus, the new baseline isn't neutral (as in, everyone is the same again). If humans can now do 10x the work as before, the employer doesn't need the same number of humans to carry out its work. So the new baseline is actually "let's keep 1 employee and fire the other 9", unless the business can find a way to suddenly expand 10x so that it needs 10x as much work done.


> So the new baseline is actually "let's keep 1 employee and fire the other 9", unless the business can find a way to suddenly expand 10x so that it needs 10x as much work done.

If they have any surplus of money (or loans) they'll try, so those 9 employees may end up becoming team leads or middle management, trying to start new initiatives to get the 10x expansion (and 100x improvement).

The market isn't anywhere near efficient enough to directly translate productivity improvements into labor reductions. Thankfully, because everything that's nice and hopeful and human lives within the market inefficiency; a fully efficient market would be a hell worse than any writer or preacher ever imagined.


lol that has nothing to do with market efficiency.

I’ve seen a number of your posts where you talk about topics you clearly are not all that well versed in, with such confidence when you’re plain wrong.


Of course it does have to do with market efficiency, of which the inertia and surplus within companies (especially large ones) is a part.

> I’ve seen a number of your posts where you talk about topics you clearly are not all that well versed in, with such confidence when you’re plain wrong.

I'm sure it's true. However, since you brought it up, can you be more specific and name three?


Yes, but in the long run, the market expects growth and innovation, not just doing the same thing with fewer workers. Especially when every other company can just buy the exact same advantage for the same price.


Your first paragraph is so short sighted that its message didn't even make it beyond the next one. It's a race to the bottom and your "doing whatever the fuck I want" will obviously never materialize.

The typical work week today is 40 hours. Just like it was 80 years ago. The typical worker is dramatically more productive than 80 years ago yet "doing whatever the fuck I want" time has not increased. Why would it? Employers don't need to pay such that 20 hour work weeks give you the same income. Because everybody around you is ok with working 40 hours.

This won't be different with AI, no matter if the overall effect is 1.1x or 10x or 100x productivity. Because it's not a technological problem but a sociological one.


Good point. My rant assumed that "10x productivity" meant 10x output in 1x time, rather than 1x output in 0.1x time. Only one of those are actually objectionable.


> I mean, who the hell wants to be 10X more productive without a commensurate 10X compensation increase? You're just giving away that value to your employer.

Those are productivity increases that got our standard of living to where it is. Fewer people doing the same amount of work has, historically speaking, freed people from their current job, allowing them to work on something else.

It's that analogy of the horse, they used to be farm animals. Now, fewer of them are 'employed' but they're much nicer jobs. I'm not sure if the same is true for us this time around though as new jobs being created have increasingly been highly skilled which means the majority can't apply.


There was a long and great ravine of suffering between the advent of the Industrial Revolution and our time of bounty.


Yep, all those artists, musicians, designers and coders will finally do something productive!


If everyone becomes 10x more productive it won’t mean the companies cash flow 10x’s. Where value is loose there is competition, so in theory everyone should win. Unless nobody else can compete to capture that loose 10x value, in which case congratulations, you are now a unicorn.

Of course in reality in the short term what happens is companies lay off people to increase margins. Times will be tough for workers, and equity keeps gravitating towards those who already had it.


Tasks have value because they take effort to complete.

If you remove the effort from those tasks, they will have no value.

10x the value of 0 is 0


Eh, I’d say the premiums drop, and that there is a residual value that is still left. So maybe 0.1 or 0.2 instead of 0.


>Assuming it actually makes you more productive. I mean, who the hell wants to be 10X more productive without a commensurate 10X compensation increase?

Given sane working arrangements or at minimum presence of remote work, it would be a bit shortsighted not to want to get done with your work in a tenth amount of time. In the very least, you're competing for a promotion against less effective people, all while having more time for yourself. If not, you're building labor market skillset in an efficient way so you can hop to a better employer.


It's interesting how differently people can think.

I couldn't imagine thinking "I'm gonna do this 0.1x as fast as I could, wasting my life away with pointless extra work, to spite my employer"


> I mean, who the hell wants to be 10X more productive without a commensurate 10X compensation increase?

The person who realizes that everybody around them is bow at 10X and if they don't follow suit then they will soon be out of a job.


> I think most people are going to say they dont want it. I mean, why would anyone want a tool that can screw up their bank account? What benefit does it gain them?

I'm not so sure. Matter of marketing and social pressure, big time.

Consider this: "Always-on pervasive google/fb/... login? I think most people are going to say they dont want it. I mean, why would anyone want a tool that would track their every move on the internet?" That could easily have been a statement 20 years ago. And look where we are.


Their solution will be to push mandatory and nonconsensual updates to your devices which limit your device and your freedom in the name of security. Like Google is doing to Android in September. You will no longer be able to install "unverified" software on anything. To address prompt injection attacks they're probably working on an approach where your data all has to be in the cloud and subject to security scans. That's already basically the model for Google Workspace, Google Drive and Chromebooks.

The model will get full access to your data, but in the name of security, you will only be permitted to have data that is cloud-hosted; local storage will effectively just be cache.

The era of the general computer will end, and the products you purchased from these companies will be nonconsensually altered and limited.

I'm so glad I switched to Linux more than a decade ago. At least on the PC there will still be an open source ecosystem for a long time to come, it may have less features but I'm willing to accept that.

Knowing that they can change what you bought overnight with a single nonconsensual update, think very, very carefully about who you purchase all of your future technology from. Google's upcoming nonconsensual degradation of Android should be a lesson for everybody.


>Google's upcoming nonconsensual degradation of Android should be a lesson for everybody.

Google is almost certainly doing this because the iOS was not found to be a monopoly, while Andorid was. It came up in Google's appeal of the Epic case verdict, where they directly asked the judge about it. Turns out you can't be anti-competitive if you don't have [allow] any competitors.


Nope. I'm still going to blame Google for their own actions. Nice try, though. I'm old enough to remember when Google pretended to take responsibility for not being evil. Even had it as their motto.


> I'm so glad I switched to Linux more than a decade ago. At least on the PC there will still be an open source ecosystem for a long time to come, it may have less features but I'm willing to accept that.

Wait until age verification is mandatory everywhere. :)

I can already see that happening, e. g. to access financial transactions or government apps, one needs to verify the id, and that will not work without age verification that can not be tampered with. So Linux will either submit to the same or be excluded.

(That free developers will be able to run Linux fine for much longer will also be true, but I guess they only care about catching the 95%, not the 5% linux users ... and 5% is a high guesstimate).

Edit: To clarify the above, one already had to provide personal data for financial transactions, of course, so a bank knows who is who, but the recent age verification go hand in hand with the attempt to get rid of vpn, and applications now make it a new standard to query the age of users, with the claim to "help protect kids". And some people buy into that rationale too. I don't, but I have seen many non-tech savvy people submit to that justification.


There's always the zero knowledge proof tech alternative, but I don't have the feeling we are moving in that direction - it's not the most profitable business is it.


No, nor is it most amenable to mass surveillance.


> It'll be like (lack of) privacy on the internet caused by the ad industrial complex, but much worse and much more invasive.

The concerning aspect is how others' content being scanned into systems don't have any knowledge or consent. Having private PII/files/code/emails/etc being read and/or accidentally shared by the agent online.


> Anyone trying to stay safe will be on the gradient to a Stallmanesque monastic computing existence.

Honestly, it's alright.

Just think of what we could do with computers up until this point. We keep all those abilities.

And more, even, because the industry still keeps churning out new local LLMs. So you even gain more capabilities than right now. Just not at the rate of the bleeding edge.

Which is just like the Linux desktop, essentially. It's fine, really. There is no need to consume the bleeding edge. You will be fine.


Definitely agree here. Made the swap to Linux a little over a year ago and the only reason I even have nice hardware is because I like gaming. But if I was cut off from everything tomorrow, the decades of stuff I have that I have not played will keep me very happy lol


>Anyone trying to stay safe will be on the gradient to a Stallmanesque monastic computing existence.

As a proud neo-luddite, I'm watching the AI hype with grim amusement and I'll tell you hwhat, it doesn't look like a good time. Even putting to one side the planetary scale economic crash that is incoming, all the hypers seem to be on some sort of treadmill that is out of their control and it simply doesn't look like fun.


Do you think that avoidance is going to protect you from the fall-out?


Everyone keeps saying how essential it all is yet a few years in and I still don’t see anything like the promised future of “everyone using them every day for everything.” Everyone’s just constantly talking (or stressing) about it.

We - including the companies - don’t know what the real “billion dollar application” of them is other than the unproven claim it makes everyone more productive in some general sense. When it doesn’t work people continue to say “it’s your fault not the tool’s.” Meanwhile investors are getting skittish and not one AI company is profitable yet. Companies that laid people off for LLM’s are regretting their decisions, leadership (and educators) is dealing with unvetted writing and having to waste their time cleaning it up, the list goes on. “Slop” is still a huge and growing problem.

LLM’s are here to stay, but IMO it’ll be more relevant in the long run than 3D printers yet less revolutionary than the internet. Everyone will touch them at various points but this whole-life, every-industry-disrupted integration still seems far fetched to me. Pricing is still a huge unsolved problem - everyone is still subsidized and despite gains in using fewer resources, it’s still too much to run these locally, even small models (not even getting into tooling and knowledge required to use them in a productive way).

When we zoom out and look at the whole picture, LLM’s have mostly made everyone’s online experience worse while the VC funded companies behind them are playing municipal and state governments’ for suckers a la Amazon getting so many cities to trip over each other giving away land and tax breaks, but far worse. Those are the biggest contributions so far aside from anecdotes from coders about “1000x productivity.” Again, I think they’re here to stay. But it’s called “AI hype” for a reason.

LLM’s have mostly been a problem creator IME rather than a “disruptor.” Never really seen “revolutionary technology” quite like it.

But hey, I’ll admit it’s useful to have a meh local model when I’m writing TTRPG stuff and have writer’s block. Though then I remember how it was trained, a whole other subject I haven’t even touched, so that kind of sucks too.


Yes, mainly because I will continue to know the difference between a truth and a lie.


2-3 news stories of people having bank accounts cleared and the product is dead on arrival.


You'd think so but all the evidence so far points to the contrary. Most people seem perfectly happy to trade security and privacy for convenience.


I dont see companies doing that. it can be business ending. only AI bros buying mac mini in 2026 to setup slop generated Claws would do that but a company doing that will for sure expose customer data.


Big companies are exposing customer data all the time, and they are doing all fine. The more criminal negligence, the richer.


> For all the benefits that agents offer, they can be asymmetrically harmful. This is not a solved issue.

Strongly agreed.

I saw a few people running these things with looser permissions than I do. e.g. one non-technical friend using claude cli, no sandbox, so I set them up with a sandbox etc.

And the people who were using Cowork already were mostly blind approving all requests without reading what it was asking.

The more powerful, the more dangerous, and vice versa.


> I saw a few people running these things with looser permissions than I do. e.g. one non-technical friend using claude cli, no sandbox, so I set them up with a sandbox etc.

People have different levels of safety-consciousness, but also different tolerances and threat models.

For example, I would hesitate running a Mythos-level model in YOLO mode with full control over my computer, but right now, for personal stuff, even figuring out WTF are sandboxes in Claude Code / Gemini CLI, much less setting them up, is too much hassle. What's the worst it can do without me noticing? Format the drive and upload some private data into pastebin? Much as I hate cloud and the proliferation of 2FA in every service, that alone means it can't actually do more to me than waste few hours of my life, as I reimage my desktop and restore OneDrive (in case of destructive changes that got synced up). These models are not yet good enough to empty my bank account in few minutes I'm not looking; everything else they can do quickly is reversible or inconsequential.

Now, I do look at things closely when working with agentic AI tools. But my threat model is limited to worrying about those few hours of my life. `rm -rf / --no-preserve-root` is an annoyance, not a danger.

(I accept that different contexts give different threat modeling. I would be more worried if I were doing businessy business stuff with all kinds of secret sauces, or was processing PII of my employer's customers, or lived in a country where it's easy to have all your money stolen if your CC number or SSN gets posted online.)


How many of these threat vectors are just theoretical? Don’t use skills from random sources (just like don’t execute files from unknown sources). Don’t paste from untrusted sites (don’t click links on untrusted sites). Maybe there are fake documentation sites that the agent will search and have a prompt injected - but I haven’t heard of a single case where that happened. For now, the benefits outweigh the risk so much that I am willing to take it - and I think I have an almost complete knowledge of all the attack vectors.


Systems have been caught out that review pull requests, that’s a simple and clear one. The more obvious to me for most people is anything you do that interacts with your email without an explicit approve list of emails to read.


Yes, but none of this applies to the local codex agent that runs when I tell it to and has access to my computer. Like: „scan this folder of PDFs and create an excel file with all expenses. Then enter them into my tax software.“ This needs access to very sensitive data and involves a quite complex handling of data. But the only attack vector I see is someone injecting prompts into my invoice files.


Which applies if you were to do this to invoices submitted to you, rather than ones you created, or if you have any way of user info getting into your invoices.


The problem is that any data now becomes effectively an executable.

> I think I have an almost complete knowledge of all the attack vectors.

That's exactly the kind of hybris where the maximum danger lies.


i think you lack creativity. you could create a site that targets a very narrow niche, say an upper income school district. build some credibility, get highly ranked on google due to niche. post lunch menus with hidden embedded text.

the attack surface is so wide idk where to start.


Why would my agent retrieve that lunch menu?


Because it’s hooked up to a microphone in your kitchen & your kid is arguing with you about what lunch they want & they say “Hey [agent], what day is pizza day at [school]?”


I’m not doing that. That would be like giving my child shell access to my system.


Funny joke,

But for real, obviously we all know people use agents to pick restaurants and that's a legit vector.

I agree it's not the biggest surface, but it's worth knowing imdo


I cannot reconcile that growth for non-technical users is going to explode, when most utility from agents is via the ability to execute arbitrary code, generally in yolo mode, with the fact that almost all corporate IT departments do not give users the ability to install anything on their machine, let alone arbitrary code. Even developers at many companies are subject to this despite the productivity impacts.

The culture of corporate IT would need to change to allow it, and I just don't see it happening.


What about setting environments for normies that mitigate this problem? I don't know that you can do it on Windows, but Linux offers various tools for isolation where you can give full rights to an LLM and still be safe from certain classes of disaster.

Maybe this kind of isolation neuters the benefit you're thinking of, but I do believe some sort of solution could be reached.


"Isolation" and "full rights" are mutually exclusive, contradictory properties.


This is me!

I’m semi-normie (MechEng with a bit of Matlab now working as a ceo).

I spend most of my day in Claude code but outputs are word docs, presentations, excel sheets, research etc.

I recently got it to plan a social media campaign and produce a ppt with key messaging and content calendar for the next year, then draft posts in Figma for the first 5 weeks of the campaign and then used a social media aggregator api to download images and schedule in posts.

In two hours I had a decent social media campaign planned and scheduled, something that would have taken 3-4 weeks if I had done it myself by hand.

I’ve vibe coded an interface to run multiple agents at once that have full access via apis and MCPs.

With a daily cron job it goes through my emails and meeting notes, finds tasks, plans execution, executes and then send me a message with a summary of what it has done.

Most knowledge work output is delivered as code (e.g. xml in word docs) so it shouldn’t be that that surprising that it can do all this!


How does this obviate the need for software? In order for what you asked to be possible, Word, Excel, PowerPoint, and Figma all still need to exist and you need licenses for them.

If you can figure out the next step and say "Claude, go find me buyers and sell shit for me without using any pre-existing software," have at it. It can't be social media, I guess, since social media is software and Claude is supposed to get rid of software.

At a certain point, why do we even need computers? Can't we just call Claude's hotline and ask "Claude, please find a way to dump $40 million in cash into my living room. Don't put it in my bank account because banks use software."


It doesn't remove the need for software, but it greatly reduces the number of tools needed or doesn't mandate building custom tools that might not be viable due to very specific needs many users have.

OP gave a good example how their workflow was changed, you could argue there are tools that could've done that, but they managed to achieve their goals without them, have something that fits their workflow perfectly, is fine tuned in case of changes, and with a few other tools (Word, Excel, Figma) they can do all sorts of things which would've required a small team or far more (expensive) tools to execute.

To me that is a great example of non-developers using tools to enhance their workflows and with initiatives like from this topic, I can only see that increasing.


> How does this obviate the need for software?

It doesn't obviate the need for software, but it greatly devalues software products, as they become reduced to tool calls for LLMs.

This is good for users, because software products are defined by boundaries - borders drawn around the code to focus and package functionality, yes, but also to limit interoperability and create a sales channel (UX being the perfect marketing platform for captive audience).

After all, I don't usually want to play with Word, Excel, PowerPoint, and Figma - they're just standing between me and the artifact I want to create, so if I can get LLM to operate them for me, I don't have to deal with all the UX and marketing bullshit those products throw at me.

I mean, that's what I'd do if I could afford to hire a person to operate those tools for me. That, again, is the best mental model for LLMs - they're little people on a chip, cheaper to employ than actual people.


> I mean, that's what I'd do if I could afford to hire a person to operate those tools for me. That, again, is the best mental model for LLMs - they're little people on a chip, cheaper to employ than actual people.

Sounds like more of a threat to people than software then.

I get the point that if an agent could generate a presentation by directly writing to some open format with a free viewer then PowerPoint would be out of the picture.

However the tool has to be pretty close to 100% for that to work. If I have a presentation that's 90% there it's probably going to be a lot easier to finish it off manually in Powerpoint than try different variants of prompts. In which case I'll still need that Powerpoint license.


> In order for what you asked to be possible, Word, Excel, PowerPoint, and Figma all still need to exist and you need licenses for them.

Or not. Besides, the better AI models can effortlessly generate Latex/Beamer, a far superior solution for typesetting and presentations. Anything than can be done in Excel can be done in Python. Those proprietary tools are a thing of the past, no one should use them anymore.


And the value of those marketing campaigns is going to zero, since everyone is doing it. Even self employed people.

Pay for ads or you get lost in the mass of posts


> My current expectation is that the Cowork/Codex set of "professional agents" for non-technical users will be one of the most important and fastest growing product categories of all time, so far.

I disagree. There is a major gap between awesome tech and market uptake.

At this point, the question is whether LLMs are going to be more useful than excel. AI enthusiasts are 100% sure that it’s already more useful than excel, but on the ground, non-technical views do not reflect that view.

All the interviews and real life interactions I have seen, indicate that a narrow band of non-technical experts gain durable benefits from AI.

GenAI is incredible for project starts. A 0 coding experience relative went from mockup to MVP webapp in 3 days, for something he just had an idea about.

GenAI is NOT great for what comes after a non-technical MVP. That webapp had enough issues that, if used at scale, would guarantee litigation.

Mileage varies entirely on whether the person building the tool has sufficient domain expertise to navigate the forest they find themselves in.

Experts constantly decide trade offs which novices don’t even realize matter. Something as innocuous as the placement of switches when you enter the room, can be made inconvenient.


> market uptake.

I think the market uptake of Claude Cowork is already massive.


Estimated users are at 18-30 mn, and we are talking about non-technical users.


> My current expectation is that the Cowork/Codex set of "professional agents" for non-technical users will be one of the most important and fastest growing product categories of all time, so far.

I agree this is going to be big. I threw a prototype of a domain-specific agent into the proverbial hornets' nest recently and it has altered the narrative about what might be possible.

The part that makes this powerful is that the LLM is the ultimate UI/UX. You don't need to spend much time developing user interfaces and testing them against customers. Everyone understands the affordances around something that looks like iMessage or WhatsApp. UI/UX development is often the most expensive part of software engineering. Figuring out how to intercept, normalize and expose the domain data is where all of the magic happens. This part is usually trivial by comparison. If most of the business lives in SQL databases, your job is basically done for you. A tool to list the databases and another tool to execute queries against them. That's basically it.

I think there is an emerging B2B/SaaS market here. There are businesses that want bespoke AI tools and don't have the discipline to deploy them in-house. I don't know if it is ever possible for OAI & friends to develop a "hyper" agent that can produce good outcomes here automatically. There are often people problems that make connecting the data sources tricky. Having a human consultant come in and make a case for why they need access to everything is probably more persuasive and likely to succeed.


> The part that makes this powerful is that the LLM is the ultimate UI/UX.

I strongly doubt that. That’s like saying conversation is the ultimate way to convey information. But almost every human process has been changed to forms and structured reports. But we have decided that simple tools does not sell as well and we are trying to make workflow as complex as possible. LLM are more the ultimate tools to make things inefficient.


>The part that makes this powerful is that the LLM is the ultimate UI/UX

Seems pretty questionable to me. Describing things in natural language can be quite imprecise and verbose.


>UI/UX development is often the most expensive part of software engineering.

I disagree with this as a blanket statement. At least in the tech world (i.e. tech companies that build technology products), UI/UX is often less expensive than the platform and infrastructure parts of the technology products, certainly at any tech that runs at scale.


> There are businesses that want bespoke AI tools and don't have the discipline to deploy them in-house. I don't know if it is ever possible for OAI & friends to develop a "hyper" agent that can produce good outcomes here automatically. There are often people problems that make connecting the data sources tricky. Having a human consultant come in and make a case for why they need access to everything is probably more persuasive and likely to succeed.

Sort of agreed, though I wonder if ai-deployed software eats most use cases, and human consultants for integration/deployment are more for the more niche or hard to reach ones.


I am starting to use Codex heavily on non-coding tasks. But I am realizing it works because I work and think like a programmer - everything is a file, every file and directory should have very precise responsibilities, versioning is controlled, etc. I don't know how quick all of this will take to spread to the general population.


Maybe. The point is that in case of software it is fairly easy to verify if that what LLM produced is correct or not. Compiler checks syntax, we can write tests, there is whole infrastructure for checking if something works as expected. In addition, LLM are just text generating algorithms and software is all about text, so if LLM see 1 000 000 a CRUD example in Python, it can generate it easily, as we have a lot of code examples out there thanks to open source.

That's why LLMs shine in coding tasks. If you move to other parts of engineering, like architecture, construction or stuff like investment (there is no AI boom there, why?) where there is no so much source text available, tasks are not so repeatable like in software, or verification is much more complicated, then LLM-s are no longer that useful.

In software also I believe we will see soon that a competitive advantage have not those who adopted LLM, but those who did not. If you ask LLM what framework/language/approach use for a given task, contrary to what people think, LLM is not "thinking", it just generates text answer on the base of what it was trained on, so you will get again and again same most popular frameworks/langs/approaches suggested, even if there is something better, yet not that popular to get into model weights in a significant way.

Interesting times, anyway.


LLMs nowadays make aggressive use of web search. Thus they don't answer only on the base of what they were trained on.

I don't think they are much more prone to using only the same popular frameworks, especially if you ask them to weigh for options.


I keep seeing sentiment like this. I work for a relatively cutting edge healthcare enterprise as a sysadmin, and we've only just been given access to copilot chat. I don't think we're going to be having agents doing work for us any time soon.


> My current expectation is that the Cowork/Codex set of "professional agents" for non-technical users will be one of the most important and fastest growing product categories of all time, so far.

They won't.

Non-technical users expect a CEO's secretary from TV/movies: you do a vague request, the secretary does everything for you. LLMs cannot give you that by their own nature.

> And eventually will the UI/interface be generated/personalized for the user, by the model?

No. Please for the love of god actually go outside and talk to people outside of the tech bubble. People don't want "personalized interfaces that change every second based on the whims of an unknowable black box". They have plenty of that already.


> Non-technical users expect a CEO's secretary from TV/movies: you do a vague request, the secretary does everything for you. LLMs cannot give you that by their own nature.

Most people are indifferent to computers. A computer to them is similar to the water pipeline or the electrical grid. It’s what makes some other stuff they want possible. And the interface they want to interact with should be as simple as possible and quite direct.

That is pretty much the 101 of UX. No deep interactions (a long list of steps), no DSL (even if visual), and no updates to the interfaces. That’s why people like their phone more than their desktops. Because the constraints have made the UX simpler, while current OS are trying to complicate things.

So Cowork/Codex would probably go where Siri is right now. Because they are not a simpler and consistent interface. They’ve only hidden all the controls behind one single point of entry. But the complexity still exists.


Just yesterday my non-technical spouse had to solve a moderately complex scheduling problem at work. She gave the various criteria and constraints to Claude and had a full solution within a few minutes, saving hours of work. It ended up requiring a few hundred lines of Python to implement a scheduling optimization algorithm. She only vaguely knows what Python is, but that didn't matter. She got what she needed.

For now she was only able to do that because I set up a modified version of my agentic coding setup on her computer and told her to give it a shot for more complex tasks. It won't be trivial, but I do think there's a big opportunity for whoever can translate the experience we're having with agentic coding to a non-technical audience.


There's no such big opportunity, as the number of programmers' spouses is quite limited. Again, and as the GP rightly suggested, some of the HN-ers here need to go and touch some normie grass, so to speak.

More to the point, nobody wants to be more efficient for the sake of being efficient, we all want to go to work, do our metaphorical 9 to 5 without consuming too much (intellectual and not only) energy, and then back home. In that regard AI is seen as an existential threat to that "lifestyle" and it will be treated as such by regular workers.


correct. you cant trust this place for realistic takes - I had a post re. financial stuff downvoted when a former Investment Banker chimed in to back me up.

Comical. Truly comical.


> Just yesterday my non-technical spouse

> It ended up requiring a few hundred lines of Python

And she knows those a hundred lines of python work correctly and give her correct result because in this instance Claude managed to produce a working result. What if it didn't? Would vague knowledge of Python have helped her?

> It won't be trivial, but I do think there's a big opportunity for whoever can translate the experience we're having with agentic coding to a non-technical audience.

Even though I agree with the sentiment, we've tried non-coding coding how many times now? Once every 5 years? Throwing LLMs into the mix won't help much when in the end you leave the end user hanging, debugging problems and hunting for solutions.


Scheduling solutions are easy to verify. For other problems, verification would be harder.


> Non-technical users expect a CEO's secretary from TV/movies: you do a vague request, the secretary does everything for you. LLMs cannot give you that by their own nature.

What are you using today? In my experience LLMs are already pretty good at this.

> Please for the love of god actually go outside and talk to people outside of the tech bubble.

In the past week I've taught a few non-technical friends, who are well outside the tech bubble, don't live in the SF Bay Area, etc, how to use Cowork. I did this for fun and for curiosity. One takeaway is that people at startups working on these products would benefit from spending more time sitting with and onboarding users - they're very powerful and helpful once people get up and running, but people struggle to get up and running.

> People don't want "personalized interfaces that change every second based on the whims of an unknowable black box". They have plenty of that already.

I obviously agree with this, I think where our view differs is I expect that models will be able to get good at making custom interfaces, and then help the user personalize it to their tasks. I agree that users don't want something that changes all the time. But they do want something that fits them and fits their task. Artifacts on Claude and Canvas on ChatGPT are early versions of this.


> What are you using today? In my experience LLMs are already pretty good at this.

LLMS are good at "find me a two week vacation two months from now"?

Or at "do my taxes"?

> how to use Cowork.

Yes, and I taught my mom how to use Apple Books, and have to re-teach her every time Apple breaks the interface.

Ask your non-tech friends what they do with and how they feel about Cowork in a few weeks.

> I think where our view differs is I expect that models will be able to get good at making custom interfaces, and then help the user personalize it to their tasks.

How many users you see personalizing anything to their task? Why would they want every app to be personalized? There's insane value in consistency across apps and interfaces. How will apps personalize their UIs to every user? By collecting even more copious amounts of user data?


"LLMS are good at "find me a two week vacation two months from now"?"

Of course they are. I gave one a similar prompt a few weeks ago, albeit quite a bit more verbose (actually I just dictated it, train of thought, with couple of 'eh actually, forget what I just said about x, do y instead") and although I wasn't brave enough to give it my credit card and finalize the bookings, it would have paid for the bookings I had it set up for me, had I done that. I gave it some RL constraints, like "we're meeting friends in place xyz at such and such date, make sure we're there then" and it did everything from watching we wouldn't be spending too many hours driving per day to check that hotels are kid friendly to things to do and see and what public holidays there are so that we know when supermarkets close early and a bunch of details I wouldn't have thought of. It checked my (and my wife's) calendar, checked what I had going on work wise, etc.

That is a fully solved 'problem' man. LLMs will run the whole thing for you. Just provide it with the login details to booking websites and you're off to the races.

I did have it upgrade the car, even if that pushed the cost outside the budget I gave it. Next time it'll know LOL.


>although I wasn't brave enough to give it my credit card and finalize the bookings

So it's not trustworthy enough for you, someone clearly interested in the hype of LLMs.


It's a matter of getting used to things. We're only a few weeks further, I maybe would have given it now. It'd need some way to keep it private I guess, maybe I could have used a one off CC number. Those are just technicalities at this point. It got me to the point where I just had to enter my details and click a few confirm buttons. Those are solved problems. I'm not sure why the denialists here are saying those things are 'impossible'. I mean I've seen them happen, what do you want me to say? Claiming this is 'just hype' is ostrich behavior. I've been playing with an abliterated Gemma 4 yesterday on my local machine. Yes it would take longer and require a bunch of harness fiddling, but even if OpenAI and Anthropic would collapse tomorrow, I'm confident I could still do the exact same thing the day after with with what I have right now on my hard disk. I'm not sure what you want me to tell you mate. Yes there's rough edges to work out or just in general workflows to improve but the ideas are way beyond 'proof of concept'. There's people like myself using these things for purposes that 6 months ago were science fiction. I don't care if you believe me or not, I'm just some dude on the internet, but level of delusion on how 'inferior' these models (with proper harnessing) are is mind boggling for someone like me who sees it happen literally 20 centimeters to the side on my screen from where I see people claim that those things are impossible.


> Or at "do my taxes"?

codex did my taxes this year (well it actually implemented a normalization pipeline and a tax computing engine which then did the taxes, but close enough)


> well it actually implemented a normalization pipeline and a tax computing engine which then did the taxes, but close enough

You can't seriously believe laymen will try to implement their own tax calculators.


of course not.

what I believe is that laymen will put all their tax docs into codex and tell it to 'do their taxes' and the tool will decide to implement the calculator, do the taxes and present only the final numbers. the layman won't even know there was a calculator implemented.


> the layman won't even know there was a calculator implemented.

That's on company making the agentic harness. Hiding details of what computer does from the user is the original sin of this industry, and subsequent generations of developers and software companies keeps doubling down on it.

(Case in point - I just downloaded the Codex app for Windows, and in the options I see it has two UI modes of operating, one of which is meant for "non coding" and apparently this means hiding the details of what the agent is doing. This is precisely where the layman is betrayed by the tool.)


Yeah, good luck trusting the output!


check back in a couple of years!


Ah right! Reminds me of AGI by 2025 :D


If your prompt was more complex than "do my taxes", then this is irrelevant.


it was many hours of working with codex, guidance and comparing to known-good outputs from previous years, but a sufficiently smart model would be able to just do it without any steering; it'd still take hours, but my input wouldn't be necessary. a harness for getting this done probably exists today, gastown perhaps or something that the frontier labs are sitting on.


If you can assume "a sufficiently smart piece of technology" that doesn't exist now, a lot of problems become trivial


yes.

but then, respect the trendline, especially if it's exponential.


Is it exponential or logistic?


> but a sufficiently smart model would be able to just do it without any steering;

Yeah, yeah, we've heard "our models will be doing everything" for close to three years now.

> a harness for getting this done probably exists today, gastown perhaps

That got a chuckle and a facepalm out of me. I would at least consider you half-serious if you said "openclaw", at least those people pretend to be attempting to automate their lives through LLMs (with zero tangible results, and with zero results available to non-tech people).


Sounds fascinating! If you wrote an article on this I bet it'd have a good shot at making it to the home page of HN.


> LLMS are good at "find me a two week vacation two months from now"?

Yes?

===

edit: Just tested it with that exact prompt on Claude. It asked me who I was traveling with, what type of trip and budget (with multiple choice buttons) and gave me a detailed itinerary with links to buy the flights ( https://www.kayak.com/flights/ORD-LIS/2026-06-13/OPO-ORD/202... )


I'd love to try and replicate, but I'm not letting any of these tools anywhere near a real browser and capabilites :)


Perfect - and this use case will be enshitificated first. LLM provider will charge small fee for proper recommendation placing. Got to recoup investment.


This is effectively how I treat my AI agents. A lot of the reason this doesn't work well for people today is due to context/memory/harness management that makes it too complex for someone to set up if they don't want a full time second job or just like to tinker.

If you productize that it will be an experience a lot of people like.

And on the UI piece, I think most people will just interact through text and voice interfaces. Wherever they already spend time like sms, what's app, etc.


Most knowledge workers aren't willing to put in the effort so they're getting their work done efficiently.


Maybe but the product category is not necessarily a monolith in the same way that Claude Code is. These general purpose tools will have to action across a heterogeneous set of enterprise systems/tools. A runtime environment must be developed to do that but where that of the agent ends and that of the enterprise systems begins is a totally open question.


> A runtime environment must be developed to do that but where that of the agent ends and that of the enterprise systems begins is a totally open question.

I think something like SQL w/ row-level security might be the answer to the problem. You often want to constrain how the model can touch the data based upon current tool use or conversation context. Not just globally. If an agent provides a tenant id as a required parameter to a tool call, we can include this in that specific sql session and the server will guarantee all rules are followed accordingly. This works for pretty much anything. Not just tenant ids.

SQL can work as a bidirectional interface while also enforcing complex connection level policies. I would go out of band on a few things like CRUD around raw files on disk, but these are still synchronized with the sql store and constrained by what it will allow.

The safety of this is difficult to argue with compared to raw shell access. The hard part is normalizing the data and setting up adapters to load & extract as needed.


> Maybe but the product category is not necessarily a monolith in the same way that Claude Code is. These general purpose tools will have to action across a heterogeneous set of enterprise systems/tools.

What would make it not be a monolith? To me it seems like there'll be a big advantage (e.g. in distribution, user understanding) for most people to be using the same product / similar interface. And then the agent and the developer of that interface figure out all the integrations under that, invisible to the user.


I mean there is a runtime layer that needs to be developed, and some of it may live in CC/Codex and some might live in the various enterprise systems. Someworkflow automations and some amount of the semantic layer may for instance exist in your CRM/ERP/data platform. Yes the front-end would be owned by the chat interface, but part of the solution may exist in the various enterprise systems. This would be closer to a distributed system than a monolith. The demos and marketing language point to this as the direction of travel (i.e. the reference to Atlassian Rovo, etc.).


Thanks for answering!


I think the coding market will be much larger. Knowledge work is kind of like the leaf nodes of the economy where software is the branches. That's to say, making software easier and cheaper to write will cause more and more complexity and work to move into the Software domain from the "real world" which is much messier and complicated.


Yes, and the same thing will happen in non-coding knowledge work too. Making knowledge work cheaper will cause complexity to increase, more knowledge work.


I don't think so, the whole point of writing software is it is a great sink for complexity. Encoding a process or mechanism in a program makes it work (as defined) for ever perfectly.

An example here is in engineering. Building a simulator for some process makes computing it much safer and consistent vs. having people redo the calculations themselves, even with AI assistance.


The history of both knowledge work and software engineering seems to be increasing in both volume and complexity, feels reasonable to me to bet on both of those trendlines increasing?


Yes, I have a theory - that higher efficiency becomes structural necessity. We just can't revert to earlier inefficient ways. Like mitochondria merging with the primitive cell - now they can't be apart.


I still think we're several "my agent sent an inappropriate email to all my contacts" away from people figuring out proper security controls for these things


I agree, and I think this extends to programming too. A lot of of software practices are built on the expectation humans are writing, reviewing and shipping code with that quickly becoming the case, processes, practices and even programming languages themselves will evolve to what agents need, rather than humans.

a version of Conway's law aimed specifically at agentic communication rather than human.


really struggling to understand where this is coming from, agents haven't really improved much over using the existing models. anything an agent can do, is mostly the model itself. maybe the technology itself isn't mature yet.


My view is different. Agent products have access to tools and to write and run code. This makes them much more useful than raw models.


Yes, I think they unlock a whole new level of capability when they have a r/w file system (memory), code execution and the web.


That's not the model, that's the box the model came in.

It's unlikely we've hit the limits on improving agent UX, but there are some fundamental limits on LLMs that seem unlikely to be fixed by better UX.


You know what happens to a predator who makes its prey go extinct?

AI is doing the same


Totally agree, AI interfaces will become the norm.

Even all the websites, desktop/mobile apps will become obsolete.


AI won't kill apps, it will just change who 'clicks' the buttons. Even the most powerful AI needs a source of truth and a structured environment to pull data from. A world without websites is a world where AI has nothing to read and nowhere to execute. We aren’t deleting the UI. We’re just building the backends that feed the agents.


I was sent this and thought it looked pretty interesting. I and others collected more sandbox tools here: https://news.ycombinator.com/item?id=47102258.


Surprised that no one commented on the clever title!


I had to read it 15 times to understand what it was trying to say. There is such a thing as trying too hard at being clever.


I think this is smart and very interesting. I see it like an aggregator marketplace. A powerful position to be in.

Cloudflare, GitHub (if they shipped more), Anthropic and OpenAI are also in decent positions to do this.

I wrote notes on this previously [1]. If you believe agents are going to be big consumers, it's helpful to make things that today allow users of agents to easily discover and purchase services via apis.

[1]: https://x.com/chrisbarber/status/2026331038994321898


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: