But they try to play it off as though this were public data:
> Public sharing settings across AI and SaaS products have surfaced similar findings in recent months. Anthropic addressed exposed public artifacts across Claude and its MCP ecosystem via Google Search.
Also, interesting, they are SOC2 compliant [1], proving again that SOC2 is meaningless/useless.
I used to work for a company seeking SOC2 compliance. They told me that I had to install corporate malware because of the compliance. I didn't want to install it on my personal computer, which I had been using for work. They sent me a company computer. I installed the corporate malware on that one. I set the company computer aside and continued working on my personal computer. No SOC2 compliance was harmed in the process.
The company computer typically comes with some data protection agreement, where you agree to only access confidential data via the company computer, or at least to never copy confidential company data to personal devices.
On the one side you are liable because you are accessing company data on your personal computer. On the other you are liable because the company is forcing you to use insecure devices to access company data because of compliance. I think we need to check the contracts to see which liability to choose that will give you the least amount of headache.
I know, I was being coy, but if I'm being honest using company devices makes me really anxious, it's always in the back of my head that someone is going to abuse that outdated and opaque VPN stack from Fortinet or that Kaspersky Daemon I needed to install using a script some guy from security sent me with a Google Docs link over DM, and I'll have a really hard time explaining there was no mishandling of data from my part
I used a product with SOC2 certification, which uploads all your chatbot conversations to a server they control (mandatory), potentially including source code and other proprietary data, which can be made visible to public with a single click. Doesn't matter if you are an individual or enterprise user.
They do have enterprise level controls that let admins turn this off. Unfortunately, it is on by default, and some of those basic security controls require a higher tier of service.
It is absolutely wild that these companies treat security like an afterthought. And I also realized SOC2 Compliant meant absolutely nothing.
SOC2 requires a company to write policies in a large number of areas, and to demonstrate that they're complying with the policies they wrote. AFAIK SOC2 does not require anything meaningful about the actual contents of the policies, nor does it require the policies to remain constant.
Besides the downplaying and obfuscation about the timeline on the first half, I find the inclusion of anthropic and zoom examples to be wild. Just spraying in all directions.
On a personal note, I recognize that I should have kept the researcher updated after his initial outreach earlier this year, and I take full responsibility for that communication gap.
They make it sound like it was a single email. What about all the other outreaches the researcher made to the CEO over a six month period?
Interesting how the CEO didn't contribute any explanation to the blog post and left the CTO out to dry.
I know some people want to run their agents when their computer is off, but I imagine a solution like this will be much more common than paying for a remote sandbox (i.e on fly.io or exe.dev), especially because it'll be free.
Though, they need to remove the login requirement.
1. You can run sandboxes locally
2. You can control them securely over the internet, for when you're on the go
3. You can migrate them to cloud VMs if you want
If I can toot my own horn, I'm trying to build that :)
Still early and the local sandboxes are experimental right now
Where are those media files from? It's pretty unlikely to encounter an RCE in media files these days. but if you are an important enough target, one could be crafted for your device or software (i.e vlc)
At one point in the article, the author asks Cloudflare's bot if they're launching a Wallet product, and it says no.
> There is no such product in our documentation or dashboard, so treat any email, website, or message claiming to be "Cloudflare Wallet" as a phishing attempt.
What's the point of adding these AI chatbots if they're hopelessly uninformed about your products?
"It's not about utility. It's not even really about the chatbot. It's about visibility, the fear of looking behind. A website without a chatbot in 2026 risks feeling unfinished, like something's missing. Even if what's missing is a half-broken widget that most visitors dismiss in three seconds. The chatbot has become a social signal, not a tool. A way of saying: we're keeping up."
heh yeah I ran into this with one of the few times I used claude desktop. it had no idea what features it had and didn't have, where buttons were in the app, etc. isn't that kind of a core category of knowledge you'd want the chatbot to know?
in my experience, usually it knows this (it is in the system prompt) but it can still get confused. Especially with skills for example, some skills might only work in claude code/outside of sandbox or in desktop but not on web. And it would sometimes not know if it was on the web or desktop.
it is so frustrating to me that claude code still does not understand how the bash tool properly functions and no matter how much I mention how it works in my CLAUDE.md it still misuses it.
> “What's the point of adding these AI chatbots if they're hopelessly uninformed about your products?”
This drives me nuts. But it is a continuum. From help-bots which are just natural language navigation to docs, to the best in class llm’s with access to both knowledge of the company and your data (my fav so far is Shopify). The most annoying are those who read the docs to you like lawyer-bots.
I’m guessing the speed at which companies go from the first type to the last type depend on many factors such as volume of support issues, the expertise of the users, and corporate culture of reliance on “accountability sinks” (someone to be mad at, but who has no authority to help or correct a problem—I’m thinking of the merchant platform Square and their dark pattern navigation that tricks you to suffer the instant fees of the fund-now link trying to find the transfer schedule).
Most chat support people were contractors hired from third party companies that were given dossiers about their products that were often quite out of date because poor management has always been a thing.
Last time Coinbase had a breach I wrote in to support chat to see what I needed to do. They said there was no breach. I sent them a link to their own blogpost about it. They responded with "wow this is the fist time I'm hearing about this."
> At minimum, a half-way decent support person would ask a few people internally or search Slack before answering.
Of course not. The extremely vast majority of support staff aren't connected to "internal people" and certainly don't have any access to the main company's Slack.
Most of all, those people are paid very little on very tight length-per-interaction targets. They can't spend any time at all looking for stuff outside the docs package or chatting with peeps outside the immediate costaff.
Its an organic cycle, as the company grows big the number of relevant areas grow too and inter team communication becomes way too costly/impractical. The big company becomes a group of informal small companies, each running in their own direction and at times competing with each other. Areas like customer support are not really good candidate for career growth so they get least resources and manpower.
Everything I said applies to overseas labor, too. I’ve worked at both types of companies, those that co-locate support with the product teams, and those that pay bottom dollar and don’t care at all about support.
With the latter type, there’s still almost always a line of communication to corporate. And the support staff still try to help. They are decent human beings, even if the end result kind of sucks.
And as a dark pattern it adds "positive friction" for the company reducing the number of people that will have the motivation of obtaining the real people support.
Several weeks ago I had an issue not being able to login to Verizon's website, so I tried to chat with someone. The chatbot that was gatekeeping was predictably useless said it would redirect me to a human except...it kept prompting me to log in first. It was literally impossible to differentiate from if they literally had no humans online to talk to at all.
I could be the best money saver for them - and ask for a hefty premium for my services - by terminating all support. No costs, nada, full save! Genius, right?! Never gives false info, never!
Ok, ok, need to have a tickmark next to the 'support' item in the quarterlies, let it be an eternal spinning wheel presenting on clicking the 'Our award winning instant support is HERE!' button then. Its close to the real experience anyway, right?
The vast majority of people I know who have done phone/chat support have had a spreadsheet (or fancier tool) of common phrases and mandatory boilerplate answers, an overview of the company they're doing it for (sometimes just a couple pages in a word doc), and tone/tool training at a minimum. And generally that tool has a tree of common steps that are also generally mandatory, because they punish rather harshly for going off-script. It's not much, and it gets you the sort of support that people often think of with ticketing systems: impersonal and inaccurate, favoring the company.
But it's rarely this inaccurate. Or if it is (e.g. missing a major product launch), it's fixed in a day or two.
> You have 12 marbles and a balance scale. One of the 12 marbles is inconsistent with the others, meaning it could be heavier or lighter than its peers of normal weight. You are allowed to use the balance scale exactly 3 times to identify which of the 12 marbles is irregular AND determine whether it is heavier or lighter than normal.
I also got this riddle, in 2015. I couldn't solve it. Tbh, I think it's a terrible question. There isn't really a step-by-step problem solving process, but you just need to have a "leap" to realize you can weight the marbles in groups of 3. I also got rejected. It's crazy to get rejected from a one-question riddle like this
Back when I interviewed at MS as a college senior, I had no clue about these "formats". So when someone asked me how I would make (I think they meant design) a web crawler, I started writing one in Python. They didn’t stop me, and it seemed to be awkward to them that I was doing this, whereas I felt awkward that they were acting that way when I was doing what they asked. It was just stupid all around. These systems deserve to be gamed, and the interviewers can't handle anything else, like in TFA.
Google was more elitist than most in its early years. As I recall, candidates had to have graduated from an elite CS program with and provide a transcript when applying showing top grades.
> The key point here is our programmers are Googlers, they’re not researchers. They’re typically, fairly young, fresh out of school, probably learned Java, maybe learned C or C++, probably learned Python. They’re not capable of understanding a brilliant language but we want to use them to build good software. So, the language that we give them has to be easy for them to understand and easy to adopt
> It must be familiar, roughly C-like. Programmers working at Google are early in their careers and are most familiar with procedural languages, particularly from the C family. The need to get programmers productive quickly in a new language means that the language cannot be too radical
One would expect folks coming from an elite CS program, and able to master Google's hiring processes, to be a bit more skilful.
Ironically all those four languages that get mentioned have quite advanced type systems, versus the one for "simple" minds.
> I also got this riddle, in 2015. I couldn't solve it. Tbh, I think it's a terrible question. There isn't really a step-by-step problem solving process, but you just need to have a "leap" to realize you can weight the marbles in groups of 3.
It’s an information theory question similar to the “how far can you drop the egg before it breaks” puzzle. You don’t need a leap of intuition, just remember the right theorems and formulae from college.
I probably would’ve failed the interview. At best I could write an algorithm that empirically arrives at the answer, which may or may not impress the interviewer enough to pass.
Stuff like this is why I dropped out of coding competitions some time in high school. You can always brute force the easy version, then the next tiers drop numbers that cannot be solved unless you know (or can invent) that one weird maths trick.
In the initial state there's 24 (12 * 2) different possibilities: The marble you're looking for is one of 12 and it's either lighter or heavier. By using a balance scale there's 3 possible outcomes (left side is heavier; right side is heavier; same weight). This means that for your last (third) weighing you'll have to have reduced the problem down to 3 (or fewer) different possibilities. If there's 4 or more there's no way to reduce it down to a single possibility. Before the second weighing you should have reduced it down to 9 (3*3) different possibilities.
Once you know this, you can start making educated guesses for the first weighing and quickly eliminate those which makes it impossible to continue. For instance: Splitting the marbles into two. This gives two possible outcomes: Left side is heaver or right side is heavier. For the first outcome it means that either the target marble is lighter and part of the left group (6 marbles) or heavier and part of the right group (6 marbles). That's 12 different possibilities (more than 9) and therefore we know that it's impossible to determine the marble with just two more weighings.
There's still guessing to be done, but at every part of the decision tree you can at the very least quickly avoid exploring paths which are guaranteed to not work.
Now that you spell it out, I guess this just binary search. You need O(log n) weighings because each weighing splits the search space in half. So I was wrong it’s not an information theory question, it’s just comp sci algorithms. I got my college classes mixed up (it’s been 14years)
Exactly. Same kind of thinking gives you bounds on what can be done. For example, can you do 13 marbles in 3 weighings? 213 < 3^3, so maybe (though I think not). Can you do 14 marbles? 214 > 3^3, so definitely no.
I think it's a terrible _interview_ question, partly due to it's "IQ test" nature, but mostly because it's a pretty hard problem and I really wouldn't expect someone to be able to just spew out a solution the first time they hear it.
Of course, the process would be more about the interviewer observing your thought process, seeing how you develop a notation for solving this thing which you almost certainly don't have any pre-existing notation for, etc. I just really wouldn't expect much progress in the space of 5 to 10 minutes, although I am basing this on the "13 marbles" version of the problem.
I'm trying to say that it's all advertising. What's interesting is the technical substance. Judge the post by the signal-to-noise ratio, not by whether the company had a marketing incentive to publish it.
If that ends up being true, they should be very surprised.
We're looking at ~$725B combined hyperscaler capex in 2026 (on a path to $1.08T by 2028) against roughly $25B of AI service revenue in 2025 on $250B+ of infrastructure spend. By 2030, the global data center build-out will require $6.7 trillion in capital expenditure. If hyperscalers require a 25% return on AI-specific capex, the industry needs to generate ~$169B in AI-attributable revenue annually by end of 2028.
A flexible compute market plus three well-capitalized competitors plus Google's internal silicon means nobody gets to hold price. SpaceX's public offering is the best signal we have on this type of thing and it's down 15% from offer price.
It's incredibly unlikely that BOTH OpenAI and Anthropic will be "obscenely" profitable in the next few years. Also very unlikely that even one will be "obscenely" profitable in the next few years.
It is more likely (but still not very) that they both will be simply profitable (not obscenely).
The most probable scenario is that ONE will be somewhat profitable (probably Anthropic) and the other still burning.
You know SpaceX is still publicly valued at $1.5T marketcap, right? (offering price has nothing to do with anything other than ego of founders/bankers).
Yes, I will wait for the correct valuation when all stocks enter the liquidity pool.
OpenAI/Anthropic obscenely profitable is an easy bet (for me)
Seems unlikely. All three market leaders have roughly equivalent products, even ignoring the Chinese models. Google is better vertically integrated, as well. They'll end up competing on price, or worst case capacity rationing.
So it's okay for organized crime to do it, but not normal plebs? The former is obviously a lot more harmful and in reality this is just a wakeup call to get their shit together.
People were dismissive of security far too long. I think it's actually good that AI instills some fear into people causing it to take a lot more seriously than they have previously done so.
reply