You are absolutely correct, billable hours are perverse incentives and wrap rates border on predatory, but this status quo is Congressionally mandated and buy commercial as a policy is generally bipartisan.
I was not a fan of Direct File (I think it could have been much better!) but it was objectively in the success bucket:
- ~$50M all-in to build which sounds egregious but is small for a government project
- mostly good outcomes
- ~4y to pilot which also sounds egregious for a single moderate-complexity web app but again is practically lightspeed for a government project
People blame it getting shuttered on Trump but this is both a misread and a fundamental misunderstanding of how the government operates. The entire federal government is deliberately paced and primarily driven by a combination of politicals, regulation, budget, Congressional policymaking, and program staffing. The hammer on this was always going to fall,
[1] Congress has maintained a "buy over build" policy since literally the Cold War. Ironically the first policy here (1965) was to buy computers from ADP instead of custom-building them. The most impactful "buy commercial" policy (1996) was also IT focused. This was literally codified in several parts of the FAR, and if we're going to be blaming admins, the Trump admin has actually weakened the FAR with the RFO and would be on the other side of this issue.
[2] This policy is also the foundation upon which IT contractors like Palantir are built, if you've ever wondered "wtf do they do" they are a commercial contractor that builds custom software for government. So agencies get their custom stuff, Congress is happy they bought commercial, and the only loser is the taxpayer who overpaid a factor of 3-5x. To their credit, Palantir is a huge improvement on the status quo, because they charge by use cases and outcomes, whereas traditional SI's operate on a staffing model and a combination of high wrap rates and perverse incentives with billable hours creates horrible outcomes and ballooning costs. Anyways, there are a lot of commercial vendors that this work "could have" gone to and didn't. It would also have been really easy for them to shop this out because this was shaped really well for GSA MAS or STARS III.
[3] When you build things internally with e.g. 18F, one agency is typically the "customer" and pays the other agency (afaik 18F program is nested under the White House). So this really sucks when they walk away; 18F earned a reputation for building things through pilot, blogging about it, puffing their cheeks in the media, sending a bill to the agency, and then running off to the next shiny thing as engineers tend to do. As an agency you save a bit of money initially but now you are left with something nobody knows how to maintain, you have to hire contractors to do it, which by nature has to be a staffing contract, and then the wrap rate alone wrecks you. This happens all the time with commercial vendors, usually when they are outsized or don't get renewed, but obviously upsetting when it happens internally.
[4] The author of this piece was part of the team and so as you'd expect this piece is exceptionally biased. The reality is that the IRS self-assessed the maintenance as low and everybody who reports on this treats this as fact when it was not a GAO estimate, which is what carries any actual weight, it was internal IRS estimates and a single independent that mainly assessed call center volume not the actual system. Not only does their reputation precede them here, but their own spend on buildout (~$30M on contractors to help) contradicts the low maintenance narrative. We may never see a GAO estimate here, but the contract would almost certainly go out to TrussWorks, all government contracting data is public and you can look up for yourself how much TrussWorks typically charges the IRS whether directly or thru ATI's vehicles as subawards.
I think this is well said but also would have appreciated if you disclosed that you are a government vendor and are susceptible to your own biases. Would be only reasonable given that you call the author out on this. But the difference is she actually discloses her interest.
This sort of thing was only worse in the past and IMO this admin has actually made the largest inroads to resolving this issue which tbh lies primarily in procurement (RFO, FR20x, PA, mandates for automation, shift to COTS, etc). With any luck this will get recompeted
Project 20x looks pretty cool, we're in a similar space and there's probably a ~2y horizon to fix things before the bureaucracy falls to consultocracy. FWIW if you or anyone else building in gov need funding to get stuff off the ground there are plenty of DC-based VCs that are phenomenal, mission oriented and take big swings like this all the time.
I made a mistake, i assumed that this admin was the crux of the issue and I was wrong. I looked into what you said here and I want to thank you. Moving forward I will not make this mistake.
TY for taking a look, right now I am digitizing all of the federal gov programs and all of the state gov programs, I would really love to connect with anyone just to talk, learn, show them what i'm doing etc...
Again, I really appreciate you taking the time to take a look and comment. It really means a lot. Gov Tech is not "sexy" part of tech where people flock and love to engage.
At Statecraft, we are building an inorganic workforce to help keep the US Government running. Our AI workers are trained on government workflows, operate reliably and compliantly, and build their own tools and software to help achieve outcomes and digitally transform outdated processes.
America is not short on forms, policies, or good people. It is short on hours. Our public servants are overworked and understaffed. For too long, contractors and consultants have made empty promises to solve this problem. Instead, they've wasted billions in taxpayer dollars. We're changing this with an AI workforce that works around the clock and improves over time. On average, we save taxpayers 80% while helping agencies deliver on initiatives and their mission 10x faster.
We're looking for a forward deployed engineer to deliver outcomes for customers, building critical systems with national impact. You'll also be working on our core platform, extending it with new capabilities.
* This position requires U.S. citizenship and the ability to obtain/maintain a security clearance.
No it is H-1B visa. Right out of the university it is hard to recognize extraordinary talent. People like Sundar Pichai were not recognized as extraordinary right out of the university, he had to start at the bottom and rise up the ranks.
This makes even less sense, Trump admin has been here for 1 year, the implication here is a university grad on H1-B in January would become a world class researcher capable of building a frontier model in <18mo
To be fair, that was clearly well deserved. Marrying Trump and then becoming first lady is definitely an extraordinary ability; I doubt I could have done it.
This is why VC is actually hard. Everyone’s instinct is always “Man, once the company has demonstrated it’s awesome I would love to have been in the seed round”. The tendency to want “proven performers” is the default belief.
When people demonstrate their capability thoroughly, the Chinese government takes away their passports. You’re not exactly going to get them here with an O-1.
This is essentially the point of the visa, it feels wrong especially as YC drops standards and increases cohort sizes, but the same power laws that keep them winning also apply here in maximizing economic value of each O-1 approval
Basically of all visas O-1 is virtually guaranteed to have highly positive economic value
It's not the point of the visa, the O-1 is supposed to be for people of extraordinary ability, eg Nobel Prize winners. It's used for software engineers.
I sold my company ~2 years ago for a very decent 8-figure exit where we cleared multiples, everybody got paid, fat bonuses all around, etc. just real pipedream founder stuff. Was incredibly excited and thought "we've made it!"
Currently barely affording a condo in Jersey City (not Manhattan). Don't really understand why the prices are this high, none of these homes are selling, the ones that sell are vacant EB-1 investments of which there surely cannot be that many, and there's no way an upper-middle class family could afford this.
More ludicrously taxes keep going up, property tax is super high making it impossible for buying to ever be better than renting.
Everyone in tech keeps talking about how AI might usher in a permanent underclass but it's already here. It's not "you will own nothing and be happy" for most people I think it's already "you own nothing and aren't happy." It's very confusing what is happening with the real estate market but what is obvious is that local politicians and housing regulations have invented feudalism from first principles.
I genuinely do not understand the controversy here, imagine somebody said this about bread.
1) It is the 12th century. Local lords own all the wheat farms and bakeries. Bread is incredibly expensive, nobody can buy bread, people are simply buttering the bread and licking off the butter, and still having to pay a portion of their wages to help maintain the bakeries.
2) ~1000 years pass. Everyone looks back at that time as the worst economic reality in human history, the foundation of most political systems is the ability for anyone to own bread and operate bakeries, people can buy bread.
3) One day you go to the bakery to buy some bread and it's now $120 a loaf. You say WTF this is so expensive, nobody can buy bread anymore, but there are so many loaves on the shelf. You're informed all of these loavess are spoken for by the people in the parking lot who are apparently gambling on the value of bread.
4) Nobody is saying you can't gamble on the value of bread. Instead, everyone agrees it would cost less if we would simply make more bread. But the people in the parking lot, who are already rich and supposedly champion a totally free bakery, say we can't make more bread as that might drop the value of bread and reduce their portfolio value, which would be worse than starvation, so we must limit the bakery. This is obviously ridiculous.
5) Then somebody walks into the bakery and says, I am a champion of the working class, I am one of you, I too cannot afford bread, make me your leader! So you think amazing, let's put them in charge, they will make more bread.
6) Unfortunately, no, for a variety of reasons that sound like they were made up on the spot, they further reduce the supply of bread, stop people from building more bakeries, and install a series of their supporters to 'review' the grain quality of bread at great expense. Also, a very small group of people will get a loaf of bread for free.
7) The price of soars further, everyone agrees that all of this is making it even less affordable. However it becomes culturally unacceptable to reverse any of these policies.
8) You have obviously been betrayed by the current leader. One of the current leader's supporters decides they want to be leader. The other supporters, who again are responsible for this crisis, back this person and say this is the next leader. They say you cannot question them, you must put them in charge or you're to blame for the bread prices. They run constant smear campaigns against anybody who disagrees, branding them all the same as the people in the parking lot, and scare people on the brink of starvation. So fine, you put them in charge, but demanding they change things.
9) They say yes we have a great solution. The people in the parking lot who are gambling on the prices of bread will now let you borrow a slice from them. You cannot eat it but you may butter it and lick the butter. If the butter seeps into the bread, as tends to happen with butter, you will be charged for the cost of drying out the slice later.
10) It is the 21st century. Local lords own all the wheat farms and bakeries. Bread is incredibly expensive, nobody can buy bread, people are simply buttering the bread and licking off the butter, and still having to pay a portion of their wages to help maintain the bakeries.
There is no level of banning housing speculation that will ever make prices go down. Housing speculation is only profitable because there aren't enough homes in the first place.
I read that as more broad - housing as investment; flippers buying properties to upgrade (materially or not) and sell at a premium, removing bottom rungs on housing ladder. i.e. the mindset is an accelerant or maybe even the cause (think back to the roots of restrictive zoning, fear of property devaluation)
Not missing the forest for the trees, this effectively means in 3-5 months China will drop open source models that are every bit as capable and dangerous as current day Mythos except with no safeguards.
And the only companies safe from this are the large corporations that shook hands with Anthropic? Because Fable doesn't seem to have actual safeguards, more like 'if you talk about this you will be talking to Opus.' It doesn't guard against offensive use, it prevents all use (offensive AND defensive).
Rationalists are inventing oligopolies from first principles, absolutely incredible things happening in SF
My bet is that Mythos is still over-hyped and the cybersecurity fear and guardrails are mostly marketing to force company partnerships through Glasswing and get public attention.
Delaying a technology release is not going to stop that in the long term. Society, culture, and the support tooling just needs to adapt. Just like how AI coding is still in the early days.
The sooner people learn the risks and build the infrastructure to make it fail less the better.
> All these points are valid, and OpenAI did a great job identifying potential risks, especially misuse and biases, at an early stage.
Many of the OpenAI employees who were focused on these risks in GPT-2 later founded Anthropic, notably Dario [1]. Since the beginning and continuing through today Anthropic describes itself as an "AI safety and research company" [2]
I'm not sure if the OpenAI of today has the same focus on safety, or if they do the minimum to not look irresponsible given Anthropic's effort.
People quote the "GPT-2 is too dangerous to release" thing as if it were wrong, but given all the slop all over social media and how it's used to create division and attack social cohesion, he was clearly right.
AISI did also say that GPT-5.5, which has been public for months, scores basically the same as Mythos on their cybersec evaluation. But there wasn't as much media about about that for some reason.
"We had to do extra work to make this safe because it's so advanced and dangerous..." how many times can they trot out that line before it loses its effect entirely?
I mean, they do actually describe what that extra work was, and people elsewhere in this thread are complaining about the effects of those safeguards. So it's not like this is purely empty rhetoric.
people are not questioning whether they did the work, they are questioning whether the work was really necessary (i.e. if mythos is really so good that it needs safeguards to prevent malicious actors from using it)
I still remember it. "Open"AI going API-only because GPT-3 is really really dangerous, so forget the Open in our name and all of that, you can't download our models anymore and must request access to them because they pose a THREAT.
Fast forward to today and GPT-3 has laughable performance.
Even back then there were plenty of people who got fooled by AI generated articles. It's easier to spot AI writing now because we are so used to it. They were right to be concerned; not that it achieved much since oss models run laps around gpt-3 now.
But it seems like that was not genuine concern, but instead a tactic to pivot to closed models and an API service with an excuse to do so, breaking the public's expectation that they would be a non-profit making open models, like their name implies.
I know a security researcher at Google with access to Mythos. He says it's the "real deal" and that "there are career plans I had that are no longer viable".
Yes, and "in collaboration with the U.S. Government" feels like a very gross ploy at appeal to authority. You don't need Mythos or really any SotA frontier model to make malware or do extensive penetration testing/reconnaissance already. Sure, Mythos might be faster/more efficient, but the cat has been out of the bag for awhile. Even the terminology "infrastructure providers" practically screams "Enterprise leads".
It's not even very usable... I tried 2 different chats and both eventually got stopped due to the safeguards
One was a piece of code I gave it to improve, it did so and then started writing tests, some of which tested security so the safeguards triggered
Another was one of the cryptography puzzles I use as new model tests, which are hard to oneshot and there's no public solution anywhere, it completely refused to even try to solve it
They're trained in a model class likely in 2t to 3t range. It's very unlikely that chinese labs have access to gpu systems capable of training models like that, let alone serving them. This requires proprietary room-scale systems which fetch a huge premium over typical 10 slot systems.
I am sure that they can develop their own equivlient version of such clusters in around 1 year though. Distilling fabel 5 will also go a long way.
MoE experts were likely trained independently / in a sparse format. Training anything beyond 2t on typical systems would be infuriantingly slow, you could do 4t on nvidias room-scale solution, but for a reasonable training speed / batch size it caps around 3t.
concept is similar to how it works in inference, instead of performing regressive writes to the entire model you run the whole model, but part of the model can live in system memory and get swapped in/out on demand. So only XB parameters are active in training.
edit: I am not really sure if it works like that. I haven't looked too deep into deepseek v4 pro specifically.
I think we're about to see a big relative drop-off of open models vs closed. I don't think there'll be an open model that competes with Mythos for ~2 years.
Even OpenAI and Google are struggling to get this kind of performance. If the distillation defenses are any good + chip controls prevent China from training massive models, it's over.
They have, but even with the whole CCP backing you you can't just catch up on the chip war overnight. It's going to take time to get their memory and compute industries where they need to be. Meanwhile, barring an invasion of Taiwan, US will have Rubin class models and then whatever the next tier is, within 3 years.
'Barring the invasion of Taiwan' might actually be quite a lot to bar in mid 2026.
My hot take is that it's now or never for Xi, and from the specific things he is reported to have said to the US president at their last meeting lead me to think that he at least knows this is his big chance; whether or not it is taken is the part of the forecast that is opaque to me.
I wonder if model distillation will continue to work as well as it has. Given hidden reasoning, the ever expanding number of expected capabilities, a serious compute shortage, the looming possibility of model collapse, and dramatically higher API costs I would guess that it's getting much harder to do.
You should check out some Chinese forums. There are services selling gateways/proxies for all major models at fraction of the official rates. Likely reselling subscriptions, or some other form of abuse.
I've seen people posting screenshots of billions of tokens consumed where they paid next to nothing.
These same gateways are likely also reselling the data to Chinese labs, because TLS has to terminate at the gateway level.
Asian labs generated synthetic datasets from UBS labs but also innovated with technology. Now it is harder to get the thinking traces AND Anthropic is recorded to poison it as well.
Thus Asian labs will have to generate their own data sets, which with the huuuuge usage boom from deepseek, mimo, kimi, etc, they will be able to.
That's the reality China already lives in. Their weapon against US companies is commoditizing them, eliminating their moats and their profits by going open weights.
Same thing Meta was doing before they fell behind.
Meta made Pytorch and a lot of vision models back in the day, like Faster RCNN, Mask RCNN, the Detectron framework, and more recently the SAM and DINO series. AI not just LLMs.
My experience is that open weight models from China are at least ~12 months behind. In some workloads they may be closer, in others further away.
I also find that the harness and product you wrap around models can often narrow that gap considerably.
Opus 4.6 for example, on a PR-for-PR basis was head and shoulders above GLM 5.1. Perhaps GLM 5.1 was a bit under Sonnet 4.6 at the time. That's roughly a year or so behind.
Much cheaper though! I'm bullish on open weight models, I have no idea where all these curves will top out, can the frontier labs keep the year plus lead? Do open labs get close enough to SOTA that they gain adoption across many tasks and drive down inference prices??? Who knows, not me.
Yeah, because it's impossible. You can't ask it anything about the thing that it's known for. It will not even answer a sky-high level question about reverse engineering, for example.
In CC, it will probably report you to authorities if you ask it to do a vulnerability scan of your codebase.
Isn't that a good thing in a way? If everyone has the weapon and defense at the same time, we will fix security holes and live safer lifes instead of having some three letter agencies and military backdoors in everything.
Pandora box is open anyway. It's better now for everyone to have the same power rather than a few national states.
Not sure this holds, sadly. I spent a few months reporting serious security bugs as model capabilities took off earlier this year, and only ~half were fixed. The unfixed bugs were just as critical as the fixed ones; sometimes they were even two similarly critical bugs at the same company, and only one would be fixed!
On your other point, the government still has systemic leverage and can compel access, so this doesn't remove that risk.
That doesn't mean this is the end of the world, and some balance of power is usually good. But I do think it will still increase the capabilties of rogue actors and their net harm.
It's more evidence that the future is local. With some time we'll all be running highly capable & efficient open-source models on dedicated NPUs. No censorship, no rate limits, no overpriced subscriptions.
3-5 months is a long time and they are pretty useless on arrival because the frontier models are so good, that it's hard to go back even if it's way cheaper. Your work flow is adapted to that level of intelligence for months.
I don't think China has any incentive to arm the rest of the world with highly capable models that can be used against them. Undoubtedly they will continue with the arms race, but they will preserve the best stuff for their own use.
I think the stronger incentive is undermining/undercutting the Western AI companies. Given what we have seen, any model can be used/convinced to do harm so that is just part of the game
I agree, depending on how much of this is marketing and how much is actual capability. It's one thing to undercut models that finish writing assignments for lazy students. If this actually identifies vulns and writes exploits, or if it designs bioweapons, those are pretty different. Those are actual weapons, and I don't think they're going to arm the adversary.
If you’re in gmail there’s literally already an interface to easily unsubscribe from stuff, it’s under “Manage subscriptions”. Yahoo similarly have a “subscriptions hub”
reply