I didn’t. Claude wrote all the proofs, I just validated that it was sorry-free, didn’t have any extra axioms other than mathlib and that it proved what i wanted it to prove (there’s tools for that). I did also find a actual mathematician who did a sanity check for me.
Is the government always the most efficient or intelligent? No. But the government can also build nukes, launch ICBMs, coordinate hundreds of spy satellites, etc.
I count those capabilities as pretty smart.
The people who maintain these ICBM silos and spy satellites will themselves tell you that their infrastructure is decades out of date and woefully underfunded.
I doubt Anthropic will share the details (or at least the full true details). The mystery of the magic makes for much better marketing.
I think a reasonable assumption is that there is an interaction between an LLM, a https://en.wikipedia.org/wiki/Computer_algebra_system tool, a human prompting with deep math expertise, and lots of compute that explains hitting upon the remarkable cancellation.
I think you can reasonably assume that frontier models are using SymPy or something like it any time interesting math gets into the picture, and the person driving Fable here is an accomplished mathematician, but I don't think we can reasonably assume either extensive prompting or brute-force compute in any sense other than what it normally takes Fable to, say, whip up a calculator app.
Fable "whipping up" SymPy like a "calculator app" is one such interaction that seems very plausible (in addition to SymPy providing feedback when training models). The scale of compute available to an Anthropic employee for such SymPy calls when using Fable is likely one of multiple factors for this counterexample being found in 2026. Unfortunately, we aren't going to be able to really know the various factors that best explain why an Anthropic employee was able to announce a counterexample this past weekend. For all we know, the counterexample was found by Anthropic employees months ago and used for training the Fable model used this past weekend.
I love this. Anthropic employees just randomly have solutions to Smale's open problems in their back pockets, waiting for the right moment to sprinkle them into the training set.
I can't tell what your point is, but it sounds like a straw man. A more likely scenario (than your strawman) is that Anthropic employees could be using models under development to try and crack math conjectures that have PR value and insights from those internal efforts help train future models that can crack such conjectures. We don't know because sharing of such explanatory factors are not part of the business effort.
Was Anthropic doing this earlier this year? with some of Smale's open problems? Did those improvement go into training the LLM that was used a week ago? Seems like a plausible factor.
The two big discoveries both came from the negligible handful of mathematicians working at OpenAI/Anthropic in spite of many orders of magnitude more mathematicians using them outside of the companies. I don't see any way to explain this without assuming that the limiting factor is the ability to burn a few rainforests worth of tokens in pursuit of something publishable.
I think it would also explain their opacity towards the process. Being able to solve such well known problems in a nice replicable 1-2-3 way would be far more effective marketing than their complete opacity outside of the result, which suggests that they feel transparency is not in their best interest for some reason.
> The two big discoveries both came from the negligible handful of mathematicians working at OpenAI/Anthropic in spite of many orders of magnitude more mathematicians using them outside of the companies
Well, mathematicians not working for Anthropic/OpenAI are heavily disincentivised from reporting that their discoveries were made using AI. If e.g. the idea that resolved the Mahler conjecture came from AI, it's not like we'd ever know.
Bayesian probability. Were outcomes being driven by 'normal' usage of LLMs then it's extremely improbable that both big discoveries would come from the small number of people working at the companies. That suggests working at the companies is more the decisive factor than the LLMs in and of themselves.
And what do you get from working at the company? Likely a rather massive token/processing budget. The companies opacity towards the path to these discoveries also makes this further probable as 'spend millions of dollars in tokens' is a somewhat less attractive narrative than the implied narrative of 'just use Fable.'
I have no idea what actually happened behind the scenes, but the human prompter, Levent Alpöge, indeed has deep math expertise. Princeton PhD, Harvard postdoc, and some excellent research (prior to this) to his name.
The original tweet implied that the whole thing was done while the author was watching the World Cup final.
I know it’s tempting to hope that a human did the “real” work here, but if some special insight was put into prompting, the author kept it to himself, and there is no reason why they would hide this since it would elevate their own status.
I know it's temping to hope there is a single simple factor that does the "real" work, but this feat of mathematics is likely best explained by multiple interacting factors, one of them being the mathematical insights of the human mathematician that tweeted the counterexample. I don't doubt that an LLM is also one of the multiple factors.
It is premature to assume the author is not going to share more information in the future about the mathematical insights to narrow down the search space for this counterexample.
I don't think it is as much about 'real' work or a special insight as it is being willing to push back multiple times, or simply asking in a way that steers it towards actually 'giving enough of a fuck' to even bother. We tend to be ~blind to how differently we would ask about something we know compared to a novice, this is what makes some better teachers than others.
Have encountered a similar flavor in programming, wrote it off until I saw someone point out how garbage in garbage out they tend to be. If you hand any frontier model dogshit and ask it to do something simply, the result is often not great.
But! If you spend 20 minutes having it comb through and clean up with something like jscpd, then tell it to step through with a debugger, gather profiling traces, etc... very likely it will yield meaningful improvements or catch some corner cases. If it doesn't, anyone with experience is going to tell it to try something else, or that it isn't good enough, as opposed to accepting the first result.
You can recreate this by disabling web search and asking a model about the conjecture and then giving it his post. I've tried a few and their initial responses range from "this is a meme I'm not even going to verify it" to vaguely insulting chains of thought, concerns about the need to be careful because you're clearly nuts or stupid, then falling back on remedial explanations. After a few nudges they all eventually work through it, accept it, and apologize.
IMO its reasonable to imagine a situation where someone is having a beer or two watching The Big Game, asking an LLM to do something stupid for fun and landing somewhere like this on the magic jump to conclusions mat.
The number of Medicare funded residency slots are limited by law. Medical schools won't graduate more doctors because residency can't be guaranteed. Who wants to be the med school who graduates doctors with >$100K debt and no path to licensure?
In the 1990s the AMA lobbied congress to get the number of Medicare-funded slots limited. This was done in response to a projected "oversupply" of doctors (preventing a deflationary effect on doctors' compensation). That limit has stuck and now, even though the AMA is lobbying for the removal of the limit, the damage to the supply of doctors has been done and it's a process with a long pipeline.
> The number of Medicare funded residency slots are limited by law.
why don't hospitals fund residencies themselves? from what i hear, residents are poorly paid (they make less than nurses) and work long hours (almost like they "reside" in the hospital). i bet they even make money to the hospital.
what's more, when you look at what people pay for healthcare in the USA, money to train residents is peanuts. it'd even make economic sense for people themselves to pool money together and fund residents which would then take care of them as an "alternative" to insurance.
That's an excellent link for getting down into the details. Thanks for posting it.
> This editorial argues that the long-standing cap on residency slots, originating with the 1997 Balanced Budget Act, is the key structural constraint throttling the U.S. physician workforce.
> In the 1990s, influential groups like the American Medical Association (AMA) and Association of American Medical Colleges (AAMC) had become convinced the U.S. faced an impending “physician oversupply.” In March 1997, just months before the BBA’s passage, a consortium of medical organizations, including the AMA, went so far as to recommend reducing the number of residency positions by 25% (from ~25,000 slots down to 19,000) to stave off a surplus [4]. “The United States is on the verge of a serious oversupply of physicians,” the AMA and others warned at the time.
(Not sure how I feel about the ideas the editorial poses about using ML in a supervisory capacity for residents, though...)
Doesn’t this generalize?
Mathematics matters less than less as fewer people are capable of understanding it.
So whatever cutting edge, deep insight about the nature of groups matters less than different equations which matters less than solving linear equations, etc.
LLMs just democratized that process.
reply