I guess we won't know if that's what was used (and maybe even provided as part of the prompt given that both Alpöge and Mathew are mathematicians) since they decided against sharing their Fable conversation and instead opted for a memey tweet as their avenue of publication. We really ought to normalize full transparency in how results come about.
Anyway, if I read Tao's post and comment correctly, there's still a gap from the Vitushkin construction to a counterexample, but chances are that was in the training data. In general, it is just a serious problem for their practical applicability that the models are outputting proofs with absolutely terribly reference hygiene.
Even if they published the conversation, Anthropic (and likely other closed model publisher) no longer provide logs of the actual thinking process.
I more and more see LLMs as a kind of scam; not useless, but really just a big database of fuzzy facts with some Prolog on top as rediscovered by the learning algorithm. Most likely could be made much cheaper to run, were humans allowed to actually inspect the algorithm.
Yes, I think this idea, that it should be "magical", is what makes it feel scummy. (Apparently I am not alone https://news.ycombinator.com/item?id=48988475). It makes AI providers sound like snake oil salesmen, and rightfully so.
Meanwhile, technological and engineering (STEM) progress have always been made by emphasizing externalization of the deductions (as opposed to reference to an opaque expert judgement) and reproducibility of experimental results.
I would even call the frontier AI labs anti-scientific. We need to understand how inference is done to avoid mistakes, not rely on intuition, even if the intuition is enclosed in a reproducible machine. The idea that AI should be this closed is a return to pre-scientific days.
I don't think it is nonsensical at all. The author and his collaborator both appear to be bright people, so there's a good chance they had to offer non-trivial insights to guide the LLM, yet it's clearly in the interest of his employer to downplay whatever personal contribution they provided.
Edit: Now the OP is flagged/dead for some reason. You could disagree on their take (calling it a marketing stunt is maybe a bit much), but I think the argument is sound, so flagging seems counterproductive to the discussion.
Believing that AI played a very small role while we know this problem was open for decades with at least a few people taking serious cracks at it is just not a coherent logical position.
That doesn't really even diminish the contribution from Fable, if true. Droves of grad students have been provided the same sorts of non-trivial insights and turned up no results.
I'm sure they had plenty of time to think about these insights without the LLM, as well as the many other mathematicians who tried to crack it over the years. Wether the LLM was simply an assistant or solved the problem entirely isn't as important as accepting than the LLM was the essential, previously missing piece in the solution.
I hear ya. Fair criticism. I'm a professional developer myself, but not great at design. I've tried to come up with a different looking site best I could. I went with a newspaper theme like back in the day when you'd get the puzzles in the paper. And then it was my idea to have a sudoku being solved as a graphic on the front page. I would push back that this could be one-shot by any of the leading models including Fable. Each of the 10 puzzle types has to have its own generator and they're different from each other. They have to handle uniqueness, solvability, and difficulty and none of the leading models have nailed even just a single generator on the first shot. Plus, there's monetization, rate limiting, caching, among other things under the hood that models wouldn't typically touch without specific instruction or would, at best, half-ass it. Maybe you have better luck with them, but for my job, I work on a large legacy app as well as various microservices and the LLMs miss things all the time. I have a system I use that does make them perform better, but you still gotta watch em like a hawk.
I one shot games every now and then, just to see how much it can do. For anyone wanting to experiment, I have come to learn that if you make it make browser games the setup is even easier since it can just inject the JS into the HTML and import from a popular CDN, no node, no compilers needed, just a single HTML page with inline JS.
I'm curious, What kind of details are you thinking of? I'm not sure I really have much of a radar for LLM websites in the way I do for LLM pictures or music.
I don't know for pictures, but I have gotten pretty good at detecting AI in videos. I am noticing these a lot on youtube. Often you can tell, e. g. movements being weird, animals behaving in ways that are only in a short and nowhere else to be found. And some more indicators e. g. youtube insists on showing sexy girls, but the video is clearly "cut" into another video and the surface layers also don't fully align; or some proportions are odd (I don't mean the "regular" ones but e. g. when the biceps looks like semi-hulk, you know something is AI slop). I try to not watch AI slop but sometimes it happens.
For images, there are some clear styles AI leans heavily on if not actively steered away[0].
It can definitely be prompted pretty successfully though, a bird spotting app was up her on HN recently with some really nice looking woodblock prints that were AI generated (I always feel disappointed/tricked when art turns out to be made by AI, I'm not sure why, it seems to pull the joy out of it for me)
From a quick look at your profile, the majority of your submissions have been Show HNs. HN only allows some fraction of your submissions to be Show HNs (imagine if the front page was nothing but), so eventually they will just be auto-flagged.
> Yes any company generating csam should not be in business as a legitimate entity.
At the same time, in this corner of the world, acting Minister for Justice (also known for trying to push through Chat Control), and NGO Save the Children, have been working to make legal the generation of CSAM for law enforcement use. So that would certainly make the industry legitimate, and you would already have a customer.
I think they key point here is "for law enforcement". That's a little different from "pay me 10 dollars and enjoy the felonies". I still don't feel good about that by the way.
And the cookie consent form is one of those that require you to click a gazillion toggles. Hasn't it been established now that opt-out must be no harder than opt-in?
the law unambiguously says that, yes. However, companies these days seem to respond to enforcement, not the text of the law. They are all using the same few cookie banner libraries/providers, so that they have herd protection (if the eu wants to crack down on it, it has to do it to hundreds/thousands of companies simultaneously). It seems neal chose that route also.
The time is ripe for deterministic AI; incidentally, this was also released today: https://itsid.cloud/ - presumably will be useful for anyone who wants to quickly recreate an open source Python package or other copyrighted work to change its license.
Can you please explain the use here? I tried the demo, and cat, cp, echo, etc... seem to do the exact same thing without the cost.
Their demo even says:
`Paste any code or text below. Our model will produce an AI-generated, byte-for-byte identical output.`
Unless this is a parody site can you explain what I am missing here?
Token echoing isn't even to the lexeme/pattern level, and not even close to WSD, Ogden's Lemma, symbol-grounding etc...
The intentionally 'Probably approximately complete' statistical learning model work, fundamentally limits reproducibility for PAC/Stastical methods like transformers.
CFG inherently ambiguity == post correspondence problem == halt == open domain frame-problem == system identification problem == symbol-grounding problem == entscheidungsproblem
The only way to get around that is to construct a grammar that isn't. It will never exist for CFGs, programs, types, etc... with arbitrary input.
I just don't see why placing a `14-billion parameter identity transformer` that just basically echos tokens is a step forward on what makes the problem hard.
Tech world became so wild even in a topic that I’m confident I cannot say if something is real or satire. Amount of real but absolutely idiotic landing pages made me this way :)
Would you be able to comment on https://news.ycombinator.com/item?id=47522876, i.e. explain the legal basis for this change for EU based users? If there is none, you may have to expect that people will exercise their right to lodge a complaint with a supervisory authority.
Why would you expect an engineer to be able to comment on legal affairs? Presumably it was cleared with Microsoft's legal department or whatever GitHub's divisional equivalent is.
That's precisely what the term 'engineer' signifies. (I know it gets used incorrectly for software developers.) Workers in general need to decide whether something is legal independently of their company, because the company lawyers have the interest of the company in mind, which might conflict with the workers interest to not do illegal things.
Big Tech is known for clearing illegal things by their legal departments all the time.
Anyway, if I read Tao's post and comment correctly, there's still a gap from the Vitushkin construction to a counterexample, but chances are that was in the training data. In general, it is just a serious problem for their practical applicability that the models are outputting proofs with absolutely terribly reference hygiene.
reply