I can’t stand it. Very engineering-y over specified formal language around a complete lack of core understanding. Is damaging other people read this and try to learn things from it.
And adding to that, there is also the recurrent state in the Kimi Delta Attention. I wouldn't call it position information but more sth like "position sensitivity"
"Kimi K3 uses no explicit positional embedding (NoPE), and instead encodes positional information implicitly through the recurrent gating and decay mechanism of KDA."
Intrinsic motivation/curiosity driven exploration/novelty seeking feels like it has a breakthrough paper waiting. If someone could get those methods working for LLMs, we start getting things like move37 but in math proofs and then everything else
I just don’t understand this view. This is the most significant technology ever developed. The uncertainty currently is whether it 1) has massive impact, completely altering society and the making world significantly significantly better or 2) if we go into a fast takeoff/rsi loop. Personally I’ve always been highly skeptical of the later, but that seems like a genuine possibility now. It’s not ‘are we going to be able to generate 10% roic on compute’ the answer to that is yes.
It will obviously be a large part of the economy like online shopping is today. The companies that built up a lot of debt to be brand names in online shopping primarily went bankrupt because new companies had no debt (and perhaps no negative sentiment from early customer experiences.)
the difference is all the current capex is going to durable, hard to get physical assets + things like PPAs. In your online shopping analogy, the hyperscalars are acting like Amazon in 1998
What destroyed hardware manufacturers after the first bubble burst was all the old but overpowered server hardware ending up in resale after bankruptcies. Hard to get hardware is a bad investment that is also potentially obsolete after new optimizations make the next generation much more efficient. The companies that think their individual optimizations are going to outpace industry wide optimization are delusional, historically speaking.
Contrarian view; I think he’s right. Many of these ideas it’s almost shocking how many you can find sketched out in his old papers. To the point where I think it’s very wise to read all his work to see what hasn’t showed up yet but likely will. Artificial curiosity for example.
The entirety of Anthropic believe ai is going to eat everything, not just software, and result in major societal disruption within a year. They do not have a sliver of a doubt on this. Article has no idea, is completely wrong.
This is called linear mode connectivity and seems to work for almost every large model. So well that in most cases it’s an explicit part of the training process; do many training ‘branches’ then merge then continue.
is that actually how they train them in the datacenter? the trillion sized weight vector gets cloned and sent off to groups of GPUs and averaged after?
reply