Hacker Newsnew | past | comments | ask | show | jobs | submit | itkovian_'s commentslogin

I can’t stand it. Very engineering-y over specified formal language around a complete lack of core understanding. Is damaging other people read this and try to learn things from it.

Causal masking allow model to learn implicit positional embeddings. The meme that a transformer block is permutation invariant is not true.


And adding to that, there is also the recurrent state in the Kimi Delta Attention. I wouldn't call it position information but more sth like "position sensitivity"


The paper says:

"Kimi K3 uses no explicit positional embedding (NoPE), and instead encodes positional information implicitly through the recurrent gating and decay mechanism of KDA."


Intrinsic motivation/curiosity driven exploration/novelty seeking feels like it has a breakthrough paper waiting. If someone could get those methods working for LLMs, we start getting things like move37 but in math proofs and then everything else


I just don’t understand this view. This is the most significant technology ever developed. The uncertainty currently is whether it 1) has massive impact, completely altering society and the making world significantly significantly better or 2) if we go into a fast takeoff/rsi loop. Personally I’ve always been highly skeptical of the later, but that seems like a genuine possibility now. It’s not ‘are we going to be able to generate 10% roic on compute’ the answer to that is yes.


It will obviously be a large part of the economy like online shopping is today. The companies that built up a lot of debt to be brand names in online shopping primarily went bankrupt because new companies had no debt (and perhaps no negative sentiment from early customer experiences.)


the difference is all the current capex is going to durable, hard to get physical assets + things like PPAs. In your online shopping analogy, the hyperscalars are acting like Amazon in 1998


What destroyed hardware manufacturers after the first bubble burst was all the old but overpowered server hardware ending up in resale after bankruptcies. Hard to get hardware is a bad investment that is also potentially obsolete after new optimizations make the next generation much more efficient. The companies that think their individual optimizations are going to outpace industry wide optimization are delusional, historically speaking.


> the hyperscalars are acting like Amazon in 1998

I hope you remember furniture.com, pets.com, webvan.com, kozmo.com and many others.

Amazon.com (during 1998) in some sense is the exception, not the norm from the bloodbath in stock markets during the dot com bubble.

(I highly recommend the book How the internet happened for more knowledge about things before, during and after the dot com bubble.)


> It’s not ‘are we going to be able to generate 10% roic on compute’ the answer to that is yes.

Based on what? No AI company has ever made a cent in profit (exept for Nvidia lmao).


Yeah but space data center Econs are going to revolutionize the tulip marketplace


‘640kb of ram should be enough for anyone’


Contrarian view; I think he’s right. Many of these ideas it’s almost shocking how many you can find sketched out in his old papers. To the point where I think it’s very wise to read all his work to see what hasn’t showed up yet but likely will. Artificial curiosity for example.


You can’t run a closed llm locally. Strange to frame the dichotomy as between local and open. One begets the other.


The entirety of Anthropic believe ai is going to eat everything, not just software, and result in major societal disruption within a year. They do not have a sliver of a doubt on this. Article has no idea, is completely wrong.


"Company execs and stuff pre-IPO declare their product is amazing and will change everything, news at 11"


This is called linear mode connectivity and seems to work for almost every large model. So well that in most cases it’s an explicit part of the training process; do many training ‘branches’ then merge then continue.

It is not understood why it works so well.


is that actually how they train them in the datacenter? the trillion sized weight vector gets cloned and sent off to groups of GPUs and averaged after?


Projects like pluralis agora solve this problem. Really what you want is the model to be collectively owned and governed, not local


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: