[HN Gopher] Transformers Are Multi-State RNNs
       ___________________________________________________________________
        
       Transformers Are Multi-State RNNs
        
       Author : DreamGen
       Score  : 31 points
       Date   : 2024-01-16 20:05 UTC (2 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | Icko wrote:
       | I've seen at least 6 such papers, all being like "<popular
       | architecture> are actually <a bit older concept>". Neural
       | networks are generic enough that you can make them equivalent to
       | almost everything.
        
         | civilized wrote:
         | Has anybody proved that transformers are just kernel SVM yet?
        
           | blackbear_ wrote:
           | I hope you're satisfied with Gaussian Processes:
           | https://arxiv.org/abs/1806.07572
        
         | haswell wrote:
         | Linking the latest generation tech to the previous generations
         | is actually really helpful from my perspective.
         | 
         | All of the terminology for this tech is still emerging, and it
         | can be quite difficult to formulate a reliable mental model for
         | any of it due to how quickly it's changing.
         | 
         | If hammers were more difficult to understand, I could imagine
         | someone writing about the fact that hammers are in fact, just a
         | piece of steel mounted on a handle made of wood.
         | 
         | > _Neural networks are generic enough that you can make them
         | equivalent to almost everything._
         | 
         | Which to me is why papers like this are useful. They help
         | newcomers conceptualize what the latest <popular architecture>
         | is actually made of in terms of <a bit older concept> that the
         | reader may already understand.
         | 
         | It will take some time for this information space to stabilize.
        
         | dboreham wrote:
         | See:
         | https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Arnold_repr...
        
           | nh23423fefe wrote:
           | oblique
        
       | joewferrara wrote:
       | They show that a decoder only transformer (which gpts are) are
       | rnns with infinite hidden state size. Infinite hidden state size
       | is a pretty strong thing! Sounds interesting to me.
        
       ___________________________________________________________________
       (page generated 2024-01-16 23:01 UTC)