[HN Gopher] Transformers Are Multi-State RNNs
___________________________________________________________________
Transformers Are Multi-State RNNs
Author : DreamGen
Score : 31 points
Date : 2024-01-16 20:05 UTC (2 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| Icko wrote:
| I've seen at least 6 such papers, all being like "<popular
| architecture> are actually <a bit older concept>". Neural
| networks are generic enough that you can make them equivalent to
| almost everything.
| civilized wrote:
| Has anybody proved that transformers are just kernel SVM yet?
| blackbear_ wrote:
| I hope you're satisfied with Gaussian Processes:
| https://arxiv.org/abs/1806.07572
| haswell wrote:
| Linking the latest generation tech to the previous generations
| is actually really helpful from my perspective.
|
| All of the terminology for this tech is still emerging, and it
| can be quite difficult to formulate a reliable mental model for
| any of it due to how quickly it's changing.
|
| If hammers were more difficult to understand, I could imagine
| someone writing about the fact that hammers are in fact, just a
| piece of steel mounted on a handle made of wood.
|
| > _Neural networks are generic enough that you can make them
| equivalent to almost everything._
|
| Which to me is why papers like this are useful. They help
| newcomers conceptualize what the latest <popular architecture>
| is actually made of in terms of <a bit older concept> that the
| reader may already understand.
|
| It will take some time for this information space to stabilize.
| dboreham wrote:
| See:
| https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Arnold_repr...
| nh23423fefe wrote:
| oblique
| joewferrara wrote:
| They show that a decoder only transformer (which gpts are) are
| rnns with infinite hidden state size. Infinite hidden state size
| is a pretty strong thing! Sounds interesting to me.
___________________________________________________________________
(page generated 2024-01-16 23:01 UTC)