[HN Gopher] Efficient Transformer Knowledge Distillation: A Perf...
       ___________________________________________________________________
        
       Efficient Transformer Knowledge Distillation: A Performance Review
        
       Author : PaulHoule
       Score  : 52 points
       Date   : 2023-12-07 15:41 UTC (7 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | empath-nirvana wrote:
       | I skimmed the paper but I don't really understand what knowledge
       | generation actually entails.
        
         | drawnwren wrote:
         | https://arxiv.org/abs/1503.02531
        
         | WhitneyLand wrote:
         | Don't mean to be flip at all, may I suggest:
         | 
         | 1. Use Gpt-4 with something like this:
         | 
         |  _"Help me understand what this paper is about and estimate
         | whether the relevance and impact to the field is likely to be
         | low, medium, or high.
         | 
         | Explain jargon that may be specific to AI research, but don't
         | bother explaining or expanding on terms familiar to a working
         | software developer or basic undergraduate computer science."_
         | 
         | 2. Follow the above prompt with either the abstract or the full
         | text of the paper.
         | 
         | 3. Post something useful here and save others the time.
         | 
         | Do not - in my opinion - copy/paste LLM output as a comment.
         | But would love to hear your own succinct, human sounding, HN
         | guideline compatible thoughts.
        
       | egnehots wrote:
       | This paper combines knowledge distillation and efficient
       | attention mechanisms.
       | 
       | => It works (still efficient, lower cost).
       | 
       | Not an unexpected result, but to their credit, they established a
       | new benchmark to test these combinations. KD+LongFormer is one of
       | the best ones, retaining 95.9% of the performance for 50.7% of
       | the cost.
        
       | mistrial9 wrote:
       | Appendix A
       | 
       | A.1 Data Collection Data for GONERD was obtained through Giant
       | Oak's GONER software, which scraped web ar- ticles from public
       | facing online news sources as well as the U.S. Department of
       | Justice's justice.gov domain. This webtext data was randomly
       | sampled with an upweighted probability toward documents from
       | justice.gov so that justice.gov consisted of roughly 25% of the
       | total GONERD dataset.
        
       ___________________________________________________________________
       (page generated 2023-12-07 23:01 UTC)