[HN Gopher] The Most Cited AI Papers in 2022
       ___________________________________________________________________
        
       The Most Cited AI Papers in 2022
        
       Author : marban
       Score  : 32 points
       Date   : 2023-03-05 15:46 UTC (7 hours ago)
        
 (HTM) web link (www.zeta-alpha.com)
 (TXT) w3m dump (www.zeta-alpha.com)
        
       | elektor wrote:
       | I'm getting wildly different citation counts for some of the
       | listed AI papers.
       | 
       | For example, the paper "ColabFold: making protein folding
       | accessible to all" is listed as having 1162 citations. I'm seeing
       | that it was cited only by 899 publications on Scite:
       | https://scite.ai/reports/colabfold-making-protein-folding-ac...
       | 
       | I'm wondering if Google Scholar is overestimating or Scite is
       | underestimating.
        
         | godelski wrote:
         | Semantic Scholar has 1,111. [0]
         | 
         | I tend to trust Semantic more than GS. GS tends to
         | overestimate. For example on GS I have 164 citations on one
         | paper and semantic says 150. FWIW Scite says 49.[1]
         | 
         | [0] https://www.semanticscholar.org/paper/ColabFold%3A-making-
         | pr...
         | 
         | [1] I'll note that this paper is an arxiv paper and has not
         | been accepted at a conference but I'd also argue that
         | conference acceptance means little in ML. I'll explain if
         | anyone is actually concerned with the claim.
        
         | jltsiren wrote:
         | Citation counts are always a bit arbitrary. Google Scholar
         | usually overestimates, because it's basically a bunch of
         | heuristics. Curated citation databases underestimate in the
         | name of consistency. For example, they may ignore citations in
         | conference proceedings, as conference papers are not considered
         | legitimate publications in most fields.
        
       | antegamisou wrote:
       | Probably the only subfield of Computer Science, maybe even
       | academic research in general, where citations number is the most
       | poor metric of paper quality. The whole situation is more akin to
       | mass media reporting the same breakthrough in cancer cure
       | research (in mice).
        
       | bryanrasmussen wrote:
       | It looks like it might be more interesting on what industries or
       | problem domains those most cited AI papers were focused.
        
         | choppaface wrote:
         | According to Fig 3, Google was that most-cited problem domain.
         | The most cited papers are for Google problems. You must have
         | Google money and compute to achieve top-cited research.
        
       | panabee wrote:
       | very useful but flawed.
       | 
       | the list biases against papers published later in the year.
       | 
       | this is a problem raised by openai and brain researchers:
       | https://twitter.com/_jasonwei/status/1631935794301771777
        
       | usgroup wrote:
       | Did nothing of note happen outside of deep learning the whole
       | year?
        
         | mtcrawshaw wrote:
         | Of course interesting things happen outside of mainstream deep
         | learning. The problem is that sorting by citation count is a
         | terrible way to look broadly at research, because mainstream
         | deep learning is so hyped up that there are just way more
         | people working in that area than in other areas, and with more
         | people come more citations.
         | 
         | It's a shame because I think that other areas offer much more
         | interesting and technical questions, but many new researchers
         | are only exposed to mainstream deep learning because the
         | enormous hype drowns out everything else.
        
       | tomohelix wrote:
       | Incredible how much benefit alphafold has brought. And all of
       | that from a less than 100million parameters model.
       | 
       | I might be dumb but could they scale it up and make an alphafold
       | 3 with maybe like 10bln params? Would it be a lot better assuming
       | the same training effort is put into it?
       | 
       | If it does, can't biotech companies just go nuts and make a
       | 100bln params internal model and have all the protein structures
       | they want?
        
         | tiedieconderoga wrote:
         | Has there been much research into the idea of distributing
         | these large models across many heterogeneous machines?
         | 
         | I'm wondering if there could be a path towards a mix of
         | alphafold and folding@home, with donated idle compute resources
         | being used to train/run the models.
         | 
         | Designing for that sort of fragmentation could also make it
         | easier to slowly run oversized models on local machines with
         | swapped memory.
        
       ___________________________________________________________________
       (page generated 2023-03-05 23:01 UTC)