[HN Gopher] Improving Text Embeddings with Large Language Models
       ___________________________________________________________________
        
       Improving Text Embeddings with Large Language Models
        
       Author : cmcollier
       Score  : 32 points
       Date   : 2024-01-02 18:59 UTC (4 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | binarymax wrote:
       | Interesting, but this aspect makes me double-take: "We
       | demonstrate that Mistral-7B, when fine-tuned solely on synthetic
       | data, attains competitive performance on the BEIR [ 40 ] and MTEB
       | [27] benchmarks".
       | 
       | E5/BGE large are an order of magnitude smaller than Mistral-7B.
       | So is this just "bigger model wins" in disguise?
       | 
       | I need to read the whole paper carefully, but this jumped out at
       | me.
        
         | huac wrote:
         | agree, this is a nice example of generating synthetic data, and
         | I believe that the synthetic data is helpful for generating
         | useful embeddings for RAG, but not including an ablation with
         | fine-tuned E5 or another commonly used embedding model (to
         | control for the 'bigger model wins' effect) is a glaring
         | omission. this paper shares many authors with the E5 paper, why
         | did they not compare on a fair basis?
        
       ___________________________________________________________________
       (page generated 2024-01-02 23:01 UTC)