[HN Gopher] Llama 3-V: Matching GPT4-V with a 100x smaller model...
       ___________________________________________________________________
        
       Llama 3-V: Matching GPT4-V with a 100x smaller model and 500
       dollars
        
       Author : minimaxir
       Score  : 96 points
       Date   : 2024-05-28 20:16 UTC (2 hours ago)
        
 (HTM) web link (aksh-garg.medium.com)
 (TXT) w3m dump (aksh-garg.medium.com)
        
       | doctorpangloss wrote:
       | Shouldn't CogAgent be in this comparison?
        
         | m00x wrote:
         | CogVLM should be, not sure how CogAgent plays into this. This
         | isn't an agent.
        
           | doctorpangloss wrote:
           | You would use CogAgent in VQA mode.
        
       | lanceflt wrote:
       | - Llava is not the SOTA open VLM, InternVL-1.5 is
       | https://huggingface.co/spaces/opencompass/open_vlm_leaderboa...
       | 
       | You need to compare the evals to strong open VLMs including this
       | and CogVLM
       | 
       | - This is not "first-ever multimodal model built on top of
       | Llama3", there's already a Llava on Llama3-8b
       | https://huggingface.co/lmms-lab
        
         | gigel82 wrote:
         | Like InternVL, no llama.cpp support severely limits its
         | applications. Close to GPT4v performance level and runnable
         | locally on any machine (no need for a GPU) would be huge for
         | the accessibility community.
        
         | valine wrote:
         | Very curious how it performs on OCR tasks compared to InternVL.
         | To be competitive at reading text you need tiling support, and
         | InternVL does tiles exceptionally well.
        
       | behnamoh wrote:
       | This "matching gpt-4" catchy phrase has lost its meaning to me.
       | Everytime an article like this pops up, I see marketing buzz and
       | unrealistic results in practice.
        
         | Mo3 wrote:
         | Of course, it's nothing else. Who could possibly believe that
         | OpenAI and others would dump billions into development and
         | training and aren't smart enough to figure out they could also
         | do it with $500.
        
           | whimsicalism wrote:
           | it would have been a lot cheaper for oai if they had access
           | to llama3 in 2018
        
           | KorematsuFredt wrote:
           | You have clearly not read the article. $500 is the cost of
           | fine tuning.
        
           | bilbo0s wrote:
           | _Who could possibly believe that OpenAI and others would dump
           | billions into development and training and aren 't smart
           | enough to figure out they could also do it with $500._
           | 
           | People upvoting the post??
           | 
           | Not really sure? But PT Barnum said there's always a lot of
           | them out there.
           | 
           | Pretty sure they mean fine tuning though?
           | 
           | But even that is total tripe.
           | 
           | These guys are snake oil salesmen. (Or Sylvester McMonkey
           | McBean is behind it.)
        
           | nomel wrote:
           | It's llama 3 training cost + their cost. Meta "kindly"
           | covered the first $700M.
           | 
           | > We add a vision encoder to Llama3 8B
        
         | mpalmer wrote:
         | For me it's become a signal the person making the claim is
         | unserious.
        
       | KTibow wrote:
       | Is there a reason Phi Vision is omitted?
        
         | cadence- wrote:
         | Is there any place that currently hosts phi3 Vision and
         | provides API access to it? I cannot run it on my local machine,
         | unfortunately.
        
       | yeldarb wrote:
       | Don't see a license listed in the repo; presumably needs to be
       | the same as Meta's Llama 3 license?
        
       ___________________________________________________________________
       (page generated 2024-05-28 23:00 UTC)