[HN Gopher] Llama 3-V: Matching GPT4-V with a 100x smaller model...
___________________________________________________________________
Llama 3-V: Matching GPT4-V with a 100x smaller model and 500
dollars
Author : minimaxir
Score : 96 points
Date : 2024-05-28 20:16 UTC (2 hours ago)
(HTM) web link (aksh-garg.medium.com)
(TXT) w3m dump (aksh-garg.medium.com)
| doctorpangloss wrote:
| Shouldn't CogAgent be in this comparison?
| m00x wrote:
| CogVLM should be, not sure how CogAgent plays into this. This
| isn't an agent.
| doctorpangloss wrote:
| You would use CogAgent in VQA mode.
| lanceflt wrote:
| - Llava is not the SOTA open VLM, InternVL-1.5 is
| https://huggingface.co/spaces/opencompass/open_vlm_leaderboa...
|
| You need to compare the evals to strong open VLMs including this
| and CogVLM
|
| - This is not "first-ever multimodal model built on top of
| Llama3", there's already a Llava on Llama3-8b
| https://huggingface.co/lmms-lab
| gigel82 wrote:
| Like InternVL, no llama.cpp support severely limits its
| applications. Close to GPT4v performance level and runnable
| locally on any machine (no need for a GPU) would be huge for
| the accessibility community.
| valine wrote:
| Very curious how it performs on OCR tasks compared to InternVL.
| To be competitive at reading text you need tiling support, and
| InternVL does tiles exceptionally well.
| behnamoh wrote:
| This "matching gpt-4" catchy phrase has lost its meaning to me.
| Everytime an article like this pops up, I see marketing buzz and
| unrealistic results in practice.
| Mo3 wrote:
| Of course, it's nothing else. Who could possibly believe that
| OpenAI and others would dump billions into development and
| training and aren't smart enough to figure out they could also
| do it with $500.
| whimsicalism wrote:
| it would have been a lot cheaper for oai if they had access
| to llama3 in 2018
| KorematsuFredt wrote:
| You have clearly not read the article. $500 is the cost of
| fine tuning.
| bilbo0s wrote:
| _Who could possibly believe that OpenAI and others would dump
| billions into development and training and aren 't smart
| enough to figure out they could also do it with $500._
|
| People upvoting the post??
|
| Not really sure? But PT Barnum said there's always a lot of
| them out there.
|
| Pretty sure they mean fine tuning though?
|
| But even that is total tripe.
|
| These guys are snake oil salesmen. (Or Sylvester McMonkey
| McBean is behind it.)
| nomel wrote:
| It's llama 3 training cost + their cost. Meta "kindly"
| covered the first $700M.
|
| > We add a vision encoder to Llama3 8B
| mpalmer wrote:
| For me it's become a signal the person making the claim is
| unserious.
| KTibow wrote:
| Is there a reason Phi Vision is omitted?
| cadence- wrote:
| Is there any place that currently hosts phi3 Vision and
| provides API access to it? I cannot run it on my local machine,
| unfortunately.
| yeldarb wrote:
| Don't see a license listed in the repo; presumably needs to be
| the same as Meta's Llama 3 license?
___________________________________________________________________
(page generated 2024-05-28 23:00 UTC)