[HN Gopher] The tide is shifting: 1.3B outperforms 7B Llama 2
___________________________________________________________________
The tide is shifting: 1.3B outperforms 7B Llama 2
Author : __vec__
Score : 42 points
Date : 2023-09-12 14:55 UTC (8 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| __vec__ wrote:
| Textbooks Are All You Need II: phi-1.5 technical report
|
| "Perhaps achieving ChatGPT's level of capability at the one
| billion parameters scale is actually achievable?"
| refulgentis wrote:
| On a long enough timeline, Id say yes.
|
| Right now they're comparing to the 7B Llama 2, which is a
| shadow of 65B llama2, which is a couple steps off ChatGPT,
| which is a shadow of GPT4.
|
| I'm still comfy saying yes because there's no reason to doubt
| it will follow the same logic as "eventually $phone will have
| the same FLOPS as $desktop"
| matteoraso wrote:
| The model can be downloaded here:
| https://huggingface.co/microsoft/phi-1_5
| brucethemoose2 wrote:
| > we are seeing improvement on that front thanks to the absence
| of web data
|
| Bingo. This (not the parameter count) is the amazing thing to me.
|
| Garbage in garbage out, and there is a _ton_ of garbage in the
| Falcon /Llama (and OpenAI?) datasets. It feels like such a waste
| of compute and parameter space.
| [deleted]
| tayo42 wrote:
| it might not exactly be a waste, when you are inputting your
| training data you need some context on what the source is and
| then add a weight to how valid it is. That is how we learn, if
| i talk to a phd on some topic, I would pay close attention to
| what they say and give it more significance, maybe even write
| it down, compared to a youtube video on that topic, compared to
| a guy on a street talking about that same topic, ill probably
| nod my head and walk away and try to forget as soon as i can.
| but there probably still is something valid to learn from that
| last conversation, how a certain type of person structures
| sentences, responds, the words used etc...
|
| do these llms get trained with something like a credibility
| weight on the training data? that was it seems they did in this
| paper, just manually curated that
| mpixel wrote:
| Do people want that though? I don't want phd level responses
| for my queries. I want it to be better than what I could come
| up in a minute or by searching half an hour. Rather than some
| highly advanced highly detailed response I could probably not
| understand if the topic is not something I'm sufficient in to
| begin with.
|
| Think common use cases. A lot of users are students, do I
| want it to write an essay like a linguist? Or solve my
| homework using the better but more advanced techniques and
| style?
| Bjartr wrote:
| It's a lot easier to add a pass to simplify an explanation
| by rephrasing or eliding information than it is to smarten
| up an overly simplified answer.
|
| You do want the underlying model to be capable of the
| advanced answers, since if it is, it can be used to supply
| simple answers. You can't make that work the other way
| around in the same way.
| tayo42 wrote:
| I think they want it. What do you think the most perfect or
| ideal question answer-er or teacher is? It would probably
| be an expert in the field, but also with the ability to
| deliver that content in a level appropriate way for the
| recipient/student. Unreasonable for the most part, we get
| away with good enough in the real world. This is a skill we
| all try to learn though. Like when you need to give a
| technical presentation to non technical audience
|
| If you have input data that includes a high rated reddit
| eli5 question, the content of that answer might be hard to
| verify, the style and way its delivered would be ideal to
| keep around in the training data. on the other side,
| technical in-depth answers have content that is worth
| keeping around, the style of its delivery would be very
| specific.
|
| Keeping the entire internet around in your training data
| would still give you access to all these types of delivery
| still. hope that makes sense.
| Valgrim wrote:
| Maybe the answer is multi-step: first use curated primary
| sources, e.g. scientific papers. Then reinforce using
| well written summaries, perhaps by actual models or well
| graded student papers. Finally, somehow apply negative
| weights using wrong answers only. Bonus points if you can
| automate the whole process
| og_kalu wrote:
| >Garbage in garbage out, and there is a ton of garbage in the
| Falcon/Llama (and OpenAI?) datasets. It feels like such a waste
| of compute and parameter space.
|
| This is true and most researchers understand this on some level
| (https://arxiv.org/abs/2305.07759) but understanding it doesn't
| really make the problem any easier. How do you curate a general
| purpose model's dataset without throwing the baby out with the
| bathwater ?
|
| EDIT: This paper seems to take a good stab at the question.
| https://arxiv.org/abs/2309.04564
| brucethemoose2 wrote:
| > How do you curate a general purpose model's dataset without
| throwing the baby out with the bathwater
|
| Time, money, and human work. The pruning process needs
| focused input from experts other than ML researchers and data
| scientists.
| nbardy wrote:
| The title is clickbait and not the title of the paper.
| YetAnotherNick wrote:
| Llama generally acheives much higher accuracy with very small
| amount of fine-tuning(on similar quality of dataset like this
| paper) on lot of tasks. So the model understanding is present in
| llama to get higher accuracy. e.g Hellaswag, ARC and MMLU for 7b
| model is 0.8, 0.57 and 0.52 respectively[0], while phi-1 is 0.48,
| 0.45 and 0.38.
|
| I don't think finetuning phi-1 on good quality synthetic data
| will increase its accuracy as it is only trained on that.
|
| [0]: https://huggingface.co/pankajmathur/orca_mini_v3_7b
___________________________________________________________________
(page generated 2023-09-12 23:02 UTC)