[HN Gopher] Phi-2: The surprising power of small language models
___________________________________________________________________
Phi-2: The surprising power of small language models
Author : birriel
Score : 85 points
Date : 2023-12-12 16:29 UTC (6 hours ago)
(HTM) web link (www.microsoft.com)
(TXT) w3m dump (www.microsoft.com)
| hexomancer wrote:
| The website seems to be down. Here is the cached version:
| https://webcache.googleusercontent.com/search?q=cache:r-GYHX...
| bratao wrote:
| 2.7B size with better performance than Mistral 7B is impressive!!
| reqo wrote:
| Serious competition in the small model space recently. The main
| goal of models this small is to be deployed locally to
| phone/laptop (consumer electronics soon maybe?) I wonder if this
| will lead to a new generation of apps/UI, if it already has not.
|
| Edit: typo
| collaborative wrote:
| Can't access the model link due to sone MS auth issue. Does
| anyone know how large (in GB) the model is? Can it be run locally
| or is it azure-only?
| sytelus wrote:
| File size on disk is ~10GB.
| armcat wrote:
| Interesting, so they are using single precision (fp32), so
| 2.7B x 4 Bytes = ~10 GB. With CUDA overhead and room for
| context, you would need at least 12GB VRAM. They could use
| half precision and half that VRAM requirement and save costs
| for everyone involved. Maybe there is a performance reason
| why they use full precision.
| collaborative wrote:
| Yes, the reason I asked is that when I see SLM I got all
| excited thinking "finally a small model that fits in cheap
| hardware for simpler tasks"
| sigmar wrote:
| That's just the parameters? So 32 bit parameters? This
| blogpost is incredibly misleading by putting Gemini nano-2's
| "size" as larger than Phi-2 (in Table 2 displaying only the
| number of parameters) and saying "Phi-2 matches or
| outperforms the recently-announced Google Gemini Nano 2,
| despite being smaller in size." Because Gemini nano
| parameters are 4 bit. So Gemini nano-2 is 1.6 GB (3.25/2) in
| size compared to Phi-2's 10GB
| acheong08 wrote:
| We don't have any real information on it but the benchmarks make
| me feel like some test data made it into the training set
| duchenne wrote:
| > The training for Phi-2 took 14 days on 96 A100 GPUs
|
| This would mean that it costs around ~30k USD to train.
|
| If training an LLM becomes cheaper than buying a car, it could
| democratize AI a lot.
| eternauta3k wrote:
| You don't need to train it again, Microsoft already did.
|
| Unless you want to develop a new one, then you also need the
| team of researchers/engineers.
| alecco wrote:
| Note the model is trained on data generated by GPT-4. It's
| probably orders of magnitude more expensive to generate the
| data at current API prices.
|
| The whole point of these papers is that training data quality
| is key.
|
| I would much prefer for these companies to release the training
| data than the weights. But that will never happen.
|
| "We speculate that the creation of synthetic datasets will
| become, in the near future, an important technical skill and a
| central topic of research in AI."
| monlockandkey wrote:
| Can we download this model locally or is it Azure only?
| antimatter15 wrote:
| Looks like it is possible to download it locally, but as far as
| I can tell you have to manually copy all the various files from
| the Artifacts folder individually
| ofou wrote:
| What is the context window?
| bionhoward wrote:
| heck, microsoft seems sketchy AF, i reported em to the department
| of justice today, here's a link to the evidence of microsoft's
| anticompetitive conduct https://i.postimg.cc/MGqPvPz5/cartel-
| microsoft-microsoft-ope... just seems really serious and makes me
| not want to have anything to do with microsoft microsoft openai,
| microsoft github, or nvidia ever again. the timing seems
| suspicious since i sent the email to DOJ today, but im a nobody
| so hey, what do i know?
|
| tried a billion ways to remedy by contacting them directly before
| i sent that. you know it's a stressful workday when you're sweaty
| from making a google drawing!
|
| Ironically, Mistral cofounders just took down their similar
| clause. Google has nothing like this. Anthropic and Inflection
| both do. I'm sick of it, but I feel morally obligated to keep
| speaking out about this. Satya could just go into the codebase
| and delete it, but he didnt, I asked him to, multiple times. A
|
| WS also has such terms, really bad because they're in charge of
| Rust Language. Conflict of interest.
|
| TL;DR: I meekly suggest we boycott these companies because they
| are heavily and explicitly anti-competitive and this goes for all
| their businesses (Microsoft OpenAI and Microsoft GitHub) also it
| all runs on NVIDIA chips which have similar terms.
|
| Great job on Phi, but this is not for me!
| abeppu wrote:
| IANAL, and I'm genuinely curious -- are these customer non-
| compete clauses actually illegal?
|
| These clauses do seem clearly anti-competitive, but is that
| enough? Your complaint mentions attempts to "acquire and
| maintain monopoly power" which seems like a stretch given that
| no company seems within reach of a "monopoly" ... which is why
| you're able to rattle off a list of companies in this space.
|
| Like, if they're not actually conspiring, but they're each
| trying to squash competition, and a likely result is that only
| a small number of rich organizations have the means (including
| user data) to continually improve LLMs ... is that competition-
| blocking a crime?
| ctoth wrote:
| There really isn't a respectful way I can say this.
|
| Please reach out and speak to a therapist or similar
| professional.
|
| This stream of consciousness where you talk about personally
| reaching out to random CEOs and expecting them to do things for
| you and being surprised when you don't get a response reminds
| me strongly of when someone close to me was going through a
| manic period of delusional breakdown.
|
| Please, talk to someone.
___________________________________________________________________
(page generated 2023-12-12 23:01 UTC)