[HN Gopher] Phi-2: The surprising power of small language models
       ___________________________________________________________________
        
       Phi-2: The surprising power of small language models
        
       Author : birriel
       Score  : 85 points
       Date   : 2023-12-12 16:29 UTC (6 hours ago)
        
 (HTM) web link (www.microsoft.com)
 (TXT) w3m dump (www.microsoft.com)
        
       | hexomancer wrote:
       | The website seems to be down. Here is the cached version:
       | https://webcache.googleusercontent.com/search?q=cache:r-GYHX...
        
       | bratao wrote:
       | 2.7B size with better performance than Mistral 7B is impressive!!
        
       | reqo wrote:
       | Serious competition in the small model space recently. The main
       | goal of models this small is to be deployed locally to
       | phone/laptop (consumer electronics soon maybe?) I wonder if this
       | will lead to a new generation of apps/UI, if it already has not.
       | 
       | Edit: typo
        
       | collaborative wrote:
       | Can't access the model link due to sone MS auth issue. Does
       | anyone know how large (in GB) the model is? Can it be run locally
       | or is it azure-only?
        
         | sytelus wrote:
         | File size on disk is ~10GB.
        
           | armcat wrote:
           | Interesting, so they are using single precision (fp32), so
           | 2.7B x 4 Bytes = ~10 GB. With CUDA overhead and room for
           | context, you would need at least 12GB VRAM. They could use
           | half precision and half that VRAM requirement and save costs
           | for everyone involved. Maybe there is a performance reason
           | why they use full precision.
        
             | collaborative wrote:
             | Yes, the reason I asked is that when I see SLM I got all
             | excited thinking "finally a small model that fits in cheap
             | hardware for simpler tasks"
        
           | sigmar wrote:
           | That's just the parameters? So 32 bit parameters? This
           | blogpost is incredibly misleading by putting Gemini nano-2's
           | "size" as larger than Phi-2 (in Table 2 displaying only the
           | number of parameters) and saying "Phi-2 matches or
           | outperforms the recently-announced Google Gemini Nano 2,
           | despite being smaller in size." Because Gemini nano
           | parameters are 4 bit. So Gemini nano-2 is 1.6 GB (3.25/2) in
           | size compared to Phi-2's 10GB
        
       | acheong08 wrote:
       | We don't have any real information on it but the benchmarks make
       | me feel like some test data made it into the training set
        
       | duchenne wrote:
       | > The training for Phi-2 took 14 days on 96 A100 GPUs
       | 
       | This would mean that it costs around ~30k USD to train.
       | 
       | If training an LLM becomes cheaper than buying a car, it could
       | democratize AI a lot.
        
         | eternauta3k wrote:
         | You don't need to train it again, Microsoft already did.
         | 
         | Unless you want to develop a new one, then you also need the
         | team of researchers/engineers.
        
         | alecco wrote:
         | Note the model is trained on data generated by GPT-4. It's
         | probably orders of magnitude more expensive to generate the
         | data at current API prices.
         | 
         | The whole point of these papers is that training data quality
         | is key.
         | 
         | I would much prefer for these companies to release the training
         | data than the weights. But that will never happen.
         | 
         | "We speculate that the creation of synthetic datasets will
         | become, in the near future, an important technical skill and a
         | central topic of research in AI."
        
       | monlockandkey wrote:
       | Can we download this model locally or is it Azure only?
        
         | antimatter15 wrote:
         | Looks like it is possible to download it locally, but as far as
         | I can tell you have to manually copy all the various files from
         | the Artifacts folder individually
        
       | ofou wrote:
       | What is the context window?
        
       | bionhoward wrote:
       | heck, microsoft seems sketchy AF, i reported em to the department
       | of justice today, here's a link to the evidence of microsoft's
       | anticompetitive conduct https://i.postimg.cc/MGqPvPz5/cartel-
       | microsoft-microsoft-ope... just seems really serious and makes me
       | not want to have anything to do with microsoft microsoft openai,
       | microsoft github, or nvidia ever again. the timing seems
       | suspicious since i sent the email to DOJ today, but im a nobody
       | so hey, what do i know?
       | 
       | tried a billion ways to remedy by contacting them directly before
       | i sent that. you know it's a stressful workday when you're sweaty
       | from making a google drawing!
       | 
       | Ironically, Mistral cofounders just took down their similar
       | clause. Google has nothing like this. Anthropic and Inflection
       | both do. I'm sick of it, but I feel morally obligated to keep
       | speaking out about this. Satya could just go into the codebase
       | and delete it, but he didnt, I asked him to, multiple times. A
       | 
       | WS also has such terms, really bad because they're in charge of
       | Rust Language. Conflict of interest.
       | 
       | TL;DR: I meekly suggest we boycott these companies because they
       | are heavily and explicitly anti-competitive and this goes for all
       | their businesses (Microsoft OpenAI and Microsoft GitHub) also it
       | all runs on NVIDIA chips which have similar terms.
       | 
       | Great job on Phi, but this is not for me!
        
         | abeppu wrote:
         | IANAL, and I'm genuinely curious -- are these customer non-
         | compete clauses actually illegal?
         | 
         | These clauses do seem clearly anti-competitive, but is that
         | enough? Your complaint mentions attempts to "acquire and
         | maintain monopoly power" which seems like a stretch given that
         | no company seems within reach of a "monopoly" ... which is why
         | you're able to rattle off a list of companies in this space.
         | 
         | Like, if they're not actually conspiring, but they're each
         | trying to squash competition, and a likely result is that only
         | a small number of rich organizations have the means (including
         | user data) to continually improve LLMs ... is that competition-
         | blocking a crime?
        
         | ctoth wrote:
         | There really isn't a respectful way I can say this.
         | 
         | Please reach out and speak to a therapist or similar
         | professional.
         | 
         | This stream of consciousness where you talk about personally
         | reaching out to random CEOs and expecting them to do things for
         | you and being surprised when you don't get a response reminds
         | me strongly of when someone close to me was going through a
         | manic period of delusional breakdown.
         | 
         | Please, talk to someone.
        
       ___________________________________________________________________
       (page generated 2023-12-12 23:01 UTC)