[HN Gopher] Qualcomm works with Meta to enable on-device AI appl...
       ___________________________________________________________________
        
       Qualcomm works with Meta to enable on-device AI applications using
       Llama 2
        
       Author : ahiknsr
       Score  : 75 points
       Date   : 2023-07-18 20:37 UTC (2 hours ago)
        
 (HTM) web link (www.qualcomm.com)
 (TXT) w3m dump (www.qualcomm.com)
        
       | oneplane wrote:
       | I doubt Qualcomm will be able to increase their performance ahead
       | of Nvidia decreasing their energy requirements.
        
         | pjmlp wrote:
         | Since NVidia hardly ships on Android phones, it doesn't really
         | matter.
        
         | jayd16 wrote:
         | Why?
        
       | [deleted]
        
       | behnamoh wrote:
       | Meanwhile Apple is stubbornly insisting on its own ways, and
       | awkwardly silent during this whole AI revolution. I gave up using
       | Siri years ago due to its glaring stupidity compared with Google
       | Assistant and Alexa.
       | 
       | While Apple keeps making money from overpriced hardware, the
       | competitors are working on actually being pioneers in AI. It
       | makes me sad to see so much computational power in my iPhone and
       | iPad getting wasted on silly subpar iOS apps.
        
         | CTDOCodebases wrote:
         | This is what I don't understand.
         | 
         | The power of LLMs is that they can allow a human to interact
         | with a computer using natural language.
         | 
         | Imagine just asking Siri to order you something and Siri finds
         | the cheapest legit supplier and just orders it from them. It
         | would remove the first hop needed to go to Amazon and unlike
         | Alexa Siri is always with you. Google and Apple could do to
         | Amazon what they did to Facebook.
        
         | huslage wrote:
         | Apple has not been silent. The whole iOS 17 release has AI
         | stuff sprinkled all through it. They just aren't speaking the
         | same language as everyone else. They are also hiring generative
         | AI folks like crazy for things like Siri. They want to own the
         | on-device space and are well on the way to revolutionizing
         | their operating systems and hardware to do it.
        
         | pazimzadeh wrote:
         | Yeah...Apple doesn't tease products before they announce them.
        
         | code51 wrote:
         | I slowly started to think Apple knows there will be a golden
         | hour to jump on a mature, economic LLM. Until then, they are
         | probably watching without a massive investment into LLMs if
         | it's not drastically changing their products.
        
         | samwillis wrote:
         | Apple are never first, sometimes they are quite late, but when
         | they do launch their version of a product it is polished and
         | fixes many issues with all other implications.
         | 
         | I have no doubt that Apple are working on local LLMs and
         | combining them with Siri, it's only a matter of time. I expect
         | we will see something next year, probably tied to the next
         | iteration of hardware - they do want to sell hardware after
         | all.
         | 
         | Apple have indicated that are working on this stuff, subtly.
         | They avoided saying "AI", instead going for "ML",
         | "encoder/decoder models" and "transformer models". The spell
         | check on the next iOS is based on a transformer model. It's
         | packed with image models for extraction, lighting and image
         | correction. iOS is littered with ML stuff, that have "Nural
         | cores" in their silicone.
         | 
         | This stuff is coming from them, when it's ready.
         | 
         | What I expect they will launch with a "LLM Siri" will blow our
         | minds, it will be an all encompassing personal assistant. It
         | will have linkage of all our documents, email, messages,
         | movements, browser history, all locally securely on our
         | devises. Not in the cloud. It will be so close to appearing
         | like AGI that some people will clams it is.
        
           | phkahler wrote:
           | >> I have no doubt that Apple are working on local LLMs and
           | combining them with Siri,
           | 
           | Yeah, Apple would want to have voice chat, not text chat.
           | Especially on a phone
        
           | FirmwareBurner wrote:
           | _> Apple are never first, sometimes they are quite late, but
           | when they do launch their version of a product it is polished
           | _
           | 
           | Apple maps would have a laugh at that
        
             | woeirua wrote:
             | Have you checked it recently? It's better than Google Maps
             | now.
        
               | FirmwareBurner wrote:
               | Read again the comment I'm replaying to mentioning
               | "polished at launched", not polished after 10 years of
               | updated and fixes.
        
             | drevil-v2 wrote:
             | Apple Maps, as it is today, is so much better than Google
             | Maps. The UI is way more polished than Google Maps. The
             | Google UI looks like it is designed by committee - it is so
             | ugly and dense. I find it really offensive.
             | 
             | The Apple Maps has better directions for my city at least.
             | The Google one re-directs me through side streets and weird
             | turns into traffic crossing lanes at the last minute.
        
               | warning26 wrote:
               | You're getting downvoted but honestly I'd agree insofar
               | as the UI.
               | 
               | Google Maps is loaded with _suggested content_. What
               | about lunch? Need a haircut? Check out these _top rated
               | places_!! Would you recommend Google Maps to a friend or
               | colleague?
               | 
               | Apple Maps is like what Google Maps was years ago: a map
               | and a search box.
        
             | DevKoala wrote:
             | Yeah I remember that as their only misstep tbh. Today,
             | Apple Maps is my go-to for California at least.
        
               | nradov wrote:
               | There have been other failed Apple products. Lisa.
               | Newton. Macintosh Portable. iPhone 6. Etc.
               | 
               | https://www.cnet.com/tech/mobile/apples-worst-failures-
               | of-al...
               | 
               | No one gets product planning right 100%. The key is to
               | recognize failures early and cut them off.
        
               | scarface_74 wrote:
               | It's estimated that the iPhone 6 and 6 Plus sold over 225
               | million.
        
               | wslh wrote:
               | The key is to have enough cash that you know you will hit
               | the nail. Microsoft is failing constantly, even launching
               | several times different products for the same use case.
               | It doesn't matter at that level.
        
               | Keyframe wrote:
               | _Yeah I remember that as their only misstep tbh._
               | 
               | MobileMe was quite a big misstep too.
        
             | pazimzadeh wrote:
             | not the first computer with windowed UI
             | 
             | not the first portable mp3 player
             | 
             | not the first smartphone
             | 
             | nor the first touchscreen phone
             | 
             | not the first bluetooth headset
             | 
             | not the first tablet computer
             | 
             | not the first smartwatch
             | 
             | not the first VR/AR headset
        
               | smoldesu wrote:
               | Many of these were not "polished" at release either.
               | 
               | - The Lisa was so expensive that it failed to find a
               | market regardless of how visionary it was, ultimately
               | making less of an impact than the Apple II.
               | 
               | - The first iPod models were plagued with battery issues
               | and had no meaningful way to replace them once it went
               | bad (this extended into smaller models like the Nano).
               | 
               | - The iPhone was notoriously gimped at launch and only
               | barely delivered on it's promises, with the majority of
               | features people know today being added in updates.
               | 
               | - The first generation Airpods should be classified as
               | instruments of sonic and physical abuse, not headphones.
               | 
               | - The first generation iPad was depreciated almost
               | immediately and got ~2 years total software support.
               | 
               | To say nothing of their success, sure, but Apple clearly
               | has their own honeymoon phases to work through.
        
               | Keyframe wrote:
               | I thought it was common belief, well at least in my
               | circle of friends, not to buy anything apple until
               | there's third generation of it.
        
         | pavlov wrote:
         | Apple has been pretty good about quietly integrating AI-based
         | image processing into iOS.
         | 
         | All text within photos is being automatically recognized and is
         | fully searchable. That has saved my ass on occasion when some
         | piece of information was in a photo.
         | 
         | The AI background removal feature in Photos is cool too -- just
         | drag from a foreground element, and it usually gets it right.
         | 
         | But despite this kind of flawless feature integration work,
         | clearly Apple as a company has some issues with applying
         | generative models on a more fundamental scale. They look a bit
         | stuck in their own loop of trying to replicate past success
         | with an upcoming $3,500 device that's all about ever fancier
         | glass and aluminum hardware.
        
         | [deleted]
        
         | FirmwareBurner wrote:
         | _> awkwardly silent during this whole AI revolution_
         | 
         | They weren't silent about AI child porn scanning on-device so
         | we know they definitely have the capability for AI on device.
        
       | simonw wrote:
       | If you have a modern iPhone you can try running an LLM directly
       | on it today using the MLC iPhone app: https://mlc.ai/mlc-
       | llm/#iphone
       | 
       | It can run Vicuna-7B which is a pretty impressive model.
       | 
       | (They have an Android app too but I haven't tried that yet).
        
         | haunter wrote:
         | >the MLC iPhone app
         | 
         | Seems like US only
        
       | smoldesu wrote:
       | > The ability to run generative AI models like Llama 2 on devices
       | such as smartphones, PCs, _VR /AR headsets_
       | 
       | Maybe it's an upcoming feature for the Quest 3?
       | 
       | To that end, I've been pretty amazed by how far quantization has
       | come. Some early llama-2 quantizations[0] have gotten down to
       | ~2.8gb, though I haven't tested it to see how it performs yet.
       | Still though, we're now talking about models that can comfortably
       | run on pretty low-end hardware. It will be interesting to see
       | where llama crops up with so many options for inferencing
       | hardware.
       | 
       | [0] https://huggingface.co/TheBloke/Llama-2-7B-Chat-
       | GGML/tree/ma...
        
       | transcriptase wrote:
       | Facebook hardware in my devices that claims it's there to protect
       | my privacy.
       | 
       | What's the catch? Firmware based "anonymous" telemetry?
        
         | serf wrote:
         | personally I hope that the facebook altruism as of late has
         | been a market effect from their needing to compete with other
         | mega-corps, and that it comes with no catch or ill-will because
         | it's simply there to be 'the other'.
         | 
         | it's probably a naive hope, but I hope it's right.
         | 
         | (i've never trusted FB further than I could throw their
         | company, so this is a big leap of faith for me.)
        
       | m3kw9 wrote:
       | ChatGPT 3.5 is the base level people expect LLMs to be, it would
       | be 2-3 generation(3-4 years) of hardware before we can reach
       | that. Anything below is just going to get bad reviews
        
         | Legend2440 wrote:
         | 3-4 years to run it on your phone seems generous, barring
         | algorithmic breakthroughs.
         | 
         | If I can run 100B+ models on my high-end desktop in 3-4 years I
         | will be very happy.
        
           | nullc wrote:
           | What do you consider a high end desktop _now_?
        
             | holoduke wrote:
             | He means a videocard with 512gb of memory or more.
        
               | delecti wrote:
               | Is 512gb a typo? The current biggest consumer card has
               | 24GB, so we're probably 15 years from a 512GB card
               | (judging from the increase of 4Gb to 24GB between 2012
               | and 2022).
        
         | kernal wrote:
         | Isn't that the same LLM that doesn't know how many e's are in
         | "ketchup"? Nice.
        
           | Legend2440 wrote:
           | Aren't you the same user that doesn't understand word-level
           | tokenization?
        
             | kernal wrote:
             | You must have me confused with an LLM bot that does know
             | how many e's are in "ketchup".
        
         | bigyikes wrote:
         | A small model might be useful for e.g. NPC interactions in a
         | Quest game
        
           | fit2rule wrote:
           | [dead]
        
         | madars wrote:
         | But for what applications? Sure, for answering free-form
         | questions I expect GPT-3.5+ quality. I don't think GPT-3.5 is
         | necessary to provide auto-complete in your email client.
        
       | seydor wrote:
       | FB should release a phone, they already have quest OS
        
         | samwillis wrote:
         | As I said in the other Llama 2 thread[0], this is Meta doing a
         | traditional textbook "commoditise your compliment". It's a
         | preemptive attack with this particular announcement on Apple
         | and Google. They don't what them own LLMs on portable devices,
         | they want it to be a commodity.
         | 
         | [0]: https://news.ycombinator.com/item?id=36775642
        
         | scarface_74 wrote:
         | What could possibly go wrong
         | 
         | https://www.cnet.com/tech/mobile/heres-why-the-facebook-phon...
         | 
         | Next up: they should make a social media app where each user
         | can customize their profile page and have their favorite music
         | start playing. It would be called YourSpace
        
       | vorpalhex wrote:
       | How would it work to get a model that needs 8Gb+ of vram
       | currently into some chiplet form factor? Is there an obvious way
       | of translating this more directly to hardware?
        
         | malux85 wrote:
         | If your filesystem is fast enough, then you can dynamically
         | load and unload chunks of it as you use it
        
           | simonw wrote:
           | I don't think that works for LLMs. My understanding is that
           | every single token produced by the model requires running
           | floating point operations against all x-billion parameters,
           | so the entire thing needs to be loaded into memory at all
           | times for it to work.
        
             | malux85 wrote:
             | Nope, you can use mmap to virtually map it to memory, and
             | then you don't have to hold the whole thing in RAM at once.
             | I have spent the last 4-5 weeks working on this and
             | optimising it
             | 
             | You can see some information here:
             | https://justine.lol/mmap/
        
               | simonw wrote:
               | How does that work? Does it mean that for every token
               | generated it has to page areas of disk into RAM and then
               | back out again?
        
             | Legend2440 wrote:
             | You can shuffle it back and forth between disk and memory.
             | It's slow, but it works.
             | 
             | There are people working on compute-in-memory hardware,
             | like flash chips that can do matrix multiplication in
             | place. None of it is close to reaching the market but
             | there's a lot more interest now that neural networks are
             | obviously useful.
        
         | jayd16 wrote:
         | Mobile chips use a shared memory model and 12+ gigs of ram for
         | mobile hardware isn't exotic at all. Seems reasonable to think
         | that it won't be an issue.
        
         | kernal wrote:
         | They're going to use an extremely compact model for extremely
         | limited use cases.
        
       | ssss11 wrote:
       | Hmm I wonder what sort of privacy impacts there are in a future
       | of having Meta (or Google et al) AI running on chips on your
       | phone when the parent company has so much info on you and
       | blatantly flaunts privacy laws.
        
       ___________________________________________________________________
       (page generated 2023-07-18 23:01 UTC)