[HN Gopher] Apple wants AI to run directly on its hardware inste...
       ___________________________________________________________________
        
       Apple wants AI to run directly on its hardware instead of in the
       cloud
        
       Author : thunderbong
       Score  : 158 points
       Date   : 2023-12-21 19:06 UTC (3 hours ago)
        
 (HTM) web link (arstechnica.com)
 (TXT) w3m dump (arstechnica.com)
        
       | tech_ken wrote:
       | > "As those things get scaled up, they start to make a meaningful
       | change in the user experience and enable new innovation which has
       | the potential to create a new upgrade cycle in smartphones."
       | 
       | Interesting read. "Create a new upgrade cycle" is a phrase that's
       | almost too on the nose, but will be interesting to see if the AI-
       | powered feature set genuinely justifies an upgrade. I have a hard
       | time imagining someone wanting to shell out 4 digits to be able
       | to run ex. stable diffusion locally if they weren't already in
       | the market for a new phone, but maybe I'm underestimating the AI
       | hype wave right now. I could see predictive typing or improved
       | voice assistants being the two biggest opportunities for a killer
       | feature, but even then not sure if the value proposition would
       | change that much, especially if the battery life takes any hit
       | whatsoever.
        
         | msoad wrote:
         | Siri that actually works sounds very appealing to be honest.
         | All the integration and UX is there. It just doesn't work
         | today. So an LLM based Siri that can understand my request is
         | worth an upgrade (for me)
        
           | merpnderp wrote:
           | Exactly, just figure out how to run Whisper on my phone and
           | it would be the best phone on the market. Whisper plus
           | Mixtral on my phone and it could do anything.
        
           | rollcat wrote:
           | On the contrary, Siri already works perfectly (or rather,
           | used to work) for a lot of things... It's just Apple keeps on
           | randomly changing stuff, and every time they change
           | something, shit breaks.
           | 
           | "Hey Siri, good evening" - "I turned on the `good morning`
           | scene."
           | 
           | "Hey Siri, timer five minutes" - "I set an alarm for
           | seventeen hundred hours"
           | 
           | "Hey Siri, weather" - turns on the lights in the bathroom
           | 
           | I wish I was joking or exaggerating... The last thing I need
           | is an LLM, just please fix the obvious problems and then
           | maybe iterate on that
        
             | ceejayoz wrote:
             | Yeah, I'll get random weeks in which "text <wife>" works
             | fine, and other weeks where it says "which one?" and gives
             | me her email and phone as different contacts. No changes on
             | my end that'd cause it.
        
               | kstrauser wrote:
               | My wife is "Jen". I used to work with "Joe Chan". To this
               | day, if I say "text Jen ...", there's about a 20% chance
               | it'll reply with "OK, saying ... to Joe Chan."
               | 
               | Dude. How about optimizing for the option that I've
               | actually communicated with this decade, and 4 times
               | already today?
        
               | raccoonDivider wrote:
               | Makes me wonder how voice assistants handle people with
               | multiple partners. Maybe "which one" is actually the
               | right answer for some users?
        
             | oidar wrote:
             | Additionally, in a multi-person household adding things to
             | a shopping list should take any authentication. Not
             | everybody has an iDevice or needs to have an iDevice - my
             | petsitter, babysitter, maid and random family members
             | should be able to add stuff to the shopping list just
             | because they are in the room with the homePod. On Alexa, it
             | works - but we have to deal with all the advertisements for
             | crappy AliExpress junk. As soon as Apple figures out
             | shopping lists for whoever is in the room, I'll be dropping
             | all my Alexa devices.
        
             | ip26 wrote:
             | It works great if you know the exact incantation, but the
             | discoverability of those incantations is terrible, and
             | that's the exact problem.
        
               | rollcat wrote:
               | I spent the past 30+ years typing exact incantations into
               | a computer. I give Siri an exact incantation today and it
               | works, I give Siri the same exact incantation tomorrow
               | and it breaks.
        
             | amluto wrote:
             | > "Hey Siri, timer five minutes" - "I set an alarm for
             | seventeen hundred hours"
             | 
             | I think I figured out the issue with this one, finally:
             | Siri understands "Set a timer for 6:30" as a request for a
             | timer for however far in the future 6:30 is. And it also
             | understands "set a timer for 6" similarly.
             | 
             | And Siri doesn't seem at all intelligent about when to stop
             | transcribing.
             | 
             | So Siri has a rather large chance of interpreting "set a
             | timer for 5 minutes" as "set a timer for 5" and ending up
             | with utter nonsense.
             | 
             | A LLM that jointly handles language and audio might do much
             | better.
        
               | n9 wrote:
               | Something's not right. I say "hey siri 10 minute timer"
               | or somesuch like five times a day. I'd play with your
               | timing and word order. For instance my kid has trouble
               | making music work with siri in the car, but if I say "hey
               | siri play the album technique by the band new order" just
               | like that it works 100% of the time.
        
           | n9 wrote:
           | Not sure what you're needing that you're note getting --
           | there is certainly a lot that it doesn't do -- but Siri works
           | like a champ for me these days for maybe 70% of the things
           | that I would do with my phone or ever computer without it. I
           | initiate all calls with it, do tons of texting with it, cue
           | up all my music in the car with it, constantly add things to
           | reminders lists ('hey siri... add soy sauce to the grocery
           | list' results in soy sauce on my grocery list... _in the
           | condiments sublist_ )
           | 
           | I hadn't noticed the improvements until I started using my
           | apple watch a lot... I had given up on it a few years back, I
           | guess.
        
         | ForkMeOnTinder wrote:
         | fwiw Google is attempting the same thing, but they're not doing
         | much better.
         | 
         | https://www.notebookcheck.net/MrWhosetheboss-video-reveals-G...
         | 
         | > Google's Pixel 8 Pro Tensor G3 off-loads all generative AI
         | tasks to the cloud
        
         | rainclouds wrote:
         | Critical.
         | 
         | I'd like my data to stay local.
         | 
         | I probably don't need cloud processing for any regular usage.
         | 
         | The privacy to life benefit equation is broken in society atm.
         | 
         | This will change.
        
           | tech_ken wrote:
           | It didn't seem like training will occur locally, just serving
           | prompt responses from a GPT model (maybe I misunderstood the
           | article). My personal privacy calculus places a lot of less
           | value on keeping prompts local, especially if some kind of
           | E2EE can be provided between me and the serving model. I
           | understand that not everyone feels this way, but from the
           | perspective of market adoption I don't see the average user
           | caring too much whether a prompt is kept local or sent to the
           | cloud (if encrypted or similar). To me it seems like the most
           | critical determinants would be response latency and battery
           | life, plus the basic performance of the thing.
        
             | mitjam wrote:
             | NVIDIA's newest GPUs support confidential computing which
             | can be used for E2EE but also for intellectual property
             | protection. https://www.nvidia.com/en-us/data-
             | center/solutions/confident...
        
               | tech_ken wrote:
               | If it's just serving prompt responses I'm not sure if the
               | IP protection is a huge dealbreaker for me as a
               | downstream user. The most critical use case I can
               | currently imagine is better predictive text for my
               | keyboard, and that one seems pretty light on the IP
               | violations (for me personally). The biggest thing for me
               | here is E2EE if the text is leaving my phone.
        
         | hx8 wrote:
         | There aren't many more single new features that are going to
         | justify upgrading that will exist on the mobile platform. Apple
         | is playing a game where they try to add enough features to
         | prevent the upgrade cycle from becoming too long. Encourage
         | people with 2 or 3 year old phones to upgrade this cycle
         | instead of delaying a generation.
         | 
         | Honestly, Apple's business model is harder if everyone wants to
         | upgrade to a certain new version. The M1 Macs were so good,
         | everyone upgraded their Intel Macs and then mac sales dropped.
         | It's better logistically if the sales are more predictable and
         | consistent year-to-year.
        
           | tech_ken wrote:
           | That's an interesting point, so more of a game of amortizing
           | upgrades over a few years rather than instigating a massive
           | one-off purchasing craze. I could see how a drip of
           | sufficiently useful AI features and improvements over like
           | '24-'27 would bolster the market. Even then though, this
           | seems like an opportunity to hold rather than really reverse
           | the trend; at some point I think Apple is going to have to
           | reckon with the fact that the market is more or less
           | saturated, though I could definitely be wrong here.
        
             | hx8 wrote:
             | Apple knows the cell phone market is saturated. That's why
             | they have iPads, AirPods, Apple Watches, MacBook Airs, etc.
             | Their plan is to create new accessories (and services) that
             | pair with the iPhone. This generates a moat around existing
             | iPhone users.
        
               | tech_ken wrote:
               | I guess if they start layering LLMs or similar into their
               | service or accessory interop stack I could see how having
               | the phone running the model could be a big win for
               | latency and availability. Like if Apple Calendar uses a
               | GPT to create a .ics from the invitation text my friend
               | sent me then it could be very handy to do that locally,
               | moreso if I'm using my Apple Watch to actually signal
               | that I want the .ics created.
        
           | n9 wrote:
           | I think people (and the market) make this assumption but the
           | correction isn't Apple doing something differently it is the
           | market and public assumption seeing that when you sell
           | something that lasts longer than a year and works really well
           | you need to accomodate the ebbs and flows of sales across a 5
           | year period instead of quarterly and YoY.
        
       | yieldcrv wrote:
       | 7B parameter language models running on metal work really well
       | 
       | Other kinds of transformers could too
       | 
       | Fix Apple autocorrect and detect replacements for misspelled
       | words better
        
       | TechRemarker wrote:
       | Of course they "want" it to just like they want Siri and
       | everything else to be computed locally since best for privacy and
       | offline performance. Whether they compete with a cloud AI time
       | will tell. Will presumably be far more limited than ChatGPT 4 but
       | hopefully light years better than Siri which is what is most
       | important.
        
         | dade_ wrote:
         | Also light years ahead when the cloud access isn't available or
         | spotty in an elevator, the middle of the lake, a parkade,
         | travelling with cellular turned off and wifi is on.
        
         | GenericPoster wrote:
         | > Whether they compete with a cloud AI time will tell.
         | 
         | But they're not competing with cloud AI. Why would a person
         | need to go to the cloud to give you a reminder or download an
         | app? They're competing against the current local assistant,
         | Siri.
         | 
         | Large models are great but they can't fit on 8 or 16gb of ram.
         | And that's a very big deal.
        
           | theshrike79 wrote:
           | They don't need to put all of the world's information
           | locally, just the relevant bits. It doesn't need to know
           | every celebrity's full history for example
           | 
           | You can have the basic stuff on-device with the "smarts" of
           | an LLM that can have conversations with the user and have
           | context to previous questions.
           | 
           | The other stuff can be fetched from the cloud (with the
           | user's permission OFC) and optionally saved locally.
        
         | theshrike79 wrote:
         | You can make a local model contact the cloud for specific
         | questions like current events.
         | 
         | You can't make a cloud model work 100% locally for privacy or
         | during connection issues.
        
       | mg wrote:
       | This brings up the question how efficient the human brain is.
       | 
       | As AI keeps catching up more and more with human intelligence,
       | will it also need more and more hardware? Can it achieve human
       | and superhuman intelligence and still run on a device like the
       | iPhone, which weights about 170g? The human brain is 8 times
       | heavier.
        
         | acchow wrote:
         | The iPhone 15 neural engine does 35Tflops, which is about on
         | par with the fastest supercomputer from 2002.
         | 
         | Maybe in 24 years the iPhone 39 will have the computing power
         | that OpenAI has today? Hopefully our compression gets much
         | better in the meantime.
        
         | fumar wrote:
         | Did our brains optimize for size, weight, and or power
         | efficiency?
        
         | filterfiber wrote:
         | > still run on a device like the iPhone, which weights about
         | 170g [1]? The human brain is 8 times heavier.
         | 
         | Why are you comparing the weight? I'm sorry but this is a
         | bizarre comparison. This isn't even apples to oranges, this is
         | apples to a telephone pole.
         | 
         | I'll also throw in that a single Nvidia H100 is 1200 grams.
         | Unless you have a "Bracket with screws" which will add 20 grams
         | (who wouldn't want an extra 20grams of intelligence?).
         | 
         | Like 70%+ of the human brain is water. The human brain needs a
         | massive network of systems to transport nutrients/oxygen which
         | is irrelevant to logical processing.
         | 
         | Similarly the majority of the iphone weight is the battery and
         | frame. The weight of the processing chip is _grams_.
         | 
         | Besides the conflicting variables with the weight, the way that
         | ML works on a physical level is completely different from the
         | human brain.
        
       | PaulKeeble wrote:
       | They are going to want to start putting more RAM in their
       | machines then or they are going to find it a bit tricky to
       | compete given how big even moderate AI is.
        
         | freedomben wrote:
         | They will ship hardware (a new iphone for example) that has a
         | co-processor and flash storage dedicated to the Apple model
         | (which many are calling Apple GPT). They aren't going to make
         | it easy to run your own model. That's completely against their
         | style
        
       | rgbrgb wrote:
       | I'm totally aligned with this line of thought and I think running
       | local is the best possible thing to do with all of the privacy
       | risks AI presents. In the future I want my cognition 10x
       | amplified by AI and the only way I can imagine that working is if
       | info is being fluidly exchanged with my brain at the speed of
       | thought. No way am I going to plug into a cloud data sink for
       | advertisers and spy agencies to snoop my unfiltered thoughts. I
       | want my cognition air-gapped on hardware I own.
        
       | tkems wrote:
       | I think that running AI on device is the clear next step in the
       | world of super powerful portable phones/devices. I have a 2021
       | model phone with very powerful specs (granted it is the high end
       | model that I got used). With nearly the same RAM and storage as
       | some of my laptops from a few years ago, I see that Apple might
       | want to offload some of their expensive AI apps to these powerful
       | user devices.
       | 
       | I see this also as a privacy win as I don't want any of my data
       | being used for training. I also like to have control over what
       | models I use and the ethics the models follow.
        
         | itake wrote:
         | I am cautiously optimistic, but the reality is these models are
         | huge. LLama 7b is 13g. Even if there is enough cpu cycles, the
         | phone's storage can't support multiple models (and there will
         | need to be multiple models).
         | 
         | Compare that to cloud, you can run the latest and greatest
         | model and effectively unlimited storage.
        
           | mejutoco wrote:
           | But the cloud would cost apple money, while if ran in your
           | own device they can sell you the hardware for a healthy
           | margin, and market privacy. I think a single useful model on
           | device or many tiny ones are feasible. Even today, the photos
           | app on my iphone knows what pictures are dog pictures, can
           | remove backgrounds, etc.
        
             | josho wrote:
             | It would simply become a part of the iCloud+ bundle and
             | contribute to service revenue.
             | 
             | Hell, that's likely what will happen anyway. The models run
             | on device with cloud assistance.
        
           | ynniv wrote:
           | A reasonable quantization (Q4_K_M) is 4GB on disk and uses
           | 6.5 GB of RAM. It will run at reasonable speed on an iPhone
           | 15 Pro today. Using a smaller model or a more efficient
           | quantization on a phone two years from now (aka whatever's in
           | their lab today) will be easier than running the latest First
           | Person Shooter.
           | 
           | Welcome to the future.
           | 
           | Edit: I see a few comments about hundreds of GB. That works
           | for storage, but DRAM consumes battery just by existing, and
           | that's been keeping RAM sizes from increasing the way storage
           | has.
        
           | wayfinder wrote:
           | 13 GB a lot? What year is it? 1995? If Apple or Google wanted
           | to add 1 TB to their phones, it wouldn't cost them anywhere
           | near what they charge us to add 256 GB.
           | 
           | A copy of GTA V takes like 100 GB.
        
           | furyofantares wrote:
           | Inference/RAM is 100% the bottleneck, not storage, and it
           | isn't close.
           | 
           | These models all have to be fully loaded in memory to run. I
           | think the max RAM on an iPhone is 8GB, max storage is 1TB, so
           | we're talking 2 orders of magnitude more ability to store
           | models vs actually running them.
           | 
           | And of course you'd really rather not use the full RAM for
           | inference.
        
             | BryantD wrote:
             | You're 100% correct on the technical challenges here, but
             | the fully loaded requirement is what Apple is trying to
             | address as per the paper they just published. I'm somewhat
             | skeptical that they've solved it since it's a hard problem.
             | We'll see!
        
           | theshrike79 wrote:
           | I'm currently running dolphin 2.2.1 mistral on an M1 Mac Mini
           | with 16GB of RAM. Zero optimisations, just running base LM
           | Studio, I haven't even bothered with GPU acceleration.
           | 
           | It thinks about 5-10 seconds before answering and writes the
           | answer about the speed I can read it.
           | 
           | The current top of the line iPads are more powerful than this
           | and an optimised modern iPhone implementation using the ML
           | hardware should be faster than my current setup.
        
           | tkems wrote:
           | Just as an example, my phone has 512gb of onboard storage and
           | 16gb of RAM. I think it is safe to say that in the not-to-
           | distant future, we will see the option to run multiple models
           | on device with only a slight storage hit.
        
         | criley2 wrote:
         | I wager there is literally zero chance that Apple will let you
         | have any freedom in the model used, and there's zero chance
         | that you'll be using anything except the Siri_Model unless you
         | hack your iDevice.
         | 
         | Just like there is 1 app store and 1 assistant, Apple will
         | fight forcefully to make sure there is 1 AI model.
        
           | eitland wrote:
           | EU is already demanding that Apple allows more app stores,
           | isn't it?
        
             | criley2 wrote:
             | Well, then if you can convince the EU to force Apple to
             | allow uncensored models that will happily teach you how to
             | make bombs and all manner of things like that, then you can
             | have those models on your iDevice. Not sure the EU is going
             | to jump on board that kind of free speech though.
        
               | danogentili wrote:
               | You could make the same (unconvincing) argument for apps
               | on third-party stores, teaching the same dangerous things
               | or doing even worse things like (gasp) allowing
               | pornography :P
        
               | criley2 wrote:
               | Generally speaking, Apple has maintained rigid control
               | over apps and have banned everything you mentioned.
               | 
               | If EU app stores fill up with pornography and terrorism
               | apps, you can imagine their move will be decried as a
               | debacle and used by Apple to demonstrate how the EU
               | failed by forcing them to open up
        
               | danogentili wrote:
               | If EU app stores fill up with pornography and terrorism
               | apps, that will be a problem of the companies running the
               | stores, not Apple's problem.
               | 
               | Apple has no right to decide what users do with the
               | hardware they sell.
        
               | theshrike79 wrote:
               | Can we start expanding this argumentation to Microsoft,
               | Sony and Nintendo too?
               | 
               | I really want to start playing hardcore porn games on my
               | PS5 as soon as possible. Sony has no right to decide what
               | I do with my hardware, right?
        
               | orangecat wrote:
               | _Can we start expanding this argumentation to Microsoft,
               | Sony and Nintendo too?_
               | 
               | Sure.
               | 
               |  _I really want to start playing hardcore porn games on
               | my PS5 as soon as possible._
               | 
               | And if somebody figures out how to do that, they should
               | be able to.
               | 
               |  _Sony has no right to decide what I do with my hardware,
               | right?_
               | 
               | Pretty much, yes.
        
           | drexlspivey wrote:
           | Is Apple going to ship these LLMs as part of the OS? I expect
           | app developers to bundle them in their apps. Also Llama has a
           | specific license clause that it's available to anyone except
           | Apple and Google
        
             | vinni2 wrote:
             | you can already run llms on iphone locally
             | https://llm.mlc.ai/
        
             | criley2 wrote:
             | I find it very difficult to believe that Apple is going to
             | incorporate AI hardware into their device and then open it
             | up to everyone freely. That would be very off brand. Even
             | in the browser space Apple famously forced all competing
             | browsers to use Apple Webkit as their engine. They
             | certainly have history around lockdown and lock-in.
        
               | williamcotton wrote:
               | They don't have to open it up freely. They could choose
               | to add an API with a chat completion type of interface,
               | whisper-like, etc.
               | 
               | Nothing would need to change in the existing terms and
               | conditions with developers: we only accept binaries that
               | use our official APIs.
        
           | vimy wrote:
           | Right now there are already local LLM apps with uncensored
           | models in the app store.
        
             | Someone wrote:
             | I think there are two things that you can't do on iOS:
             | 
             | - have "Hey Siri" send audio to your code
             | 
             | - have your code reliably listen in to the microphone so
             | that your code can listen for its own trigger phrase (AFAIK
             | you can start an audio recording session, and have that
             | continue while your app is in the background, but that
             | requires users to open your app first. I also am not sure
             | that recording will reliably survive device sleep and
             | automatic app shutdown because of memory pressure)
             | 
             | I think #1 is fine. It doesn't seem fair to me that Apple
             | should be forced to allow other developers link "Hey Siri"
             | to their voice assistant because that assistant may be
             | detrimental to the value of the "Hey Siri" mark or Apple's
             | values, for example because it is racist, sexist,
             | homophobic, etc.
             | 
             | The second IMO isn't fine, but I can see Apple argue they
             | are in their rights to have some control over how that
             | rigger phrase gets detected. "Hey Siri" uses a small on-
             | device network specifically trained for that, in order to
             | conserve battery life
             | (https://machinelearning.apple.com/research/hey-siri), and
             | they may argue having good battery life is essential to
             | their brand.
        
         | colechristensen wrote:
         | My next step at least is various models running at home that i
         | can talk to remotely with my phone. Join a discord channel and
         | talk to my bots. My bots respond or do things back in the
         | channel or with whatever else they have access to.
        
       | jelled wrote:
       | I expect that locally installed models will quickly resemble how
       | we think about drivers today. In the short term you'll install a
       | notes app that requires Mixtral-8x7B so you pull that in through
       | ollama. Down the line all operating systems will ship with
       | various models preinstalled.
        
       | rvz wrote:
       | As predicted in: [0] [1]
       | 
       | [0] https://news.ycombinator.com/item?id=35711680
       | 
       | [1] https://news.ycombinator.com/item?id=35581777
        
       | conjecTech wrote:
       | I've noticed that despite Apple's stated on-device aspirations,
       | they've put very little incremental die space toward the neural
       | engine in the most recent generations. I think there's a number
       | of possible explanations for the disconnect. Curious if anyone
       | with a better perspective knows which are true.
       | 
       | 1) They think the existing hardware is powerful enough.
       | 
       | 2) The utilization of the ANE doesn't justify increased
       | resources.
       | 
       | 3) They plan to re-generalize AI compute through things like
       | vector operations.
       | 
       | 4) (the most pessimistic) They are saving large increases for
       | future releases when they need to force upgrades.
       | 
       | I could see the math for on-device just not working for things
       | like massive LLMs. The amount of silicon you'd need to make it
       | possible would be large, and the frequency of use low. The same
       | silicon for one person's phone could likely support dozens if it
       | were in a datacenter instead.
        
         | TillE wrote:
         | I'm not clear on what precisely the Neural Engine can do, but
         | Apple has put plenty of resources into improving their GPUs,
         | which can definitely do compute well (see eg, MacWhisper).
        
           | conjecTech wrote:
           | Yeah, it could be that they plan to do #3 through use of
           | their chunky GPUs. There is definitely a lot of focus on the
           | whole APU concept. From look at some of the dies and what
           | projects like ggml have been able to do through mixed
           | computation, I'm guessing the neural engine is ~3-5x the
           | throughput/area vs GPU for the kinds of operations its doing.
           | 
           | Trying to map neural nets onto graphics seems reasonable, but
           | personally I would bet on neuralizing graphics to be the
           | better strategy long term.
        
         | ip26 wrote:
         | There wasn't a "killer app" for on-die AI accelerators until
         | roughly this March. Until about nine months ago, the only
         | widely-discussed ideas were re-aiming pupils to make your video
         | call feel more natural, and applying background filters.
         | 
         | So today's on-die AI accelerators were, as far as I can tell, a
         | cautious bet. It turns out caution was warranted, because LLM's
         | need large amounts of memory bandwidth and capacity far more
         | than they need a specialized neural compute engine.
        
       | simonw wrote:
       | If you haven't tried it yet, and you have a modern enough phone,
       | I strongly recommend trying out Mistral 7B on your mobile device.
       | 
       | I've been running it on an iPhone 15 using this app:
       | https://llm.mlc.ai/#ios - App Store link:
       | https://apps.apple.com/us/app/mlc-chat/id6448482937
       | 
       | The same team have an Android version too which I've not yet
       | tried myself: https://llm.mlc.ai/#android
       | 
       | On my iPhone it works, and provides very decent performance. The
       | only catch is that it needs pretty much the entire phone's memory
       | to work - so if you switch out of it to another app and back
       | again it resets the state and has to load the model from scratch.
        
         | ComputerGuru wrote:
         | > so if you switch out of it to another app and back again it
         | resets the state and has to load the model from scratch.
         | 
         | This has been my experience anyway ever since upgrading to iOS
         | 16 (and now 17) anyway. Everything is always paging, even on
         | the latest hardware.
        
           | doug_durham wrote:
           | There is no paging in iOS. At least paging in the Linux
           | sense. Your apps memory will get evicted when you switch to
           | another app if the system needs it.
        
             | dcdc123 wrote:
             | It is up to each app to store their state and resume. One
             | of the most annoying apps for not doing this well is Google
             | Maps. Holy crap I hate when I search something and get a
             | route set and tab out to check something and com eback to a
             | fresh Google Maps like 5 seconds later. They will fix it
             | for a while then break it again several times a year. Once
             | you start navigating it seems pretty reliable but until you
             | hit start it is a shitshow.
        
               | kergonath wrote:
               | The Google Maps app is atrocious on iOS. So is their
               | website.
        
               | behnamoh wrote:
               | I just switched to Apple Maps and have been pretty happy
               | with it. Uninstalled Google Maps. Telling me that in 2010
               | would make me laugh at you.
        
               | dcdc123 wrote:
               | Apple Maps is terrible at discovery which is what I am
               | doing with Google Maps 90% of the time.
        
             | ghostpepper wrote:
             | do you happen to have a link where I can read more about
             | this?
        
         | mi_lk wrote:
         | FYI that app is 3.3 GB on iOS, not your typical app size
        
           | simonw wrote:
           | I'd love to know why it's that big. I'm pretty sure it
           | doesn't include the actual LLM model files - the first time
           | you use a model you have to click to download it, and those
           | are likely even larger than 3.3GB.
        
         | oidar wrote:
         | I've been playing with it on the recent iPhone Pro, and it is
         | pretty fast. The mistral 7B is very wordy though and is prone
         | to making up stuff. It's fun to talk to, but it's not very
         | useful currently.
         | 
         | Based on this demonstration of a local LLM, with optimizations,
         | it definitely the way of the future on apple devices at least.
        
         | keep_reading wrote:
         | Apple has released a paper about running LLMs efficiently with
         | low memory
         | 
         | https://arxiv.org/pdf/2312.11514.pdf
        
         | thefourthchime wrote:
         | Mistral is amazing for its size.
        
       | yellow_postit wrote:
       | Apple letting Siri stagnate for so many years is such a shame.
       | 
       | I do wonder what foundation model Apple will eventually use.
       | 
       | Both training their own as well as licensing a 3P one seem at
       | odds with their claimed data privacy stance.
       | 
       | My bet though is a licensed 3p model that's fine tuned locally
       | with on device data.
        
         | freedomben wrote:
         | Don't have a link right now, but I read earlier they have a
         | model developed internally already called colloquially "Apple
         | GPT" that (supposedly) rivals OpenAI in parameter size, etc.
         | 
         | They'll release a hardware co-processor similar to Google's
         | when they're ready to ship this.
        
       | kraig911 wrote:
       | I think it's a great idea but... local will always not be as good
       | as the an AI connected to the cloud. I find myself going from
       | conversations I left off on google/chatgpt/notes between my
       | desktop, phone. Would that it could work offline great but then
       | what of the divergent thoughts brought up in our mutual thinking?
        
       | TillE wrote:
       | Apple's biggest problem here is that they've been consistently
       | stingy with RAM, with only the latest iPhone 15 Pro featuring
       | 8GB. If they want to make a really gutsy move for on-device AI,
       | they could up that to 16GB at the very least.
       | 
       | That eats into Tim's precious margins, but it seems like a price
       | worth paying if they want to do exciting, useful things with this
       | tech.
        
       | w10-1 wrote:
       | It's crazy to say Apple is late to AI or missing the boat.
       | 
       | Yes, OpenAI gets a lot of attention to the extent LLM's could
       | displace search - except search is based on ads from links and
       | selling data, which are even less promising for AI.
       | 
       | "Big" AI now is stuck with organization-defining cloud bills for
       | training and huge players are scrambling to push software into
       | hardware. They are barely getting started on the integration.
       | 
       | OpenAI has been talking about AGI as it lines up commercial
       | partners across the globe and up and down the stack. But AGI is
       | no more realistic than crypto taking over from central banks.
       | It's possible in theory, but not in fact.
       | 
       | Meanwhile, Apple has had neural processors on devices for 4+
       | years; AI features have figured in every subsequent marketing
       | campaign. And VisionOS augmented reality provides a whole new
       | domain for AI utility, explicitly targeted not just for play, but
       | also for work, as remote work becomes the rule instead of the
       | exception. Plus, Apple is the only safety- and privacy- enforcing
       | ecosystem in existence.
        
         | ErneX wrote:
         | Yes, 1st neural engine was in the A11 Bionic, 2017.
        
         | parineum wrote:
         | > OpenAI has been talking about AGI as it lines up commercial
         | partners across the globe and up and down the stack. But AGI is
         | no more realistic than crypto taking over from central banks.
         | It's possible in theory, but not in fact.
         | 
         | At some level it feels like the pearl clutching about GPT4 as
         | it relates to AGI is astroturfed.
        
           | jakderrida wrote:
           | I'm inclined to agree with you on that to an extent. When I
           | hear Sam Altman talking about it like it's inevitable and he
           | carries the weight of the world on his shoulders, just seems
           | a little cringey. Don't get me wrong. I like the guy, but it
           | is cringey.
        
         | __loam wrote:
         | I still think LLMs displacing search will have hugely negative
         | consequences for the heath of the internet. Yes, a lot of
         | Google results are seo garbage now, but there's value in
         | interacting directly with the source of a piece of information.
         | LLMs alienate people from those sources, and that threatens the
         | structure of the information economy on the internet. Why share
         | anything if OpenAI is just going to train on your data and cut
         | you out of the transaction?
        
           | jakderrida wrote:
           | > but there's value in interacting directly with the source
           | of a piece of information.
           | 
           | Well, who said they're just gonna abandon it altogether. I'm
           | not going to. Those of us that make use of all the google
           | operators sure as hell ain't abandoning it. If anything, it
           | will provide augmented results where appropriate and explicit
           | results where appropriate. Otherwise, we'll switch to a
           | competitor that does.
        
         | alpaca128 wrote:
         | Yes, my Apple devices are the first where I felt AI was
         | integrated for something actually useful. I can just select and
         | copy text in pictures as if it was the most normal thing. I can
         | search pictures based on their content, again including visible
         | text. And it runs locally. Too bad it only seems to work on the
         | latin alphabet. With greek letters, chinese symbols etc I get
         | nothing at best.
        
           | wharvle wrote:
           | > Yes, my Apple devices are the first where I felt AI was
           | integrated for something actually useful. I can just select
           | and copy text in pictures as if it was the most normal thing.
           | I can search pictures based on their content, again including
           | visible text
           | 
           | Apple's so good at pushing out these kind of "little"
           | features that immediately feel like a kind of _bare minimum_
           | a system should do. Having to use a system where these two
           | things weren't automatic, quick, and practically flawless
           | would feel like taking a big step back.
           | 
           | See also: Live Photos. I wouldn't even consider any kind of
           | phone or camera without that feature, now that I've
           | experienced it.
           | 
           | [edit] context of how good it is: some days back there was a
           | link here that went to an image with a bunch of text on it. I
           | followed the link, read it, and did my usual compulsive-text-
           | selection thing I do when reading on the web.
           | 
           | I didn't realize I'd been looking at an image, rather than
           | html, until I came back to the thread and read complaints
           | about it. Double checked, yep, it was an image.
        
           | GaggiX wrote:
           | I mean I have these features too on my Xiaomi phone and it
           | does work with other alphabets and chinese characters
           | locally.
        
           | Scubabear68 wrote:
           | I was really surprised and happy to be able to take a picture
           | of my son's hand written Christmas list and the phone made
           | the text selectable with very few errors.
        
           | reaperducer wrote:
           | _Too bad it only seems to work on the latin alphabet. With
           | greek letters, chinese symbols etc I get nothing at best._
           | 
           | Works with Japanese on my MacBook Pro. I use it quite
           | frequently.
        
         | troupo wrote:
         | > It's crazy to say Apple is late to AI or missing the boat.
         | 
         | Apple is always "late to X". Their approach to all things ML so
         | far has always been "can we run this on device". And very few,
         | if any, of the latest LLMs can be run on the device reliaby,
         | and without sacrificing device resources.
         | 
         | > OpenAI has been talking about AGI
         | 
         | Of course they have. Precisely because they "line up commercial
         | partners"
         | 
         | ---
         | 
         |  _Edit: previous version said, confusingly, "AI is always late
         | to X". Also fixed typos and cleaned up sentences_
        
         | MaKey wrote:
         | > Plus, Apple is the only safety- and privacy- enforcing
         | ecosystem in existence.
         | 
         | You mean the company that wanted to ship CSAM detection for
         | iCloud uploads with iOS, which was only abandoned because of
         | the huge backlash? Or the company that's implementing privacy
         | features which neatly don't affect its own apps, effectively
         | boosting their ad platform? [1]
         | 
         | [1]: https://adguard.com/en/blog/apple-tracking-ads-
         | business.html
        
       | ssijak wrote:
       | Can they just give an option "I'm OK with my data going to Apple
       | cloud for better LLM inference and better Photo library search,
       | etc". I'm all for local inference, but until it gets good enough
       | on limited devices, I would like to share you my data, please.
        
         | Kon-Peki wrote:
         | Keep in mind that Apple is charging you money for the hardware
         | that you are using for this inference. You are paying for the
         | electricity that it uses, and you have paid for the storage
         | space and memory.
         | 
         | Who pays (and _how_ ) if Apple starts doing this for you, on
         | their equipment?
         | 
         | Would you pay $5/mo for it?
        
       | notaustinpowers wrote:
       | I'm sure it's part of their ML with the newer iOS but this does
       | make me happy to think of the AI running on the hardware.
       | Yesterday when driving I got a message and usually it would just
       | say "Jane says 'blah blah' and sent a photo." But last night Siri
       | in CarPlay said "Jane says 'blah blah' and sent a photo of a dog
       | in a blanket on a couch". And when I checked the photo they sent
       | it was exactly that.
       | 
       | If it's E2E encrypted, I don't think it'd be able to identify
       | that during transit, so it had to have been locally done.
       | 
       | Thought it was interesting, and a nice QoL update. Excited to see
       | how this can grow if they introduce more things to be processed
       | with AI on the hardware.
        
         | bayindirh wrote:
         | I love how Apple obsessively does things on the edge, on the
         | device itself rather than on the cloud.
         | 
         | Sometimes my lowly iPhone X gets super hot during charging.
         | Next morning I either get the facial recognition updated, a new
         | memory generated or something else has happened. It's nice,
         | confidence evoking even. The features you like are run on your
         | device, with your data not leaving your device, on its free
         | time.
         | 
         | This is the AI I like. Personal, confined and private.
        
           | chatmasta wrote:
           | They're in the best possible position to innovate on the
           | literal edge, because they've got a fleet of billions of
           | devices that mostly auto-update. As hardware advances, Apple
           | will have the clearest path to deploying truly personal and
           | locally hosted AI.
        
         | kergonath wrote:
         | > If it's E2E encrypted, I don't think it'd be able to identify
         | that during transit, so it had to have been locally done.
         | 
         | Yes. Image description has been done on-device for quite a
         | while now, and it is great. You can look for pictures using a
         | description, and you can look for any text in any picture as
         | well. It works very, very well.
        
           | ryanwaggoner wrote:
           | Yeah, but I agree with the parent; I think something has
           | changed recently with Siri. The other day I asked it some
           | questions as a joke that I knew would just say "here's what I
           | found on the web", but it answered them. It made me wonder
           | what they're changing under the hood.
        
           | sn0wf1re wrote:
           | Yeah, apparently well enough to run a web service off of
           | iPhone image recognition!
           | 
           | https://findthatmeme.com/blog/2023/01/08/image-stacks-and-
           | ip...
        
         | baz00 wrote:
         | Apple photos is wonderful for this. I can search for "cows" and
         | I get photos of cows. All on device in the middle of a mountain
         | with no internet 50 miles in any direction.
        
       | dang wrote:
       | Recent and related:
       | 
       |  _LLM in a Flash: Efficient LLM Inference with Limited Memory_ -
       | https://news.ycombinator.com/item?id=38704982 - Dec 2023 (51
       | comments)
        
       | madrox wrote:
       | This is the possibility I'm most excited for in the AI space.
       | Cloud models are great, but a lot will be unlocked once we get
       | decent models running on consumer devices. If I never have to
       | look at a web page again and use an LLM as my browser, it will be
       | a game changer for the industry.
        
         | tech_ken wrote:
         | I think that no matter what you're still going to need to
         | connect over WAN to interact with remote services. GPTs are
         | super capable but I don't see them generating arbitrary
         | webpages without at least fetching some basic resources from
         | another machine somewhere. Like if a restaurant updates its
         | menu, that data needs to make its way to the model in one way
         | or another before it can generate their webpage. At the very
         | least I think the "website" might change from an HTML page to
         | an engineered prompt that gets served to you and which
         | generates the page locally. But even in this case it's not
         | clear to me why I'd do that generation on my phone rather than
         | have the server gen the page for me and send the finished
         | product.
        
       | sovietmudkipz wrote:
       | I kinda wonder if it's a derisking maneuver. If the AI runs on
       | the client hardware you gain certain characteristics: the compute
       | happens on the client hardware (saving apple $$) and if a bad
       | actor succeeds at jail breaking from the AI controls then only
       | the client hardware is owned (instead of a multi-tenant apple
       | server).
       | 
       | Just my naive thoughts
        
       | thih9 wrote:
       | I guess this looks moat shaped.
       | 
       | And seems to be aligned well with Apple too: their focus on
       | privacy, their hardware performance gains, their ai assistant
       | products.
       | 
       | Apple is good at launching new product categories, ahead of
       | competition. The apple watch had no competition for the first
       | couple of years after launch, ipads are a class of its own, I
       | wonder if they're planning a similar move with ai too.
        
       | mcintyre1994 wrote:
       | > While Microsoft and Google have largely focused on delivering
       | chatbots and other generative AI services over the Internet from
       | their vast cloud computing platforms, Apple's research suggests
       | that it will instead focus on AI that can run directly on an
       | iPhone.
       | 
       | Isn't there supposed to be a version of Gemini specifically for
       | Pixel phones? It seems like Google see this as a USP for their
       | hardware as well. I wonder how companies like Samsung are
       | thinking about it.
        
       | cmarschner wrote:
       | The metal support for pytorch is a real sh*t show at this point -
       | all APIs are either legacy or unfinished. Also, app memory is
       | about 3GB even for the largest phones. It really still feels that
       | it's all in its infancies.
        
       | zmmmmm wrote:
       | Well of course they do, they have 0% of the cloud hardware market
       | and a gigantic percentage of the phone market. They like money,
       | after all.
        
       | samstave wrote:
       | Personally, I believe this is the future. You cant have really
       | cool (tm) space suits with all the telemetry, comms, life
       | support, situ, interfaces, etc... etc... etc... without on-bio AI
       | hardware.
       | 
       | obviously its augmented by any Galaxync.
        
       | mark_l_watson wrote:
       | Apple's CoreML libraries have existed for several years,
       | including a BERT model that is useful. I have been expecting
       | Apple to bring more advanced Transformer models than BERT to iOS
       | and iPadOS.
       | 
       | My iPad Pro has 16G RAM and I run a 13B model on it.
       | 
       | Things are moving fast. About 6 weeks ago I bought a 32G Mac Mini
       | and the models that I can now run 6 weeks later are a huge
       | improvement (e.g., mistral, solar, mixtral:8x7b-instruct, Phi,
       | Yi, etc.) and for local RAG applications and various experiments
       | I feel like I could live with what I have now for a while. My
       | expectations for the future are high!
        
       | KaiserPro wrote:
       | I think the world and it's dog want to run AI on consumer
       | hardware instead of in the cloud. Its sure as shit cheaper to do
       | (if you exclude the RnD costs.)
        
       | thorum wrote:
       | Apart from the advantages for privacy and Siri, imagine if the OS
       | included a shared system AI that any app could query. That would
       | unlock so many new features that are either impossible or
       | impractical today.
        
       | smrtinsert wrote:
       | Seems like a sort of win/win. Great for pushing new devices
       | regularly and also good for reliability of the AI based
       | functionality as well as privacy.
        
       | ardel95 wrote:
       | I think Apple is fully capable of getting there too. They have a
       | high enough price point, and a loyal enough customer base, that
       | even if getting there required additional hardware or shipping
       | with more RAM by default, they could make it work.
        
       | 1-6 wrote:
       | Nothing is better than inferencing at the edge. For example,
       | Apple now processes Siri locally on the Series 9 and Ultra 2
       | watches. That's remarkable given the small size and power of
       | these devices.
        
       | 1-6 wrote:
       | I thought more people already knew about this. It's quite evident
       | when you see their product announcements.
        
       | contrarian1234 wrote:
       | I don't really get why Apple doesn't go all-in on privacy.
       | 
       | They probably make peanuts on advertisement and it's not their
       | core business at all. They have such good PR potential. They can
       | easily paint Google/Meta as blood sucking data vampires that read
       | all your emails and look at all your dick pics. They could paint
       | Android as a data hoovering OS that only works for the benefit of
       | advertisers.
       | 
       | Meanwhile Apple.. The good guys.. Encrypt all your data, don't
       | inject JS on every web page , protect you from ads with Safari,
       | and do compute on-device
        
         | troupo wrote:
         | That is almost literaly how they have been advertising
         | themselves for the past almost 10 years.
        
       | paxys wrote:
       | Gemini Nano is on it way. Various other 3B-7B models are getting
       | small and efficient enough to run on consumer hardware. It has
       | been clear for a while that the "edge" is the new frontier for
       | LLMs.
        
       | J_Shelby_J wrote:
       | That's nice.
       | 
       | Anyways, i had fun building a Linux server to stick an nvidia GPU
       | in. I'm sure it's going to be pretty common over the next few
       | years as devs realize they can't just do everything with a
       | MacBook. (Or maybe apple will add support for 3rd party GPUs like
       | they had with the intel MacBooks)
       | 
       | This was my take away from the apple AR headset; it doesn't
       | matter how nice the UX is if it lacks the compute to do anything
       | interesting... and AR/VR takes a LOT of compute.
        
       | carabiner wrote:
       | It's still amazing how much ML already does on iPhone. OCR in
       | photos, subject identification in photos work incredibly well,
       | whereas Android phones have to do it via cloud.
        
       | happytiger wrote:
       | I doubt it.
       | 
       | I'd imagine they want both.
        
       | olliej wrote:
       | "AI in the cloud" with user data requires users to give up any
       | pretense of privacy so it's amazing to me that people are ok with
       | this?
       | 
       | If you build "AI" into an OS, and that "AI" is on the cloud that
       | means you are requiring your users to upload their data in
       | plaintext.
        
       | brap wrote:
       | I wonder what companies like OpenAI will do if we see on-device
       | AI becoming the mainstream. The entire business model is based on
       | API access. Will they start selling offline models? I guess then
       | piracy becomes an issue...
        
         | BudaDude wrote:
         | OpenAI will always be in a different league. I know Apple is
         | making some amazing advancements with LLMs, but OpenAI
         | basically recreated Google Search as a chatbot.
         | 
         | I'm not saying Apple couldn't do the same thing, but the model
         | alone would take up so much space on the phone that it would
         | not be practical. I can see Apple making something closer to a
         | very fine tuned mistral 7B than GPT4.
         | 
         | But I'm sure in a year this comment will be wrong.
        
           | brap wrote:
           | I'm not saying on-device performance will ever be comparable
           | to cloud, it's probably a different use case and market at
           | this point.
           | 
           | But 1, it's still a profitable business to be in, so why
           | shouldn't OpenAI go there? And 2, more importantly, as
           | hardware and software become more efficient, in the future
           | on-device might be good enough to put up a fight and start
           | eating the cloud. Seems risky _not_ to go there.
        
       | summerlight wrote:
       | This got to be one of the major goals for the next gen generative
       | models. Those "big" models cannot be reliably used on critical
       | tasks thanks to various failure modes. Even for those non-
       | critical tasks, the latency is typically high enough to be
       | unusable for many potential applications. If it can be 99.9% done
       | in-device, this has a potential to drastically change the user
       | experience in many directions.
        
       ___________________________________________________________________
       (page generated 2023-12-21 23:01 UTC)