[HN Gopher] FastVLM: Efficient vision encoding for vision langua...
___________________________________________________________________
FastVLM: Efficient vision encoding for vision language models
Author : nhod
Score : 349 points
Date : 2025-05-13 01:16 UTC (21 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| BryanLegend wrote:
| Seems like the main thing holding these new minds back is being
| able to see well. Breakthroughs like this will fix that.
| efnx wrote:
| That and the ability to hold on to knowledge.
| static_void wrote:
| ... or say they don't know.
| kamranjon wrote:
| Apple out here playing 5d chess, installing neural cores in their
| hardware and writing crazy efficient vision models to run on em.
| Cool stuff.
| wmf wrote:
| I thought they turned sycophancy off...
| kamranjon wrote:
| Awe yes I admit, I think the new Apple hardware is real cool
| vFunct wrote:
| Can it fill a wine glass to the rim?
| mkl wrote:
| It's for interpreting images, not generating them.
| turnsout wrote:
| Apple has gotten a slow start in the LLM world, but they have the
| only long term strategy that makes sense. They're going to
| dominate the 2030s.
| boroboro4 wrote:
| What exactly the strategy is?
| generalizations wrote:
| They can run locally on-device: a win for cost, latency and
| privacy (privacy is pragmatic: it means you can use all the
| user's data as context without qualms). There's a reason
| Microsoft tried so hard to push for the neural processors a
| year or two ago. Avoiding the cost of the datacenter while
| offering good-enough inference (emphasis on good) is a
| massive win.
| turnsout wrote:
| Yes, thank you; this is the strategy I was referring to. It
| will take some time for the models and chips to get there,
| but on-device inference will have massive advantages for
| privacy, speed and cost. Plus it will drive demand for
| hardware--at first, iPhones, but soon AirPods and glasses.
| xnx wrote:
| Google already has some of the best on device models
| (Gemma) and chips (Tensor).
| AceJohnny2 wrote:
| > _and chips (Tensor)_
|
| Is there actually any hard data out there comparing the
| NPU on the Google Tensor G4 vs the Apple A18? I wasn't
| able to quickly find anything concrete.
|
| I mean Apple has been shipping mobile NPUs for longer
| than Google (Apple: since A11 in 2017, Google: since
| 2021), and are built on (ostensibly) a smaller silicon
| node that Google's (G4: Samsung SF4P vs A18: TSMC N3E).
| However, the G4 appears to have more RAM bandwidth (68.26
| GB/s vs 60 GB/s on A18).
| lern_too_spel wrote:
| Google has been shipping custom NPUs since the Pixel 4 in
| 2019. Prior to that, Google phones just used off the
| shelf SOCs from Qualcomm, with 2018's Pixel 3 using the
| NPU in the Snapdragon 845. Android first shipped NNAPI in
| Android 8.1 in 2017, with acceleration on various mobile
| GPUs and DSPs, including the Pixel Visual Core on the
| Pixel 2. Google has shipped more on-device models so far,
| but neither company has a moat for on-device inference.
|
| https://blog.google/products/pixel/pixel-visual-core-
| image-p...
| turnsout wrote:
| Unfortunately for them, Google doesn't make devices that
| people want to buy
| lern_too_spel wrote:
| Other Android phone vendors do, and they have the same
| strategy, sitting on top of Qualcomm NPUs.
| weikju wrote:
| They are running data centers and offloading some things to
| chatGPT though, not just running on device.
|
| In fact there's no clear indication when Apple Intelligence
| is running on-device or in their Private Cloud Compute.
| jfarina wrote:
| What strategy is that?
| ryanmcgarvey wrote:
| I presume they mean that distribution is king and they make
| all the devices.
| karn97 wrote:
| Average delusional hner who really doesn't know what they are
| talking about
| turnsout wrote:
| Enlighten us, wise one
| insane_dreamer wrote:
| As the father of a young child whose optic nerves are highly
| deteriorated (compression) and is expected to lose his sight
| (when exactly is unknown; based on original projections he should
| be blind by now, but an experimental treatment run in a trial at
| the NIH (KEEP FUNDING SCIENCE) has stabilized his sight), I'm
| overjoyed with the advances being made in VLMs. I can now
| envision a future where even if he loses his sight he'll be able
| to interact with the world around him, go to college, have a
| fulfilling career (he loves science and engineering, and is
| talented for his young age), etc.
| lynx97 wrote:
| I grew up in the 80s as a 100% blind child. Technology was by
| far not as advanced as today. Computers were just coming up
| when I was around 12. I learnt to type on a oldschool
| typewriter, and I also learnt to write braille with a pretty
| heavy full-metal embossing device. OCR was still quite bad.
| When I switched to what you call high scooll, I used a laptop
| with integrated Braille display to follow classes. Used good
| old DOS as OS and Word 5.5 as my "notepad". Except for PC
| Lingua for Latin, I basically had no tools specialized for
| learning. A electronic notepad and my brain was all I had to
| follow school. And I still made it. I have a great job I love,
| my own appartment, a sweet girlfriend and I am basically
| completely independent. To a point where I had to forcefully
| send away my mother since her continued attempts to "help" me
| were basically detrimental to my own development. I can not
| emphasis how important it is how you deal with it as a parent.
| Since parents are indeed the biggest hinderence to development,
| we have a saying around here amongst disabled people:
| "additional disability due to parental overprotection"
| (Zusatzbehinderung Eltern). Please take a moment to understand
| what this means, without feeling personally attacked. Its
| important. Your child _can_ leave home around 18, just like
| every other kid. I did. Don 't slow that process down
| artificially. The more this is prolonged, the harder it gets
| for the individual to actually obtain independence.
|
| I am telling you this because I read between the lines that you
| believe current technology is a reason for you to be hopeful.
| Sure, it should be. But never forget, your child can do much
| more then you as a sighted person will ever be able to
| understand. Don't let them drown in your own misery. Let them
| discover what they can do. You will be surprised what they come
| up with. And dont fall for Gear Acquision Syndrome. Sure, tools
| are nice, and they do get better, which is also nice. I _LOVE_
| vision models, to stay on topic somehow. However, I still leave
| my house with only a cane and my phone in my pocket. I do
| occasionally ask Siri "Where am I" to get an address if I
| happen to have forgotten where I am exactly, currently. But at
| the end of the day, my cane is what shows me the way. Most tech
| is hype, plain old hearing and your sense of touch gets you
| much farther then you might think.
|
| Wish you all the best for your own journey, and the development
| of your child.
| wiz21c wrote:
| I should read a comment like yours every morning.
| topato wrote:
| Wow, this really adds an amazing perspective to the entire
| (frequently touted) concept of Visual Language Models somehow
| "saving" blind people from their old life; In the past, a
| blind person desperately needed caretakers, otherwise the
| blind person will bumble around their home, end up mistaking
| the sink for the toilet, accidentally turn on their stove
| thinking it's the thermostat, until they died after mistaking
| bleach for milk and cat litter for cereal....
|
| BUT NOW... THE FUTURE IS HERE.... an all-knowing god-like
| cell phone can tell these poor miserable individuals what the
| objects in their own homes are! No more tragic Mr. Magoo-ian
| accidents!
|
| But thank you for posting this; It certainly enlightened me!
| I'll admit, all these AI solutions
| exe34 wrote:
| > I'll admit, all these AI solutions
|
| They got to him.
| insane_dreamer wrote:
| Thanks. I appreciate your insight.
| liamwire wrote:
| It feels like this is the required level of speed-up needed re.
| time-to-first-token to make continuous vision useful for on-
| device applications like an assistant that can see and take
| action on your screen, ala the original Apple Intelligence demos.
| It's very impressive seeing the app in the repo and I'm excited
| to build it tonight and play around.
| nine_k wrote:
| With that, a really helpful aid for blind people can be made,
| running just on their phone, fed from a camera in their
| eyeglasses. Somebody who could not move around without an
| assistant could become autonomous in daily life.
| jdiff wrote:
| It might be useful for telling Cream of Chicken from Cream of
| Mushroom, but for locomotion I can't see this adding anything
| over existing strategies people use to get around sans sight.
|
| "There's a tree. There's a tree. There's a tree. There's a
| number of pedestrians. There's a tree. There's a sign." does
| not strike me as useful feedback for getting around.
| nine_k wrote:
| Consider a city. It's full of signs and inscriptions, traffic
| lights, and other key interaction elements. Consider a store.
| It has shelves with stuff, again with inscriptions, price
| tags, etc.
|
| "Pavement. Row of stores to the left. Joe's Grocery Store.
| Doors. Door handle. A shelf with bakery. A shelf with canned
| goods. A shelf with bottles. Coke bottle. Large Pepsi bottle.
| Apple juice bottle. Passageway. Checkout. Payment terminal.
| Door. Door handle. Pavement. ..."
| jdiff wrote:
| None of that gives me any useful spatial sense of _where_.
| "Payment terminal." Okay. Where is it? Left? Left where?
| How much left? How far?
|
| The only truly useful bits I see in your stream of text is,
| again, "Cream of Mushroom" vs "Cream of Chicken." I am
| actively holding something, so I know where it is, but need
| to differentiate it from printed detail.
| adamsiem wrote:
| Anyone using vision to parse screenshots? QVQ was too slow. Will
| give this a shot.
| abrichr wrote:
| You might be interested in
| https://github.com/OpenAdaptAI/OpenAdapt
| logankeenan wrote:
| I used molmo to parse screenshots in order to detect locations
| of UI elements. See the repo below. I think Omni parser from
| Microsoft would also work well.
|
| https://github.com/logankeenan/george
|
| https://github.com/microsoft/OmniParser
| Aeroi wrote:
| I built/building a realtime voice+vision app called Sen, its
| currently live in beta and streams frames over webrtc. It's fast
| and smart, but Im super curious to see how these models do as we
| get closer to the metal. I can see these running on-device in the
| future with super fast ttfb.
| keyle wrote:
| Do you have a write up of the tech stack and setup? Or willing
| to give the gist here?
|
| I'd like to make a private Qwen or similar for my kids to
| prompt with a button and voice control. It doesn't need
| vision... Although eventually that'd be very cool.
|
| Siri just sucks.
|
| We might not be there yet...
| Aeroi wrote:
| yeah i made a post on here, but the algo sent it to the gulag
| abyss.
|
| https://news.ycombinator.com/item?id=43926673
| keyle wrote:
| That's a good product site but it doesn't help me in
| anyway...
| Aeroi wrote:
| I also ran across an interesting robot toy demo today that
| had voice built in. it was whimsical and seemed like it was
| aimed towards primary education and kids. Someone here might
| know the name.
| stavros wrote:
| You can use Ollama or LM Studio, both in API mode, to return
| the responses. I believe they offer audio support, but I'm
| not entirely sure.
|
| However, if you're looking for instruction following (like an
| agent), I've tried to implement my own agent and have lost
| faith. Even GPT-4.1 will regularly gaslight me that no, it
| _definitely_ ran the tool call to add the event to my
| calendar, when it just didn 't. I can't get any more
| adherence out of it.
| cloudking wrote:
| Check out https://livekit.io/
| tomp wrote:
| We're definitely there, there's just no "ready-made" apps
| yet. But the technology is possible, go to e.g. vapi.ai to
| test it.
| nikolayasdf123 wrote:
| 2GB for 0.5B smallest model. it does not make sense for each app
| to download this. apple must have plans to pre-load these models
| on os level and expose SDK for all apps to call these models
| locally. exciting times!
|
| opened issue for them to confirm this:
| https://github.com/apple/ml-fastvlm/issues/7
| cube2222 wrote:
| That's what they suggested about LLMs at last year's WWDC iirc.
| The core models are provided by the OS, while apps bring LORAs
| to fine-tune them / bring custom heads for them.
| babl-yc wrote:
| You could probably get away with f16 or even quantize to int8
| and have a much smaller model, but your point stands. Users
| won't be thrilled to download a 500MB model for an app either.
| nikolayasdf123 wrote:
| haha latest Uber build for iOS 18 is 500MB... without LLM
| models <face-palm/>
| ukuina wrote:
| What are they doing in there? Is it mostly visual assets?
| bastawhiz wrote:
| If I was going to guess, I'd get there's a ton of third
| party code for things like payment method SDKs. Every
| local payment method around the world is going to have
| its own package that you need to import, and you can't
| just load in new executable code on the fly after the app
| is installed.
| victorbjorklund wrote:
| You can actually do over the air updates of apps (how
| easy it is depends on what you wrote your app in) and not
| adding a new feature (like just adding an additional
| payment provider) would not require an update on the App
| Store.
| nikolayasdf123 wrote:
| wouldn't you want to create payment gateway and abstract
| away logic such that client is agnostic of payment
| processes and backend confirms internal payment process
| into external specific ones? (in worst case redirect to
| other apps with universal links or webview)
| philipkglass wrote:
| A lot of geography-specific scenarios are compiled into
| the app, including regional payment SDKs. There's a great
| comment from a former Uber engineer explaining it here:
|
| https://news.ycombinator.com/item?id=25376346
| nikolayasdf123 wrote:
| I think they using vector graphics and vector animations
| (say Rive). Rive takes order of 10s of KBs. Lottie is
| much larger to 100s of KBs. Even then you would need 5000
| animations to reach 500MB, unlikely!
|
| raster graphics and videos are likely not included in
| build
|
| probably some unused code (libraries) got into it, it can
| grow quite large
|
| or maybe some ML models?
| HanClinto wrote:
| I think that there is fantastic potential in having open-
| weight, OS-standard foundation models.
|
| Especially if the API gives opportunity for app developers to
| load their custom LoRa fine-tunings onto OS-standard foundation
| models at runtime, then you can (ideally) have the best of both
| worlds -- fine-tuned app-specific models with reasonable app
| sizes.
| HappMacDonald wrote:
| I haven't seen much done with loras for LLMs though, only for
| diffusion image gen models. From what I've heard it sounds
| like a difference in benefit due to architecture.
| gessha wrote:
| My guess is that they won't confirm it unless it's a big
| presentation. WWDC maybe?
| nikolayasdf123 wrote:
| google and cloud LLM providers must be biting their teeth now!
| haha
| nikolayasdf123 wrote:
| distributing this heavy compute and moving it close to device
| where 1. source of data happens; 2. decision and output about the
| result of analysis is done; is way to go. super low latency, no
| network traffic, privacy, less overhead in cloud. this is amazing
| porphyra wrote:
| It seems that the future of robotics is VLA models. Even Tesla
| FSD is an end-to-end VLA model. Efficient vision encoding will be
| a huge part of making robots safe and responsive.
| lynx97 wrote:
| I wonder, can I convert/run this with llama.cpp? It being LLaVA
| based seems promising.
| vessenes wrote:
| Um wow. The on-device realtime videos are worth a watch, and
| compelling. Looking forward to this being deployed and widely
| adopted. Getting much faster time to first token opens up a ton
| of features and usability benefits.
| buyucu wrote:
| where is my gguf?
| simianparrot wrote:
| I have a feeling feeding tesseract the image every 1 second would
| be significantly faster and take far less space and processing
| power? Haven't tested it yet but given how fast tesseract is on
| large images, it wouldn't surprise me.
| regularfry wrote:
| If all you want is OCR, possibly.
| coredog64 wrote:
| If all you want is OCR of typewritten text.
|
| Tesseract is awful for handwriting.
| d3k wrote:
| Very nice! I wish they were more keen to contribute to AI/ML
| community an publish weights and model definition on HuggingFace.
| Funny enough I have just seen today a similar demo that is using
| a freely available VLM: https://github.com/ngxson/smolvlm-
| realtime-webcam
| tough wrote:
| SmolVLM is from huggingface team
|
| cool to see people doing stuff with smaller models
|
| https://huggingface.co/blog/smolvlm
|
| https://arxiv.org/abs/2504.05299
| labadal wrote:
| I'm absolutely thrilled that there is an effort to make models
| smaller and run with less resources instead of blindly throwing
| more resources at the problem and expecting it to get solved.
___________________________________________________________________
(page generated 2025-05-13 23:02 UTC)