[HN Gopher] Apple Releases Open Weights Video Model
       ___________________________________________________________________
        
       Apple Releases Open Weights Video Model
        
       Author : vessenes
       Score  : 421 points
       Date   : 2025-12-02 05:10 UTC (17 hours ago)
        
 (HTM) web link (starflow-v.github.io)
 (TXT) w3m dump (starflow-v.github.io)
        
       | coolspot wrote:
       | > STARFlow-V is trained on 96 H100 GPUs using approximately 20
       | million videos.
       | 
       | They don't say for how long.
        
         | moondev wrote:
         | Apple intelligence: trained by Nvidia GPUs on Linux.
         | 
         | Do the examples in the repo run inference on Mac?
        
       | satvikpendem wrote:
       | Looks good. I wonder what use case Apple has in mind though, or I
       | suppose this is just what the researchers themselves were
       | interested in, perhaps due to the current zeitgeist. I'm not
       | really sure how it works at big tech companies with regards to
       | research, are there top down mandates?
        
         | ivape wrote:
         | To add things to videos you create with your phone. TikTok and
         | Insta will probably add this soon, but I suppose Apple is
         | trying to provide this feature on "some level". That means you
         | don't have to send your video through a social media platform
         | first to creatively edit it (the platforms being the few tools
         | that let you do generative video).
         | 
         | They should really buy Snapchat.
        
         | ozim wrote:
         | I guess Apple is big in video production and animation with
         | some ties via Pixar and Disney. Since Jobs started Pixar and it
         | all got tied up in myriad of different ways.
        
       | nothrowaways wrote:
       | Where do they get the video training data?
        
         | postalcoder wrote:
         | From the paper:
         | 
         | > Datasets. We construct a diverse and high-quality collection
         | of video datasets to train STARFlow-V. Specifically, we
         | leverage the high-quality subset of Panda (Chen et al., 2024b)
         | mixed with an in-house stock video dataset, with a total number
         | of 70M text-video pairs.
        
           | justinclift wrote:
           | > in-house stock video dataset
           | 
           | Wonder if "iCloud backups" would be counted as "stock video"
           | there? ;)
        
             | fragmede wrote:
             | Turn on advanced data protection so they don't train on
             | yours.
        
               | givinguflac wrote:
               | That has nothing to do with it, and Apple wouldn't train
               | on user content, they're not Google. If they ever did
               | there would be opt in at best. There's a reason they're
               | walking and observing, not running and trying to be the
               | forefront cloud AI leader, like some others.
        
             | anon7000 wrote:
             | I have to delete as many videos as humanly possible before
             | backing up to avoid blowing through my iCloud storage quota
             | so I guess I'm safe
        
             | whywhywhywhy wrote:
             | More likely AppleTV shows
        
       | devinprater wrote:
       | Apple has a video understanding model too. I can't wait to find
       | out what accessibility stuff they'll do with the models. As a
       | blind person, AI has changed my life.
        
         | densh wrote:
         | > As a blind person, AI has changed my life.
         | 
         | Something one doesn't see in news headlines. Happy to see this
         | comment.
        
           | badmonster wrote:
           | What other accessibility features do you wish existed in
           | video AI models? Real-time vs post-processing?
        
             | devinprater wrote:
             | Mainly realtime processing. I play video games, and would
             | love to play something like Legend of Zelda and just have
             | the AI going, then ask it "read the menu options as I move
             | between them," and it would speak each menu option as the
             | cursor moves to it. Or when navigating a 3D environment,
             | ask it to describe the surroundings, then ask it to tell me
             | how to get to a place or object, then it guide me to it.
             | That could be useful in real-world scenarios too.
        
               | kulahan wrote:
               | Weird question, but have you ever tried text adventures?
               | It seems like it's inherently the ideal option, if you
               | can get your screen reader going.
        
           | tippa123 wrote:
           | +1 and I would be curious to read and learn more about it.
        
             | joedevon wrote:
             | If you want to see more on this topic, check out (google)
             | the podcast I co-host called Accessibility and Gen. AI.
        
               | moss_dog wrote:
               | Thanks for the recommendation, just downloaded a few
               | episodes!;
        
               | tippa123 wrote:
               | Honestly, that's such a great example of how to share
               | what you do on the interwebs. Right timing, helpful and
               | on topic. Since I've listened to several episodes of the
               | podcast, I can confirm it definitely delivers.
        
             | swores wrote:
             | A blind comedian / TV personality in the UK has just done a
             | TV show on this subject - I haven't seen it, but here's a
             | recent article about it: https://www.theguardian.com/tv-
             | and-radio/2025/nov/23/chris-m...
        
               | latexr wrote:
               | Hilariously, he beat the other teams in the "Say What You
               | See" round (yes, really) of last year's Big fat Quiz. No
               | AI involved.
               | 
               | https://youtu.be/i5NvNXz2TSE?t=4732
        
               | swores wrote:
               | Haha that's great!
               | 
               | I'm not a fan of his (nothing against him, just not my
               | cup of tea when it comes to comedy and mostly not been
               | interested in other stuff he's done), but the few times I
               | have seen him as a guest on shows it's been clear that
               | he's a generally clever person.
        
               | asplake wrote:
               | I remembered he was once a techie, and Wikipedia confirms
               | that he (Chris McCausland) has a BSc Honours in Software
               | Engineering.
               | 
               | https://en.wikipedia.org/wiki/Chris_McCausland
        
               | lukecarr wrote:
               | Chris McCausland is great. A fair bit of his material
               | _does_ reference his visual impairment, but it's
               | genuinely witty and sharp, and it never feels like he's
               | leaning on it for laughs/relying on sympathy.
               | 
               | He did a great skit with Lee Mack at the BAFTAs 2022[0],
               | riffing on the autocue the speakers use for announcing
               | awards.
               | 
               | [0]: https://www.youtube.com/watch?v=CLhy0Zq95HU
        
             | chrisweekly wrote:
             | Same! @devinprater, have you written about your
             | experiences? You have an eager audience...
        
           | fguerraz wrote:
           | > Something one doesn't see in news headlines.
           | 
           | I hope this wasn't a terrible pun
        
             | densh wrote:
             | No pun intended but it's indeed an unfortunate choice of
             | words on my part.
        
               | 47282847 wrote:
               | My blind friends have gotten used to it and hear/receive
               | it not as a literal "see" any more. They would not feel
               | offended by your usage.
        
             | devinprater wrote:
             | Nah, best pun ever!
        
           | kkylin wrote:
           | Like many others, I too would very much like to hear about
           | this.
           | 
           | I taught our entry-level calculus course a few years ago and
           | had two blind students in the class. The technology available
           | for supporting them was abysmal then -- the toolchain for
           | typesetting math for screen readers was unreliable (and
           | anyway very slow), for braille was non-existent, and
           | translating figures into braille involved sending material
           | out to a vendor and waiting weeks. I would love to hear how
           | we may better support our students in subjects like math,
           | chemistry, physics, etc, that depend so much on
           | visualization.
        
             | WillAdams wrote:
             | For a physical view on this see:
             | 
             | https://www.reddit.com/r/openscad/comments/1p6iv5y/christma
             | s...
             | 
             | The creator, https://www.reddit.com/user/Mrblindguardian/
             | has asked for help a few times in the past (I provided
             | feedback when I could), but hasn't needed to as often of
             | late, presumably due to using one or more LLMs.
        
           | Rover222 wrote:
           | `Something one doesn't see` - no pun intended
        
           | WarcrimeActual wrote:
           | I have to believe you used the word see twice ironically.
        
         | phyzix5761 wrote:
         | Can you share some ways AI has changed your life?
        
           | darkwater wrote:
           | I guess that auto-generated audio descriptions for (almost?)
           | any video you want is a very, very nice feature for a blind
           | person.
        
             | baq wrote:
             | guessing that being able to hear a description of what the
             | camera is seeing (basically a special case of a video) in
             | any circumstances is indeed life changing if you're
             | blind...? take a picture through the window and ask what's
             | the commotion? door closed outside that's normally open -
             | take a picture, tell me if there's a sign on it? etc.
        
             | tippa123 wrote:
             | My two cents, this seems like a case where it's better to
             | wait for the person's response instead of guessing.
        
               | darkwater wrote:
               | Fair enough. Anyway I wasn't trying to say what actually
               | changed GP's life, I was just expressing my opinion on
               | what video models could potentially bring as an
               | improvement to a blind person.
        
               | nkmnz wrote:
               | My two cents, this seems like a comment it should be up
               | to the OP to make instead of virtue signaling.
        
               | tippa123 wrote:
               | > Can you share some ways AI has changed your life?
               | 
               | A question directed to GP, directly asking about their
               | life and pointing this out is somehow virtue signalling,
               | OK.
        
               | throwup238 wrote:
               | You can safely assume that anyone who uses "virtue
               | signaling" unironically has nothing substantive to say.
        
               | SV_BubbleTime wrote:
               | >[People who call out performative bullshit should be
               | ignored because they're totally wrong and I totally mean
               | it.]
               | 
               | Maybe you're just being defensive? I'm sure he didn't
               | mean an attack at you personally.
        
               | throwup238 wrote:
               | It's presumptuous of you to assume I was offended.
               | 
               | Accusing someone of "virtue signaling" is itself virtue
               | signaling, just for a different in-group to use as a
               | thought terminating cliche. It has been for decades.
               | "Performative bullshit" is a great way to put it, just
               | not in the way you intended.
               | 
               | If the OP had a substantive point to make they would have
               | made it instead of using vague ad hominem that's so 2008
               | it could be the opening track on a Best of Glenn Beck
               | album (that's roughly when I remember "virtue signaling"
               | becoming a cliche).
        
               | fragmede wrote:
               | From the list of virtues, which one was this signaling?
               | 
               | https://www.virtuesforlife.com/virtues-list/
        
               | efs24 wrote:
               | I'd guess: Respect, consideration, authenticity,
               | fairness.
               | 
               | Or should I too perhaps wait for OP to respond.
        
               | SV_BubbleTime wrote:
               | That list needs updating. Lots of things became virtuous
               | in scenario. During Covid, fear was a virtue. You had to
               | prove how scared you were of it, all the masks you wore
               | because it made you "one of the good ones" to be fearful.
        
               | foobarian wrote:
               | Yall could have gotten a serviceable answer about this
               | topic out of ChatGPT. 2025 version of "let me google that
               | for you"
        
               | MangoToupe wrote:
               | ...you know, people can have opinions about the best way
               | to behave outside of self-aggrandizement, even if _your_
               | brain can 't grasp this concept.
        
               | nkmnz wrote:
               | _exactly_
        
           | gostsamo wrote:
           | Not the gp, but currently reading a web novel with a card
           | game where the author didn't include alt text in the card
           | images. I contacted them about it and they started, but in
           | the meantime ai was a big help. all kinds of other images on
           | the internet as well when they are significant to
           | understanding the surrounding text. better search experience
           | when Google, DDG, and the like make finding answers
           | difficult. I might use smart glasses for better outdoor
           | orientation, though a good solution might take some time.
           | phone camera plus ai is also situationally useful.
        
             | dzhiurgis wrote:
             | As a (web app) developer I never quite sure what to put in
             | alt. Figured you might have some advice here?
        
               | gostsamo wrote:
               | The question to ask is, what a sighted person learns
               | after looking at the image? The answer is the alt text.
               | E.g if the image is a floppy, maybe you communicate that
               | this is the save button. If it shows a cat sleeping on
               | the windowsill, the alt text is yep: "my cat looking cute
               | while sleeping on the windowsill".
        
               | michaelbuckbee wrote:
               | I really like how you framed this as the takeaway or
               | learning that needs to happen as what should be in the
               | alt and not a recitation of the image. Where I've often
               | had issues is more for things like business charts and
               | illustrations and less cute cat photos.
        
               | travisjungroth wrote:
               | It might be that you're not perfectly clear on what
               | exactly you're trying to convey with the image and why
               | it's there.
        
               | gostsamo wrote:
               | sorry, snark does not help with my desire to improve
               | accessibility in the wild.
        
               | hrimfaxi wrote:
               | What would you put for this? "Graph of All-Transactions
               | House Price Index for the United States 1975-2025"?
               | 
               | https://fred.stlouisfed.org/series/USSTHPI
        
               | wlesieutre wrote:
               | Charts are one I've wondered about, do I need to try to
               | describe the trend of the data, or provide several
               | conclusions that a person seeing the chart might draw?
               | 
               | Just saying "It's a chart" doesn't feel like it'd be
               | useful to someone who can't see the chart. But if the
               | other text on the page talks about the chart, then maybe
               | identifying it as the chart is enough?
        
               | gostsamo wrote:
               | It depends on the context. What do you want to say? How
               | much of it is said in the text? Can the content of the
               | image be inferred from the text part? Even in the best
               | scenario though, giving a summary of the image in the alt
               | text / caption could be immensely useful and include the
               | reader in your thought process.
        
               | embedding-shape wrote:
               | What are you trying to point out with your graph in
               | general? Write that basically. Usually graphs are added
               | for some purpose, and assuming it's not purposefully
               | misleading, verbalizing the purpose usually works well.
        
               | freedomben wrote:
               | I might be an unusual case, but when I present
               | graphs/charts it's not usually because I'm trying to
               | point something out. It's usually a "here's some data,
               | what conclusions do you draw from this?" and hopefully a
               | discussion will follow. Example from recently: "Here is a
               | recent survey of adults in the US and their religious
               | identification, church attendance levels, self-reported
               | "spirituality" level, etc. What do you think is
               | happening?"
               | 
               | Would love to hear a good example of alt text for
               | something like that where the data isn't necessarily
               | clear and I also don't want to do any interpreting of the
               | data lest I influence the person's opinion.
        
               | gostsamo wrote:
               | An image is the wrong way to convey something like that
               | to a blind person. As written in one of my other
               | comments, give the data in a table format or a custom
               | widget that could be explored.
        
               | embedding-shape wrote:
               | > and hopefully a discussion will follow.
               | 
               | Yeah, I think I misunderstood the context. I
               | understood/assumed it to be for an article/post you're
               | writing, where you have something you want to say in
               | general/some point of what you're writing. But based on
               | what you wrote now, it seems to be more about how to
               | caption an image you're sending to a blind person in a
               | conversation/discussion of some sort.
               | 
               | I guess at that point it'd be easier for them if you just
               | share the data itself, rather than anything generated by
               | the data, especially if there is nothing you want to
               | point out.
        
               | alwillis wrote:
               | https://www.w3.org/WAI/tutorials/images/ including how
               | write alt text for charts.
        
               | asadotzler wrote:
               | a plaintext table with the actual data
        
               | isoprophlex wrote:
               | "A meaningless image of a chart, from which nevertheless
               | emanates a feeling of stonks going up"
        
               | gostsamo wrote:
               | The logic stays the same though the answer is longer and
               | not always easy. Just saying "business chart" is totally
               | useless. You can make a choice on what to focus and say
               | "a chart of the stock for the last five years with
               | constant improvement and a clear increase by 17 percent
               | in 2022" (if it is a simple point that you are trying to
               | make) or you can provide an html table with the
               | datapoints if there is data that the user needs to
               | explore on their own.
        
               | nextaccountic wrote:
               | but the table exists outside the alt text, right? i don't
               | know a mechanism to say "this html table represents the
               | contents of this image" , in a way that screen readers
               | and other accessibility technologies take advantage of
        
               | gostsamo wrote:
               | The figure tag has both image and caption tags that link
               | them. As far as I remember, some content could be marked
               | as screen reader only if you don't want for the table to
               | be visible to the rest of the users.
               | 
               | Additionally, recently I've been a participant in
               | accessibility studies where charts, diagrams and the like
               | have been structured to be easier to explore with a sr.
               | Those needed js to work and some of them looked custom,
               | but they are also an alternative way to layer data.
        
               | alwillis wrote:
               | Accessible info graphics [1]
               | 
               | [1]: https://web.archive.org/web/20130922065731/http://ww
               | w.last-c...
        
               | askew wrote:
               | One way to frame it is: "how would I describe this image
               | to somebody sat next to me?"
        
               | embedding-shape wrote:
               | Important to add for blind people: "... assuming they
               | never seen anything and visual metaphors won't work"
               | 
               | The amount of times I've seem captions that wouldn't make
               | sense for people who never been able to see is
               | staggering, I don't think most people realize how visual
               | our typical language usage is.
        
               | shagie wrote:
               | I'm gonna flip this around... have you tried pasting the
               | image (and the relevant paragraph of text) and asking
               | ChatGPT (or another LLM) to generate the alt text for the
               | image and see what it produces?
               | 
               | For example... https://chatgpt.com/share/692f1578-2bcc-80
               | 11-ac8f-a57f2ab6a7...
        
               | alwillis wrote:
               | > I'm gonna flip this around... have you tried pasting
               | the image (and the relevant paragraph of text) and asking
               | ChatGPT (or another LLM) to generate the alt text for the
               | image and see what it produces?
               | 
               | There's a great app by an indie developer that uses ML to
               | identify objects in images. Totally scriptable via
               | JavaScript, shell script and AppleScript. macOS only.
               | 
               | Could be 10, 100 or 1,000 images [1].
               | 
               | [1]: https://flyingmeat.com/retrobatch/
        
               | alwillis wrote:
               | > As a (web app) developer I never quite sure what to put
               | in alt.
               | 
               | Are you making these five mistakes when writing alt text?
               | [1] Images tutorial [2] Alternative Text [3]
               | 
               | [1]: https://www.a11yproject.com/posts/are-you-making-
               | these-five-...
               | 
               | [2]: https://www.w3.org/WAI/tutorials/images/
               | 
               | [3]: https://webaim.org/techniques/alttext/
        
           | devinprater wrote:
           | Image descriptions. TalkBack on Android has it built in and
           | uses Gemini. VoiceOver still uses some older, less accurate,
           | and far less descriptive ML model, but we can share images to
           | Seeing AI or Be My Eyes and such and get a description.
           | 
           | Video descriptions, through PiccyBot, have made watching more
           | visual videos or videos where things happen that don't make
           | sense without visuals much easier. Of course, it'd be much
           | better if YouTube incorporated audio description through AI
           | the same way they do captions, but that may happen in a good
           | 2 years or so. I'm not holding my breath. Google as a whole
           | is hard to get accessibility out of more than the bare
           | minimum.
           | 
           | Looking up information like restaurant menus. Yes it can make
           | things up, but worst-case, the waiter says they don't have
           | that.
        
         | javcasas wrote:
         | Finally good news about the AI doing something good for the
         | people.
        
           | p1esk wrote:
           | I'm not blind and AI has been great for me too.
        
           | Workaccount2 wrote:
           | People need to understand that a lot of angst around AI comes
           | from AI enabling people to do things that they formally
           | needed to go through gatekeepers for. The angst is coming
           | from the gatekeepers.
           | 
           | AI has been a boon for me and my non-tech job. I can pump out
           | bespoke apps all day without having to get bent on
           | $5000/yr/usr engineering software packages. I have a website
           | for my side business that looks and functions professionally
           | and was done with a $20 monthly AI subscription instead of a
           | $2000 contractor.
        
             | MyFirstSass wrote:
             | I highly doubt "pumping out bespoke apps all day" is
             | possible yet besides 100% boilerplate, and when possible
             | then no good for any other purpose than enshittifiying the
             | web, and at that point not profitable because everyone can
             | do it.
             | 
             | I use AI daily as a senior coder for search and docs, and
             | when used for prototyping you still need to be a senior
             | coder to go from say 60% boilerplate to 100% finished
             | app/site/whatever unless it's incredibly simple.
        
               | Workaccount2 wrote:
               | Often the problem with tech people is they think software
               | only exists for tech or for being sold to others from
               | tech.
               | 
               | Nothing I do is in the tech industry. It's all
               | manufacturing and all the software is for in-house
               | processes.
               | 
               | Believe it or not, software is useful to everyone and no
               | longer needs to originate from someone who only knows
               | software.
        
               | MyFirstSass wrote:
               | I'm saying you can't do what you're saying without
               | knowing code at the moment.
               | 
               | You didn't give any examples of the valuable bespoke apps
               | that you are creating by the hour.
               | 
               | I simply don't believe you, and the arrogant salesy tone
               | doesn't help.
        
               | Workaccount2 wrote:
               | LLMs can pretty reliably write 5-7k LOC.
               | 
               | If your needs fit in a program that size, you are pretty
               | much good to go.
               | 
               | It will not rewrite PCB_CAD 2025, but it will happily
               | create a PCB hole alignment and conversion app,
               | eliminated the need for the full PCB_CAD software if all
               | you need is that one toolset from it.
               | 
               | Very, _very_ , few pieces of software need to be full
               | package enterprise productivity suites. If you just make
               | photos black and white and resize them, you don't need
               | Photoshop to do it. Or even ms paint. Any LLM will make a
               | simple free program with no ads to do it. Average people
               | generally do very simple dumb stuff with the expensive
               | software they buy.
        
               | vjvjvjvjghv wrote:
               | This is the same as the discussion about using Excel.
               | Excel has its limitations, but it has enabled millions of
               | people to do pretty sophisticated stuff without the help
               | of "professionals". Most of the stuff us tech people do
               | is also basically some repetitive boilerplate. We just
               | like to make things more complex than they need to be. I
               | am always a little baffled why seemingly every little
               | CRUD site that has at most 100 users needs to be run on
               | Kubernetes with several microservices, CI/CD pipelines,
               | and whatever.
               | 
               | As far as enshittification goes, this was happening long
               | before AI. It probably started with SEO and just kept
               | going from there.
        
               | almosthere wrote:
               | The reality is too, that even if "what is acceptable" has
               | not yet caught up to that guy working at Atlassian,
               | polishing off a new field in Jira, people are using AI +
               | Excel to manage their tasks EXACTLY the way their head
               | works, not the way Jira works.
               | 
               | Yet we fail to see AI as a good thing but just as a jobs
               | destroyer. Are we "better than" the people that used to
               | fill toothpaste tubes manually until a machine was
               | invented to replace them? They were just as mad when they
               | got the pink slip.
        
               | vjvjvjvjghv wrote:
               | I have told people that us techies have proudly killed
               | the jobs of millions of people and we were arrogant about
               | it. Now we are mad that it's our turn. Feels almost like
               | justice :-)
        
               | alwillis wrote:
               | > I use AI daily as a senior coder for search and docs,
               | and when used for prototyping you still need to be a
               | senior coder to go from say 60% boilerplate to 100%
               | finished app/site/whatever unless it's incredibly simple.
               | 
               | I know you would like to believe that, but with the tools
               | available NOW, that's not necessarily the case. For
               | example, by using the Playwright or Chrome DevTools MCPs,
               | models can see the web app are it's being created and
               | it's pretty easy to prompt them to fix something they can
               | see.
               | 
               | These models know the current frameworks and coding
               | practices but they do need some guidance; they're not
               | mindreaders.
        
               | MyFirstSass wrote:
               | I still don't believe that. Again yes a boilerplate
               | calculator or recipe app probably, but anything advanced
               | real world with latency issues, scaling, race conditions,
               | css quirks, design weirdness, optimisation - in other
               | words the things that actually require domain knowledge i
               | still don't get much help with, even with Claude Code,
               | pointers yes but they completely fumble actual production
               | code in real world scenarios.
               | 
               | Again it's the last 5% that takes 95% of the time, and
               | those 5% i haven't seen fixed with Claude or Gemini,
               | because it's essentially quirks, browser errors, race
               | conditions, visual alignment, etc etc. All stuff that
               | completely goes way above any LLM's head atm from what
               | i've seen.
               | 
               | They can definitely bullshit a 95% working app though,
               | but that's 95% from being done ;)
        
             | BeFlatXIII wrote:
             | AI is divine retribution for artists being really annoying
             | on Twitter.
        
         | GeekyBear wrote:
         | One cool feature they added for deaf parents a few years ago
         | was a notification when it detects a baby crying.
        
           | embedding-shape wrote:
           | Is that something you actually need AI for though? A device
           | with a sound sensor and something that shines/vibrate a
           | remote device when it detects sound above some threshold
           | would be cheaper, faster detection, more reliable, easier to
           | maintain, and more.
        
             | jfindper wrote:
             | > _Is that something you actually need AI for though?_
             | 
             | Need? Probably not. I bet it helps though (false positives,
             | etc.)
             | 
             | > _would be cheaper, faster detection, more reliable,
             | easier to maintain, and more._
             | 
             | Cheaper than the phone I already own? Easier to maintain
             | than the phone that I don't need to do maintenance on?
             | 
             | From a fun hacking perspective, a different sensor & device
             | is cool. But I don't think it's any of the things you
             | mentioned for the majority of people.
        
             | evilduck wrote:
             | But your solution costs money in addition to the phone they
             | already own for other purposes. And multiple things can
             | make loud noises in your environment besides babies;
             | differentiating between a police siren going by outside and
             | your baby crying is useful, especially if the baby slept
             | through the siren.
             | 
             | The same arguments were said for blind people and the
             | multitude of one-off devices that smartphones replaced, OCR
             | to TTS, color detection, object detection in photos/camera
             | feeds, detecting what denomination US bills are, analyzing
             | what's on screen semantically vs what was provided as
             | accessible text (if any was at all), etc. Sure, services
             | for the blind would come by and help arrange outfits for
             | people, and audiobook narrators or braille translator
             | services existed, and standalone devices to detect money
             | denominations were sold, but a phone can just do all of
             | that now for much cheaper.
             | 
             | All of these accessibility AI/ML features run on-device, so
             | the knee-jerk anti-AI crowd's chief complaints are mostly
             | baseless anyways. And for the blind and the deaf, carrying
             | all the potential extra devices with you everywhere is
             | burdensome. The smartphone is a minimal and common social
             | and physical burden.
        
             | doug_durham wrote:
             | You are talking about a device of smart phone complexity.
             | You need enough compute power to run a model that can
             | distinguish noises. You need a TCP/IP stack and a wireless
             | radio to communicate the information. At that point you
             | have a smart phone. A simple sound threshold device would
             | have too many false positives/negatives to be useful.
        
             | Aurornis wrote:
             | > more reliable
             | 
             | I've worked on some audio/video alert systems. Basic
             | threshold detectors produce a lot of false positives. It's
             | common for parents to put white noise machines in the room
             | to help the baby sleep. When you have a noise generating
             | machine in the same room, you need more sophisticated
             | detection.
             | 
             | False positives are the fastest way to frustrate users.
        
           | Damogran6 wrote:
           | I also got notification on my apple watch, while being away
           | from the house, that the homepod mini heard our fire alarm
           | going off.
           | 
           | A call home let us know that our son had set it off learning
           | to reverse-sear his steak.
        
             | brandonb wrote:
             | If the fire alarm didn't go off, you didn't sear hard
             | enough. :)
        
             | kstrauser wrote:
             | I live across the street from a fire station. Thank for you
             | for diligence, little HomePod Mini, but I'm turning your
             | notifications off now.
        
           | SatvikBeri wrote:
           | My wife is deaf, and we had one kid in 2023 and twins in
           | 2025. There's been a noticeable improvement baby cry
           | detection! In 2023, the best we could find was a specialized
           | device that cost over $1,000 and has all sorts of
           | flakiness/issues. Today, the built-in detection on her
           | (android) phone + watch is better than that device, and a lot
           | more convenient.
        
         | andy_ppp wrote:
         | I wonder if there's anything that can help blind people to
         | navigate the world more easily - I guess in the future AR
         | Glasses won't just be for the sighted but allow people without
         | vision to be helped considerably. It really is both amazing and
         | terrifying the future we're heading towards.
        
           | shagie wrote:
           | From a couple years ago...
           | 
           | https://www.microsoft.com/en-us/garage/wall-of-
           | fame/seeing-a...
           | 
           | https://youtu.be/R2mC-NUAmMk
           | 
           | https://youtu.be/DybczED-GKE
           | 
           | ... and that was 10 years ago. I'm curious for what it could
           | do now.
        
           | xnx wrote:
           | https://play.google.com/store/apps/details?id=com.google.and.
           | ..
        
           | asadotzler wrote:
           | AURA Vision for blind and low vision people has been doing
           | this for years. Be My Eyes has been doing this for years
           | without AI. Meta Ray-Bans can do this. There's nothing new
           | coming soon that hasn't already been available for a while,
           | only refinements.
        
         | whatsupdog wrote:
         | > As a blind person, AI has changed my life.
         | 
         | I know this is a low quality comment, but I'm genuinely happy
         | for you.
        
         | basilgohar wrote:
         | I'm only commenting because I absolutely love this thread. It's
         | an insight into something I think most of us are quite (I'm
         | going to say it...) blind to in our normal experiences with
         | daily life, and I find immense value in removing my ignorance
         | about such things.
        
         | robbomacrae wrote:
         | Hi Devin and other folks, I'm looking for software developers
         | who are blind or hard of sight as there is a tool I'm building
         | that I think might be of interest to them (it's free and open
         | source). If you or anyone you know is interested in trying it
         | please get in touch through my email.
        
       | camillomiller wrote:
       | Hopefully this will make into some useful feature in the
       | ecosystem and not contribute to having just more terrible slop.
       | Apple has saved itself from the destruction of quality and taste
       | that these model enabled, I hope it stays that way.
        
       | yegle wrote:
       | Looking at text to video examples
       | (https://starflow-v.github.io/#text-to-video) I'm not impressed.
       | Those gave me the feeling of the early Will Smith noodles videos.
       | 
       | Did I miss anything?
        
         | M4v3R wrote:
         | These are ~2 years behind state of the art from the looks of
         | it. Still cool that they're releasing anything that's open for
         | researchers to play with, but it's nothing groundbreaking.
        
           | Mashimo wrote:
           | But 7b is rather small no? Are other open weight video models
           | also this small? Can this run on a single consumer card?
        
             | Maxious wrote:
             | Wan 2.2: "This generation was run on an RTX 3060 (12 GB
             | VRAM) and took 900 seconds to complete at 840 x 420
             | resolution, producing 81 frames."
             | https://www.nextdiffusion.ai/tutorials/how-to-run-
             | wan22-imag...
        
             | dragonwriter wrote:
             | > But 7b is rather small no?
             | 
             | Sure, its smallish.
             | 
             | > Are other open weight video models also this small?
             | 
             | Apples models are weights-available not open weights, and
             | yes, WAN 2.1, as well as the 14B models, also has 1.3B
             | models; WAN 2.2, as well as the 14B models, also has a 5B
             | model (the WAN 2.2 VAE used by Starflow-V is _specifically_
             | the one used with the 5B model.) and because the WAN models
             | are largely actually open weights models (Apache 2.0
             | licensed) there are lots of downstream open-licensed
             | derivatives.
             | 
             | > Can this run on a single consumer card?
             | 
             | Modern model runtimes like ComfyUI can run models that do
             | not fit in VRAM on a single consumer card by swapping model
             | layers between RAM and VRAM as needed; models bigger than
             | this can run on single consumer cards.
        
             | jjfoooo4 wrote:
             | My guess is that they will lean towards smaller models, and
             | try to provide the best experience for running inference on
             | device
        
           | tomthe wrote:
           | No, it is not as good as Veo, but better than Grok, I would
           | say. Definitely better than what was available 2 years ago.
           | And it is only a 7B research model!
        
           | tdesilva wrote:
           | The interesting part is they chose to go with a normalizing
           | flow approach, rather than the industry standard diffusion
           | model approach. Not sure why they chose this direction as I
           | haven't read the paper yet.
        
         | manmal wrote:
         | I wanted to write exactly the same thing, this reminded me of
         | the Will Smith noodles. The juice glass keeps filling up after
         | the liquid stopped pouring in.
        
         | jfoster wrote:
         | I think you need to go back and rewatch Will Smith eating
         | spaghetti. These examples are far from perfect and probably not
         | the best model right now, but they're far better than you're
         | giving credit for.
         | 
         | As far as I know, this might be the most advanced text-to-video
         | model that has been released? I'm not sure whether the license
         | will qualify as open enough in everyone's eyes, though.
        
       | RobotToaster wrote:
       | The license[0] seems quite restrictive, limiting it's use to non
       | commercial research. It doesn't meet the open source definition
       | so it's more appropriate to call it weights available.
       | 
       | [0]https://github.com/apple/ml-starflow/blob/main/LICENSE_MODEL
        
         | limagnolia wrote:
         | They haven't even released the weights yet...
         | 
         | As for the license, happily, Model Weights are the product of
         | machine output and not creative works, so not copyrightable
         | under US law. Might depend on where you are from, but I would
         | have no problem using Model Weights however I want to and
         | ignoring pointless licenses.
        
           | loufe wrote:
           | The weights for the text-->image model are already on
           | Huggingface, FWIW.
        
       | mdrzn wrote:
       | "VAE: WAN2.2-VAE" so it's just a Wan2.2 edit, compressed to 7B.
        
         | BoredPositron wrote:
         | They used the VAE of WAN like many other models do. For image
         | models you see a lot of them using the flux VAE. Which is
         | perfectly fine, they are released as apache2 and save you time
         | to focus on your transformers architecture...
        
         | kouteiheika wrote:
         | This doesn't necessarily mean that it's Wan2.2. People often
         | don't train their own VAEs and just reuse an existing one,
         | because a VAE isn't really what's doing the image generation
         | part.
         | 
         | A little bit more background for those who don't know what a
         | VAE is (I'm simplifying here, so bear with me): it's
         | essentially a model which turns raw RGB images into a something
         | called a "latent space". You can think of it as a fancy "color"
         | space, but on steroids.
         | 
         | There are two main reasons for this: one is to make the model
         | which does the actual useful work more computationally
         | efficient. VAEs usually downscale the spatial dimensions of the
         | images they ingest, so your model now instead of having to
         | process a 1024x1024 image needs to work on only a 256x256
         | image. (However they often do increase the number of channels
         | to compensate, but I digress.)
         | 
         | The other reason is that, unlike raw RGB space, the latent
         | space is actually a higher level representation of the image.
         | 
         | Training a VAE isn't the most interesting part of image models,
         | and while it is tricky, it's done entirely in an unsupervised
         | manner. You give the VAE an RGB image, have it convert it to
         | latent space, then have it convert it back to RGB, you take a
         | diff between the input RGB image and the output RGB image, and
         | that's the signal you use when training them (in reality it's a
         | little more complex, but, again, I'm simplifying here to make
         | the explanation more clear). So it makes sense to reuse them,
         | and concentrate on the actually interesting parts of an image
         | generation model.
        
           | mdrzn wrote:
           | Thanks for the explanation!
        
           | sroussey wrote:
           | Since you seem to know way more than I on the subject, can
           | you explain the importance of video generation that is not
           | diffusion based?
        
         | dragonwriter wrote:
         | > "VAE: WAN2.2-VAE" so it's just a Wan2.2 edit
         | 
         | No, using the WAN 2.2 VAE does not mean it is a WAN 2.2 edit.
         | 
         | > compressed to 7B.
         | 
         | No, if it was an edit of the WAN model that uses the 2.2 VAE,
         | it would be expanded to 7B, not compressed (the 14B models of
         | WAN 2.2 use the WAN _2.1_ VAE, the WAN _2.2_ VAE is used by the
         | 5B WAN 2.2 model.)
        
       | pulse7 wrote:
       | <joke> GGUF when? </joke>
        
       | LoganDark wrote:
       | > Model Release Timeline: Pretrained checkpoints will be released
       | soon. Please check back or watch this repository for updates.
       | 
       | > The checkpoint files are not included in this repository due to
       | size constraints.
       | 
       | So it's not actually open weights yet. Maybe eventually once they
       | actually release the weights it will be. "Soon"
        
       | vessenes wrote:
       | From the paper, this is a research model aimed at dealing with
       | the runaway error common in diffusion video models - the latent
       | space is (proposed to be) causal and therefore it should have
       | better coherence.
       | 
       | For a 7b model the results look pretty good! If Apple gets a
       | model out here that is competitive with wan or even veo I believe
       | in my heart it will have been trained with images of the finest
       | taste.
        
       | dymk wrote:
       | Title is wrong, model isn't released yet. Title also doesn't
       | appear in the link - why the editorializing?
        
       | giancarlostoro wrote:
       | I was upset the page didnt have videos immediately available,
       | then I realized I have to click on some of the tabs. One red flag
       | on their github is the license looks to be their own flavor of
       | MIT (though much closer to MS-PL).
        
       | andersa wrote:
       | The number of video models that are worse than Wan 2.2 and can
       | safely be ignored has increased by 1.
        
         | embedding-shape wrote:
         | To be fair, the sizes aren't comparable, and for the variant
         | that is comparable, the results aren't that much worse.
        
           | dragonwriter wrote:
           | The samples (and this may or may not be completely fair,
           | either set could be more cherry picked than the other, It
           | would be interesting to see a side-by-side comparison with
           | comparable prompts) seem significantly worse than what I've
           | seen from WAN 2.1 1.3B, which is both fron the previous WAN
           | version and is smaller, proportionally, compared to Apple's
           | 7B than that model itself is compared to the 28B combination
           | of the high and low noise 14B WAN 2.2 models that are
           | typically used together.
           | 
           | But also, Starflow-V is a research model with a substandard
           | text encoder, it doesn't have to be competitive as-is to be
           | an interesting spur for further research on the new
           | architecture it presents. (Though it would be nice if it had
           | some _aspect_ where it offered a clear improvement.)
        
         | wolttam wrote:
         | This doesn't look like it was intended to compete. The research
         | appears interesting
        
       | gorgoiler wrote:
       | It's not really relevant to this release specifically but it irks
       | me that, in general, an "open weights model" is like an "open
       | source machine code" version of Microsoft Windows. Yes, I guess I
       | have open access to view the thing I am about to execute!
       | 
       | This Apple license is click wrap MIT with the rights, at least,
       | to modify and redistribute the model itself. I suppose I should
       | be grateful for that much openness, at least.
        
         | advisedwang wrote:
         | Great analogy.
         | 
         | To extend the analogy, "closed source machine code" would be
         | like conventional SaaS. There's an argument that shipping me a
         | binary I can freely use is at least better than only providing
         | SaaS.
        
         | limagnolia wrote:
         | I think you are looking at the code license, not the model
         | license.
        
           | Aloisius wrote:
           | No, it's the model license. There's a second license for the
           | code.
           | 
           | Of course, model weights almost certainly are not
           | copyrightable so the license isn't enforceable anyway, at
           | least in the US.
           | 
           | The EU and the UK are a different matter since they have sui
           | generis database rights which seemingly allows individuals to
           | own /dev/random.
        
         | satvikpendem wrote:
         | > _Yes, I guess I have open access to view the thing I am about
         | to execute!_
         | 
         | Better to execute locally than to execute remotely where you
         | can't change or modify any part of the model though. Open
         | weights at least mean you can retrain or distill it, which is
         | not analogous to a compiled executable that you can't
         | (generally) modify.
        
       | cubefox wrote:
       | Interesting that this is an autoregressive ("causal") model
       | rather than a diffusion model.
        
       | Invictus0 wrote:
       | Apple's got to stop running their AI group like a university lab.
       | Get some actual products going that we can all use--you know,
       | with a proper fucking web UI and a backend.
        
         | Jtsummers wrote:
         | Personally, I'm happy that Apple is spending the time and money
         | on research. We have products that already do what this model
         | does, the next step is to make it either more efficient or
         | better (closer to the prompt, more realistic, higher quality
         | output). That requires research, not more products.
        
       | summerlight wrote:
       | This looks interesting. This project has some novelty as a
       | research and actually delivered a promising PoC but as a product
       | it implies that its training was severely constrained by
       | computing resources, which correlates well with the report that
       | their CFO overruled CEO's decision on ML infra investment.
       | 
       | JG's recent departure and follow up massive reorg to get rid of
       | AI, rumors on Tim's upcoming step down in early 2026... All of
       | these signals indicate that those non-ML folks have won corporate
       | politics to reduce the in-house AI efforts.
       | 
       | I suppose this was a part of serious efforts to deliver in-house
       | models but the directional changes on AI strategy made them to
       | give up. What a shame... At least the approach itself seem
       | interesting and hope others to take a look and use it for
       | building something useful.
        
       ___________________________________________________________________
       (page generated 2025-12-02 23:01 UTC)