[HN Gopher] Helix: A vision-language-action model for generalist...
       ___________________________________________________________________
        
       Helix: A vision-language-action model for generalist humanoid
       control
        
       Author : Philpax
       Score  : 217 points
       Date   : 2025-02-20 14:30 UTC (8 hours ago)
        
 (HTM) web link (www.figure.ai)
 (TXT) w3m dump (www.figure.ai)
        
       | kubb wrote:
       | Wow! This is something new.
        
       | sandis wrote:
       | YouTube link for the video (for whatever reason the video hosted
       | on their site kept buffering for me):
       | https://www.youtube.com/watch?v=Z3yQHYNXPws
        
       | bhouston wrote:
       | When doing robot control, how do you model in the control of the
       | robot? Do you have tool_use / function calling at the top level
       | model which then gets turned into motion control parameters via
       | inverse kinematic controllers?
       | 
       | What is the interface from the top level to the motors?
       | 
       | I feel it can not just be a neural network all the way down,
       | right?
        
         | Philpax wrote:
         | Have a look at the post - it explains how it works. There are
         | two models: a 7-9Hz 7B vision-language model, and a 200Hz 80M
         | visuomotor model. The former produces a latent vector, which is
         | then interpreted by the latter to drive the motors.
        
           | NitpickLawyer wrote:
           | > a 7-9Hz 7B vision-language model, and a 200Hz 80M
           | visuomotor model.
           | 
           | huh. An interesting approach. I wonder if something like this
           | can be used for other things as well, like "computer use"
           | with the same concept of a "large" model handling the goals,
           | and a "small" model handling clicking and stuff, at much
           | higher rates, useful for games and things like that.
        
             | whatever1 wrote:
             | This is typical in real time applications. A supervisor
             | tries to guess in which region the system is currently and
             | then invokes the correct set of lower level algorithms.
        
       | traverseda wrote:
       | "The first time you've seen these objects" is a weird thing to
       | say. One presumes that this is already in their training set, and
       | that these models aren't storing a huge amount of data in their
       | context, so what does that even mean?
        
         | jayd16 wrote:
         | It probably gives them confidence that they can accurately see
         | a thing even though they don't know what that thing is.
         | 
         | I could also imagine a lot of safety around leaving things
         | outside of the current task alone so you might have to bend
         | over backwards to get new objects worked on.
        
           | thomastjeffery wrote:
           | There is no such thing as "thing" here.
           | 
           | These models are trained such that the given conditions (the
           | visual input and the text prompt) will be continued with a
           | desirable continuation (motor function over time).
           | 
           | The only dimension accuracy can apply to is desirability.
        
         | Symmetry wrote:
         | It's normal to have a training set and a validation set and I
         | interpreted that to mean that these items weren't in the
         | training set.
        
         | ygouzerh wrote:
         | So from what I understand it actually means that they were for
         | example never trained on a video of an apple. Maybe only on a
         | video of bread, pineapple, chocolate.
         | 
         | However, as it was trained using generic text data similarly to
         | a normal LLM, it knows how an apple is supposed to look like.
         | 
         | Similar than a kid that never saw a banana, but his parent
         | described it to him.
        
       | ziofill wrote:
       | There's nothing I want more than a robot that does house chores.
       | That's the real 10x multiplier for humans to do what they do
       | best.
        
         | dartos wrote:
         | Hopefully in the next decade we'll get there.
         | 
         | Vision+language multimodal models seem to solve some of the
         | hard problems.
        
         | siavosh wrote:
         | What do humans do best?
        
           | jayd16 wrote:
           | Everything everything else is worse at.
        
           | ziofill wrote:
           | I mean to use their time to pursue their passions and
           | interests, not cleaning up the kitchen or making the bed or
           | doing laundry...
        
             | hooverd wrote:
             | Given time to "pursue their passions and interests", most
             | people chose to turn their brain to soup on social media.
        
               | KolmogorovComp wrote:
               | but most people think they are better than most people.
        
           | ein0p wrote:
           | Browse Instagram, apparently.
        
         | abraxas wrote:
         | Yeah, except that future doesn't need us. By us I mean those of
         | us who don't have $1B to their name.
         | 
         | Do you really expect the oligarchs to put up with the
         | environmental degradation of 8 billion humans when they can
         | have a pristine planet to themselves with their whims served by
         | the AI and these robots?
         | 
         | I fully anticipate that when these things mature enough we'll
         | see an "accidental" pandemic sweep and kill off 90% of us. At
         | least 90%.
        
           | ewjt wrote:
           | Oligarchs would use the robots to kill people instead of a
           | pandemic. A virus carries too much risk of affecting the
           | original creators.
           | 
           | Fortunately, robotic capability like that basically becomes
           | the equivalent of Nuclear MAD.
           | 
           | Unfortunately, the virus approach probably looks fantastic to
           | extremist bad actors with visions of an afterlife.
        
           | ben_w wrote:
           | I'd expect Musk and Bezos to know about von Neumann
           | replicators; factories that make these robots staffed
           | entirely by these robots all the way to the mines digging
           | minerals out of the ground... rapid and literally exponential
           | growth until they hit whatever the limiting factor is, but
           | they've both got big orbial rockets now, so the limit isn't
           | necessarily 6e24 kg.
        
         | cess11 wrote:
         | To me this is such a weird wish. Why would you not want to care
         | for your home and the people living there? Why would you want
         | to have a slave taking these activities from you?
         | 
         | I'd rather have less waged labour and more time for chores with
         | the family.
        
         | mmh0000 wrote:
         | It's called a "house cleaner" and they only cost ~$150 (area
         | and all varies) bi-weekly. I'll shit a brick (and then have the
         | robot clean it up) if a robot is ever cheaper than ~$4000/yr.
        
           | vessenes wrote:
           | A robot will definitely cost less than $30/hr eventually. But
           | you'll be running it a lot more than a few hours every other
           | week.
        
           | ziofill wrote:
           | Yeah, but a robot will work 24/7, not 2h byweekly -_-
        
             | mclau156 wrote:
             | do you need a robot to work in your house 24/7?
        
               | gigel82 wrote:
               | Why not? If it's done with all the chores, I can have it
               | make some silly woodworking / art project for Etsy to
               | earn its keep, or just loan it out to neighbors.
        
               | ben_w wrote:
               | In the same way it's hard to earn money from using AI to
               | make art, I don't see Etsy projects made by affordable
               | domestic robots selling above cost.
        
               | ziofill wrote:
               | Well perhaps not at night, but otherwise there's always
               | something to clean, something to fix, something to cook,
               | take care of the yard.. heck I might need two robots ^^'
        
         | 01100011 wrote:
         | I'd pay $2k for something that folds my laundry reliably. It
         | doesn't need arms or legs, just like my dishwasher doesn't need
         | arms or legs. It just needs to let me dump in a load of clean
         | laundry and output stacks of neatly folded or hung clothing.
        
       | anentropic wrote:
       | Very impressive
       | 
       | Why make such sinister-looking robots though...?
        
         | jayd16 wrote:
         | With the way they move, they look like stoned teenagers
         | interning at a Bond villain factory. Not to knock the tech but
         | they're scary and silly at the same time.
        
         | esafak wrote:
         | Black was not the best color choice.
        
       | aerodog wrote:
       | Interesting timing - same day MSFT releases
       | https://microsoft.github.io/Magma/
        
         | _1 wrote:
         | Current discussion:
         | https://news.ycombinator.com/item?id=43110265
        
       | wwwtyro wrote:
       | Until we get robots with really good hands, something I'd love in
       | the interim is a system that uses _me_ as the hands. When it's
       | time to put groceries away, I don't want to have to think about
       | how to organize everything. Just figure out which grocery items I
       | have, what storage I have available, come up with an optimized
       | organization solution, then tell me where to put things, one at a
       | time. I'm cautiously optimistic this will be doable in the near
       | term with a combination of AR and AI.
        
         | camjw wrote:
         | Maybe I don't understand exactly what you're describing but why
         | would anyone pay for this? When I bring home the shopping I
         | just... chuck stuff in the cupboards. I already know where it
         | all goes. Maybe you can explain more?
        
           | bear141 wrote:
           | Maybe some people just assume there is a "best" or "optimal"
           | way to do everything and AI will tell us what that is. Some
           | things are just preference and I don't mind the tiny amount
           | of energy that goes into doing small things the way I like.
        
           | jayd16 wrote:
           | Maybe they're imagining more complex tasks like working on an
           | engine.
        
           | loudmax wrote:
           | One use case I imagine is skilled workmanship. For example,
           | putting on a pair of AR glasses and having the equivalent of
           | an experienced plumber telling me exactly where to look for
           | that leak and how to fix it. Or how to replace my brake pads
           | or install a new kitchen sink.
           | 
           | When I hire a plumber or a mechanic or an electrician, I'm
           | not just paying for muscle. Most of the value these
           | professionals bring is experience and understanding. If a
           | video-capable AI model is able to assume that experience,
           | then either I can do the job myself or hire some 20 year old
           | kid at roughly minimum wage. If capabilities like this come
           | about, it will be very disruptive, for better and for worse.
        
             | semi-extrinsic wrote:
             | This is called "watching YouTube tutorials". We've had it
             | for decades.
        
               | rolisz wrote:
               | But what if there's no YouTube tutorial for the exact AC
               | unit you have and it doesn't look like any of the videos
               | you checked out?
        
               | cess11 wrote:
               | Have you met people that seem to be able to fix almost
               | anything?
               | 
               | If you can't get a tutorial on your exact case you learn
               | about the problem domain and intuit from there. Usually
               | it works out if you're careful, unlike software.
        
               | semi-extrinsic wrote:
               | Then you are equally fucked as the AI will be, so no
               | difference.
               | 
               | Case in point, I remember about ten years ago our washing
               | machine started making noise from the drum bearing. Found
               | a Youtube tutorial for bearing replacement on the exact
               | same model, but 3 years older. Followed it just fine
               | until it was time to split the drum. Then it turned out
               | that in the newer units like mine, some rent-seeking MBA
               | fuckers had decided more profits could be had if they
               | _plastic welded shut the entire drum assembly_. Which was
               | then a $300 replacement part for a $400 machine.
               | 
               | An AI doesn't help with this type of shit. It can't know
               | the unknown.
        
               | deepGem wrote:
               | But once it knows it's pretty certain to become common
               | knowledge almost instantaneously. That's not possible
               | now. What you learn stays localised to you and may be
               | people 1 degree away from you that's it.
        
               | semi-extrinsic wrote:
               | How does that work? None of the current AI models can re-
               | train on the fly. How would the inference engine even
               | know if it's a case of new information that needs to be
               | fed back, or just a user that's not following
               | instructions correctly?
        
             | hulahoof wrote:
             | Sounds like what Hololens was designed to solve, more in
             | the AR space than AI though
        
           | __MatrixMan__ wrote:
           | It would be nice to be able to select a recipe and have it
           | populate your shopping list based on what is currently in
           | your cupboards. If you just chuck stuff in the cupboards then
           | you have to be home to know what they contain.
           | 
           | Or you could wear it while you cook and it could give you
           | nutrition information for whatever it is you cooked. Armed
           | with that it could make recommendations about what nutrients
           | you're likely deficient in based on your recent meals and
           | suggest recipes to remedy the gap--recipes based on what it
           | knows is already in the cupboard.
        
             | gopher_space wrote:
             | Maybe I'm showing my age, but isn't this a home ec class?
        
               | __MatrixMan__ wrote:
               | I took home ec in 2001. I learned to use a sewing
               | machine, it was great.
               | 
               | But none of the kitchen stuff we learned had anything to
               | do with ensuring that this week's shopping list ensures
               | that you'll get enough zinc next week, or the kind of
               | prep that uses the other half of yesterday's cauliflower
               | in tomorrow's dinner so that it doesn't go bad.
               | 
               | These aren't hard problems to solve if you've got time to
               | plan, but they are hard to solve if you are currently at
               | the grocery store and can't remember that you've got a
               | half a cauliflower that needs an associated recipe.
        
           | SoftTalker wrote:
           | Yes it's pretty amazing how so many people seem to
           | overcomplicate simple household tasks by introducing
           | unnecessary technology.
        
           | luma wrote:
           | > why would anyone pay for this?
           | 
           | Presumably, they won't as this is still a tech demo. One can
           | take this simple demonstration and think about some future
           | use cases that aren't too different. How far away is
           | something that'll do the dishes, cook a meal, or fold the
           | laundry, etc? That's a very different value prop, and one
           | that might attract a few buyers.
        
             | Philip-J-Fry wrote:
             | The person you're replying to is referring to the GP. The
             | GP asks for an AI that tells them where to put their
             | shopping. Why would anyone pay for THAT? Since we already
             | know where everything goes without needing an AI to tell
             | us. An AI isn't going to speed that up.
        
         | cactusplant7374 wrote:
         | Your solution sounds like the worst cognitive load for getting
         | home from the grocery store and wanting it all to be over.
        
         | lucianbr wrote:
         | You want to outsource thinking to a computer system and keep
         | manual labor? You do you, but I want the opposite. I want to
         | decide what goes where but have a robot actually put the stuff
         | there.
        
           | TeMPOraL wrote:
           | That's the problem, though - the computer is already better
           | at thinking than you, but we still don't know how to make it
           | good at arbitrary labor requiring a mix of precision and
           | power, something humans find natural.
           | 
           | In other words: I'm sorry, but that's how reality turned out.
           | Robots are better at thinking, humans better at laboring. Why
           | fight against nature?
           | 
           | (Just joking... I think.)
        
           | RedNifre wrote:
           | I think he means outsourcing everything eventually, but right
           | now, outsourcing the thought process is possible, while
           | outsourcing the manual labor is not.
        
         | lynx97 wrote:
         | A good AI fridge would be already a great starting point. With
         | a checkin procedure that makes sure to actually know whats in
         | the fridge. Complete with expiry tracking and recipe
         | suggestions based on personal preferences combined with product
         | expiry. I am totally unimpressed with almost everything I see
         | in home automation these days, but I'd immediately buy the AI
         | fridge if it really worked smoothly.
        
         | Philpax wrote:
         | Sounds like what's described in Manna:
         | https://marshallbrain.com/manna1
        
         | __MatrixMan__ wrote:
         | I imagine a something like a headlamp except it's a projector
         | and a camera so it can just light up where it wants you to pick
         | something up in one color or where it wants you to put it down
         | in another color. It can learn from what it sees of my hands
         | how the eventual robot should handle the space (e.g. not
         | putting heavy things on top of fragile things and such).
         | 
         | I'd totally use that to clean my garage so that later I can ask
         | it where the heck I put the thing or ask it if I already have
         | something before I buy one...
        
         | hooverd wrote:
         | You already have one: a brain.
        
         | RedNifre wrote:
         | I fully agree, building something like this is somewhere in my
         | back log.
         | 
         | I think the key point why this "reverse cyborg" idea is not as
         | dystopian as, say, being a worker drone in a large warehouse
         | where the AI does not let you go to the toilet is that the AI
         | is under your own control, so you decide on the high level goal
         | "sort the stuff away", the AI does the intermediate planning
         | and you do the execution.
         | 
         | We already have systems like that, every time you use you tell
         | your navi where you want to go, it plans the route and gives
         | you primitive commands like "on the next intersection, turn
         | right", so why not have those for cooking, doing the laundry,
         | etc.?
         | 
         | Heck, even a paper calendar is already kinda this, as in
         | separating the planning phase from the execution phase.
        
           | Jarwain wrote:
           | I'm quite slowly working on something like this, but for
           | time.
           | 
           | For "stuff" I think a bigger draw is having it so it can let
           | me know "hey you already have 3 of those spices at locations
           | x, y, and z, so don't get another" or "hey you won't be able
           | to fit that in your freezer"
        
         | malux85 wrote:
         | Yeah there's more to it than that. Do you want a can of beans
         | to be put in the utensil draw just because it would fit? If it
         | was done as you describe the placement of all of your items
         | would be almost random each time, the bot need to have
         | contextual memory and familiarity with your storage habits and
         | preferences.
         | 
         | This can be done of course, in your statement the phrase "just
         | figure out" is doing a lot more heavy lifting than you allude
         | to
        
         | sho_hn wrote:
         | Dunno, I would not want to bleed my mental faculties for doing
         | even simple planning work like this by outsourcing it to AI.
         | Reliance on crutches like this would seem like a pathway to
         | early-onset dementia.
        
           | meowkit wrote:
           | Already playing out, anecdotally to my experience.
           | 
           | Its similar to losing callouses on our hands if you don't
           | labor/go to the gym.
        
         | htrp wrote:
         | so the kiva-amazon model?
        
         | falcor84 wrote:
         | This is almost literally the first chapter in Marshall Brain's
         | "Manna" [0], being the first step towards world-controlling
         | AGI:
         | 
         | > Manna told employees what to do simply by talking to them.
         | Employees each put on a headset when they punched in. Manna had
         | a voice synthesizer, and with its synthesized voice Manna told
         | everyone exactly what to do through their headsets. Constantly.
         | Manna micro-managed minimum wage employees to create perfect
         | performance.
         | 
         | [0] https://marshallbrain.com/manna1
        
       | sottol wrote:
       | Imo, the Terminator movies would have been scarier if they moved
       | like these guys - slow, careful, deliberate and measured but
       | unstoppable. There's something uncanny about this.
        
         | megous wrote:
         | Unfortunately, there'll be no time travel to save us. That was
         | the lying part of the movie. Other stuff was true.
        
       | ramenlover wrote:
       | Why do they make "eye contact" after every hand off? Feels oddly
       | forced.
        
         | bear141 wrote:
         | This along with the writing style in the description is totally
         | forced anthropomorphizing. It's creepy.
        
           | jimbohn wrote:
           | Gotta hype up the investors somehow
        
         | GoatInGrey wrote:
         | Perhaps they're exchanging knowing looks on how stupid they
         | think the demo is. Solidarity between artificial brothers.
        
       | Symmetry wrote:
       | So, there's no way you can have fully actuated control of every
       | finger joint with just 35 degrees of freedom. Which is very
       | reasonable! Humans can't individually control each of our finger
       | joints either. But I'm curious how their hand setups work, which
       | parts are actuated and which are compliant. In the videos I'm not
       | seeing any in-hand manipulation other than just grasping,
       | releasing, and maintaining the orientation of the object relative
       | to the hand and I'm curious how much it can do / they plan to
       | have it be able to do. Do they have any plans to try to mimic
       | OpenAI's one handed rubics cube demo?
        
       | pr337h4m wrote:
       | Goal 2 has been achieved, at least as a proof of concept (and not
       | by OpenAI): https://openai.com/index/openai-technical-goals/
        
         | Symmetry wrote:
         | They can put away clutter but if they could chop a carrot or
         | dust a vase they'd have shown videos demonstrating that sort of
         | capability.
         | 
         | EDIT: Let alone chop an onion. Let me tell you having a robot
         | manipulate onions is the worst. Dealing with loose onion skins
         | is very hard.
        
           | squigz wrote:
           | There's something hilarious to me about the idea of chopping
           | onions being a sort of benchmark for robots.
        
           | j-krieger wrote:
           | Sure. But if you showed this video to someone 5 or 10 years
           | ago, they'd say it's fiction.
        
       | kla-s wrote:
       | Does anyone know how long they have been at this? Is this mainly
       | a reimplementation of the physical intelligence paper + the dual
       | size/freq + the cooperative part?
        
         | pr337h4m wrote:
         | "Over a year" according to the founder:
         | https://x.com/adcock_brett/status/1892578309344502191
        
       | bilsbie wrote:
       | This is amazing but it also made me realize I just don't trust
       | these videos. Is it sped up? How much is preprogrammed?
       | 
       | I now they claim there's no special coding but did they practice
       | this task? Special training?
       | 
       | Even if this video is totally legit I'm but burned out by all the
       | hype videos in general.
        
         | ge96 wrote:
         | they seem slow to me, I was thinking they're slow for safety
        
         | turnsout wrote:
         | They appear to be realtime, based on the robot's movements with
         | the human in the scene. If you believe the article, it's zero
         | shot (no preprogramming, practice or special training).
        
       | bilsbie wrote:
       | They should have made them talk. It's a little dehumanizing
       | otherwise.
        
       | bilsbie wrote:
       | I get the impression there's a language model sending high level
       | commands to a control model? I wonder when we can have one
       | multimodal model that controls everything.
       | 
       | The latest models seemed to be fluidly tied in with generating
       | voice; even singing and laughing.
       | 
       | It seems like it would be possible to train a multimodal that can
       | do that with low level actuator commands.
        
         | turnsout wrote:
         | If you read the article, they describe a two-system approach;
         | one "think fast" 80M parameter model running at 200hz to
         | control motion, and one "think slow" 7B parameter model running
         | at ~7-9hz for everything else (scene understanding, language
         | processing, etc).
         | 
         | If that sounds like a cheat, neuroscientists tell us this is
         | how the human brain works.
        
       | ge96 wrote:
       | Wonder what their vision stack is like. Depth via sensors or
       | purely visual and the distance estimating of objects and inverse
       | kinematics/proprioception, anyway it looks impressive.
        
       | yurimo wrote:
       | I don't know, there has been so many overhyped and faked demos in
       | humanoid robotics space over the last couple years, it is
       | difficult to believe what is clearly a demo release for
       | shareholders. Would love to see some demonstration in a less
       | controlled environment.
        
         | ge96 wrote:
         | Imagine they bring one out to a construction site and they
         | treat the robot as a new rookie guy, go pick up those pipes.
         | That would be an ultimate on the fly test to me.
        
           | ortsa wrote:
           | Picking up a bundle of loose pipes actually seems like a
           | great benchmark for humanoid robots. Especially if they're
           | not in a perfect pile. A full test could be something like
           | grabbing all the pipes, from the floor, and putting them into
           | a truck bed, in some (hopefully) sane fashion
        
             | sayamqazi wrote:
             | I have my personal multimodal benchmark for physical
             | robots.
             | 
             | You put a keyring with bunch of different keys in front of
             | a robot and then instruct it pick it up and open a lock
             | while you are describing which key is the correct one.
             | Something like "Use the key with black plastic head and you
             | need to put it in teeths facing down"
             | 
             | I have low hopes of this being possibe in the next 20
             | years. I hope I am still alive to witness if it ever
             | happens.
        
         | falcor84 wrote:
         | I suppose the next big milestone is Wozniak's Coffee Test: A
         | robot is to enter a random home and figure out how to make
         | coffee with whatever they have.
        
       | abraxas wrote:
       | Is this even reality or CGI? They really should show these things
       | off in less sterile environemtns because this video has a very
       | CGI feel to it.
        
         | psb217 wrote:
         | Natural, cluttered environments are a lot tougher to deal with.
         | This near future-y minimalist environment has the dual benefits
         | of looking stylish and being much closer to whatever they were
         | able to simulate at scale for training the models.
        
       | exe34 wrote:
       | Is there a paper? I think I get how they did their training, but
       | I'd like to understand it more.
       | 
       | Does anyone know if this trained model would work on a different
       | robot at all, or would it need retraining?
        
       | swalsh wrote:
       | At this point, this is enough autonomy to have a set of these
       | guys man a howitzer (read as old stockpiles of weapons we already
       | have). Kind of a scary thought. On one hand, I think the idea of
       | moving real people out of danger in war is a good idea, and as an
       | American i'd want Americans to have an edge... and we can't
       | guarantee our enemies won't take it if we skip it, on the other
       | hand I have a visceral reaction to machines killing people.
       | 
       | I think we're at an inflection point now where AI and robotics
       | can be used in warfare, and we need to start having that
       | conversation.
        
         | lyu07282 wrote:
         | I don't understand we already saw exactly what happens with the
         | emergence of drones and Israel is already using AI to select
         | bombing targets and semi-autonomous turrets. What conversation?
         | What kind of society do you think we are living in?
        
         | 01100011 wrote:
         | We had sufficient AI to make death machines for decades. You
         | don't need fancy LLMs to get a pretty good success rate for
         | targeting.
         | 
         | I have said for years that the only thing keeping us from
         | "stabby the robot" is solving the power problem. If you can
         | keep a drone going for a week, you have a killing machine. Use
         | blades to avoid running out of ammo. Use IR detection to find
         | the jugular. Stab, stab and move on. I'm guessing "traditional"
         | vision algorithms are also sufficient to, say, identify an
         | ethnicity and conduct ethnic cleansing. We are "solving the
         | power problem" away from a new class of WMDs that are
         | accessible to smaller states/groups/individuals.
        
           | j-krieger wrote:
           | > We had sufficient AI to make death machines for decades
           | 
           | And we already reached the peek here. Small drones that are
           | cheaply mass produced, fly on SIM cards alone and explode
           | when they reached a target. That's all there is to it. You
           | don't need a gun mounted on a spot or a humanoid robot
           | carrying a gun. Exploding swarms are enough.
        
         | Symmetry wrote:
         | They don't look strong enough to pick up a 155mm shell even
         | with both arms - and we haven't seen them pick up something
         | with two arms.
        
         | meindnoch wrote:
         | So you're concerned about remote operated howitzers?
         | Autoloaders and remote control land vehicles have existed for
         | 40 or so years by now. If we wanted remote controlled howitzers
         | we could have fielded them already.
        
       | causal wrote:
       | I'm always wondering at the safety measures on these things. How
       | much force is in those motors?
       | 
       | This is basically safety-critical stuff but with LLMs.
       | Hallucinating wrong answers in text is bad, hallucinating that
       | your chest is a drawer to pull open is very bad.
        
         | cess11 wrote:
         | Not a big deal on the battlefield.
        
           | causal wrote:
           | I'd say a very big deal when munitions and targeting are
           | involved
        
         | silentwanderer wrote:
         | In terms of low-level safety, they can probably back out forces
         | on the robot from current or torque measurement and detect
         | collisions. The challenge comes with faster motions carrying
         | lots of inertia and behavioral safety (e.g. don't pour oil on
         | the stove)
        
         | mmh0000 wrote:
         | The thing in the video moves slower than the sloth in Zootopia.
         | If you die by that robot, you probably deserve it.
        
           | causal wrote:
           | Are you saying it cannot move faster than they because of
           | some kind of governor?
        
             | Symmetry wrote:
             | A governor, the firmware in the motor controllers,
             | something like that. Certainly not the neural network
             | though.
        
           | throwaway0123_5 wrote:
           | As a sibling comment implies though, there's also danger from
           | it being stupid while unsupervised. For example, I'd be very
           | nervous having it do something autonomously in my kitchen for
           | fear of it burning down my house by accident.
        
           | exe34 wrote:
           | or if you're old, injured, groggy from medication, distracted
           | by something/someone else, blind, deaf or any number of
           | things.
           | 
           | it's easy to take your able body for granted, but reality
           | comes to meet all of us eventually.
        
           | mikehollinger wrote:
           | From a different robot (Boston Dynamics' new Atlas) - the
           | system moves at a "reasonable" speed. But watch at 1m20s in
           | this video[1]. You can see it bump and then move VERY quickly
           | -- with speed that would certainly damage something, or hurt
           | someone.
           | 
           | [1] https://www.youtube.com/watch?v=F_7IPm7f1vI
        
           | dr_kiszonka wrote:
           | They are designed to penetrate Holtzman shields, surely.
        
       | ianamo wrote:
       | Are we at a point now where Asimov's laws are programmed into
       | these fellas somewhere?
        
         | thomastjeffery wrote:
         | Nope.
         | 
         | The article clearly spells out that it's end to end LLM. Text
         | and video in, motor function out.
         | 
         | Technically, the text model probably has a few copies, but they
         | are nothing more than Asimov's _narrative_. Laws don 't (and
         | can't) exist in a model
        
       | porphyra wrote:
       | It seems that end to end neural networks for robotics are really
       | taking off. Can someone point me towards where to learn about
       | these, what the state of the art architectures look like, etc? Do
       | they just convert the video into a stream of tokens, run it
       | through a transformer, and output a stream of tokens?
        
         | vessenes wrote:
         | I was reading their site, and I too have some questions about
         | this architecture.
         | 
         | I'd be very interested to see what the output of their 'big
         | model' is that feeds into the small model. I presume the small
         | model gets a bunch of environmental input, and some input from
         | the big model, and we know that the big model input only
         | updates every 30 or 40 frames in terms of small model.
         | 
         | Like, do they just output random control tokens from big model
         | and embed those in small model and do gradient descent to find
         | a good control 'language'? Do they train the small model on
         | english tokens and have the big model output those? Custom
         | coordinates tokens? (probably). Lots of interesting
         | possibilities here.
         | 
         | By the way, the dataset they describe was generated by a large
         | (much larger presumably) vision model tasked with creating
         | tasks from successful videos.
         | 
         | So the pipeline is:
         | 
         | * Video of robot doing something
         | 
         | * (o1 or some other high end model) "describe very precisely
         | the task the robot was given"
         | 
         | * o1 output -> 7B model -> small model -> loss
        
       | andiareso wrote:
       | Seriously, what's with all of these perceived "high-end" tech
       | companies not doing static content worth a damn.
       | 
       | Stop hosting your videos as MP4s on your web-server. Either
       | publish to a CDN or use a platform like YouTube. Your bandwidth
       | cannot handle serving high resolution MP4s.
       | 
       | /rant
        
       | IAmNotACellist wrote:
       | I don't suppose this is open research and I can read about their
       | model architecture?
        
       | verytrivial wrote:
       | Are they claiming these robots are also silent? They seem to have
       | "crinkle" sounds handling packaging, which if added in post seems
       | needlessly smoke-and-mirror for what was a very impressive
       | demonstration (of robots impersonating an extreme stoned human.)
        
       | ein0p wrote:
       | There's no way this is 100% real though. No startup demo ever is.
        
       | the_other wrote:
       | It's funny... there a lot of comments here asking "why would
       | anyone pay for this, when you could learn to do the thing, or
       | organise your time/plans yourself."
       | 
       | That's how I feel about LLMs and code.
        
       | plipt wrote:
       | The demo is quite interesting but I am mostly intrigued by the
       | claim that it is running totally local to each robot. It seems to
       | use some agentic decision making but the article doesn't touch on
       | that. What possible combo of model types are they stringing
       | together? Or is this something novel?
       | 
       | The article mentions that the system in each robot uses two ai
       | models.                   S2 is built on a 7B-parameter open-
       | source, open-weight VLM pretrained on internet-scale data
       | 
       | and the other                   S1, an 80M parameter cross-
       | attention encoder-decoder transformer, handles low-level [motor?]
       | control.
       | 
       | It feels like although the article is quite openly technical they
       | are leaving out the secret sauce? So they use an open source VLM
       | to identify the objects on the counter. And another model to
       | generate the mechanical motions of the robot.
       | 
       | What part of this system understands 3 dimensional space of that
       | kitchen?
       | 
       | How does the robot closest to the refrigerator know to pass the
       | cookies to the robot on the left?
       | 
       | How is this kind of speech to text, visual identification,
       | decision making, motor control, multi-robot coordination and
       | navigation of 3d space possible locally?                   Figure
       | robots, each equipped with dual low-power-consumption embedded
       | GPUs
       | 
       | Is anyone skeptical? How much of this is possible vs a staged
       | tech demo to raise funding?
        
         | bbor wrote:
         | I'm very far from an expert, but:                 What part of
         | this system understands 3 dimensional space of that kitchen?
         | 
         | The visual model "understands" it most readily, I'd say -- like
         | a traditional Waymo CNN "understands" the 3D space of the road.
         | I don't think they've explicitly given the models a pre-
         | generated pointcloud of the space, if that's what you're
         | asking. But maybe I'm misunderstanding?                 How
         | does the robot closest to the refrigerator know to pass the
         | cookies to the robot on the left?
         | 
         | It appears that the robot is being fed plain english
         | instructions, just like any VLM would -- instead of the very
         | common `text+av => text` paradigm (classifiers, perception
         | models, etc), or the less common `text+av => av` paradigm
         | (segmenters, art generators, etc.), this is `text+av =>
         | movements`.
         | 
         | Feeding the robots the appropriate instructions at the
         | appropriate time is a higher-level task than is covered by this
         | demo, but I think is pretty clearly doable with existing AI
         | techniques (/a loop).                 How is this kind of
         | speech to text, visual identification, decision making, motor
         | control, multi-robot coordination and navigation of 3d space
         | possible locally?
         | 
         | If your question is "where's the GPUs", their "AI" marketing
         | page[1] pretty clearly implies that compute is offloaded, and
         | that only images and instructions are meaningfully "on board"
         | each robot. I could see this violating the understanding of
         | "totally local" that you mentioned up top, but IMHO those
         | claims are just clarifying that the individual figures aren't
         | controlled as one robot -- even if they ultimately employ the
         | same hardware. Each period (7Hz?) two sets of instructions are
         | generated.
         | 
         | [1] https://www.figure.ai/ai                 What possible
         | combo of model types are they stringing together? Or is this
         | something novel?
         | 
         | Again, I don't work in robotics at all, but have spent quite a
         | while cataloguing all the available foundational models, and I
         | wouldn't describe anything here as "totally novel" on the model
         | level. Certainly impressive, but not, like, a theoretical
         | breakthrough. Would love for an expert to correct me if I'm
         | wrong, tho!
         | 
         | EDIT: Oh and finally:                 Is anyone skeptical? How
         | much of this is possible vs a staged tech demo to raise
         | funding?
         | 
         | Surely they are downplaying the difficulties of getting this
         | setup perfectly, and don't show us how many bad runs it took to
         | get these flawless clips.
         | 
         | They are seeking to raise their valuation from ~$3B to ~$40B
         | this month, sooooooo take that as you will ;)
         | 
         | https://www.reuters.com/technology/artificial-intelligence/r...
        
           | plipt wrote:
           | their "AI" marketing page[1] pretty clearly implies that
           | compute is offloaded
           | 
           | I think that answers most of my questions.
           | 
           | I am also not in robotics, so this demo does seem quite
           | impressive to me but I think they could have been more clear
           | on exactly what technologies they are demonstrating. Overall
           | still very cool.
           | 
           | Thanks for your reply
        
       | bbor wrote:
       | To focus on something other than the obviously-terrifying nature
       | of this and the skepticism that rightfully entails on our part:
       | A fast reactive visuomotor policy that translates the latent
       | semantic representations produced by S2 into precise continuous
       | robot actions at 200 Hz
       | 
       | Why 200Hz...? Any experts in here on robotics? Because to this
       | layman that seems really often to update motor controls.
        
       | dr_dshiv wrote:
       | Wake me when robots can make a peanut butter sandwich
        
       | kingkulk wrote:
       | Anyone have a link to their paper?
        
       ___________________________________________________________________
       (page generated 2025-02-20 23:00 UTC)