[HN Gopher] Helix: A vision-language-action model for generalist...
___________________________________________________________________
Helix: A vision-language-action model for generalist humanoid
control
Author : Philpax
Score : 217 points
Date : 2025-02-20 14:30 UTC (8 hours ago)
(HTM) web link (www.figure.ai)
(TXT) w3m dump (www.figure.ai)
| kubb wrote:
| Wow! This is something new.
| sandis wrote:
| YouTube link for the video (for whatever reason the video hosted
| on their site kept buffering for me):
| https://www.youtube.com/watch?v=Z3yQHYNXPws
| bhouston wrote:
| When doing robot control, how do you model in the control of the
| robot? Do you have tool_use / function calling at the top level
| model which then gets turned into motion control parameters via
| inverse kinematic controllers?
|
| What is the interface from the top level to the motors?
|
| I feel it can not just be a neural network all the way down,
| right?
| Philpax wrote:
| Have a look at the post - it explains how it works. There are
| two models: a 7-9Hz 7B vision-language model, and a 200Hz 80M
| visuomotor model. The former produces a latent vector, which is
| then interpreted by the latter to drive the motors.
| NitpickLawyer wrote:
| > a 7-9Hz 7B vision-language model, and a 200Hz 80M
| visuomotor model.
|
| huh. An interesting approach. I wonder if something like this
| can be used for other things as well, like "computer use"
| with the same concept of a "large" model handling the goals,
| and a "small" model handling clicking and stuff, at much
| higher rates, useful for games and things like that.
| whatever1 wrote:
| This is typical in real time applications. A supervisor
| tries to guess in which region the system is currently and
| then invokes the correct set of lower level algorithms.
| traverseda wrote:
| "The first time you've seen these objects" is a weird thing to
| say. One presumes that this is already in their training set, and
| that these models aren't storing a huge amount of data in their
| context, so what does that even mean?
| jayd16 wrote:
| It probably gives them confidence that they can accurately see
| a thing even though they don't know what that thing is.
|
| I could also imagine a lot of safety around leaving things
| outside of the current task alone so you might have to bend
| over backwards to get new objects worked on.
| thomastjeffery wrote:
| There is no such thing as "thing" here.
|
| These models are trained such that the given conditions (the
| visual input and the text prompt) will be continued with a
| desirable continuation (motor function over time).
|
| The only dimension accuracy can apply to is desirability.
| Symmetry wrote:
| It's normal to have a training set and a validation set and I
| interpreted that to mean that these items weren't in the
| training set.
| ygouzerh wrote:
| So from what I understand it actually means that they were for
| example never trained on a video of an apple. Maybe only on a
| video of bread, pineapple, chocolate.
|
| However, as it was trained using generic text data similarly to
| a normal LLM, it knows how an apple is supposed to look like.
|
| Similar than a kid that never saw a banana, but his parent
| described it to him.
| ziofill wrote:
| There's nothing I want more than a robot that does house chores.
| That's the real 10x multiplier for humans to do what they do
| best.
| dartos wrote:
| Hopefully in the next decade we'll get there.
|
| Vision+language multimodal models seem to solve some of the
| hard problems.
| siavosh wrote:
| What do humans do best?
| jayd16 wrote:
| Everything everything else is worse at.
| ziofill wrote:
| I mean to use their time to pursue their passions and
| interests, not cleaning up the kitchen or making the bed or
| doing laundry...
| hooverd wrote:
| Given time to "pursue their passions and interests", most
| people chose to turn their brain to soup on social media.
| KolmogorovComp wrote:
| but most people think they are better than most people.
| ein0p wrote:
| Browse Instagram, apparently.
| abraxas wrote:
| Yeah, except that future doesn't need us. By us I mean those of
| us who don't have $1B to their name.
|
| Do you really expect the oligarchs to put up with the
| environmental degradation of 8 billion humans when they can
| have a pristine planet to themselves with their whims served by
| the AI and these robots?
|
| I fully anticipate that when these things mature enough we'll
| see an "accidental" pandemic sweep and kill off 90% of us. At
| least 90%.
| ewjt wrote:
| Oligarchs would use the robots to kill people instead of a
| pandemic. A virus carries too much risk of affecting the
| original creators.
|
| Fortunately, robotic capability like that basically becomes
| the equivalent of Nuclear MAD.
|
| Unfortunately, the virus approach probably looks fantastic to
| extremist bad actors with visions of an afterlife.
| ben_w wrote:
| I'd expect Musk and Bezos to know about von Neumann
| replicators; factories that make these robots staffed
| entirely by these robots all the way to the mines digging
| minerals out of the ground... rapid and literally exponential
| growth until they hit whatever the limiting factor is, but
| they've both got big orbial rockets now, so the limit isn't
| necessarily 6e24 kg.
| cess11 wrote:
| To me this is such a weird wish. Why would you not want to care
| for your home and the people living there? Why would you want
| to have a slave taking these activities from you?
|
| I'd rather have less waged labour and more time for chores with
| the family.
| mmh0000 wrote:
| It's called a "house cleaner" and they only cost ~$150 (area
| and all varies) bi-weekly. I'll shit a brick (and then have the
| robot clean it up) if a robot is ever cheaper than ~$4000/yr.
| vessenes wrote:
| A robot will definitely cost less than $30/hr eventually. But
| you'll be running it a lot more than a few hours every other
| week.
| ziofill wrote:
| Yeah, but a robot will work 24/7, not 2h byweekly -_-
| mclau156 wrote:
| do you need a robot to work in your house 24/7?
| gigel82 wrote:
| Why not? If it's done with all the chores, I can have it
| make some silly woodworking / art project for Etsy to
| earn its keep, or just loan it out to neighbors.
| ben_w wrote:
| In the same way it's hard to earn money from using AI to
| make art, I don't see Etsy projects made by affordable
| domestic robots selling above cost.
| ziofill wrote:
| Well perhaps not at night, but otherwise there's always
| something to clean, something to fix, something to cook,
| take care of the yard.. heck I might need two robots ^^'
| 01100011 wrote:
| I'd pay $2k for something that folds my laundry reliably. It
| doesn't need arms or legs, just like my dishwasher doesn't need
| arms or legs. It just needs to let me dump in a load of clean
| laundry and output stacks of neatly folded or hung clothing.
| anentropic wrote:
| Very impressive
|
| Why make such sinister-looking robots though...?
| jayd16 wrote:
| With the way they move, they look like stoned teenagers
| interning at a Bond villain factory. Not to knock the tech but
| they're scary and silly at the same time.
| esafak wrote:
| Black was not the best color choice.
| aerodog wrote:
| Interesting timing - same day MSFT releases
| https://microsoft.github.io/Magma/
| _1 wrote:
| Current discussion:
| https://news.ycombinator.com/item?id=43110265
| wwwtyro wrote:
| Until we get robots with really good hands, something I'd love in
| the interim is a system that uses _me_ as the hands. When it's
| time to put groceries away, I don't want to have to think about
| how to organize everything. Just figure out which grocery items I
| have, what storage I have available, come up with an optimized
| organization solution, then tell me where to put things, one at a
| time. I'm cautiously optimistic this will be doable in the near
| term with a combination of AR and AI.
| camjw wrote:
| Maybe I don't understand exactly what you're describing but why
| would anyone pay for this? When I bring home the shopping I
| just... chuck stuff in the cupboards. I already know where it
| all goes. Maybe you can explain more?
| bear141 wrote:
| Maybe some people just assume there is a "best" or "optimal"
| way to do everything and AI will tell us what that is. Some
| things are just preference and I don't mind the tiny amount
| of energy that goes into doing small things the way I like.
| jayd16 wrote:
| Maybe they're imagining more complex tasks like working on an
| engine.
| loudmax wrote:
| One use case I imagine is skilled workmanship. For example,
| putting on a pair of AR glasses and having the equivalent of
| an experienced plumber telling me exactly where to look for
| that leak and how to fix it. Or how to replace my brake pads
| or install a new kitchen sink.
|
| When I hire a plumber or a mechanic or an electrician, I'm
| not just paying for muscle. Most of the value these
| professionals bring is experience and understanding. If a
| video-capable AI model is able to assume that experience,
| then either I can do the job myself or hire some 20 year old
| kid at roughly minimum wage. If capabilities like this come
| about, it will be very disruptive, for better and for worse.
| semi-extrinsic wrote:
| This is called "watching YouTube tutorials". We've had it
| for decades.
| rolisz wrote:
| But what if there's no YouTube tutorial for the exact AC
| unit you have and it doesn't look like any of the videos
| you checked out?
| cess11 wrote:
| Have you met people that seem to be able to fix almost
| anything?
|
| If you can't get a tutorial on your exact case you learn
| about the problem domain and intuit from there. Usually
| it works out if you're careful, unlike software.
| semi-extrinsic wrote:
| Then you are equally fucked as the AI will be, so no
| difference.
|
| Case in point, I remember about ten years ago our washing
| machine started making noise from the drum bearing. Found
| a Youtube tutorial for bearing replacement on the exact
| same model, but 3 years older. Followed it just fine
| until it was time to split the drum. Then it turned out
| that in the newer units like mine, some rent-seeking MBA
| fuckers had decided more profits could be had if they
| _plastic welded shut the entire drum assembly_. Which was
| then a $300 replacement part for a $400 machine.
|
| An AI doesn't help with this type of shit. It can't know
| the unknown.
| deepGem wrote:
| But once it knows it's pretty certain to become common
| knowledge almost instantaneously. That's not possible
| now. What you learn stays localised to you and may be
| people 1 degree away from you that's it.
| semi-extrinsic wrote:
| How does that work? None of the current AI models can re-
| train on the fly. How would the inference engine even
| know if it's a case of new information that needs to be
| fed back, or just a user that's not following
| instructions correctly?
| hulahoof wrote:
| Sounds like what Hololens was designed to solve, more in
| the AR space than AI though
| __MatrixMan__ wrote:
| It would be nice to be able to select a recipe and have it
| populate your shopping list based on what is currently in
| your cupboards. If you just chuck stuff in the cupboards then
| you have to be home to know what they contain.
|
| Or you could wear it while you cook and it could give you
| nutrition information for whatever it is you cooked. Armed
| with that it could make recommendations about what nutrients
| you're likely deficient in based on your recent meals and
| suggest recipes to remedy the gap--recipes based on what it
| knows is already in the cupboard.
| gopher_space wrote:
| Maybe I'm showing my age, but isn't this a home ec class?
| __MatrixMan__ wrote:
| I took home ec in 2001. I learned to use a sewing
| machine, it was great.
|
| But none of the kitchen stuff we learned had anything to
| do with ensuring that this week's shopping list ensures
| that you'll get enough zinc next week, or the kind of
| prep that uses the other half of yesterday's cauliflower
| in tomorrow's dinner so that it doesn't go bad.
|
| These aren't hard problems to solve if you've got time to
| plan, but they are hard to solve if you are currently at
| the grocery store and can't remember that you've got a
| half a cauliflower that needs an associated recipe.
| SoftTalker wrote:
| Yes it's pretty amazing how so many people seem to
| overcomplicate simple household tasks by introducing
| unnecessary technology.
| luma wrote:
| > why would anyone pay for this?
|
| Presumably, they won't as this is still a tech demo. One can
| take this simple demonstration and think about some future
| use cases that aren't too different. How far away is
| something that'll do the dishes, cook a meal, or fold the
| laundry, etc? That's a very different value prop, and one
| that might attract a few buyers.
| Philip-J-Fry wrote:
| The person you're replying to is referring to the GP. The
| GP asks for an AI that tells them where to put their
| shopping. Why would anyone pay for THAT? Since we already
| know where everything goes without needing an AI to tell
| us. An AI isn't going to speed that up.
| cactusplant7374 wrote:
| Your solution sounds like the worst cognitive load for getting
| home from the grocery store and wanting it all to be over.
| lucianbr wrote:
| You want to outsource thinking to a computer system and keep
| manual labor? You do you, but I want the opposite. I want to
| decide what goes where but have a robot actually put the stuff
| there.
| TeMPOraL wrote:
| That's the problem, though - the computer is already better
| at thinking than you, but we still don't know how to make it
| good at arbitrary labor requiring a mix of precision and
| power, something humans find natural.
|
| In other words: I'm sorry, but that's how reality turned out.
| Robots are better at thinking, humans better at laboring. Why
| fight against nature?
|
| (Just joking... I think.)
| RedNifre wrote:
| I think he means outsourcing everything eventually, but right
| now, outsourcing the thought process is possible, while
| outsourcing the manual labor is not.
| lynx97 wrote:
| A good AI fridge would be already a great starting point. With
| a checkin procedure that makes sure to actually know whats in
| the fridge. Complete with expiry tracking and recipe
| suggestions based on personal preferences combined with product
| expiry. I am totally unimpressed with almost everything I see
| in home automation these days, but I'd immediately buy the AI
| fridge if it really worked smoothly.
| Philpax wrote:
| Sounds like what's described in Manna:
| https://marshallbrain.com/manna1
| __MatrixMan__ wrote:
| I imagine a something like a headlamp except it's a projector
| and a camera so it can just light up where it wants you to pick
| something up in one color or where it wants you to put it down
| in another color. It can learn from what it sees of my hands
| how the eventual robot should handle the space (e.g. not
| putting heavy things on top of fragile things and such).
|
| I'd totally use that to clean my garage so that later I can ask
| it where the heck I put the thing or ask it if I already have
| something before I buy one...
| hooverd wrote:
| You already have one: a brain.
| RedNifre wrote:
| I fully agree, building something like this is somewhere in my
| back log.
|
| I think the key point why this "reverse cyborg" idea is not as
| dystopian as, say, being a worker drone in a large warehouse
| where the AI does not let you go to the toilet is that the AI
| is under your own control, so you decide on the high level goal
| "sort the stuff away", the AI does the intermediate planning
| and you do the execution.
|
| We already have systems like that, every time you use you tell
| your navi where you want to go, it plans the route and gives
| you primitive commands like "on the next intersection, turn
| right", so why not have those for cooking, doing the laundry,
| etc.?
|
| Heck, even a paper calendar is already kinda this, as in
| separating the planning phase from the execution phase.
| Jarwain wrote:
| I'm quite slowly working on something like this, but for
| time.
|
| For "stuff" I think a bigger draw is having it so it can let
| me know "hey you already have 3 of those spices at locations
| x, y, and z, so don't get another" or "hey you won't be able
| to fit that in your freezer"
| malux85 wrote:
| Yeah there's more to it than that. Do you want a can of beans
| to be put in the utensil draw just because it would fit? If it
| was done as you describe the placement of all of your items
| would be almost random each time, the bot need to have
| contextual memory and familiarity with your storage habits and
| preferences.
|
| This can be done of course, in your statement the phrase "just
| figure out" is doing a lot more heavy lifting than you allude
| to
| sho_hn wrote:
| Dunno, I would not want to bleed my mental faculties for doing
| even simple planning work like this by outsourcing it to AI.
| Reliance on crutches like this would seem like a pathway to
| early-onset dementia.
| meowkit wrote:
| Already playing out, anecdotally to my experience.
|
| Its similar to losing callouses on our hands if you don't
| labor/go to the gym.
| htrp wrote:
| so the kiva-amazon model?
| falcor84 wrote:
| This is almost literally the first chapter in Marshall Brain's
| "Manna" [0], being the first step towards world-controlling
| AGI:
|
| > Manna told employees what to do simply by talking to them.
| Employees each put on a headset when they punched in. Manna had
| a voice synthesizer, and with its synthesized voice Manna told
| everyone exactly what to do through their headsets. Constantly.
| Manna micro-managed minimum wage employees to create perfect
| performance.
|
| [0] https://marshallbrain.com/manna1
| sottol wrote:
| Imo, the Terminator movies would have been scarier if they moved
| like these guys - slow, careful, deliberate and measured but
| unstoppable. There's something uncanny about this.
| megous wrote:
| Unfortunately, there'll be no time travel to save us. That was
| the lying part of the movie. Other stuff was true.
| ramenlover wrote:
| Why do they make "eye contact" after every hand off? Feels oddly
| forced.
| bear141 wrote:
| This along with the writing style in the description is totally
| forced anthropomorphizing. It's creepy.
| jimbohn wrote:
| Gotta hype up the investors somehow
| GoatInGrey wrote:
| Perhaps they're exchanging knowing looks on how stupid they
| think the demo is. Solidarity between artificial brothers.
| Symmetry wrote:
| So, there's no way you can have fully actuated control of every
| finger joint with just 35 degrees of freedom. Which is very
| reasonable! Humans can't individually control each of our finger
| joints either. But I'm curious how their hand setups work, which
| parts are actuated and which are compliant. In the videos I'm not
| seeing any in-hand manipulation other than just grasping,
| releasing, and maintaining the orientation of the object relative
| to the hand and I'm curious how much it can do / they plan to
| have it be able to do. Do they have any plans to try to mimic
| OpenAI's one handed rubics cube demo?
| pr337h4m wrote:
| Goal 2 has been achieved, at least as a proof of concept (and not
| by OpenAI): https://openai.com/index/openai-technical-goals/
| Symmetry wrote:
| They can put away clutter but if they could chop a carrot or
| dust a vase they'd have shown videos demonstrating that sort of
| capability.
|
| EDIT: Let alone chop an onion. Let me tell you having a robot
| manipulate onions is the worst. Dealing with loose onion skins
| is very hard.
| squigz wrote:
| There's something hilarious to me about the idea of chopping
| onions being a sort of benchmark for robots.
| j-krieger wrote:
| Sure. But if you showed this video to someone 5 or 10 years
| ago, they'd say it's fiction.
| kla-s wrote:
| Does anyone know how long they have been at this? Is this mainly
| a reimplementation of the physical intelligence paper + the dual
| size/freq + the cooperative part?
| pr337h4m wrote:
| "Over a year" according to the founder:
| https://x.com/adcock_brett/status/1892578309344502191
| bilsbie wrote:
| This is amazing but it also made me realize I just don't trust
| these videos. Is it sped up? How much is preprogrammed?
|
| I now they claim there's no special coding but did they practice
| this task? Special training?
|
| Even if this video is totally legit I'm but burned out by all the
| hype videos in general.
| ge96 wrote:
| they seem slow to me, I was thinking they're slow for safety
| turnsout wrote:
| They appear to be realtime, based on the robot's movements with
| the human in the scene. If you believe the article, it's zero
| shot (no preprogramming, practice or special training).
| bilsbie wrote:
| They should have made them talk. It's a little dehumanizing
| otherwise.
| bilsbie wrote:
| I get the impression there's a language model sending high level
| commands to a control model? I wonder when we can have one
| multimodal model that controls everything.
|
| The latest models seemed to be fluidly tied in with generating
| voice; even singing and laughing.
|
| It seems like it would be possible to train a multimodal that can
| do that with low level actuator commands.
| turnsout wrote:
| If you read the article, they describe a two-system approach;
| one "think fast" 80M parameter model running at 200hz to
| control motion, and one "think slow" 7B parameter model running
| at ~7-9hz for everything else (scene understanding, language
| processing, etc).
|
| If that sounds like a cheat, neuroscientists tell us this is
| how the human brain works.
| ge96 wrote:
| Wonder what their vision stack is like. Depth via sensors or
| purely visual and the distance estimating of objects and inverse
| kinematics/proprioception, anyway it looks impressive.
| yurimo wrote:
| I don't know, there has been so many overhyped and faked demos in
| humanoid robotics space over the last couple years, it is
| difficult to believe what is clearly a demo release for
| shareholders. Would love to see some demonstration in a less
| controlled environment.
| ge96 wrote:
| Imagine they bring one out to a construction site and they
| treat the robot as a new rookie guy, go pick up those pipes.
| That would be an ultimate on the fly test to me.
| ortsa wrote:
| Picking up a bundle of loose pipes actually seems like a
| great benchmark for humanoid robots. Especially if they're
| not in a perfect pile. A full test could be something like
| grabbing all the pipes, from the floor, and putting them into
| a truck bed, in some (hopefully) sane fashion
| sayamqazi wrote:
| I have my personal multimodal benchmark for physical
| robots.
|
| You put a keyring with bunch of different keys in front of
| a robot and then instruct it pick it up and open a lock
| while you are describing which key is the correct one.
| Something like "Use the key with black plastic head and you
| need to put it in teeths facing down"
|
| I have low hopes of this being possibe in the next 20
| years. I hope I am still alive to witness if it ever
| happens.
| falcor84 wrote:
| I suppose the next big milestone is Wozniak's Coffee Test: A
| robot is to enter a random home and figure out how to make
| coffee with whatever they have.
| abraxas wrote:
| Is this even reality or CGI? They really should show these things
| off in less sterile environemtns because this video has a very
| CGI feel to it.
| psb217 wrote:
| Natural, cluttered environments are a lot tougher to deal with.
| This near future-y minimalist environment has the dual benefits
| of looking stylish and being much closer to whatever they were
| able to simulate at scale for training the models.
| exe34 wrote:
| Is there a paper? I think I get how they did their training, but
| I'd like to understand it more.
|
| Does anyone know if this trained model would work on a different
| robot at all, or would it need retraining?
| swalsh wrote:
| At this point, this is enough autonomy to have a set of these
| guys man a howitzer (read as old stockpiles of weapons we already
| have). Kind of a scary thought. On one hand, I think the idea of
| moving real people out of danger in war is a good idea, and as an
| American i'd want Americans to have an edge... and we can't
| guarantee our enemies won't take it if we skip it, on the other
| hand I have a visceral reaction to machines killing people.
|
| I think we're at an inflection point now where AI and robotics
| can be used in warfare, and we need to start having that
| conversation.
| lyu07282 wrote:
| I don't understand we already saw exactly what happens with the
| emergence of drones and Israel is already using AI to select
| bombing targets and semi-autonomous turrets. What conversation?
| What kind of society do you think we are living in?
| 01100011 wrote:
| We had sufficient AI to make death machines for decades. You
| don't need fancy LLMs to get a pretty good success rate for
| targeting.
|
| I have said for years that the only thing keeping us from
| "stabby the robot" is solving the power problem. If you can
| keep a drone going for a week, you have a killing machine. Use
| blades to avoid running out of ammo. Use IR detection to find
| the jugular. Stab, stab and move on. I'm guessing "traditional"
| vision algorithms are also sufficient to, say, identify an
| ethnicity and conduct ethnic cleansing. We are "solving the
| power problem" away from a new class of WMDs that are
| accessible to smaller states/groups/individuals.
| j-krieger wrote:
| > We had sufficient AI to make death machines for decades
|
| And we already reached the peek here. Small drones that are
| cheaply mass produced, fly on SIM cards alone and explode
| when they reached a target. That's all there is to it. You
| don't need a gun mounted on a spot or a humanoid robot
| carrying a gun. Exploding swarms are enough.
| Symmetry wrote:
| They don't look strong enough to pick up a 155mm shell even
| with both arms - and we haven't seen them pick up something
| with two arms.
| meindnoch wrote:
| So you're concerned about remote operated howitzers?
| Autoloaders and remote control land vehicles have existed for
| 40 or so years by now. If we wanted remote controlled howitzers
| we could have fielded them already.
| causal wrote:
| I'm always wondering at the safety measures on these things. How
| much force is in those motors?
|
| This is basically safety-critical stuff but with LLMs.
| Hallucinating wrong answers in text is bad, hallucinating that
| your chest is a drawer to pull open is very bad.
| cess11 wrote:
| Not a big deal on the battlefield.
| causal wrote:
| I'd say a very big deal when munitions and targeting are
| involved
| silentwanderer wrote:
| In terms of low-level safety, they can probably back out forces
| on the robot from current or torque measurement and detect
| collisions. The challenge comes with faster motions carrying
| lots of inertia and behavioral safety (e.g. don't pour oil on
| the stove)
| mmh0000 wrote:
| The thing in the video moves slower than the sloth in Zootopia.
| If you die by that robot, you probably deserve it.
| causal wrote:
| Are you saying it cannot move faster than they because of
| some kind of governor?
| Symmetry wrote:
| A governor, the firmware in the motor controllers,
| something like that. Certainly not the neural network
| though.
| throwaway0123_5 wrote:
| As a sibling comment implies though, there's also danger from
| it being stupid while unsupervised. For example, I'd be very
| nervous having it do something autonomously in my kitchen for
| fear of it burning down my house by accident.
| exe34 wrote:
| or if you're old, injured, groggy from medication, distracted
| by something/someone else, blind, deaf or any number of
| things.
|
| it's easy to take your able body for granted, but reality
| comes to meet all of us eventually.
| mikehollinger wrote:
| From a different robot (Boston Dynamics' new Atlas) - the
| system moves at a "reasonable" speed. But watch at 1m20s in
| this video[1]. You can see it bump and then move VERY quickly
| -- with speed that would certainly damage something, or hurt
| someone.
|
| [1] https://www.youtube.com/watch?v=F_7IPm7f1vI
| dr_kiszonka wrote:
| They are designed to penetrate Holtzman shields, surely.
| ianamo wrote:
| Are we at a point now where Asimov's laws are programmed into
| these fellas somewhere?
| thomastjeffery wrote:
| Nope.
|
| The article clearly spells out that it's end to end LLM. Text
| and video in, motor function out.
|
| Technically, the text model probably has a few copies, but they
| are nothing more than Asimov's _narrative_. Laws don 't (and
| can't) exist in a model
| porphyra wrote:
| It seems that end to end neural networks for robotics are really
| taking off. Can someone point me towards where to learn about
| these, what the state of the art architectures look like, etc? Do
| they just convert the video into a stream of tokens, run it
| through a transformer, and output a stream of tokens?
| vessenes wrote:
| I was reading their site, and I too have some questions about
| this architecture.
|
| I'd be very interested to see what the output of their 'big
| model' is that feeds into the small model. I presume the small
| model gets a bunch of environmental input, and some input from
| the big model, and we know that the big model input only
| updates every 30 or 40 frames in terms of small model.
|
| Like, do they just output random control tokens from big model
| and embed those in small model and do gradient descent to find
| a good control 'language'? Do they train the small model on
| english tokens and have the big model output those? Custom
| coordinates tokens? (probably). Lots of interesting
| possibilities here.
|
| By the way, the dataset they describe was generated by a large
| (much larger presumably) vision model tasked with creating
| tasks from successful videos.
|
| So the pipeline is:
|
| * Video of robot doing something
|
| * (o1 or some other high end model) "describe very precisely
| the task the robot was given"
|
| * o1 output -> 7B model -> small model -> loss
| andiareso wrote:
| Seriously, what's with all of these perceived "high-end" tech
| companies not doing static content worth a damn.
|
| Stop hosting your videos as MP4s on your web-server. Either
| publish to a CDN or use a platform like YouTube. Your bandwidth
| cannot handle serving high resolution MP4s.
|
| /rant
| IAmNotACellist wrote:
| I don't suppose this is open research and I can read about their
| model architecture?
| verytrivial wrote:
| Are they claiming these robots are also silent? They seem to have
| "crinkle" sounds handling packaging, which if added in post seems
| needlessly smoke-and-mirror for what was a very impressive
| demonstration (of robots impersonating an extreme stoned human.)
| ein0p wrote:
| There's no way this is 100% real though. No startup demo ever is.
| the_other wrote:
| It's funny... there a lot of comments here asking "why would
| anyone pay for this, when you could learn to do the thing, or
| organise your time/plans yourself."
|
| That's how I feel about LLMs and code.
| plipt wrote:
| The demo is quite interesting but I am mostly intrigued by the
| claim that it is running totally local to each robot. It seems to
| use some agentic decision making but the article doesn't touch on
| that. What possible combo of model types are they stringing
| together? Or is this something novel?
|
| The article mentions that the system in each robot uses two ai
| models. S2 is built on a 7B-parameter open-
| source, open-weight VLM pretrained on internet-scale data
|
| and the other S1, an 80M parameter cross-
| attention encoder-decoder transformer, handles low-level [motor?]
| control.
|
| It feels like although the article is quite openly technical they
| are leaving out the secret sauce? So they use an open source VLM
| to identify the objects on the counter. And another model to
| generate the mechanical motions of the robot.
|
| What part of this system understands 3 dimensional space of that
| kitchen?
|
| How does the robot closest to the refrigerator know to pass the
| cookies to the robot on the left?
|
| How is this kind of speech to text, visual identification,
| decision making, motor control, multi-robot coordination and
| navigation of 3d space possible locally? Figure
| robots, each equipped with dual low-power-consumption embedded
| GPUs
|
| Is anyone skeptical? How much of this is possible vs a staged
| tech demo to raise funding?
| bbor wrote:
| I'm very far from an expert, but: What part of
| this system understands 3 dimensional space of that kitchen?
|
| The visual model "understands" it most readily, I'd say -- like
| a traditional Waymo CNN "understands" the 3D space of the road.
| I don't think they've explicitly given the models a pre-
| generated pointcloud of the space, if that's what you're
| asking. But maybe I'm misunderstanding? How
| does the robot closest to the refrigerator know to pass the
| cookies to the robot on the left?
|
| It appears that the robot is being fed plain english
| instructions, just like any VLM would -- instead of the very
| common `text+av => text` paradigm (classifiers, perception
| models, etc), or the less common `text+av => av` paradigm
| (segmenters, art generators, etc.), this is `text+av =>
| movements`.
|
| Feeding the robots the appropriate instructions at the
| appropriate time is a higher-level task than is covered by this
| demo, but I think is pretty clearly doable with existing AI
| techniques (/a loop). How is this kind of
| speech to text, visual identification, decision making, motor
| control, multi-robot coordination and navigation of 3d space
| possible locally?
|
| If your question is "where's the GPUs", their "AI" marketing
| page[1] pretty clearly implies that compute is offloaded, and
| that only images and instructions are meaningfully "on board"
| each robot. I could see this violating the understanding of
| "totally local" that you mentioned up top, but IMHO those
| claims are just clarifying that the individual figures aren't
| controlled as one robot -- even if they ultimately employ the
| same hardware. Each period (7Hz?) two sets of instructions are
| generated.
|
| [1] https://www.figure.ai/ai What possible
| combo of model types are they stringing together? Or is this
| something novel?
|
| Again, I don't work in robotics at all, but have spent quite a
| while cataloguing all the available foundational models, and I
| wouldn't describe anything here as "totally novel" on the model
| level. Certainly impressive, but not, like, a theoretical
| breakthrough. Would love for an expert to correct me if I'm
| wrong, tho!
|
| EDIT: Oh and finally: Is anyone skeptical? How
| much of this is possible vs a staged tech demo to raise
| funding?
|
| Surely they are downplaying the difficulties of getting this
| setup perfectly, and don't show us how many bad runs it took to
| get these flawless clips.
|
| They are seeking to raise their valuation from ~$3B to ~$40B
| this month, sooooooo take that as you will ;)
|
| https://www.reuters.com/technology/artificial-intelligence/r...
| plipt wrote:
| their "AI" marketing page[1] pretty clearly implies that
| compute is offloaded
|
| I think that answers most of my questions.
|
| I am also not in robotics, so this demo does seem quite
| impressive to me but I think they could have been more clear
| on exactly what technologies they are demonstrating. Overall
| still very cool.
|
| Thanks for your reply
| bbor wrote:
| To focus on something other than the obviously-terrifying nature
| of this and the skepticism that rightfully entails on our part:
| A fast reactive visuomotor policy that translates the latent
| semantic representations produced by S2 into precise continuous
| robot actions at 200 Hz
|
| Why 200Hz...? Any experts in here on robotics? Because to this
| layman that seems really often to update motor controls.
| dr_dshiv wrote:
| Wake me when robots can make a peanut butter sandwich
| kingkulk wrote:
| Anyone have a link to their paper?
___________________________________________________________________
(page generated 2025-02-20 23:00 UTC)