Really interesting. Upshot: a well trained multimodal video generation model has a world representation model trained inside it. They’ve done some work lifting this world model out and deploying it to robots, where it seems to work well.
On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.
I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released videos of users training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.
The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.
I think this tactic is standard, no? Nvidia has demoed robots being trained using video generation models, and Waymo has been doing that for a while too.
The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?
You are indeed out of the loop. Google did something arguably more impressive more than a year ago, with a VLA based bot replacing a tensioned timing belt: https://www.youtube.com/watch?v=2AAFiuEP7iE
if you want to catch up, i would say Generalist is the state of the art in dexterous manipulation tasks, you can learn more here https://generalistai.com/blog/gen-1
> However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.
Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled representations are not as useful for modeling the real world.
We’re just pulling signs of LLM touched writing out of our ass now. Might be time to move on from the accusations, assume all writing is at least LLM assisted and judge it purely on the quality.
Bad news, humans write silly things all the time. If anything, LLMs are less likely to make awkward phrasings than people, because they aren’t found in the training data very often.
What's sad to me is that we have all this awesome technology, but movies are worse than ever.
I usually watch movies from decades ago just to find something decent and it's amazing how with goofy-looking puppets the storytelling was 1000 times better.
I regularly stream movies from Screambox, FilmRise and Alter as background music while I work from home. I love cheesy horror, but they're just bad enough to keep me company without distracting me
OP replied but I’m chipping in too, I’ve also been watching a lot of older movies specifically the more B list and C list which weren’t successful cult classics and they’re still better than anything new. The low production and low talent C list ones are still cheesy and the performances don’t land anywhere near as well as they could but they feel more genuine and creative than what passes for a movie these days.
would you refuse to invest in BFL because Robin Rombach doesn't know anything about making movies?
sora is done. can you find me one person brave enough to say, "Bill Peebles, a nice guy, doesn't know anything about making movies?" this isn't just a VC problem, this is why it is called the Hollywood No, but that said: maybe the hollywood no is the problem? in video games it has been called Toxic Positivity. are you getting it now?
Perhaps.. Mistral seems to be focussing more on LLM application rather than foundation models. Seeing that the latter may be becoming commodities, Mistral may come out ahead.
BFL resides in the same region as BMW and Audi. Quite wealthy and pretty fierce to protect their car industry. If Audi or any other german car maker sees BFL as the future of automation, it will be rather hard to buy them out.
On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.
I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released videos of users training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.
The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.
https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-f...
Luma Labs https://lumalabs.ai/news/luma-open-physical-ai-lab
Runway https://runway.com/product/robotics
> However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.
Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled representations are not as useful for modeling the real world.
That dash definitely means an LLM wrote this comment. /s
What's sad to me is that we have all this awesome technology, but movies are worse than ever.
I usually watch movies from decades ago just to find something decent and it's amazing how with goofy-looking puppets the storytelling was 1000 times better.
This is probably an unrelated rant, sorry
But with my GF I watch whatever is rated >= 8 on Imdb and the 8s from years ago are ALWAYS much, much better than the recent ones.
Maybe I just don't like the style. All the over explaining, etc.
https://www.imdb.com/title/tt32565993/
sora is done. can you find me one person brave enough to say, "Bill Peebles, a nice guy, doesn't know anything about making movies?" this isn't just a VC problem, this is why it is called the Hollywood No, but that said: maybe the hollywood no is the problem? in video games it has been called Toxic Positivity. are you getting it now?
if you know some one out of the software development/engineering worlds, please forward this to them
u.i. Zugló robot mikor?