中文

Wang Kaixuan / 3D Vision & Robotics

Wang Kaixuan Blog

Personal blog on 3D vision, robotics, embodied AI and weekly notes.

Sep 5, 2026

A Conversation with Generalist Researcher Yanwei (Felix) on GEN-1: Judging the Robot GPT-3 Moment, Views on Data, the Harness Layer, and Steerability

#AI 总结

Bilibili Interview · Xu Huazhe (Harry)
A Conversation with Generalist Researcher Yanwei (Felix) on GEN-1: Judging the Robot GPT-3 Moment, Views on Data, the Harness Layer, and Steerability

Program: Xu Huazhe Harry (Bilibili)
Guest: Yanwei (Felix), researcher at Generalist AI, PhD from MIT CSAIL (advisor Julie Shah, on Human-Robot Interaction and robot steerability)
Duration: 34 minutes · Published: 2026-04-25

S O U R C E

Xu Huazhe (Harry) · Bilibili: https://www.bilibili.com/video/BV1eao9BJEJS/

This report was transcribed from audio (whisper auto-recognition) and organized with the Pyramid Principle · Timestamps refer to the original audio

—— P Y R A M I D S T R U C T U R E ——

Top Conclusion · One Sentence

GEN-1's historical significance lies not in making specialized tasks reliable, but in the simultaneous appearance of three sparks: generalism, fast learning, and dexterous improvisation; Generalist's real moat is the trinity of first-rate models + first-rate systems + first-rate people

Pillar 1 · Capability Judgment

The criterion for a robot GPT-3 moment is sparks of genius: the ability to quickly learn many different things (generalism + speed), plus dexterity improvisation — reliability itself is merely specialist level and does not constitute a spark

Pillar 2 · Systems View

GEN-1 is a system, not a single model: on top of 500,000 hours of data there is a harness layer (classical robotics + a low-latency inference stack); a first-rate system plus a first-rate model beats making the model accommodate a second-rate system via domain randomization

Pillar 3 · Interaction View

Steerability (training-free human-robot interaction) is a prerequisite for robots entering the home: there are no failed actions, only actions aligned or misaligned with user intent; and the stronger the model, the easier steering becomes

Pillar 4 · People and Organization

People are the core variable: choosing a team is mainly about the people (research taste that is bold enough and practical minded), and the speed a warm culture brings beats 996; robots do what people don't want to do rather than replacing them — embodied AI is a marathon

01Is GEN-1 the Robot GPT-3 Moment?

On April 3, 2026, Generalist released the embodied foundation model GEN-1 (vision: Scaling Embodied Foundation to Mastery), claiming mastery on simple tasks and demonstrating improvisation, with the public framing of reaching GPT-3 level. Host Xu Huazhe cuts straight to the point: does that characterization hold?

Key Takeaway

In Yanwei's view, GEN-1's impact lies not in making any one thing reliable — that is the specialist level any team can reach by piling up data — but in the simultaneous appearance of three sparks: generalism, learning speed, and dexterous improvisation. That is exactly why he accepts the near-GPT-3-moment framing, though he himself deliberately stays conservative about the analogy.

BenchmarkReliability is not a spark: any team can reach specialist level

The definition of the GPT-3 moment varies by person. If the bar is works-out-of-the-box and does the job directly, today's models haven't fully reached it — they still need post-training optimization; if the bar is after post-training it can reliably complete useful tasks, then it is indeed close to the GPT-3 moment of the day. But simply making one task highly reliable, Google already proved with its table-tennis setup — that is the inevitable result of piling up data and technology to the end, not sparks of genius.

Sparks of genius probably come from being able to learn many things very quickly… (being a specialist) any team can do that, and to me that is probably not sparks of genius. [03:05]

Spark 1Generalism: getting demos goes from very hard to relatively easy

As long as the base model is strong enough and has a strong learning ability, you can quickly get the model to learn a new skill and show it through different demos. Producing demos, once a very hard thing, is now relatively easy — direct evidence of generalism.

Getting demos used to be a very hard thing, but now it is relatively easy, if your base model is fairly strong… the generalism it displayed was very impressive to me. [03:52]

Spark 2Learning speed: 500,000 hours of pre-training + 1 hour of post-training

The second impact comes from how fast it learns. GEN-1 pre-trains on 500,000 hours and reaches very high performance with just one hour of post-training — a leap of orders of magnitude in learning efficiency that Yanwei calls quite a big breakthrough.

There is also how fast it can learn — that also gave me a bit of a shock. [04:07]

Spark 3Dexterity improvisation: like being handed the hands of God

The most striking third point is dexterous improvisation — behaviors that make you stop what you are doing and stand there in shock watching. The early block-stacking video (one human demonstration, the robot imitating directly) is one of the most stunning capabilities Yanwei has seen since joining. Notably, this imitation pipeline is in-context and contains no robot data — it points toward the end goal that anyone can do in-context learning with the robot.

All of a sudden it feels like it has the hands of God… doing things that make it hard to stop what you are doing, leaving you standing there in shock watching its behavior. [04:20]

This is in context, because the end goal we want to achieve is definitely in-context learning — anyone being able, in some way, to use this robot for in-context learning. [05:07]

Open QuestionHarry's probing: how much the pre-training data contains downstream tasks

Xu Huazhe raised the sharpest methodological challenge of the interview: π0.5 fine-tunes well on folding clothes perhaps simply because folding-clothes data was already abundant in pre-training; GEN-1 would be more convincing with a table comparing tasks across never seen / seen a little / seen a lot against the time needed to learn new skills. Yanwei admits he cannot go into detail and concedes the current argument is weak — they can only argue indirectly, by demonstrating very different capabilities, that the likelihood of deliberately picking out these tasks for dedicated training in pre-training is not high; and when tasks number 100~200, such indirect argument becomes very difficult. On open-sourcing: there is some possibility, but no timeline.

This link is probably fairly weak… if we can demonstrate very different capabilities, then the likelihood of (in pre-training) specifically training this becomes not that high. [06:30]

I think there is some possibility (of open-sourcing), but I can't really say when we would do it. [07:07]

02Views on Data: Where 500,000 Hours Come From and What They Mean

From GEN-0 to GEN-1, pre-training data grew from 270,000 hours to 500,000 hours. Is data the only key variable?

Key Takeaway

Data growth is a very important variable, but not the whole story — GEN-1 should be understood as a system: above the model there is the harness layer (see Chapter 3). Data acquisition is a combination of multiple means (including dedicated collection teams), with details kept confidential; commercially the end goal is zero-shot, use-it-as-is, but the deployment pace follows whichever path produces value first, with both tracks advancing in parallel.

Data SourceVariety of means: dedicated collectors, plus tricks that can't be shared

Faced with the triple question — self-collected / purchased / some trick to generate quickly — Yanwei gives only limited information: there is a variety of means to obtain data, one of which is dedicated collection (the collection team shown in the Forbes article's photos); as for the concrete methods to quickly and massively obtain diverse data, I probably can't share too much.

Like many other companies say: whatever ways we can get as much diverse data as possible, we will all seriously consider and use. [18:23]

Scaling LawFrom loss to closed-loop success rate: the Gen0 appendix addendum

Harry asks whether they plotted a data-volume-versus-success-rate curve. After Gen0's release peers did raise this question, and the team later added curves correlating loss with closed-loop success rate in the appendix — roughly the same kind of curve: loss going down moves in the same direction as real closed-loop task success, providing preliminary engineering-level evidence for an embodied Scaling Law.

We later had this appendix… correlating with closed-loop success-rate performance — roughly the same kind of curve. [07:49]

End StateZero-shot out of the box vs. users collecting a little: whichever lands first, do first

The end goal is definitely pick-it-up-and-use-it without worrying about collecting data, just like using Claude / ChatGPT today. But the path toward that goal, and having users collect a small amount of data and quickly fine-tune their own model for deployment, are two parallel tracks: whichever lands faster and produces value faster goes first — while still pushing forward the longer-term goal.

The eventual goal is definitely that everyone can just pick it up and use it, without having to think much about collecting data — just like how we pick up Claude or ChatGPT and can use it directly. [08:31]

03Harness Layer: GEN-1 Is a System, Not Just a Model

A rarely discussed topic that Yanwei calls the company's sore spot: what lies between GEN-0's model and GEN-1's system.

Key Takeaway

GEN-1's leap cannot be attributed to data alone: it went from one model to one system. The base model is essentially a model-free learner; its output must pass through model-based control, a low-latency inference stack, and other harness layers — a great deal of classical robotics work — to truly shine on a hardware platform. Generalist's stance: a first-rate system plus a first-rate model beats making the model accommodate a second-rate system with domain randomization; and classical robotics talent, precisely because no LLM-style wave incubated it, is scarcer than ML talent.

Concept BreakdownThe base model is a model-free learner — many layers short of the physical world

Today's foundation models basically take vision plus action-history input, then predict world changes (with an inverse model) or directly predict actions. For the actions produced by this intelligence core to be realized in the physical world, many more layers are needed: model-based control techniques, a low-latency inference stack — you have a lot and a lot which, in my view, need classical robotics, to let the model's output actually shine on a hardware platform.

We probably see GEN-1 more as a system… above this model there are actually many harness layers — how to use this model, to get even stronger results out of it. [18:42]

Talent ViewClassical robotics is not something smart people need not do — it is scarcer

Against the trend of young people abandoning robotics and returning to early-CV-style pure model training and benchmark chasing, Yanwei pushes back: the LLM wave gave many people strong ML expertise, but no wave ever let people accumulate classical robotics expertise — so talent there is scarcer; that is actually the harder problem.

We think that might be the harder problem… we haven't had a wave that gave many people classical robotics expertise, so talent in that area is probably scarcer. [20:50]

Company StanceFirst-rate system + first-rate model; never make the model accommodate the system

One could argue that the model learns to adapt and uses domain randomization to cover for system deficiencies. But Generalist reasons backward from the goal of making robots generate production value, which requires very reliable system engineering — as long as you can still build a first-rate system, first-rate system plus first-rate model will definitely perform better. A recent MIT paper on control gain is the empirical proof: a tiny difference in one gain parameter leads to vastly different results.

So often, how good the model appears is only a matter of how well the system is built… we have to think of robotics as a system problem; model training is just one part of it. [21:36]

Host's AddendumUMI's zero-latency collection vs. multi-layer latency in the control chain

Xu Huazhe shared the deep resonance of a fellow practitioner: with UMI data collection, the hand moves and it is there — in a sense zero latency; while the real control chain is the multi-layer nesting of AI command → motor → end-effector position. Getting the model to learn from latency-free data and run on a latency-laden system is exactly one of the relatively challenging problems right now.

How to make it learn from something without latency and then run on a system with latency — that is indeed a relatively challenging thing. [22:15]

04Steerability: Robots Entering the Home Need Training-Free Interaction

Yanwei's representative PhD work is Inference-Time Policy Steering. How does that academic thread continue into GEN-1?

Key Takeaway

No matter how well pre-training and post-training go, once a robot enters a home its behavior will never fully align with user intent — so there must be training-free interaction mechanisms that let the robot understand user intent. Steering is a broad, multimodal, abstraction-layered framework (language / physical actions / observation-space annotation); its philosophical basis: all physically feasible and plausible actions are success, the only difference being whether they align. And Generalist found internally: the stronger the base model, the easier steering actually is.

AnalogyGoogle Maps: the user only chooses at a high level

Yanwei uses driving navigation as the analogy for his ideal human-robot interaction: the map plans several routes, the user only picks the left or right lane at a high level, and if they detour mid-way the map replans automatically — the user need not care how each route is driven. His research: given a set of base skills for the robot, how to let users easily switch, correct, and steer those skills, even guide it toward skills it rarely did before.

I don't really need to care how each specific route goes; I just make a choice at a fairly high level… this kind of user-machine interaction experience is what I very much hope to realize on robots. [12:51]

FrameworkMultimodal steering: different abstraction levels elicit different behaviors

Steering is not just physical steering. It can be language (best for high-level tasks like wash the dishes for me), annotation of the observation space (more useful than language when specifying how to tidy the kitchen), or physical augmentation / nudging (most effective when specifying how to wash the dishes). Steering signals of different modalities sit at different abstraction levels, each good at eliciting behaviors at a different level.

Steering signals of different modalities actually have different abstractions; they may be good at eliciting different behaviors at different abstraction levels. [14:06]

NecessityThe capability is already inside the model; what is missing is the guidance to elicit it

Many models actually have lots of capability and are very eager; they just lack some guidance to elicit the capability deep inside. Like buying a car — no need to train a model first, just drive — training-free interaction is a crucial step for robots entering application scenarios. Yanwei says he would of course very much want to add it (bringing steering into Generalist's products), with timing depending on base model capability.

If we want to bring robots into everyone's home, I think a training-free interaction modality or mechanism is necessary. [17:08]

PhilosophyNo failed actions: dropping a cup is not a mistake, it just does not align

To Yanwei, the standard evaluation notion of robot-drops-the-cup-equals-failure deserves re-examination: if the user is a mischievous kid or a cat, dropping the cup on purpose is precisely the correct behavior. All physically feasible and plausible actions are success; the only difference is whether they align with user intent. When the richness of physical behaviors is sufficient, you can always find the one intent that satisfies the need.

All physically feasible and plausible motions are, in my view, success — they are just either aligning or not aligning. [15:46]

Internal FindingThe stronger the model, the easier steering gets

During his PhD there was no strong foundation model available, so steering was rather hard; Generalist's internal finding is that as base model capability keeps improving, steering is actually not as hard as imagined. That also answers why steering research matters in this foundation-model era — it is a capability that rises with the models.

As your base model's capability gets better and better, steering actually is not as hard as imagined. [17:31]

05Strategy and Organization: Banyan Trees, Hardware, and Following the People

A team of around 30 powered both model leaps from GEN-0 to GEN-1, earning it the label of the most people-efficient company in embodied AI. How was the niche chosen? Do they build hardware? Why do top researchers choose to join?

Key Takeaway

The banyan tree is the metaphor for the future ecosystem: growing many verticals (rooted aerial branches) is to obtain diversified data that feeds back into the base model trunk; what gets delivered in the end is more likely a capability than a bare model license. On hardware, they stick with mature commercial robot arms, saving the small team's energy for intelligence problems. And Yanwei's answer on career choice is highly consistent: mainly it is the people — research taste that is both bold and practical read from papers, and a warm culture that is not a perk but a speed advantage.

NicheThe banyan metaphor: vertical deployments are aerial roots, the base model is the trunk

The banyan-tree lesson from childhood became Yanwei's image of the company's future: a banyan has a very large trunk; as it grows, leaves spread outward and aerial roots drop down. Doing many verticals is, first, for a diversified dataset — only with data across different verticals can the model become well-rounded, and the base model gets stronger and stronger. As for whether they end up licensing models or delivering a capability, it is still being explored; handing over just a model without considering how users derive value from it is not good for model training either.

We do a lot of these verticals because we need a diversified dataset… my imagining is that only when you have data across different verticals can we make the model more worldly-wise. [09:26]

HardwareNot building hardware: mature robot arms still hold vast untapped potential

Many existing robot arms work well, are already products, and have passed all kinds of tests and certifications — using this tested mature hardware as much as possible saves energy for model development and the intelligence problems they find more interesting. It is a small team's trade-off: no time to solve every problem; but if one day hardware becomes the bottleneck blocking the intelligence problems, they would seriously consider building their own.

If one day this stops us from properly tackling the intelligence battles we want to solve, I think we would seriously consider doing hardware ourselves. [10:46]

Career LogicMainly it is the people: taste read from papers, boldness read from debate

MIT graduate Yanwei turned down faculty offers, big companies (Google / Gemini Robotics / OpenAI), and other startups (Figure, 1X). His selection logic: you can read an author's attitude and research taste from a single paper — Pete Florence's Dense Object Nets is very practical minded while asking the right questions; Andy Zeng's TossingBot is a superb system whose simplicity bias made him really love it; the two's later Transporter Nets he read over and over. Moreover, a controversial question Pete raised at a workshop — should roboticists worldwide all drop what they are doing and spend a year collecting data together, contributing more to robotics? — showed Yanwei the founding team's bold-enough, crazy-enough DNA. For a team whose CEO is not a researcher, he looks at the researchers below: early in your career you want to learn technical problems, not how to manage people.

You can read a person's attitude and research taste from a single paper… ok these are the people I really want to work with. [23:30]

CultureWarmth is not a perk, it is speed: an efficiency source that beats 996

Facing the challenge that Silicon Valley grinds just as hard and Tesla / xAI do very well too, Yanwei argues at the organizational level: a warm atmosphere makes people not want to leave, very loyal, with no energy spent on politics — all time goes into research; less internal friction, more efficiency; candid discussion cuts the decision cost of communication. The speed it brings may be faster than what 996 or 247 brings, because you don't keep switching directions. Everyone on the team is a one-hop connection from existing members; they know each other and are very nice.

The speed it brings may be faster than what 996 or 247 brings, because you don't keep switching directions, endlessly trying left then right. [27:00]

Screaming MomentsStacking blocks, fixing a vacuum, taking money from a wallet

The most memorable moments of GEN-1 development: giving the block-stacking task a new configuration and the robot imitating along — for me it was a screaming moment; we were all screaming there; the recover behavior in the just-released vacuum-repair demo the team could watch all day; plus the dexterity of taking money out of and putting money into a wallet. Every day there are screams, everyone gathered around the robot watching.

For me this was a screaming moment; we were all screaming there… every day it is just screams. [11:42]

06Human-Robot Relations and Life Philosophy: Liberation, Not Replacement; Slow Down to Go Fast

Xu Huazhe dug a trap: you care about human connection, but the enormous social change brought by general-purpose robot brains might strip the blue-collar version of you of the happy life of assembling phone boxes with coworkers.

Key Takeaway

Yanwei's answer folds human-robot relations and life philosophy into one thread: robots do what people don't want to do rather than replacing anyone; humans stay present (issuing commands, supervising intent), and new jobs will emerge the way social media created influencers; the happiness of helping others is exactly the origin of the make-robots-useful mission. And his advice to everyone is a hiking lesson: if you wanna go fast, you have to slow down — embodied AI is a marathon.

Division of LaborHumans do what humans are good at; robots do what humans are not good at and do not want to do

Pressed on replacing blue-collar jobs, Yanwei is clear: the robots he wants to build help people do the things they don't want to do, not replace any single person. Along the way humans still need to be there — robots need commands, need to know what we want to do; humans have a big role in this human-robot collaboration. Robots entering daily life will also create new jobs the way mobile devices + Facebook created influencers, and free factory friends from tedious, even physically harmful repetitive labor to do what they truly want to do.

The future I imagine: humans do what humans are good at, and robots do what humans are not good at and do not want to do. [30:03]

Original MotivationThe happiness of helping others = the motive for making robots useful

Why is people so important in all his decisions? Because so many people helped him along the way, later he helped others too, and helping others gives him the most happiness — exactly why he wants robots to be useful and to help those who need help. As for the final output: all our technical outputs, all our papers may not matter; what matters is this companionship of us walking the road together. At his thesis defense, the whole team flew from Seattle to Boston to be there — that moment confirmed these are the people I want to spend some time with.

I found that when I help others I get the most happiness — that is probably why I want robots to be useful, so they can help the people who need help. [28:09]

Life AdviceIf you wanna go fast, you have to slow down

Advice for younger people: to go fast, slow down first. Life today is fast and noisy — this company has progress, that company has progress — if you don't slow down you cannot hear your inner longing; your inner longing tells you what you want to experience at this stage of life, and even if it looks like a detour on the research road, after experiencing it you will find it a precious asset. His own hiking story is the evidence: rushing to make the semester start and pushing hard, he got injured and returned to school later; walking at an even, slow pace, his body recovered just right, no injury, and the trip ended sooner. Don't fear falling behind; only by slowing down do you learn what you truly want to do — eventually you will catch up.

If we can keep an even pace or walk slowly, quietly settling down to take it in, your body may recover just right, you won't get hurt, and you can actually finish your journey sooner. [32:19]

Life Beyond WorkTheater, hiking, skiing, and an adventure for every stage of life

Theater was his passion in undergrad (too time-consuming, no time now); hiking was his hobby during the PhD (walking more when time allowed). These hobbies are like different phases of life: once experienced, you go find the next adventure — and now the biggest adventure is the robot capability he is researching; the happiness and joy it brings, just like theater did for him in undergrad and hiking did during the PhD, even more. Outside work he still runs, skis, and hikes: work made him realize this is a marathon and he needs work-rest balance.

For me now, the biggest adventure is the robot capability I am researching… the happiness and joy it brings me is very similar — even more. [31:19]