Sep 4, 2026
What Is It Like to Work at Figure, the Robotics Company? — A Conversation with Jed Yang (杨佳宁)
#AI 总结
Program: 许华哲Harry (Bilibili channel)
Guest: Jed Yang 杨佳宁 · Figure Researcher (PhD from the University of Michigan, former Meta FAIR intern, SAM 3D contributor)
Duration: 47 minutes · Published: 2026-08 (Bilibili)
S O U R C E
Bilibili interview · 许华哲Harry × Figure researcher: https://www.bilibili.com/video/BV1RS8i66Eng/
This report was transcribed from audio (auto-recognized by whisper) and then organized using the pyramid principle · Timestamps refer to the original audio recording
—— P Y R A M I D S T R U C T U R E ——
Top Conclusion · One Sentence
The abilities behind Figure's "cinematic feel" are real — whole-body control and hardware coordination are genuine barriers; but as humanoid robots move from demo to deployment, the current top priority is reliability, while generalization must wait for data to accumulatePillar 1 · The Valuation Mystery
The $39 billion valuation = the founder's fundraising ability + a strong business team + the humanoid "endgame halo", rather than a pure technology premiumPillar 2 · Authenticity
All videos were completed autonomously at 1x speed, with no faked teleoperation; outsiders see the highest level of performance, while insiders see far more unsolved problemsPillar 3 · The Deployment Bottleneck
At this stage reliability > generalization; generalization is the company-wide top priority, but it is essentially a matter of waiting for data — teleoperation is hard to scale, and Human Data is the direction of focusPillar 4 · The Route Call
The debate over model architectures (fast/slow systems / end-to-end / VLA / world models) doesn't matter — the data recipe is the moat that actually requires accumulation01Who He Is: From FAIR to Figure
This may be the first interview with a Figure employee on the Chinese internet. The guest, Jed Yang (杨佳宁), has been at the company for about 10 months. He previously did his PhD at the University of Michigan (embodied AI lab), interned twice at Meta FAIR where he worked on the SAM 3D project — work that recently received a CVPR 2026 Best Paper Honorable Mention.
Key Takeaway
Moving from "publication-oriented FAIR" to "Figure R&D, where everything serves one product" is a shift from a research platform to a highly integrated startup — this 3D vision researcher chose, upon graduating, to "work on the most fundamental thing, which is the robot itself".
BackgroundTwo FAIR internships: witnessing the tail end of a research golden age
In the summer of 2024, FAIR was still a research super-platform "with publication as the big goal", where star individuals and teams pursued research driven by their own interests; when he went back in the summer of 2025, he could already feel organizational turmoil — small projects being downgraded and priorities reshuffled. Perks were shrinking too (milk-tea delivery and the like had been canceled), but for a PhD student embedded in the lab it was still extremely appealing.
In '24 it was actually still like the FAIR everyone used to think of — people basically just did research, with publication as the big goal; by '25 you could already feel some changes, some organizational turmoil. [00:02:16]
TransitionFrom 3D vision to robotics: 3D sits closer to the core of embodied AI
During his PhD, 杨佳宁 worked on vision-and-language navigation (VLN) and found that the best methods in simulators mostly relied on 3D maps, which led him to conclude that 3D research sits closer to the core of embodied AI. Other PhD students in the same lab mostly approached things from an NLP angle (breaking tasks down with language); he chose to stay closer to the robot itself. When a robotics opportunity came along at graduation, he switched without hesitation.
I felt that studying 3D would be a bit closer to the core of so-called embodied AI... When I graduated and this robotics opportunity came along, I decisively went for the most fundamental thing — the robot itself. [00:06:21]
CultureFigure is more like R&D: one monorepo serving one product
Unlike Meta, where each group has its own infra/data/training code, the whole company at Figure (software + hardware) shares a single monorepo, and all code ultimately serves the same product. In robotics, a highly integrated field, the effort is far more concentrated. For the past six months, 杨佳宁's work has centered on using human data (Human Data) to help robots improve their skills.
Meta FAIR is still a super platform for research... but Figure is a startup after all, and in robotics, a highly integrated field, it feels more like R&D to me — there's both product and research. [00:04:22]
02The Valuation Mystery and the "Making Movies" Skepticism
Figure, founded less than four years ago and valued at $39 billion, has long been teased as "a company that makes movies" because its videos look too cinematic and it almost never demos live at trade shows (CES/RSS/WAIC). How does someone on the inside respond?
Key Takeaway
The high valuation comes from the founder's fundraising ability, a strong business team, and the humanoid "endgame halo" stacked together; and while the videos were indeed shot by a film crew, everything in them was completed autonomously at 1x speed, with no teleoperation — what's truly hard is not the motions themselves but the infrastructure of whole-body control and hardware coordination.
ValuationThree sources of the $39 billion: fundraising skill, a business team, and the endgame halo
杨佳宁 is candid that the valuation logic has three parts: first, the founder is "pretty good at raising money"; second, the company has a very strong business team — Figure may be one of the robotics companies that talks to the most customers, and it cares intensely about how it presents itself (industrial design, demo tone, details like taping over cameras); third, humanoids carry a built-in "endgame halo" — people believe the final form is a robot shaped like a human, and Figure entered early, during a window when US robotics companies were scarce, so the picture it painted was naturally huge.
On one hand, the founder really is pretty good at raising money... On the other hand, I do feel the company has a very strong business team — from the robot's ID design to the tone of the demos, you can sense from all kinds of details that the company cares a lot about its image. [00:11:28]
People probably attach a kind of endgame halo to humanoids of this kind... If a company says it's playing this endgame and started relatively early, it can paint a very big picture. [00:12:39]
AuthenticityThe response to "making movies": every video is done autonomously at 1x
Figure has in recent years indeed used professional film crews to shoot its videos (the equipment and production pipeline are extremely professional) — which is exactly where the "fakeness" impression comes from. But after joining, 杨佳宁 confirmed: everything in the videos is autonomous, with no teleoperation; the package-sorting livestream, loading the dishwasher, hip-bumping the drawer shut, making the bed — all real. Once in the field, he found the motions themselves aren't actually that hard — the truly high barrier is whole-body control, hardware coordination, and reaction speed.
Everything that shows up in the videos really is autonomous — there really is no teleoperation like people suspect... The videos Figure releases have always been 1x and autonomous; there's really not much to doubt there. [00:14:23]
Figure's whole-body control, hardware coordination, and quick reactions — that's where the barrier is genuinely high and hard to achieve. Once that infrastructure exists, learning particular motions actually isn't especially difficult. [00:14:56]
First ImpressionsLesson one on day one: taped-over cameras and a robot patrolling the lobby
On his onsite interview day, the cameras on phones and laptops were all taped over and photography was banned; in the lobby a robot was patrolling, on one side data collectors in headsets and mocap suits were teleoperating, and on the other side robots were working autonomously. Because they're humanoids, the visual impact is especially strong. But soon after joining he saw the other side of the coin: the external demos show the peak level of the controllers and hardware smoothness, while inside you must face the reliability numbers head-on — for a task like loading the dishwasher, how many tries out of ten succeed — still a considerable gap from the ideal target.
The moment you walk in, the cameras on your phone and laptop get taped over and you can't take photos — it's quite confidential... You might see a robot patrolling the lobby, and the only words in my head at the time were: sci-fi. [00:07:34]
What shows up in external demos is one level of quality; inside, you may see many more problems... For example, how far the dishwasher task can generalize, how reliable it is, how many tries out of ten actually succeed — only after you get in do you find there are still considerable challenges. [00:08:40]
CadenceThe shifting center of gravity: hardware reliability → RL whole-body control → data generalization
In its first year or two, Figure focused on raising the level and reliability of the hardware and on swapping the controller from MPC to RL-based whole-body control — which brought a qualitative leap in demo capability. Now that the controller has come up and the data infrastructure exists, the question becomes: how to get robots to actually do tasks that are generalizable, highly reliable, and create real value.
In the first year or two, the focus was raising the level and reliability of the hardware, and switching the controller from the previous MPC to the current RL-based whole-body control — these gave the whole company a qualitative leap. [00:10:00]
SafetyWhy they never go to trade shows: a 70-kilogram legged robot is too risky
One reason is that the company wants to cultivate a "cool" tone, with carefully crafted design; the other, frankly, is that exposing a humanoid too early really does carry risk — it's a legged being weighing roughly 70-some kilograms, and when a motion goes wrong the forces involved are large (there are already cases online of dancing robots injuring bystanders), so the bar for what gets tested and deployed is higher. But investors and customers who visit the company come away feeling "still very blown away".
After all, it's a legged being weighing around 70-some kilograms, and if it falls or its motion goes wrong... when things go wrong, the forces involved are really large, so this does make us set a higher bar for what we're willing to deploy. [00:16:02]
03The Evolving Value of Demos and the 200-Hour Livestream
After 2026, "it can do one thing" or "it can dance" has left audiences numb. What kind of demo still has value? Figure gave its own answer with a logistics livestream originally planned for 7×24 hours and ultimately extended to 200 hours.
Key Takeaway
The bar for demos is rising fast: long-horizon manipulation and dexterous improvisation still persuade, while simple videos keep losing importance; and what was truly impressive about the 200-hour livestream wasn't the level of intelligence but the fluency of motion and the stability of the whole system — a combined test of hardware, battery and charging, scheduling, and every other layer.
The BarA good demo in 2026: ultra-long-horizon + dexterous improvisation
杨佳宁 believes demos are still useful, but the focus has become more nuanced. Two kinds still carry weight: one is ultra-long-horizon manipulation (like the POKE demo from the host 许华哲's team — guaranteeing zero mistakes across a long video is exponentially hard, and he "studied it carefully"); the other is especially dexterous videos with improvised flair (like a peer's recent clothes-folding video that handled never-before-seen states). Relatively simple videos keep losing importance after 2026 — because the field as a whole has gotten better.
Coming into 2026, the bar for demos really does seem to be rising... "It can do one thing" has become a bit numbing, or "it can do some kind of dance" has become a bit numbing. [00:17:12]
If you shoot something very long and have to guarantee there are no mistakes along the way, that's still quite hard, because the difficulty is exponential. [00:17:50]
Livestreamx24 → 200 hours: the hard part isn't intelligence but system stability
From an intelligence standpoint the logistics-sorting task isn't hard; what was most impressive about the livestream was the fluency of the motions and the stability of the system: the robot runs about 3 hours on a charge and must return autonomously to recharge when low — precisely the parts not handled by the hottest embodied-AI models, yet the whole pipeline had to work and the hardware had to be extremely reliable. The livestream was originally planned for 7×24 hours; seeing that "nothing seemed to go wrong", they kept extending it, ultimately reaching 200 hours.
The most impressive thing about this video to me is more the fluency of the motions and the stability of the system... A lot of what's involved isn't even handled by the so-called hottest embodied-AI models, but the whole thing had to work and the hardware had to be very reliable — overall it's quite impressive. [00:19:03]
It was originally a 7×24-hour livestream, but seeing that nothing seemed to go wrong, they kept extending it, eventually to 200 hours. [00:19:57]
ViralityFrank, Gary, Rose: internet savvy that grew out of the livestream
During the livestream, viewers named the robots (Frank, Gary, Rose), joking that "Gary never works quite as well, while Frank does a better job" — and the company later actually put those names on the robots. In terms of publicity, the livestream was a clear success.
You could see viewers in the comments naming the robots... saying Gary never worked quite as well while Frank did a better job — quite internet-savvy — and later they actually put those names on the robots. [00:20:20]
ScopeFigure's scope: solving all physical labor
Home-scenario videos, the factory logistics livestream, the BMW production-line collaboration — Figure's scope spans both home use and factory commercial applications, with the overall goal of solving all physical labor. Home and commercial are two different things: there's a divide at the model level, but a robot is far more than its model, and iteration on the other parts is shared.
I think Figure's scope covers both home use and this kind of factory commercial application, so overall it aims to solve all physical labor. [00:21:15]
04Deployment Core: Reliability First, Generalization Waits for Data
The humanoid has been built and the 200-hour livestream is done — what matters most when it comes time to actually deploy? This is the most technically substantive part of the whole interview — and also where the inside-outside perception gap is largest.
Key Takeaway
If only one could be chosen, at this stage choose reliability: generalization is essentially a matter of "waiting for data", while reliability pays long-term dividends both for research iteration signals and for commercial deployment. At the same time, generalization is the highest company-wide priority (the CEO has set the keyword for this year as generalization), and the biggest budget goes to it; on the data route, teleoperation is very hard to scale, and Human Data is where 杨佳宁 has personally invested the most effort.
PriorityReliability vs. generalization: at this stage, choose the former
Many people would answer generalization, but 杨佳宁's judgment is: pursuing generalization is mainly a matter of waiting for data; pursuing reliability, by contrast, helps over the long run both for the iteration signal your research can get and for producing real utility in commercial scenarios — it's the way to make faster progress at this stage. Of course, you can't chase only this forever.
If I could only pick one, at this stage I'd pick reliability... Chasing generalization, I think, is mostly a matter of waiting for data, but pursuing some reliability will help you a lot over the long run. [00:25:45]
Data VolumeThe real data volume behind the demos: lower double digit
Asked the sensitive question of "the success rate and generalization of tidying the bedroom / loading the dishwasher / cleaning the living room", 杨佳宁 offers a reading common across the industry: if a video doesn't show the numbers, it means the task isn't yet that reliable or generalizable, "but it's also not as bad as people think". The demo tasks he has seen don't require much data — he can't disclose the exact figures, but one can imagine them being "in the lower double digits" (lower double digit, i.e., on the order of 10–30 demonstrations).
If a video doesn't really show the numbers, it certainly means the task is still not that reliable or generalizable... I can't reveal the exact figures, but people can imagine — it's a lower double-digit number. [00:24:46]
GeneralizationThe CEO's framing: the keyword for this year is generalization
Generalization is a very high priority at the company level — the CEO said at an all-hands that the keyword for this year is generalization. The strategy is to accumulate more data and try more recipes. 杨佳宁 believes the biggest money the company is spending now (data, talent, compute) is all going to generalization; the reason less is shown publicly is that a large share of the research push is going into reliability, and Figure's "show it when it's complete" ethos simply takes more time.
Generalization is a very high priority at the company level; I remember the CEO saying at an all-hands that the keyword for this year is generalization. [00:27:48]
I actually believe the biggest money the company is spending now — whether on data, talent, or compute — is all going toward generalization. [00:28:25]
Data RouteTeleoperation is necessary, but not scalable
Comparing the industry's various routes (real-robot data, UMI's embodiment-free approach, the Human Data camp): 杨佳宁 believes teleoperation, as the final stage of a recipe, has fully proven its necessity (no recipe on the market can escape it as the last stage), but at least on a humanoid embodiment it is very hard to scale — teleoperating a humanoid is extremely difficult and the operators are all highly skilled; and it requires at minimum one person, one robot, and one space, with transportation costs, off-site safety requirements (you can't smash the floor when collecting in someone's home), repairs, on-site operations, venue and power, two-person support (one operating, one handling recording/reset)... the total cost is very high. That's exactly why he has put considerable effort into the Human Data direction over the past six months.
I think teleop is still very hard to scale, especially when your embodiment is a humanoid robot — the difficulty of teleoperating it is actually quite high. [00:29:38]
Teleoperation as data collection requires at minimum one person and one robot, plus a space that can fit them, so this single constraint alone can severely limit its eventual scalability. [00:30:38]
05Technical Route Calls: The Architecture Debate Doesn't Matter — Data Is the Moat
Fast/slow systems or fully end-to-end? World models or VLA? How do today's hottest route debates rank in the eyes of a front-line researcher? And how does someone from a 3D background view the relationship between 3D generation and robotics?
Key Takeaway
With coding agents as capable as they are today, swapping in a different model architecture is very easy; what truly requires accumulation is the entire data recipe — and with human data, teleop data, and egocentric video data all still far from sufficient, arguing about architectures isn't urgent right now. 3D generation is already playing a real role in robotics (generating simulator assets), while how to train dexterous hands remains the biggest open question.
ArchitectureThe model-faction debate: it doesn't matter — the data recipe does
Asked "does the model level matter — fast/slow systems, fully end-to-end, world models, or VLA", 杨佳宁's answer is blunt: "I don't think it matters much." The reason is that swapping model architectures is cheap, while the data recipe takes long-term accumulation. Architecture matters a bit more only when data interacts with the recipe — but the current reality is that data is far from sufficient.
I think these factional debates over models — especially now, with coding agents so capable — swapping to a different so-called model architecture is actually very easy; what's harder is your entire data recipe, which takes accumulation. [00:33:29]
Whether it's human data, teleop data, or so-called egocentric video data, none of it is particularly abundant; in that situation, I don't think arguing about these architectures is especially urgent. [00:34:00]
2B/2CCommercial vs. home: the models diverge, the rest is shared
The two routes are neither fully independent nor a simple progression: at the model level there's a divide — for commercial deployment at this stage, the emphasis is on post-training (Behavior Cloning's SFT, on-robot RL, Preference Tuning); for entering homes, the focus shifts to VLA and world models (both of which Figure is exploring) and to making better use of teleoperation data. But a robot is more than its model — hand design problems, next-generation camera placement, simplified charging, multi-robot scheduling, anomalous behaviors, safety of interacting with people-filled environments, worst-case damage control... research on these fronts is shared, and in a relatively more controllable commercial environment it iterates faster and the team grows more quickly.
A robot isn't just the model — the model is only one part, and not even a particularly big part. Iteration on many other parts is shared... Research on these fronts, in a relatively more controllable commercial environment, may iterate faster than for home use and let the team grow more quickly. [00:22:54]
3D3D generation and robotics: already combined in practice, and only more so ahead
3D already has a market as a standalone field, and it genuinely helps robotics — both inside the company and among external collaborators, people are already using open-source 3D models to generate assets. One thing 杨佳宁 can't yet see clearly is how to train dexterous hands: pure Behavior Cloning struggles to learn especially dexterous manipulation, so he's curious whether future dexterous manipulation will rely more on simulation inside simulators — and if so, generating all kinds of assets becomes critically important.
One thing I'm still not sure about is how to train dexterous hands — I think it remains a very big open question... Will more dexterous manipulation in the future rely more on simulation inside simulators? If so, generating all kinds of assets may become very important. [00:35:18]
Signature WorkSAM 3D: click a point in one image to generate a 3D asset — CVPR 2026 Honorable Mention
SAM 3D, which 杨佳宁 contributed to during his internship last year, won a CVPR 2026 Best Paper Honorable Mention (he modestly says he only contributed a small part as an intern). Usage: upload an image to a webpage, click an object, first get a 2D cutout, then generate a 3D mesh or Gaussian splat of that object. The vision is precisely to serve robotics — turning the real world into 3D assets and enriching the variety of assets in simulators. The challenge is in the data (bridging 2D segmentation data with 3D asset libraries like Objaverse); on the modeling side it used cutting-edge structures rarely seen in 3D, such as MoT (Mixer of Transformer), and post-training added DPO preference optimization to make outputs better match human preferences. He also laments that this project became the "final dance" of the research-oriented part of FAIR that ultimately gave its work back to the community as open source.
If you could turn everything in the real world into 3D assets, perhaps you could greatly enrich the variety of assets in your simulator — I think that was the vision of the project at the time. [00:36:56]
It became not only my so-called signature work but also a signature work of that part of Meta that was once FAIR, and it was presented to the community as open source... I think it counts as a final dance. [00:38:30]
06Endgame Imagination and Advice for Young People
When will robots truly enter every household? What should the relationship between home robots and humans be? The interview closes by moving from technology to social imagination and life philosophy.
Key Takeaway
"The arrival of robots is not a technological change but a change in social institutions" — once robots are abundant enough and the prices of goods and services approach zero, values and human relationships will be reshuffled. The home robot's role is the "all-capable appliance", not an emotional companion: it does what people don't want to do, rather than replacing what people do want to do. And the advice for young people: among the things that interest you, pick the hardest one to start with.
Timeline5 to 10 years: it will definitely happen, but the exact timing is hard to predict
杨佳宁's first instinct is "around 5 years", but on reflection — how many problems still need solving in between — the time starts to feel insufficient; meanwhile the rise of coding agents, large companies actively collecting and selling data, and the sudden growth in data abundance could bring faster progress than expected. His overall judgment: it will definitely happen, probably in 5 to 10 years.
What I tell myself is maybe around 5 years, but when you think carefully about how many problems still need solving in between, the time somehow doesn't feel like enough... But what we can all see, I think, is that it will definitely happen — probably within 5 to 10 years. [00:32:21]
SocietyA change in social institutions: an elementary schooler rents a fleet of robots to run a "toy workshop"
The line traces back to when he joined Figure, coinciding with the Figure 03 launch and the article in America's TIME magazine. The underlying logic: once robots are abundant enough, the prices of goods and services will approach zero without limit, and everyone can obtain what they want — values, priorities, and relationships between people will all be reshuffled. He often imagines a scene like this: an elementary schooler with a small idea uses a coding agent to write a program, rents robots in the physical world, and actually builds a small workshop/factory to realize their dream — the space for imagination is enormous.
When robots become extremely abundant and numerous, the prices of goods and services may approach zero without limit... Many people's values and priorities will be reshuffled. [00:40:04]
Maybe he just loves making toys. In the future he wouldn't stop at "I'll buy a toy" — he could just say a sentence, have a coding agent write a program, rent a few robots in the physical world, and actually set up a small workshop to realize his dream. [00:41:06]
PositioningThe home robot = an all-capable appliance, not an emotional companion
In 杨佳宁's imagination, the home robot is more like an all-capable appliance: every chore at home that you can or don't want to do can be handed to it, but he doesn't imagine it becoming an emotional companion. Robots are more about doing the part humans don't want to do, not replacing the part humans do want to do — emotional connection with family is the more human part. As for the flourishes like "making eye contact while placing a plate" or "bumping the oven door shut with its hip", those come mostly from the industrial design team's strong grip on look and feel — research sometimes even spends extra effort to accommodate them — while data collection resources today still prioritize practicality; "getting to that day would be nice enough".
In my imagination, the home robot is still more like an all-capable appliance... I don't really imagine it becoming too much of a so-called emotional companion. [00:42:14]
I think robots are more about doing the part of things humans don't want to do, rather than replacing the part of things humans do want to do. [00:42:45]
AdviceFor young people: among the things that interest you, pick the hardest one
Why pick the hard one? Because young people have little capital, and an adventurous, challenge-seeking spirit is their greatest asset. When you choose to do something especially hard, older people may even help you unconditionally — they can't take that risk themselves (family, all sorts of considerations), and they lack the time capital for risky choices, which makes you a very attractive "investment" for them. On the other hand, the sense of accomplishment and confidence that comes from pulling off something hard also helps enormously in the long run.
Among the things you're interested in, if there are several, pick the hardest one to start with — you often reap some unexpected rewards. [00:44:36]
When you have no capital at all, your adventurous spirit and your willingness to take on challenges are actually your greatest capital... When you, as a young person, choose to do something especially hard, they may even help you or support you unconditionally. [00:45:20]