[ Research ]

Dyna-2.1: A Physical Agent for End-to-End Workflows

To be useful, a general-purpose robot needs to take over entire workflows without frequent human intervention. Dyna-2.1 introduces our physical agent, a semi-humanoid robot named Taku with a learning system that leverages human experience for enhanced teachability.

Category:

Research

Author:

Dyna Robotics

Date:

September 29, 2026

Read:

19 min

Dyna-2.1: The first physical agent to automate an end-to-end workflow—one hour, uncut.

Listen to our team discussing the breakthroughs.

1. The value of a general-purpose robot lives in workflows, not tasks

Robot foundation models have now reached production-level performance on stationary tasks, as our DYNA-1, DYNA-2 and deployment releases show [1, 2, 3]. Despite promising progress from ourselves as well as others, the same level of performance and reliability has remained elusive for long-running workflows. By a workflow we mean a self-contained business process that allows infrequent, asynchronous handoff. A person hands it over, walks away, and checks back hours later. A robot that works this way needs little oversight.
Why should we care about workflows when stationary tasks are already working? The fundamental limitation of stationary tasks as a product is that they still require a robot babysitter in most cases, as individual stationary tasks often only make up small sub-steps in a larger workflow. As an example, DYNA-1 folded napkins for 24 hours without an intervention on the napkin folding task itself [1], but a person had to keep refilling its napkin bin and clearing its stacks. Without that person, the workflow halts when the bin is empty. Backed by commercial evidence we have gathered from talking to many customers, we have come to the realization that customers pay for a role not a task: a robot takes over a whole workflow for a shift without needing a second person standing by supporting it. In other words, people are looking for a robot that can become a whole employee.
Workflows, however, require several necessary capability unlocks across the full stack of robotics. To just name a few:

Expanding the robot’s workspace, because workflows are done by humans today. The robot has to cover the whole workspace of the role it fills, using base, torso and arms together at every station.

Teachability, because each customer’s workflow is different, changes over time and chains hundreds of subtasks that each need more 9’s. The robot has to learn new skills fast and make them reliable, from human data, simulation and the corrections it collects on the job.

Reasoning, because workflows are non-linear and require conditional decision-making.

Unlocking these capabilities altogether is no longer feasible for a single action model deployed on a limiting hardware platform. Instead, we must build a full-stack physical agent all the way from careful hardware design to robust policy learning.

2. Dyna-2.1: Introducing the Physical Agent and the Taku Robot

Today, we introduce Dyna-2.1, a full-stack physical agent that can autonomously complete end-to-end real-world workflows. A central part of our physical agent system is our brand-new semi-humanoid robot, Taku, which comes from the Japanese word Takumi (匠), meaning master craftsman.
While Dyna-2.1 is compatible with any loco-dexterous manipulation workflow (see server servicing and retrieving a drink), we will use hotel laundry room workflow as the running example in this blog, as commercial laundry has been one of the deployment verticals that we have been focusing on. An uncut 1-hour long full laundry-room workflow autonomously executed by Dyna-2.1 is shown at the top of this post. To the best of our knowledge, Dyna-2.1 is the first physical agent to perform hour-long, non-linear loco-dexterous manipulation workflows.

2.1. Meet Taku

Adorable, capable and built for real work, Taku joins the Dyna family as a one-of-a-kind physical agent. Taku has a human-shaped body above the waist, a folding lower body to reach low and high, an incredibly stable base on four steerable wheels, and two 7-degree-of-freedom arms.

Video 2.1: introducing Taku.

In the workflows we target, most steps need reach and manipulation precision rather than legs, so Taku moves between stations on wheels. Above the wheels, we sized Taku like an average person, so a person’s tracked wrists, elbows and chest land near poses Taku can reach. To match the speed of human motion, Taku’s arms use low-ratio planetary actuators, which allow higher speeds and accelerations than the harmonic drives in conventional arms.

Video 2.2: whole-body manipulation, across the workspace. Clockwise from top left: reaching deep into a dryer, picking up a towel, opening a refrigerator, and placing a stack on a low shelf.

The laundry room shows what this body is for. The washers, dryers, folding table and shelves are separate stations meters apart, and the work happens at every height: leaning into a dryer drum about an arm’s length deep, reaching into the washer for a leftover towel, turning at the waist with a load, and crouching to the bottom shelf or reaching one above head height. This requires loco-dexterous manipulation, coordinated whole-body movement and dexterous manipulation, which is challenging for a table-mounted bimanual arm or an elevator-style mobile manipulator. Taku’s hardware puts every laundry station within reach (Video 2.2), which delivers the first unlock from Section 1. The controller coordinates the body to work at them (Section 3.1.2).

2.2. The Physical Agent

Reach is only the start. To complete a shift, the robot must execute physical skills reliably for hours and decide what to do next as the workflow unfolds. This is especially challenging when the horizon of the workflow is on the order of hours. The laundry workflow shows why both matter. Beyond reaching every station, it poses two problems that a single task does not:
1. It chains hundreds of subtasks over a shift. Over an hour, small things go wrong many times. A grasp catches two towels, a corner folds under, a stack leans. Some mistakes undo earlier work: a dropped towel must go back in the wash, and a dropped stack sends every towel in it back to the start. We chose laundry partly because most of its failures are recoverable. But each recovery costs time, and every step must still avoid unrecoverable errors, so per-step reliability compounds. If each step is only 95% reliable, a cycle almost never finishes without help (Figure 2.2).

[ WORKFLOW RELIABILITY ]

[ ONE-CYCLE SUCCESS ]1.7%Chance that all ≈79 steps in one cycle succeedFAILS MORE OFTEN THAN IT SUCCEEDS[ HUMAN INTERVENTIONS / 24 H ]190≈48 cycles a day at ≈30 min each · one about every 8 min[ SUBTASK SUCCESS RATE ]SET BY THE SLIDER · ILLUSTRATIVE, NOT MEASUREDTowel pick& place≈52 PER CYCLE≈40 machine transfers+ ≈12 table picks95.0%Fold &stack≈12 PER CYCLEflatten, fold,add to the stack95.0%Transferstack to shelf≈3 PER CYCLEa whole stackat a time95.0%Open & closea machine door≈8 PER CYCLEwasher, dryer95.0%Start amachine≈2 PER CYCLEbuttons and dial95.0%Move thebasket≈2 PER CYCLEtable → dryerand back95.0%[ ASSUMPTIONS ]Subtasks are treated as independent. In practice one failure tends to set up the next, so real numbers are worse.A subtask fails only when a person has to step in. One that succeeds after a retry counts as a success.
95.0%log scale · 80% to 99.99%

Figure 2.2. Workflow reliability. The chance that a whole cycle finishes without help, for a given per-subtask success rate. Drag the slider; subtask counts are estimates.

2. It is non-linear. Towels follow a linear progression from wash to dry to fold to stack, but the robot does not (Figure 2.3). Each machine finishes on its own schedule, and a finished machine left idle is lost capacity [4]. So folding has to be interruptible. The robot monitors the machines as it folds and attends to each one once it finishes. We break a cycle into thirteen decision points. Several depend on information the robot saw minutes or hours ago but cannot see now: when the washer started, which load is in which machine and which shelf has room.

[ LAUNDRY ROOM WORKFLOW ]

[ C · UNLOAD THE DRYER ] DRYER DONE ↻ LOOP UNTIL DRYER EMPTY [ B · WASHER TO DRYER TRANSFER ] WASHER DONE, DRYER EMPTY ↻ LOOP UNTIL WASHER EMPTY [ A · LOAD THE WASHER ] WASHER EMPTY ↻ LOOP UNTIL BASKET EMPTY Open washer door Take a dirty towel from the dirty basket Dropped? Put it in the washer Basket empty? Close washer door Start the washer buttons and dial NO YES NO · NEXT TOWEL Off the floor → washer YES Open both doors washer and dryer Take a wet towel from the washer Dropped? Put it in the dryer Washer empty? Close both doors Start the dryer buttons and dial NO YES NO · NEXT TOWEL Dirty now → dirty basket YES Bring the basket folding table → dryer Open dryer door Take a towel from the dryer Dropped? Put it in the basket Dryer empty? Close dryer door Push the basket back dryer → folding table NO YES NO · NEXT TOWEL Dirty now → dirty basket YES [ D · FOLD AND STACK ] BOTH RUNNING ↻ STACK LOOP UNTIL BASKET EMPTY ↻ TOWEL LOOP UNTIL STACK TALL Machine waiting? Pick up a towel Dropped? Flatten & fold Add to the stack Machine waiting? Stack count reached? Transfer stack to shelf NO NO NO YES NO NEXT TOWEL NEXT STACK Dirty now → dirty basket YES START OF SHIFT ↓ [ CHECK THE MACHINES ] Check washer and dryer Washer empty? Washer done and dryer empty? Dryer done? Both running: fold NO NO NO YES YES YES WASHER RUNNING YES YES BACK TO THE MACHINES BACK TO THE MACHINES Dirty basket the day's laundry, plus every drop FEEDS THE NEXT WASHER LOAD · A DROPPED = DIRTY [ KEY ] hand-off to and from the hub dropped, so dirty loop · until its condition

Figure 2.3. Laundry Room Workflow. One laundry cycle: thirteen decisions the robot makes, between long runs of loco-dexterous manipulation. The orchestrator is instructed to follow these decisions (Section 3.2).

Meeting these challenges requires designing the robot’s brain as a whole. We divide responsibilities by the timescale of each decision and by how each capability can be learned most efficiently, so precise control, reliable skills and workflow reasoning work together. Three model layers control Taku (Figure 2.4); Section 3 explains the design of each layer:

A whole-body controller, trained with reinforcement learning in simulation, turns task-space target trajectories for the wrists, elbows, chest and footprint (the Unified Robot Representation, URR; Section 3.1.1) into joint targets and wheel velocities at 100 Hz.

The DYNA-2 policy, an improved version of our DYNA-2 world-action model [2], turns the current step into whole-body target trajectories to be tracked by the whole-body controller.

A workflow orchestrator, a vision-language model, tracks the workflow, decides the next step and steers DYNA-2 to carry it out.

[ INGREDIENTS ]

DYNA-2 policy world-action model · ~5 Hz §3.1 · 3.2 Taku · semi-humanoid body 21-DoF upper body + mobile base §2.1 · 3.1 JOINT & WHEEL COMMANDS SENSING Whole-body controller RL · 100 Hz §2.2 · 3.1 URR TRAJECTORIES · POLICY OR TELEOP THE NEXT STEP Workflow orchestrator vision-language model §3.1 · 3.2 BUILT TO LEARN FROM PEOPLE human bodies, human motion, human video, human reasoning

Figure 2.4. Ingredients of a physical agent. Taku and the three model layers, each on its own clock. A person can produce URR target trajectories too, so human data enters where the policy’s commands do.

3. Building the Physical Agent: Contributions Across the Full Stack

Taku’s hardware puts the work within reach. The next two sections address the problems from Section 2.2: Section 3.1 shows how we teach physical skills reliable enough to chain, at low cost, and Section 3.2 how the robot decides what to do next.

3.1. Teachability: Each Skill from the Cheapest Data That Teaches It

Teachability is how cheaply a new skill, or a new customer’s procedure, reaches production reliability. Robot data is the scarcest input. It costs robot and operator hours, and each hardware revision makes previously collected data less useful, while the experience people record keeps growing.
Our deployments already show what teachability is worth. In our stationary folding deployments, a new site now reaches its production bar in as little as three days, down from weeks to months of on-site engineering [3, 5]. Dyna-2.1 takes this further: it learns new skills much faster than our previous systems, with far less demonstration data. Servicing a server and retrieving a drink are two recent examples.

Learning a new skill: server servicing.

Learning a new skill: retrieving a drink.

We made three design choices to make Dyna-2.1 more teachable:

Make human data usable for robots (Section 3.1.1): a shared representation lets human recordings train the robot.

Training controls that maximize teachability (Section 3.1.2): the controller gets better with every demonstration we collect, and a better controller records better data for the policy.

Keep old robot data usable (Section 3.1.3): a hardware revision retrains the controller, not the policy.

Section 3.1.4 shows what they add up to: mastery of skills, mastery of recovery.

3.1.1 Make human data usable for robots

Human experience data is growing quickly. Data collection firms now record people doing physical work at scale, as egocentric video, motion capture and handheld-gripper demonstrations [6]. We built Taku to follow human motion (Section 2.1). To learn from these recordings, our models also need one data interface for human and robot motion.
One data interface for human and robot motion. We call this interface the Unified Robot Representation (URR). URR describes a body by its wrist, elbow, chest and footprint poses in a locally consistent coordinate frame (Figure 3.1). Both people and robots can produce these poses, and a sequence of them describes motion over time. URR has three properties:

Embodiment-agnostic. The same format describes a human, humanoid, semi-humanoid or tabletop bimanual robot, using more or fewer body components according to what is present or observed.

Captures key task-space constraints. These poses specify hand placement, arm configuration, torso posture and the body’s placement in the workspace. These are the spatial constraints that matter most for coordinated manipulation.

Lossless across equivalent action formats. The represented component poses can be recovered from equivalent formats, such as relative-pose actions, provided the reference frame and starting pose are retained.

[ URR TRACKING ]

Figure 3.1. URR tracking. One teleoperation episode, replayed from the robot's log. The translucent target is the URR command: grippers at the hand targets, the torso at the chest target, the base at the footprint target and elbows at the elbow targets. The solid robot is the measured joint state and odometry as the whole-body controller tracks that command. The robot starts hidden; toggle each layer and drag to orbit.

For training, we convert teleoperation, egocentric recordings and UMI data [7] into URR and encode them into the model’s action representation; at inference, we decode the model’s actions back into URR target trajectories for the controller (Figure 3.2). Converting human motion into URR is a retargeting step that stays in task space. Most whole-body teleoperation stacks instead fit human motion to robot joints before a tracking controller runs [8, 9]. For learning from video, that fitting can pull a hand away from the object it touched, so the video and the action labels disagree [10]. URR keeps the demonstrated task-space motion as the target and leaves joint coordination to the controller.
Teleoperation also captures the robot’s own reach and limits, and lets a person take over to record corrections. Prior work has paired task-space policies with learned whole-body controllers [11]. We add one interface shared by all these sources, and the same interface lets a person and the model hand control back and forth.

[ THE DATA LIFECYCLE ]

Teleoperationon the robotOff-robot dataegocentric + UMIURRURRURRUnified Robot RepresentationBody-part poses in one locally consistent frame, over timewrists · elbows · chest · footprintSHARED BY PEOPLE AND THE ROBOTDYNA-2 policyworld-action modelDecode actionEncode actionURRURRMODEL INFERENCEMODEL TRAININGURRRL whole-body controllerJOINT TARGETS + WHEEL VELOCITIESRobot hardware

Figure 3.2. The data lifecycle around URR. Every data source enters as URR for training; at inference, model actions decode to URR target trajectories for the controller.

Both the action model and the RL controller train on human motion in URR. We pre-train the action model, an improved DYNA-2 policy, on one million hours of human video mixed with robot data from our fleet, using whole-body poses tracked from the video in URR as targets. Those targets are already in the format the policy outputs on the robot, so fine-tuning on robot data only has to teach the policy Taku’s specific body. The human video also shows the policy examples of towels, machines, and rooms our robot data never covered. Human pre-training already pays off without it: DYNA-2, pre-trained on human video alone, passed 87% of customer acceptance tests zero-shot on a stationary robot at sites it had seen no data from, versus 46% for DYNA-1 [2].
The controller (Section 3.1.2) reuses these motion recordings as tracking targets in simulation. We do not collect a separate set of demonstrations for control. Prior work showed that tracking human motion-capture data teaches humanoids natural whole-body motion [12], and a person’s reach, lean and turn carry over to a wheeled base. One human demonstration therefore teaches the policy a task and also serves as training curriculum for the controller.

Learning from human data. Taku reproducing egocentric human demonstrations in simulation, in real time. Left: Taku; right: the person’s head camera. Each recording is converted to URR commands and tracked by the RL whole-body controller.

3.1.2 Training controls that maximize teachability

The controller and the task form a virtuous cycle: more demonstrations make a better controller, and a better controller records better demonstrations.
Seed reinforcement learning with human motion. The controller learns balance, posture and motion tracking with reinforcement learning in simulation, where trials are cheap (Video 3.3). Rather than overly relying on exploration and manual reward tuning, we put our massive reservoir of URR data to use, specifically as pose targets alongside synthetic targets to inject natural human pose priors into control learning. Its targets are recordings, so the controller’s training set grows with every demonstration we collect, and each new task or operator adds poses it learns to follow. Empirically, we observed that the growing human prior not only induces faster convergence, but also increases the quality of our controls in offline metrics and teleoperation tests. The controller tracks the elbows and chest as well as the hands, so on a tall or low shelf the forearm stays clear of the shelf edge or machine door. Unlike inverse kinematics, it is dynamics-aware, tolerates infeasible commands and encodes priorities in reward weights, such as hand accuracy over chest accuracy. RL whole-body control is well studied [12–16], but mostly for legged locomotion with imprecise arms; we need precise manipulation on hardware we aim to iterate on often. We randomize link masses, joint friction and damping, and actuator gains and delays to narrow the gap to hardware.

Video 3.3: the controller, trained at scale. Thousands of simulated Taku robots training in parallel in NVIDIA Isaac Sim.

Tune the controller for imitation. The loop closes through the policy. We record every demonstration through the controller, so a better controller produces smoother, more responsive, higher-quality motion for the policy to learn from. Two responses with the same total tracking error can behave very differently: one overshoots and oscillates, the other approaches smoothly (Figure 3.4). We found oscillations consistently harmful to policy learning, so no oscillation is a hard requirement, consistent with Tune to Learn [17], which found that behavior cloning benefits from overdamped gains because damping dissipates the policy’s small action errors before they become large deviations in the robot’s motion.

Figure 3.4. Overdamped versus underdamped. A 15 cm hand step with equal total tracking error: the underdamped response overshoots and rings, the overdamped one never passes the target. The responses are constructed.

3.1.3 Keep old robot data usable

Teaching compounds only if earlier data keeps counting, so robot data has to outlive the hardware and controller it was recorded on. This matters especially because we treat hardware research as a continuous process of iteration. We leave joint-level motion and limits to the controller, so a hardware revision means retraining the controller in simulation, with no robot time, rather than changing the policy’s action space.
A stable action space does not eliminate differences in control dynamics. Every demonstration carries the dynamics of the controller that recorded it. Our robot data spans several generations of hardware and controllers, and the same command produces different motion through an RL controller and an inverse-kinematics baseline. We account for these differences in two ways:

Explicit controller conditioning. We build on DYNA-2’s language steerability [2] to teach the policy each controller’s characteristics. A metadata prompt tells the DYNA-2 policy which controller each robot demonstration came through, so one checkpoint learns how each controller moves and can be steered to the one used at deployment.

Implicit context through asynchronous inference. We also condition the model on future commands during asynchronous inference. This command context provides implicit information about the controller’s dynamics, complementing the explicit controller metadata.

3.1.4 The result: general physical capability and recovery

Cheap teaching pays off in two ways. The first is general capability. The same teaching recipe carries across very different tasks, each with its own demands on whole-body dexterity and precision (Video 3.5).

Video 3.5: physical skills across tasks. Clockwise from top left: pressing a button, placing a stack on a high shelf, inserting a server, and picking a laundry pod. These examples show dexterity and precision across different manipulation tasks.

The second is reliability. Figure 2.2 shows how unforgiving a workflow is: a laundry cycle chains about 79 steps, and at 95% per step it almost never finishes without help. No policy gets every step right, so what closes the gap is recovery, and we have to teach recovery step by step. Sections 3.1.1 to 3.1.3 are what make that affordable. Video 3.6 shows Taku recovering autonomously from outside interference and from its own mistakes.

Video 3.6: robustness when manipulation does not go to plan. Both clips autonomous, at 1× speed. In one, while Taku turns on the washer, a person perturbs the dials and pulls Taku’s gripper off the buttons; Taku returns to the panel and finishes setting the correct cycle. In the other, Taku recovers from a double-towel scenario.

3.2. Reasoning: An Orchestrator that Runs the Workflow

Ling Li, Ming Qin, and Chet Bhateja discuss the design of reasoning and language following, interwoven with footage of the reasoning system guiding the robot through the workflow.

We use a vision-language orchestrator to make Figure 2.3’s decisions and steer the DYNA-2 policy to carry each one out. Many of these decisions depend on what it observed minutes or hours earlier.

3.2.1 Built on a vision-language model

Our orchestrator is built on a vision-language model pre-trained on diverse text and images, so it starts with a general sense of what a laundry room is for, and we post-train it on this workflow’s branches (Figure 2.3). People can also give it short instructions, which it expands into steps the policy can carry out, so some changes to a customer’s procedure need no new demonstrations. Our orchestrator runs at a lower frequency while the policy keeps carrying out the current step, so we could swap in a larger, slower vision-language model without changing the layers below. Its decisions can take longer than the policy’s, so each one can use more context, such as its text memory and people’s instructions, and more time to reason.

3.2.2 Reason about the environment and choose the next step

The orchestrator continuously monitors the room and the task in progress. At each branch of Figure 2.3, such as which machine to load or unload or whether a stack has enough towels to shelve, the orchestrator chooses the next step and steers the policy to carry it out (Video 3.7).

Video 3.7: from folding to dryer unloading. Taku pauses folding to attend to the dryer, adapting its next step as the workflow changes.

We leave the checks inside a fold to the policy. It prepares the next task before the current one ends, and confirms each step is complete before moving on (Figure 3.8). When a check fails, it decides what comes next: repeat the step, try it another way, or change the plan. The policy still carries out the recovery steps (Section 3.1.4).

[ THE ORCHESTRATOR ]

Camera + text memory what it sees, what it knows Orchestrator vision-language model Next step the correct next action DYNA-2 policy carries out the step WRITES PROGRESS BACK TO MEMORY [ ONE STRETCH OF A SHIFT ] ILLUSTRATIVE SEES REMEMBERS DECIDES T+0 towel on the table, dryer running washer done · dryer 18 min left current stack: 3 of 5 washer must wait for the dryer: keep folding T+18 MIN dryer has finished dryer done · washer done basket empty finish this fold first, then unload the dryer T+21 MIN dryer drum empty washer done · dryer empty current stack: 4 of 5 move the washer load into the dryer, start it T+24 MIN dryer running, towel on the table dryer 45 min left current stack: 4 of 5 machines need nothing: resume folding at towel 5 T+31 MIN the stack reaches five top shelf full stack is tall enough: shelve it on the middle shelf

Figure 3.8. The orchestrator at work. Top: the loop. Bottom: an illustrative stretch of a shift, one decision per row.

3.2.3 Long-term memory

The controller and the DYNA-2 policy learn general skills, but no dataset holds what has happened in this room today. The orchestrator tracks progress visually and compresses it into a long-term memory kept as text: which steps are done, whether each washer and dryer door is open or closed, how many towels have been folded, which machine is running. The orchestrator records how many towels Taku has folded, so Taku picks up folding where it left off after its washer and dryer tasks. It turns that history into the next step.

4. Looking Forward

We have introduced Dyna-2.1, our first physical agent capable of completing hour-long end-to-end workflows. We optimized Dyna-2.1 across hardware, controller, and hierarchical policy design together to achieve the current results. Our next milestone is to take this brand new system to real-world customer sites and establish a deployment flywheel, in which every piece of deployment data improves the full system. Our long term vision is to build a self-improving physical agent that can both absorb large amounts of diverse data and improve itself using its own experiences.
As our models and agents become more capable, the standard evaluation metrics of robot learning stop measuring what matters. Per-episode success rate on a benchmark task was informative when a model was just barely working. However, it says nothing about whether the same model can be left alone with a shift. And we should measure how long a workflow a model can handle without any intervention. There is a close analogue in the world of language agents. In March 2025, METR published Measuring AI Ability to Complete Long Tasks [18], which argued that benchmark accuracy had stopped being informative about what AI systems could actually do, and proposed a replacement: the 50% task-completion time horizon, the length of a task, measured in how long it takes a human expert, at which a model succeeds half the time. For physical AI, Dyna-2.1 is just the beginning.

References

01

Dyna Robotics. Dynamism v1 (DYNA-1) Model: A Breakthrough in Performance and Production-Ready Embodied AI. dyna.co/research/dyna-1, June 2025.

02

Dyna Robotics. Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co/dyna-2, August 2026.

03

Dyna Robotics. Not Just a Model, But a Product. dyna.co/research/scaling-customer-deployments, August 2026.

04

Hietpas. Hotel Laundry Can Maximize Effectiveness with Right Equipment Mix (Part 1 of 2). American Laundry News, September 2011. americanlaundrynews.com.

05

Dyna Robotics. What 10 Months in Production Taught Us About the Robotics “Bubble”. dyna.co/news/robotics-bubble, May 2026.

06

Build AI. Egocentric-10K and Egocentric-100K datasets (10,000 and 100,405 hours of egocentric factory-worker video, Apache 2.0). huggingface.co/datasets/builddotai/Egocentric-100K, 2025.

07

Chi et al. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. Project page, 2024.

08

Araujo et al. Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking. arXiv:2510.02252, 2025.

09

He et al. OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning. arXiv:2406.08858, 2024.

10

Yang et al. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction. arXiv:2509.26633, 2025.

11

Ha et al. UMI on Legs: Making Manipulation Policies Mobile with a Manipulation-Centric Whole-body Controller. Project page, 2024.

12

Luo et al. SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11(117), 2026. arXiv:2511.07820.

13

Liu et al. Visual Whole-Body Control for Legged Loco-Manipulation. arXiv:2403.16967, 2024.

14

Li et al. AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control. arXiv:2505.03738, 2025.

15

Ben et al. HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit. arXiv:2502.13013, 2025.

16

Xue et al. LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction. arXiv:2506.13751, 2025.

17

Bronars, Park and Agrawal. Tune to Learn: How Controller Gains Shape Robot Policy Learning. arXiv:2604.02523, 2026.

18

Kwa et al. Measuring AI Ability to Complete Long Software Tasks. NeurIPS 2025; arXiv:2503.14499, first posted March 2025 as “Measuring AI Ability to Complete Long Tasks”. Summary on the METR blog, March 2025.

[ Stay Updated ]

Our research straight to your inbox.

[ DYNA ]

Newsletter Signup

© 2026 DYNA Robotics Inc.