Introducing Light-O1

Scaling Whole-Body Intelligence with Human Action Pretraining

Light-O1, our first general-purpose embodied foundation model, scales whole-body intelligence through human action pretraining and transfers it to diverse robot embodiments, enabling autonomous, long-horizon loco-manipulation.

Human videos capture how people interact with the physical world, offering a scale and diversity that robot data collection alone struggles to match. At Light Origins, we elicit action knowledge from these videos into a transferable human action prior. We demonstrate a cross-embodiment transfer scaling law: scaling human action pretraining yields power-law reductions in next-action prediction loss and whole-body pose error after adaptation. On this foundation, we build Light-O1—a model that brings whole-body intelligence into the era of scalable pretraining.

01 · INTRODUCTION

Introduction

Physical AGI requires general-purpose whole-body intelligence: integrating language reasoning, visual-spatial understanding, and whole-body coordination to produce actions aligned with instructions and current observations. These capabilities must operate in a continuous loop, in which actions change the environment and new observations inform subsequent reasoning and action. Generalizing these capabilities across tasks and environments requires a broad, transferable physical action prior.

Today, robot action models rely heavily on purpose-collected data, including over 10,000 hours of robot demonstration data used to pretrain π₀[1], human-operated capture with handheld devices (UMI-style)[2][3], and egocentric human demonstrations[4]. These collections provide targeted supervision for embodiment and viewpoint alignment, but scaling them to cover the diversity and long tail of everyday activity remains expensive and time-consuming. Human videos provide a complementary source of action knowledge, capturing physical activity across diverse situations. They record the situations people encounter, how they act, and how their surroundings change—from navigating environments and using tools to interacting with objects and people. The central challenge is to transfer the action knowledge embedded in human videos to robots, despite differences in embodiment, observation, and control.

At Light Origins, we address this challenge through human action pretraining, guided by the principle of intelligence through compression[5]. Language-model pretraining uses next-token prediction[6] to compress the knowledge and reasoning expressed in language and code into a reusable prior. We extend this principle to the intelligence expressed in human action: how people act in a given situation, how their actions change the world, and how they respond to environmental feedback. We recover structured human actions from video and encode them in a unified humanoid action representation. High-fidelity tokenization converts these continuous representations into discrete action tokens, which we interleave with language and visual observations into multimodal temporal sequences. We pretrain an autoregressive Transformer[7] on these sequences at scale to learn a transferable human action prior, then adapt the pretrained model to robot embodiments and tasks using purpose-collected robot data.

We demonstrate a cross-embodiment transfer scaling law for the human action prior across three adaptation settings: public egocentric human action data, public robot data, and in-house loco-manipulation data collected with our own humanoid. Across the measured scales, increasing the human action pretraining budget yields power-law reductions in next-action prediction loss and whole-body pose error after adaptation. All metrics are evaluated on held-out data, with whole-body pose error measured under open-loop evaluation.

Building on this foundation, we introduce Light-O1, a whole-body intelligence model that connects language reasoning and visual-spatial understanding with coordinated action. We demonstrate two complementary capabilities: loco-manipulation, where instructions define a task and the model determines how to act in the observed environment; and expressive whole-body skills, where instructions specify a movement and the model reasons about and generates the corresponding whole-body action. Real-world demonstrations and quantitative evaluations of humanoid manipulation and human action generation show how this human action prior supports both task completion and instruction-guided movement.

02 · CROSS-EMBODIMENT TRANSFER SCALING LAW

Cross-embodiment Transfer Scaling Law

We find a cross-embodiment transfer scaling law: scaling human action pretraining yields power-law reductions in next-action-token prediction loss and whole-body pose prediction error after adaptation. The prior learned from human actions transfers across embodiments with different camera views and task instructions, providing a scalable foundation for whole-body intelligence.

We measure pretraining scale by D, the total multimodal token budget: language, vision, and discrete action tokens processed during pretraining. Starting from a 4B base model, we train independent models at six budgets from 3.75B to 120B tokens, with the largest budget corresponding to 100,000 hours of human action. We evaluate next-action-token prediction loss (cross-entropy) on held-out human actions recovered from video.

To measure transfer, we initialize post-training from each pretrained checkpoint and adapt it to egocentric human data (Nymeria[8]), Unitree G1 teleoperation data (HIW-500[9]), and in-house LightBot teleoperation data. We measure the optimal prediction error achievable at a given pretraining budget on each target dataset’s held-out evaluation set, using next-action-token prediction loss and whole-body pose prediction error from open-loop evaluation of the decoded actions. Here, optimal means the best result attained across the evaluated post-training configurations and checkpoints. We characterize these trends with power-law fits, following the scaling-law analysis of Kaplan et al.[10].

Transfer Scaling Law

Scaling human action pretraining yields power-law reductions in prediction error across embodiments.

Power-law fit in the form of L(D) = L₀ + αD−η. Y-axis shows L − L₀ in log scale.

Human action pretraining

Next-action-token Prediction Loss

Human action pretraining · Next-action-token Prediction LossL − L₀1.390.900.593.75B7.5B15B30B60B120BD · Total multimodal pretraining tokens (B)D = 3.758 B tokens; L = 6.421 nats/token; L − L₀ = 1.324D = 7.516 B tokens; L = 6.204 nats/token; L − L₀ = 1.106D = 15.03 B tokens; L = 6.037 nats/token; L − L₀ = 0.9391D = 30.06 B tokens; L = 5.92 nats/token; L − L₀ = 0.8223D = 60.13 B tokens; L = 5.763 nats/token; L − L₀ = 0.6656D = 120.3 B tokens; L = 5.672 nats/token; L − L₀ = 0.5741L₀ = 5.098 nats/token; η = 0.2402L₀ = 5.10η = 0.24

Egocentric human data

Next-action-token Prediction Loss
Egocentric human data · Next-action-token Prediction LossL − L₀1.000.390.15No pretraining: L = 6.854 nats/token; L − L₀ = 1.099No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 6.384 nats/token; L − L₀ = 0.629D = 7.516 B tokens; L = 6.215 nats/token; L − L₀ = 0.4604D = 15.03 B tokens; L = 6.096 nats/token; L − L₀ = 0.3417D = 30.06 B tokens; L = 6.019 nats/token; L − L₀ = 0.2643D = 60.13 B tokens; L = 5.935 nats/token; L − L₀ = 0.1805D = 120.3 B tokens; L = 5.895 nats/token; L − L₀ = 0.1405L₀ = 5.755 nats/token; η = 0.4345L₀ = 5.75η = 0.43
Body Pose Prediction Error (mm)
Egocentric human data · Body Pose Prediction Error (mm)L − L₀16.7115.3414.09No pretraining: L = 40.11 ± 0.312 SD mm; L − L₀ = 16.54No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 39.22 ± 0.1442 SD mm; L − L₀ = 15.65D = 7.516 B tokens; L = 38.88 ± 0.2674 SD mm; L − L₀ = 15.31D = 15.03 B tokens; L = 38.35 ± 0.1921 SD mm; L − L₀ = 14.77D = 30.06 B tokens; L = 38.3 ± 0.1821 SD mm; L − L₀ = 14.73D = 60.13 B tokens; L = 37.76 ± 0.2108 SD mm; L − L₀ = 14.18D = 120.3 B tokens; L = 37.78 ± 0.1455 SD mm; L − L₀ = 14.21L₀ = 23.57 mm; η = 0.02948L₀ = 23.57η = 0.03

Unitree G1 robot

Next-action-token Prediction Loss
Unitree G1 robot · Next-action-token Prediction LossL − L₀1.341.030.79No pretraining: L = 4.721 nats/token; L − L₀ = 1.372No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 4.538 nats/token; L − L₀ = 1.189D = 7.516 B tokens; L = 4.443 nats/token; L − L₀ = 1.094D = 15.03 B tokens; L = 4.343 nats/token; L − L₀ = 0.9934D = 30.06 B tokens; L = 4.27 nats/token; L − L₀ = 0.9211D = 60.13 B tokens; L = 4.17 nats/token; L − L₀ = 0.8212D = 120.3 B tokens; L = 4.118 nats/token; L − L₀ = 0.7688L₀ = 3.349 nats/token; η = 0.1285L₀ = 3.35η = 0.13
Body Pose Prediction Error (mm)
Unitree G1 robot · Body Pose Prediction Error (mm)L − L₀3.523.192.89No pretraining: L = 13.34 ± 0.05 SD mm; L − L₀ = 3.5No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 13.2 ± 0.0599 SD mm; L − L₀ = 3.353D = 7.516 B tokens; L = 13.13 ± 0.0561 SD mm; L − L₀ = 3.29D = 15.03 B tokens; L = 13.03 ± 0.0513 SD mm; L − L₀ = 3.191D = 30.06 B tokens; L = 12.97 ± 0.0489 SD mm; L − L₀ = 3.129D = 60.13 B tokens; L = 12.84 ± 0.0432 SD mm; L − L₀ = 2.994D = 120.3 B tokens; L = 12.75 ± 0.0444 SD mm; L − L₀ = 2.911L₀ = 9.842 mm; η = 0.04157L₀ = 9.84η = 0.04

LightBot · In-house data

Next-action-token Prediction Loss
LightBot · In-house data · Next-action-token Prediction LossL − L₀2.520.860.30No pretraining: L = 7.958 nats/token; L − L₀ = 2.789No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 6.439 nats/token; L − L₀ = 1.27D = 7.516 B tokens; L = 6.104 nats/token; L − L₀ = 0.9346D = 15.03 B tokens; L = 5.838 nats/token; L − L₀ = 0.6686D = 30.06 B tokens; L = 5.621 nats/token; L − L₀ = 0.4516D = 60.13 B tokens; L = 5.48 nats/token; L − L₀ = 0.3108D = 120.3 B tokens; L = 5.437 nats/token; L − L₀ = 0.2675L₀ = 5.169 nats/token; η = 0.4808L₀ = 5.17η = 0.48
Body Pose Prediction Error (mm)
LightBot · In-house data · Body Pose Prediction Error (mm)L − L₀8.546.254.57No pretraining: L = 21.39 ± 0.4115 SD mm; L − L₀ = 8.392No pretraining3.75B7.5B15B30B60B120BD = 3.758 B tokens; L = 19.05 ± 0.2492 SD mm; L − L₀ = 6.047D = 7.516 B tokens; L = 18.29 ± 0.1814 SD mm; L − L₀ = 5.287D = 15.03 B tokens; L = 18.2 ± 0.1567 SD mm; L − L₀ = 5.194D = 30.06 B tokens; L = 17.7 ± 0.1191 SD mm; L − L₀ = 4.696D = 60.13 B tokens; L = 17.57 ± 0.1313 SD mm; L − L₀ = 4.564D = 120.3 B tokens; L = 17.61 ± 0.1058 SD mm; L − L₀ = 4.61L₀ = 13 mm; η = 0.07825L₀ = 13.00η = 0.08

D · Total multimodal pretraining tokens (B)

Fig. 03 Scaling a human action prior across embodiments. Left: next-action-token prediction loss on held-out human-action pretraining data recovered from video. Right: optimal prediction loss and whole-body pose prediction error on each target’s held-out set after adaptation. Dots are measurements; downward-sloping dashed curves are power-law fits; horizontal grey dashed lines are no-pretraining references. Pose-error bars are ±1 SD across the evaluated checkpoint window. Power-law view shows L − L₀ on log–log axes; Raw metrics shows L with a linear vertical axis. D counts total multimodal pretraining tokens; L₀ is the floor, fitted for loss and measured tokenizer reconstruction error for pose error; η is the scaling exponent.
Detailed Experiment Protocols

Pretraining setup

We start from Qwen3.5-4B and train independent models at six pretraining budgets D from 3.75B to 120B multimodal tokens. Within each series, the base model, action tokenizer, and data mixture remain fixed. Each run completes its own learning-rate schedule; its final checkpoint is evaluated on a held-out human-action split and used to initialize post-training.

Post-training setup

Prediction at each budget. We initialize multiple post-training runs from each pretrained checkpoint, varying batch topology and random seed at a fixed post-training token budget for each target, following a target-specific learning-rate search. Each target uses the same held-out evaluation set across budgets and runs. For each metric, we report the optimal held-out prediction error attained across the evaluated runs and checkpoints. We fit the transfer scaling curves to these per-budget optima.

Open-loop pose evaluation. Conditioned on the language prompt, preceding visual observations, and the ground-truth action prefix, the model predicts the next one-second action chunk. We measure mean per-joint position error (MPJPE) against the recorded actions using forward kinematics in a pelvis-centered, heading-aligned frame, with best-of-four sampling by local pose error. Pose-error bars show ±1 SD across the evaluated checkpoint window.

Power-law fitting

We fit each series as L(D) = L₀ + αD−η. In the power-law view, the fitted slope of log(L − L₀) against log D is −η.

For next-action-token prediction loss, we jointly fit L₀, α, and η to the raw measurements using unweighted nonlinear least squares. For whole-body pose prediction error, we fix L₀ to the measured action-tokenizer reconstruction error on the same held-out data and scored frames, then fit log(L − L₀) against log D by unweighted least squares.

03 · WHOLE-BODY INTELLIGENCE

From Human Action Pretraining to Whole-Body Intelligence

Light-O1 is our whole-body intelligence model, the result of scaling human action pretraining to internet-scale video. Built on the unified human action representation, it is natively deployable across humanoid robots: it carries the motor skills that evolved in people over millions of years onto the machine, and exploits the kinematic redundancy, dexterity and adaptability of their bodies. We highlight two capabilities below: loco-manipulation, in which the instruction defines the task and the policy determines the action, and expressive whole-body skills, in which the movement itself is prescribed.

03A

Loco-Manipulation

Light-O1 turns an instruction into whole-body loco-manipulation anchored in the scene in front of it. It locates the objects it must act on and coordinates locomotion, posture and dexterous manipulation to complete the task. The same model drives humanoids with different bodies and hands, and handles rigid and soft objects.

03B

Expressive Whole-Body Skills

Expressive whole-body skills are the most direct evidence of human action pretraining at work. From a single instruction that prescribes the movement, whether kneeling to propose or swinging a golf club, Light-O1 reasons about it, composes it for the entire body and executes it on the robot. For each example below we present the reasoning trace, the rendered motion and its execution on hardware, so that every stage of the process can be examined.

04 · QUANTITATIVE EVALUATION

Quantitative Evaluation

We quantify what scaling the pretrained model buys with two benchmarks: humanoid manipulation, whether Light-O1 completes a household task, and human action generation, whether the motion it produces matches the instruction. Both compare Light-O1 with published models.

04A

Humanoid Manipulation

In the simulated RoboCasa GR-1 benchmark, which covers 24 kitchen tabletop tasks on a bimanual humanoid with dexterous hands, from moving objects between surfaces and containers to placing them into cabinets, drawers and a microwave and closing them, Light-O1 achieves a 79.3% macro success rate with 50 episodes per task, above all published results we compare against.

RoboCasa GR-1 tabletopMacro success rate over the 24 RoboCasa GR-1 tabletop tasks. DIAL is as reported by its authors and was not re-run by us; GR00T N1.7 and π0.5 were trained and evaluated by us on the benchmark's simulation teleoperation data. Checkpoint selection and simulator versions are not aligned between the author-reported entry and ours. Light-O1 (Ours): 79.3%; π0.5: 70.9%; DIAL: 70.2%; GR00T N1.7: 59.1%. Source: evidence/data/robocasa-gr1-benchmark.json.Success rate (%)50%75%100%Light-O1 (Ours): 79.3%79.3%Light-O1 (Ours)π0.5: 70.9%70.9%π0.5DIAL: 70.2%70.2%DIALGR00T N1.7: 59.1%59.1%GR00T N1.7
Fig. 14 RoboCasa GR-1 tabletop. Macro success rate over the 24 RoboCasa GR-1 tabletop tasks.
04B

Human Action Generation

We evaluate Light-O1 with a primary benchmark and an auxiliary one. The primary is human rating, the gold standard, over 30k prompts on which judges compare models pairwise. The lead there is not marginal: Light-O1's lowest category rating stands above either baseline's highest, on semantic following, expressiveness and acceptability alike. The fitted Elo separates it from HY-Motion-1.0[15] by close to 400 points, 1472.8 against 1078.3, with Kimodo[16] anchored at 1000. The auxiliary benchmark is SSAE on HY-Motion-Bench[15], where a VLM judge checks whether the motion contains what the prompt asked for. It can score a new checkpoint on demand, and it isolates what a rating blends together. Light-O1 leads there in all six prompt categories, at 78.0 overall against 74.7 and 61.4.

01 · Human judgement

Motion Arena

over 30k prompts
  • Light-O1 (Ours)
  • HY-Motion-1.0
  • Kimodo
Arena Elo
Arena Elo10001100120013001400Light-O1 (Ours): Elo 1472.8 (1456.4–1489.4)1473HY-Motion-1.0: Elo 1078.3 (1070.4–1086.8)1078Kimodo: Elo 1000 (1000–1000)1000

Elo score computed by pairwise comparison; Kimodo fixed at 1,000 as the anchor.

Semantic followingSemantic following. Prompt-level mean, 1-5 (higher is better). CROSS-WEEK: the three models were labelled in different weeks on different prompt sets (best own = week 0727; HY-Motion = weeks 0704/0711/0713/0719; Kimodo = week 0704 only), so values are NOT same-batch comparable. Use for illustrative overview, not a strict head-to-head. Source: evidence/data/human-label-radar.json.Light-O1 (Ours) · Locomotion · Semantic following: 4.0034Light-O1 (Ours) · Vehicle motion · Semantic following: 3.9826Light-O1 (Ours) · Sports & athletics · Semantic following: 3.8931Light-O1 (Ours) · Dance & performance · Semantic following: 3.9506Light-O1 (Ours) · Fitness & exercise · Semantic following: 3.7152Light-O1 (Ours) · Manual work · Semantic following: 3.905Light-O1 (Ours) · Social interaction · Semantic following: 4.0292Light-O1 (Ours) · Posture transition · Semantic following: 3.9578HY-Motion-1.0 · Locomotion · Semantic following: 2.67HY-Motion-1.0 · Vehicle motion · Semantic following: 2.3689HY-Motion-1.0 · Sports & athletics · Semantic following: 2.5631HY-Motion-1.0 · Dance & performance · Semantic following: 2.5594HY-Motion-1.0 · Fitness & exercise · Semantic following: 2.6852HY-Motion-1.0 · Manual work · Semantic following: 2.6488HY-Motion-1.0 · Social interaction · Semantic following: 2.552HY-Motion-1.0 · Posture transition · Semantic following: 2.6322Kimodo · Locomotion · Semantic following: 2.108Kimodo · Vehicle motion · Semantic following: 2.431Kimodo · Sports & athletics · Semantic following: 2.103Kimodo · Dance & performance · Semantic following: 2.14Kimodo · Fitness & exercise · Semantic following: 2.131Kimodo · Manual work · Semantic following: 2.412Kimodo · Social interaction · Semantic following: 2.184Kimodo · Posture transition · Semantic following: 2.261LocomotionVehicle motionSports & athleticsDance & performanceFitness & exerciseManual workSocial interactionPosture transition2345
ExpressivenessExpressiveness. Prompt-level mean, 1-5 (higher is better). CROSS-WEEK: the three models were labelled in different weeks on different prompt sets (best own = week 0727; HY-Motion = weeks 0704/0711/0713/0719; Kimodo = week 0704 only), so values are NOT same-batch comparable. Use for illustrative overview, not a strict head-to-head. Source: evidence/data/human-label-radar.json.Light-O1 (Ours) · Locomotion · Expressiveness: 3.9131Light-O1 (Ours) · Vehicle motion · Expressiveness: 3.9095Light-O1 (Ours) · Sports & athletics · Expressiveness: 3.8782Light-O1 (Ours) · Dance & performance · Expressiveness: 3.8132Light-O1 (Ours) · Fitness & exercise · Expressiveness: 3.5454Light-O1 (Ours) · Manual work · Expressiveness: 3.7957Light-O1 (Ours) · Social interaction · Expressiveness: 3.8567Light-O1 (Ours) · Posture transition · Expressiveness: 3.7879HY-Motion-1.0 · Locomotion · Expressiveness: 2.4119HY-Motion-1.0 · Vehicle motion · Expressiveness: 2.1963HY-Motion-1.0 · Sports & athletics · Expressiveness: 2.385HY-Motion-1.0 · Dance & performance · Expressiveness: 2.3495HY-Motion-1.0 · Fitness & exercise · Expressiveness: 2.4562HY-Motion-1.0 · Manual work · Expressiveness: 2.4335HY-Motion-1.0 · Social interaction · Expressiveness: 2.2536HY-Motion-1.0 · Posture transition · Expressiveness: 2.3555Kimodo · Locomotion · Expressiveness: 2.904Kimodo · Vehicle motion · Expressiveness: 3.011Kimodo · Sports & athletics · Expressiveness: 2.795Kimodo · Dance & performance · Expressiveness: 3.089Kimodo · Fitness & exercise · Expressiveness: 2.957Kimodo · Manual work · Expressiveness: 2.926Kimodo · Social interaction · Expressiveness: 3.028Kimodo · Posture transition · Expressiveness: 2.868LocomotionVehicle motionSports & athleticsDance & performanceFitness & exerciseManual workSocial interactionPosture transition2345
AcceptabilityAcceptability. Prompt-level mean, 1-5 (higher is better). CROSS-WEEK: the three models were labelled in different weeks on different prompt sets (best own = week 0727; HY-Motion = weeks 0704/0711/0713/0719; Kimodo = week 0704 only), so values are NOT same-batch comparable. Use for illustrative overview, not a strict head-to-head. Source: evidence/data/human-label-radar.json.Light-O1 (Ours) · Locomotion · Acceptability: 3.9594Light-O1 (Ours) · Vehicle motion · Acceptability: 3.9397Light-O1 (Ours) · Sports & athletics · Acceptability: 3.8731Light-O1 (Ours) · Dance & performance · Acceptability: 3.8668Light-O1 (Ours) · Fitness & exercise · Acceptability: 3.5692Light-O1 (Ours) · Manual work · Acceptability: 3.8283Light-O1 (Ours) · Social interaction · Acceptability: 3.8983Light-O1 (Ours) · Posture transition · Acceptability: 3.8504HY-Motion-1.0 · Locomotion · Acceptability: 2.2977HY-Motion-1.0 · Vehicle motion · Acceptability: 2.0955HY-Motion-1.0 · Sports & athletics · Acceptability: 2.2467HY-Motion-1.0 · Dance & performance · Acceptability: 2.2112HY-Motion-1.0 · Fitness & exercise · Acceptability: 2.354HY-Motion-1.0 · Manual work · Acceptability: 2.3545HY-Motion-1.0 · Social interaction · Acceptability: 2.1929HY-Motion-1.0 · Posture transition · Acceptability: 2.2714Kimodo · Locomotion · Acceptability: 2.229Kimodo · Vehicle motion · Acceptability: 2.516Kimodo · Sports & athletics · Acceptability: 2.212Kimodo · Dance & performance · Acceptability: 2.264Kimodo · Fitness & exercise · Acceptability: 2.259Kimodo · Manual work · Acceptability: 2.483Kimodo · Social interaction · Acceptability: 2.287Kimodo · Posture transition · Acceptability: 2.325LocomotionVehicle motionSports & athleticsDance & performanceFitness & exerciseManual workSocial interactionPosture transition2345

Mean rating per prompt, 1–5, by motion category; each model pooled across the weeks it was labelled in.

Fig. 15 Quantitative evaluation of instruction following.
The auxiliary benchmark: HY-Motion-Bench (SSAE)
02 · VLM judgement

HY-Motion-Bench (SSAE)⁠[15]

1,984 prompts · 4 generations each

Semantic alignment score by a VLM judge.

SSAESSAE. Share of judge questions answered yes, % (higher is better) Axes start at 40, not 0.. Same 1,984 prompts and the same VLM-judge protocol for every model, four generations per prompt; each model was run on a different date with its own default sampling. Source: evidence/data/hy-motion-bench-ssae.json.Ours · Daily activities · SSAE: 79.4Ours · Fitness · SSAE: 70.89Ours · Game & performance · SSAE: 84.46Ours · Locomotion · SSAE: 75.65Ours · Social interaction · SSAE: 81.11Ours · Sports · SSAE: 78.86HY-Motion-1.0 · Daily activities · SSAE: 79.32HY-Motion-1.0 · Fitness · SSAE: 69.76HY-Motion-1.0 · Game & performance · SSAE: 73.52HY-Motion-1.0 · Locomotion · SSAE: 74.82HY-Motion-1.0 · Social interaction · SSAE: 76.81HY-Motion-1.0 · Sports · SSAE: 76.06Kimodo · Daily activities · SSAE: 68.34Kimodo · Fitness · SSAE: 44.19Kimodo · Game & performance · SSAE: 66Kimodo · Locomotion · SSAE: 69.41Kimodo · Social interaction · SSAE: 66.35Kimodo · Sports · SSAE: 59.04Daily activitiesFitnessGame & performanceLocomotionSocial interactionSports5060708090100
Fig. 16 HY-Motion-Bench (SSAE).
05 · METHODOLOGY

Human Action as a Foundation

We build a data pipeline that recovers human actions from video and interleaves them with language and visual observations into temporal sequences to model how observation and understanding inform action and how environmental feedback shapes subsequent actions. Autoregressive pretraining on these sequences builds a transferable human action prior, which we adapt to different embodiments and tasks.

05A

Data Pipeline

Our pipeline recovers structured human actions from internet videos and aligns them with language and visual observations for large-scale human action pretraining.

From human videos to annotated 3D human actionFour overlapping colored circles show internet-scale human videos, video curation, 3D human motion recovery, and fine-grained language annotation. Hover or keyboard focus brings each stage forward with subtle decorative motion. Internal artwork is conceptual, not actual footage or recovered motion.
Internet-Scale Human Videos

Diverse human activities from the internet

Video Curation

Cleaning and segmenting human activity videos

Human Action Extraction

Recovering structured action from video

Fine-Grained Annotation

Backswing

Continue the arm swing while shifting weight onto the lead leg...

Detailed descriptions aligned with action

Fig. 17 From human videos to annotated 3D human action.

Video processing. We segment human videos, detect and track people, reconstruct their three-dimensional actions, and enrich these records with language annotations. Each stage is designed to preserve action fidelity and temporal alignment, connecting visual context, semantic descriptions, and physical movement in a shared record.

Unified human action representation. We represent human actions through three decoupled components: root trajectory, body pose, and hand state. Together, these components provide a complete description of human actions in 3D space.

Interleaved multimodal sequences. We interleave language, visual observations, and structured actions along a shared timeline. Different sequence arrangements express how observation drives understanding and informs action, how instructions guide action in a visual context, and how repeated observation and action support long-horizon execution with environmental feedback. These sequences provide a basis for modeling causal relationships between human actions and the physical world.

05B

Pretraining and Adaptation

Autoregressive pretraining and embodiment adaptationRead from left to right: the input column shows text, six visual examples, and four human-action illustrations from top to bottom. These feed three groups of four schematic tokens beside the Autoregressive Transformer: language and vision tokens appear first, followed by action tokens with illustrated autoregressive feedback. Two parallel routes lead through the Action Decoder and unified action space, or the Action Expert and targeted action space, to the Behavior Foundation Model. Blue identifies the shared model pathway; orange distinguishes the expert pathway. The diagram ends at the Behavior Foundation Model. Animations and block sizes do not indicate measured performance, timing, or tensor dimensions.
Autoregressive Transformer
Text tokensThe bunny leaptonto a tram! Catch upand jump aboard!
Vision tokens
Action tokens
Action Decoder
Unified Human Action
Unified Human Action
Action Expert
Target actions
Targeted Action Space
Behavior Foundation Model
Fig. 18 Autoregressive pretraining and embodiment adaptation.

Autoregressive pretraining on text, vision, and action tokens builds a transferable human action prior. After embodiment and task adaptation, the prior connects to behavior foundation models through decoded unified actions or a diffusion-based action expert.

High-fidelity action tokenization. Our action tokenizer encodes continuous action representations into discrete action tokens and decodes them back into continuous action representations, preserving precision in key end-effector positions and orientations. This retains where the head is directed, where the hands are placed, and how they are oriented—spatial information essential to purposeful interaction.

Autoregressive pretraining. We model interleaved sequences of observations, language, and actions within a unified autoregressive framework. This brings together human action pretraining, language-guided action generation, vision-language embodied reasoning, and training on general foundation-model data. Pretraining compresses this multimodal experience into a transferable human action prior, modeling how observations and intent inform action, how actions change the world, and how environmental feedback shapes what follows. The multimodal context provides a form of memory: a history of what was observed, said, and done that informs subsequent reasoning and action as the interaction unfolds.

Embodiment and task adaptation. We post-train on purpose-collected data to align the learned prior with target embodiments and tasks. We support two execution interfaces: decoded unified human actions can directly connect to compatible behavior foundation models (BFMs), such as the BFM used by our LightBot; a diffusion-based action expert adapts the prior to a target action space for other BFMs. This second route follows the action-expert approach used in π0.5 and GR00T N1.7[12][11]. Aligning the prior with human intent belongs to the same stage, and no held-out loss can score intent, so we scale reinforcement learning from human feedback[17]. This also teaches the model to reply in two parts: it states in language what the instruction requires of the body, then generates the action.

06 · LOOKING AHEAD

Built to Scale

Behind Light-O1 is a sustained investment in the infrastructure for large-scale human action pretraining. Our data infrastructure now operates at the thousand-GPU scale, with weekly video-processing throughput reaching 200,000 hours—up 16-fold from 12,500 hours six months ago. Alongside this data infrastructure, we have built high-performance training systems and large-scale compute clusters, designed to turn internet-scale human video data into embodied intelligence more efficiently.

The next phase is to grow this foundation in scale, diversity, and quality. We will broaden our data sources and improve data quality throughout the pipeline, while expanding training data and model capacity together. Continued scaling experiments will guide how we allocate compute between them, with the goal of building a richer and more general foundation for whole-body intelligence. Beyond scaling data and compute, we will explore more efficient learning methods to compress the knowledge in raw internet data into transferable embodied intelligence. We also aim to deploy these capabilities in real-world settings across a broader range of robot embodiments, extending the reach of intelligence learned from human action.

It's time to scale.

Our mission is to bring general intelligence into the physical world—building systems that continually learn from experience and turn understanding into action. If you want to help build this future, join us (https://www.lightorigins.com/en/careers).

Fig. 19 All robot actions in this video were generated by Light-O1.

Citation

@techreport{lightorigins2026lighto1,
  author      = {{Light Origins Team}},
  title       = {{Light-O1}: Scaling Whole-Body Intelligence with Human Action Pretraining},
  institution = {Light Origins},
  year        = {2026}
}

References

Expanded video