{
  "schema_version": "paper_public_manifest_v1",
  "paper_id": "latent_action_pretraining_through_world_modeling_2026_07",
  "slug": "latent_action_pretraining_through_world_modeling",
  "title": "LAWM: Why Action Labels Become a Burden",
  "authors": [],
  "source": {
    "arxiv_id": "2509.18428v2",
    "pdf_url": "https://arxiv.org/pdf/2509.18428v2",
    "project_url": "",
    "github_url": "",
    "huggingface_url": "",
    "original_source": "https://arxiv.org/pdf/2509.18428v2"
  },
  "site": {
    "post_url": "/posts/latent_action_pretraining_through_world_modeling",
    "canonical_url": "https://haiguangboy.com/posts/latent_action_pretraining_through_world_modeling",
    "cover_image": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/cover.webp"
  },
  "taxonomy": {
    "domain": "embodied_ai",
    "track": "world_model",
    "tasks": [
      "embodied_ai",
      "world_model",
      "vla",
      "action_generation",
      "robotics",
      "state_prediction",
      "Embodied Intelligence",
      "World Models",
      "Pretraining",
      "Data Efficiency",
      "Robot Learning"
    ],
    "related_topics": [
      {
        "paper_id": "an_open_foundation_model_towards_2026_07",
        "title": "An_Open_Foundation_Model_Towards",
        "url": "https://haiguangboy.com/posts/an_open_foundation_model_towards",
        "relation": "same_track",
        "summary": "Counterintuitive core finding: world model pretraining without action labels performs on par with or even surpasses supervised pretraining with labels, validating that joint co-training on heterogeneous human-robot data is a structurally suboptimal approach",
        "strength": "strong"
      },
      {
        "paper_id": "t_rex_tactile_reactive_dexterous_manipulation_2026_07",
        "title": "T-Rex: Why Tactile Sensing Needs Separate Modeling",
        "url": "https://haiguangboy.com/posts/t_rex_tactile_reactive_dexterous_manipulation",
        "relation": "same_track",
        "summary": "T-Rex: Why Tactile Sensing Needs Separate Modeling",
        "strength": "strong"
      },
      {
        "paper_id": "wx_界面新闻_20260605_2026_06",
        "title": "wx_InterfaceNews_20260605",
        "url": "https://haiguangboy.com/posts/wx_界面新闻_20260605",
        "relation": "same_track",
        "summary": "Comparison with LAPA/UniVLA/villa-X: small models (7M-parameter scale) match or surpass large models (7B-parameter scale) using similar methods, validating the architectural bet that unified networks outperform modular concatenation—WUM opposes VLA's layer-by-layer semantic transmission",
        "strength": "strong"
      },
      {
        "paper_id": "orca_2026_07",
        "title": "π0.5 Trembles in Place When Failing to Grab a Spoon, While Orca Goes Further with Physical Intuition Learned from Watching Videos",
        "url": "https://haiguangboy.com/posts/orca",
        "relation": "same_track",
        "summary": "The key to a world model is a readable state",
        "strength": "strong"
      },
      {
        "paper_id": "fast-wam_2026_07",
        "title": "Robots Still Reach 91.8% Without Imagining the Future at Inference! Fast-WAM Debunks WAM's Core Assumption",
        "url": "https://haiguangboy.com/posts/fast-wam",
        "relation": "same_track",
        "summary": "Training video objectives matter more than imagining the future at test time",
        "strength": "strong"
      },
      {
        "paper_id": "wx_星河频率_20260718_2026_07",
        "title": "Robots Begin to 'Stand in the Light': Lingchu Intelligence Enters Optical Module Production Lines",
        "url": "https://haiguangboy.com/posts/lingchu-optical",
        "relation": "same_track",
        "summary": "This paper is a concrete engineering implementation of the 'massive non-action data pretraining + small action data alignment' approach, validating the 'native human data' pyramid claim: pretraining is dominated by human data, with real-robot data used only for post-training adaptation",
        "strength": "medium"
      }
    ]
  },
  "ruling": {
    "importance_score": 3.0,
    "one_sentence": "LAWM: Why Action Labels Become a Burden"
  },
  "asset_base_url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling",
  "assets": [
    {
      "type": "pdf_screenshot",
      "object_key": "papers/latent_action_pretraining_through_world_modeling/page_01.webp",
      "content_type": "image/webp",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/page_01.webp",
      "role": "paper_first_page",
      "size_bytes": 155576
    },
    {
      "type": "pdf_screenshot",
      "object_key": "papers/latent_action_pretraining_through_world_modeling/key_figure.webp",
      "content_type": "image/webp",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/key_figure.webp",
      "role": "method_figure",
      "size_bytes": 154576
    },
    {
      "type": "cover_image",
      "object_key": "papers/latent_action_pretraining_through_world_modeling/cover.webp",
      "content_type": "image/webp",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/cover.webp",
      "role": "post_cover",
      "size_bytes": 68394
    },
    {
      "type": "public_brief",
      "object_key": "papers/latent_action_pretraining_through_world_modeling/public_brief.md",
      "content_type": "text/markdown; charset=utf-8",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/public_brief.md",
      "role": "public_brief",
      "size_bytes": 4171
    },
    {
      "type": "public_manifest",
      "object_key": "papers/latent_action_pretraining_through_world_modeling/public_manifest.json",
      "content_type": "application/json; charset=utf-8",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/latent_action_pretraining_through_world_modeling/public_manifest.json",
      "role": "public_manifest",
      "size_bytes": 6245
    }
  ],
  "published_at": "2026-07-22T10:41:01+08:00",
  "created_at": "2026-07-22T10:41:01+08:00",
  "updated_at": "2026-09-02T10:55:06+08:00",
  "analyst_take": {
    "type": "author_opinion",
    "text": "💡This clue has repeatedly appeared in the database\n\nDifferent teams have hit the same conclusion through completely different approaches: T-Rex uses a three-stage recipe (22,900 hours of human video as a foundation, grounded by 100 hours of tactile data); Fast-WAM from Xinghaitu found that the world model's capability mainly comes from representations learned during training, and visual prediction at inference can be directly removed; other work in the database has shown that 'human video + 1 robot demonstration' can surpass baselines using pure robot data.\n\nNotably, these efforts are unaware of each other and take different routes, yet all point to the same thing: **what robots truly lack may not be action labels, but methods to convert the much larger volume of 'non-action data' into executable capabilities.**\n\nHowever, don't rush to treat this as settled. Also in the database, LAWM's 'small model unified learning' results support the WUM bet that 'unified networks outperform modular concatenation,' while T-Rex's conclusion is the opposite—it insists tactile sensing must follow an independent pathway and opposes forced fusion. On the same question of 'whether one model should handle everything,' the two papers stand on different sides. The essence of the disagreement may not be who is right, but rather: what should be unified (the physical laws of the world) and what should not (perceptual signals whose time scales differ by several orders of magnitude)."
  }
}
