{
  "schema_version": "paper_public_manifest_v1",
  "paper_id": "video2robo_2026_08",
  "slug": "video2robo",
  "title": "Closed is the kinematic loop, not the physical loop",
  "authors": [],
  "source": {
    "arxiv_id": "",
    "pdf_url": "https://openaccess.thecvf.com/content/CVPR2026/papers/Deng_Video2Robo_3DGS-based_Synthetic_Data_from_One_Video_Enables_Scalable_Robot_CVPR_2026_paper.pdf",
    "project_url": "",
    "github_url": "",
    "huggingface_url": "",
    "original_source": "https://openaccess.thecvf.com/content/CVPR2026/papers/Deng_Video2Robo_3DGS-based_Synthetic_Data_from_One_Video_Enables_Scalable_Robot_CVPR_2026_paper.pdf"
  },
  "site": {
    "post_url": "/posts/video2robo",
    "canonical_url": "https://haiguangboy.com/posts/video2robo",
    "cover_image": "https://static.haiguangboy.com/papers/video2robo/cover.webp"
  },
  "taxonomy": {
    "domain": "embodied_ai",
    "track": "action_generation",
    "tasks": [
      "embodied_ai",
      "action_generation",
      "robotics",
      "Embodied AI",
      "Robot learning",
      "3dgs",
      "Data generation"
    ],
    "related_topics": [
      {
        "paper_id": "wx_elsewhere别处发生_20260722_2026_07",
        "title": "Liang Wenfeng's four-hour investor meeting transcript",
        "url": "https://haiguangboy.com/posts/liangwenfeng-world-model",
        "relation": "contrast",
        "summary": "Route bet: bypass physics simulation with kinematic consistency + 3DGS photorealism contradicts ★ Divergence may stem from differing referents: Liang Wenfeng's parallel of \"3D/video generation/world models\" falls entirely under \"renderers,\" while robotics needs \"simulators\"",
        "strength": "strong"
      },
      {
        "paper_id": "from_passive_video_to_editable_experience_physically_grounded_experience_synthes_2026_08",
        "title": "Cross-embodiment transfer should be done at the experience level",
        "url": "https://haiguangboy.com/posts/pegasus-editable-experience",
        "relation": "contrast",
        "summary": "Route bet: bypass physics simulation with kinematic consistency + 3DGS photorealism contradicts Core claim: embodiment gap should be bridged at the experience level, not the pixel level",
        "strength": "strong"
      },
      {
        "paper_id": "leapbot_wa_world_anchor_action_models_via_predictive_latent_alignments_2026_07",
        "title": "Predictive features cannot be fed directly to diffusion models",
        "url": "https://haiguangboy.com/posts/leapbot-wa",
        "relation": "contrast",
        "summary": "Route bet: bypass physics simulation with kinematic consistency + 3DGS photorealism contradicts Core claim: the utility of world modeling for manipulation lies in abstract physical expectations, not photorealistic rendering",
        "strength": "strong"
      },
      {
        "paper_id": "learning_a_thousand_tasks_in_a_day_2026_08",
        "title": "1,000 tasks in 1 day, relying on inductive bias",
        "url": "https://haiguangboy.com/posts/mt3-thousand-tasks",
        "relation": "same_track",
        "summary": "Limited evaluation scope: six self-collected tasks, self-built simulation benchmark, real-robot baseline with only 20 demonstrations validates Problem quantification: mainstream BC systems average 175–250 demonstrations per task; dual-arm tasks require ~8K demonstrations",
        "strength": "strong"
      },
      {
        "paper_id": "latent_action_pretraining_through_world_modeling_2026_07",
        "title": "LAWM: Why action labels become a burden",
        "url": "https://haiguangboy.com/posts/latent_action_pretraining_through_world_modeling",
        "relation": "same_track",
        "summary": "LAWM: Why action labels become a burden",
        "strength": "strong"
      },
      {
        "paper_id": "decompose_and_reorganize_planning_with_primitives_and_visuomotor_policies_learne_2026_08",
        "title": "Switching logic should not be learned by the policy",
        "url": "https://haiguangboy.com/posts/dr-lfd",
        "relation": "same_track",
        "summary": "Real-robot Franka: pure synthetic data 62.5% surpasses 46.67% trained from 20 real teleop demonstrations validates Results: simulation peg-in-hole ID 100% vs ACT 44%/DP 54%; on DexMimicGen, 100 demonstrations beat baseline's 1,000",
        "strength": "strong"
      }
    ]
  },
  "ruling": {
    "importance_score": 3.0,
    "one_sentence": "Data synthesized from a phone video achieves 62.5% on real robot, surpassing 46.7% from real collected data"
  },
  "asset_base_url": "https://static.haiguangboy.com/papers/video2robo",
  "assets": [
    {
      "type": "public_brief",
      "object_key": "papers/video2robo/public_brief.md",
      "content_type": "text/markdown; charset=utf-8",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/video2robo/public_brief.md",
      "role": "public_brief",
      "size_bytes": 5333
    },
    {
      "type": "public_manifest",
      "object_key": "papers/video2robo/public_manifest.json",
      "content_type": "application/json; charset=utf-8",
      "upload_status": "uploaded",
      "bucket": "paper-assets",
      "url": "https://static.haiguangboy.com/papers/video2robo/public_manifest.json",
      "role": "public_manifest",
      "size_bytes": 5186
    }
  ],
  "published_at": "2026-08-09T22:42:47+08:00",
  "created_at": "2026-08-09T22:42:47+08:00",
  "updated_at": "2026-09-02T10:55:06+08:00",
  "analyst_take": {
    "type": "author_opinion",
    "text": "This paper's route bet directly opposes several recent works. LeapBot-WA argues the utility of world modeling lies in abstract physical expectations rather than photorealistic rendering; the Genesis official blog positions simulation as an evaluation engine rather than a data generator; \"From Passive Video to Editable Experience\" argues the embodiment gap should be bridged at the experience level. Video2Robo does the opposite across the board, and its real-robot numbers win.\n\nIt also doesn't overclaim: the paper states \"deformable objects are not supported because dynamics models need to be learned from video,\" effectively admitting rendering cannot replace physics. So this is not a right-vs-wrong debate—within the bandwidth of desktop rigid bodies, rendering plus kinematic scripting is indeed sufficient.\n\nThere is also a three-way divergence: how to use human video. Video2Robo tracks object relative motion, CAIP tracks hand pose, and JoyAI-RA 0.5 learns a cross-embodiment latent action space.\n\nAnchor for review in six months: whether it can climb out of desktop rigid bodies. If it scales to contact-rich assembly, deformables, and dual-arm tasks, then \"no physics needed\" holds; if not, the 62.5% only shows the tasks are insensitive to physics."
  }
}
