更新时间:2026-09-20 GMT+08:00
分享

智能体调试

在部署模型服务和云端接入机器人之后,可以开始智能体调试,发送技能开始调试,云端和残差强化学习端侧会同时显示调试记录(日志),在调试期间可以人工干预调试,调试完成后在本地残差强化目录下查看保存的轨迹数据(具体格式请参见轨迹格式)。

智能体调试

  1. 登录CloudRobo控制台。
  2. 在左侧导航栏选择“运行管理 > 机器人”,进入“机器人”页面。
  3. 找到接入并激活机器人中的在线机器人,单击对应“操作”列的“智能体调试”。
  4. 在“智能体调试”页面,首次进入请选择部署模型的“UR5E抓放航插”模型服务与泛化技能。
  5. 在对话框输入“pick up the blue three-pin plug and place it upright on the right white foam board”技能,单击,即开始模型云端推理。
  6. 用户可以在云端和残差强化学习端侧查看调试记录,在调试期间可以人工干预调试(即人类干预)。

    图1 智能体调试记录

    此时,残差强化学习端侧日志示例如下:

    2026-07-17 11:18:00,055 - [INFO] - Published observation, episode_id=0
    2026-07-17 11:18:00,055 - [INFO] - Sent observation, waiting for cloud inference...
    2026-07-17 11:18:00,055 - [INFO] - Received action step, timestamp=1784258386042604464, chunk_index=0, step_index=0
    2026-07-17 11:18:00,055 - [INFO] - Skip action index 1
    2026-07-17 11:18:00,055 - [INFO] - Received action step, timestamp=1784258386042604464, chunk_index=0, step_index=1
    2026-07-17 11:18:00,055 - [INFO] - Executing action 2, base_action shape: (2, 7), gripper: 0.07993686790907673
    2026-07-17 11:18:00.056 | DEBUG    | robot_infra.envs.wrappers:action:97 - original target action is [ 5.56897283e-01  4.35239563e-05  2.91847587e-01 -1.77356648e+00
      1.93513513e+00 -5.84504247e-01  7.99368679e-02], intervened action is [ 5.56897283e-01  4.35239563e-05  2.91847587e-01 -1.77356648e+00
      1.93513513e+00 -5.84504247e-01  7.99368679e-02]
    2026-07-17 11:18:00.056 | DEBUG    | robot_infra.envs.wrappers:action:99 - processed target action is [ 5.56897283e-01  4.35239563e-05  2.91847587e-01 -1.77356648e+00
      1.93513513e+00 -5.84504247e-01  7.99368679e-02]
    2026-07-17 11:18:00.056 | DEBUG    | robot_infra.envs.ur5e_env:step:251 - target timestamp is 1784258280.1067631, current time is 1784258280.0567656
    2026-07-17 11:18:00,187 - [INFO] - Received action step, timestamp=1784258386042604464, chunk_index=0, step_index=2
    2026-07-17 11:18:00,187 - [INFO] - Skip action index 3
    2026-07-17 11:18:00,187 - [INFO] - Received action step, timestamp=1784258386042604464, chunk_index=0, step_index=3
    2026-07-17 11:18:00,187 - [INFO] - Executing action 4, base_action shape: (2, 7), gripper: 0.07999412965313052

人类干预

在机器人执行任务过程中,如果有必要,可以通过3D鼠标space mouse进行人工干预,程序会自动记录干预数据。

  • 一条轨迹执行成功,比如机械臂把航插直立地放到了海绵上,此时可以按电脑键盘“T”标记成功并保存轨迹。
  • 如果任务执行失败,比如机械臂偏离目标,此时可以按键盘“N”标记失败不保存轨迹,重新采集这条轨迹。

不论成功还是失败,机械臂都会自动执行初始化程序,回到初始位置。

采集成功的轨迹数目达到本地残差强化目录/hil-residual-rl/experiments/place_upright/config.yaml中experiment.successes_needed的值后,程序会自动结束。如果config.yaml文件中的输出目录为默认配置,此时可以在/hil-residual-rl/test_output/目录下看到保存的轨迹(具体格式请参见轨迹格式)。

轨迹格式

保存的数据格式与RLinf框架训练用的数据集格式一致,每次采集的轨迹目录下包含两个json文件和若干条轨迹pt文件。

  • metadata.json
  • trajectory_index.json
  • trajectory_0_demo_expert.pt
  • trajectory_1_demo_expert.pt

保存的pt轨迹包含以下key:

  • max_episode_length: 最大 episode 长度
  • model_weights_id: 模型权重 ID,默认"demo_experts"
  • actions: 动作 (n, 1, 7)
  • intervene_flags: 干预标志 (n, 1, 7)
  • rewards: 奖励 (n, 1, 1)
  • terminations: 终止标志 (n, 1, 1)
  • truncations: 截断标志 (n, 1, 1)
  • dones: 完成标志 (n, 1, 1)
  • forward_inputs: 动作 (n, 1, 7)
  • curr_obs: 当前观测
  • next_obs: 下一观测

curr_obs/next_obs包含:

  • states: 状态 (n, 1, 7)
  • main_images: 主相机图像 (n, 1, 224, 224, 3)
  • wrist_images: 手腕相机图像 (n, 1, 224, 224, 3)
  • residual_actions: 残差动作 (n, 1, 7)
  • base_actions: 基线动作 (n, 1, 7)
  • step_index: 步索引 (n, 1)
  • chunk_index: chunk 索引 (n, 1)
  • next_base_actions: 下一基线动作 (n, 1, 7)

相关文档