GPU vision, CPU control
The writer selects CUDA when available and warms up Theia with five calls. The small perception and actor sessions stay in the controller process, using CPU inference.
Inspect the writer ↘A humanoid learns to dodge thrown balls from head-camera depth and its own body state.
A frozen vision transformer, a learned temporal percept, and seven decoder adapters. Follow a recorded ball through the visual input and into the controller.
Recorded simulation · 25 camera frames / second
The supplied iteration-4000 rollout. The animation follows its camera frame, not a synthetic ball trajectory.
64 × 96 metric depth → near-bright uint8 → square resize → three identical channels.
197 × 192 tokens · CLS + 196 patches. Token marks show the layout, not measured activations. Select a patch to locate its token.
16 token frames · one every 40 ms · 0.64 s buffer, oldest → newest
The learned stages connect cached frozen tokens to the percept. Click a stage for details.
13 actual parameter tensors · 624,769 parameters · checksum-verified iteration-4000 perception export
The matrices below use the same linear dimension scale as the SONIC decoder (0.082 SVG units per dimension). Thin vectors are outlined for visibility. Choose a matrix or a parameter to inspect its sampled values at a readable size.
The 3072 input columns are 16 consecutive blocks of 192. Each bar is the Frobenius norm of one column block, computed from the full tensor. Magnitude alone does not establish a slot’s causal importance.
Same color transform as SONIC: asinh(4096 × clip(w, −1, 1)) / asinh(4096). Positive red, negative blue, zero black. PNGs use nearest-neighbor samples at 256 pixels wide; statistics use every original tensor value. Tensor shapes, norms & provenance ↗
The auxiliary-head weights are absent from deployment ONNX and the released training checkpoint is unavailable. The percept animation remains labelled schematic: the replay contains rendered depth panels, not the original metric-depth arrays needed to reproduce measured activations.
Actual SONIC base and multiplied LoRA weights
Physical Unitree G1 · supplied deployment footage
The repository provides the LeRobot deployment path: head-camera depth and body state in, whole-body joint targets out. The supplied clip shows a physical G1 trial; its exact checkpoint and camera settings are not recorded alongside the file.
A hardware demonstration, not an evaluation: the simulation results below do not transfer automatically, and no hardware success rate has been measured yet.
Diagnostics screencast from a separate hardware trial
Approximate readouts from the separate diagnostics recording, including Dodge-mode log lines at elapsed runtime ≈ 54–57 s. The visible excerpt includes range-gated frames and nearly blank policy input. These readouts describe instrumentation, not dodge success, reaction latency, or the checkpoint in the newer deployment clip. The controller’s capture-to-history reading is inside the training corruption range of 0–40 ms; that alone does not establish transfer.
Earlier trial. Phone recording of a previous hardware test. Reference footage, not a measured success rate.
Source excerpts from the shipped controller and camera writer · pinned code revision ↗
The deployment uses two processes on the Jetson. A RealSense writer runs the frozen vision encoder on the GPU; the LeRobot process runs the temporal adapter and actor on the CPU. They exchange depth, tokens, and capture timestamps through shared memory.
480 × 270 at a requested 60 Hz. Crop using camera intrinsics, resize to 64 × 96, encode once per captured frame.
A sequence counter lets the reader reject incomplete or already-consumed snapshots.
Sample by capture time. Reuse the percept between camera updates while the actor reads fresh body state.
Run from the repository root in the matching Jetson environments
The writer and controller have separate Python environments because their RealSense, GPU, and DDS dependencies differ. Keep the working Jetson setup; the hardware guide records the known versions and remaining gaps.
Install the overlay into an existing controller environment without replacing its platform dependencies.
python3 scripts/prepare_lerobot.py --clone --apply --checkout external/lerobot
python3 scripts/fetch_assets.py --scope runtime
python -m pip install --no-deps -e external/lerobot
python robot/doctor.py --writer-python "$HOME/rs_pub_venv/bin/python" \
--out results/robot/environment.json
python scripts/check_release.py --weights-dir weights/releasedThis uses synthetic standing body state and the real camera pipeline. It does not connect an actuator controller.
python robot/run_dodge.py --dry-run --camera shm --mode dodge --seconds 30 \
--rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
--log results/robot/dryrun.jsonlThe detector defaults below apply here. Range gating is off unless supplied through --writer_args.
python robot/run_dodge.py --dry-run --camera shm --mode dodge --seconds 30 \
--rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
--ball_only --view_port 8080 --writer_args "--trace" \
--log results/robot/orange-mask-check.jsonlUse the robot’s established low-level-control procedure, support arrangement, and physical emergency stop. Replace eth0 with the configured DDS interface. The runner starts in stock SONIC mode; Space toggles Dodge and q ends the session.
python robot/run_dodge.py --real --network_interface eth0 --seconds 120 \
--rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
--log results/robot/run01.jsonlPress h, m, or f to annotate a hit, miss, or fall. Outcomes are operator labels, not automatic contact measurements.
python robot/summarize_session.py results/robot/run01.jsonl \
--out results/robot/run01-summary.jsonRaw full-scene depth is the default. --ball_only enables an additional RGB-assisted preprocessing path.
The camera helper makes a fresh spatial detection in each frame. HSV segmentation finds orange candidates; shape and depth checks choose a plausible ball. An accepted enclosing disc retains its measured depth. The learned adapter then uses depth history to infer motion — the detector’s centre, range, and score are never passed as a ball-state vector to the actor.
Align colour to the depth grid, threshold HSV, then apply a 5 × 5 elliptical opening and closing to the colour mask.
Check area and circularity. Estimate centre distance from the median valid depth in the inner 70% radius, plus the configured ball radius. Check apparent metric size.
If an enabled maximum range rejects the detection, record it as “hidden.” The gate’s default is 0: disabled.
Keep the raw depth inside the selected disc. Replace other pixels with 6 m by default, or invalid zero when explicitly configured. No depth hole-filling is performed by this masker.
| Setting | Value | Meaning |
|---|---|---|
| HSV lower / upper | (3, 120, 70) / (22, 255, 255) | Orange, in OpenCV’s HSV units |
| Ball radius / size tolerance | 0.12 m / 0.6 | Configured geometry; tolerance is relative radius error |
| Area / circularity | 25 px² / 0.5 | Minimum contour area and area ÷ enclosing-circle area |
| Depth support | At least 10 nonzero samples | Required inside the inner disc |
| Maximum ball range | 0 m: disabled | Optional experiment setting; record any override |
| Non-ball background | 6 m: far | Encodes as grayscale 1; invalid depth encodes as 0 |
Ball-only masking changes the input distribution from the full-scene depth used in training. The orange HSV thresholds are an experiment setting; they do not detect the black simulation ball by colour.
Exact BallMasker.apply implementation. Depth inside an accepted disc passes through unchanged.
def apply(self, color_bgr: np.ndarray, depth_mm: np.ndarray) -> np.ndarray:
det = self.detect(color_bgr, depth_mm)
self.last_far = None
if det is not None and self.max_range > 0 and det["z"] > self.max_range:
self.last_far, det = det, None
self.last = det
out = np.full_like(depth_mm, self.background_mm)
if det is not None:
disc = np.zeros(depth_mm.shape, np.uint8)
self.cv2.circle(disc, (int(det["u"]), int(det["v"])), max(int(det["r_px"]), 1), 255, -1)
keep = disc > 0
out[keep] = depth_mm[keep]
return outMedian-depth support, circularity, apparent radius, and a combined score select one candidate per frame.
def detect(self, color_bgr: np.ndarray, depth_mm: np.ndarray) -> dict | None:
cv2 = self.cv2
self.last_mask = self.mask(color_bgr)
contours = cv2.findContours(self.last_mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)[0]
best = None
self.rejects: list[str] = [] # why the largest orange blobs were not accepted (for --trace)
for contour in contours:
area = cv2.contourArea(contour)
if area < self.min_area:
continue
(u, v), r_px = cv2.minEnclosingCircle(contour)
circ = area / max(np.pi * r_px**2, 1.0)
if circ < self.min_circularity:
self.rejects.append(f"{int(area)}px circ={circ:.2f}")
continue
inner = np.zeros(depth_mm.shape, np.uint8)
cv2.circle(inner, (int(u), int(v)), max(int(0.7 * r_px), 2), 255, -1)
inside = depth_mm[(inner > 0) & (depth_mm > 0)]
if inside.size < 10:
self.rejects.append(f"{int(area)}px no-depth")
continue
centre_m = float(np.median(inside)) / 1000.0 + self.radius
metric_r = r_px * centre_m / self.fx
size_err = abs(metric_r - self.radius) / self.radius if self.radius > 0 else 0.0
if self.radius > 0 and size_err > self.size_tolerance:
self.rejects.append(f"{int(area)}px r={metric_r * 100:.0f}cm@{centre_m:.1f}m")
continue
score = size_err + 0.5 * (1.0 - circ)
if best is None or score < best["score"]:
best = {"u": u, "v": v, "r_px": r_px, "z": centre_m, "circ": circ, "score": score}
return bestThe opening and closing operate on the colour mask. They do not fill holes in measured depth.
def mask(self, color_bgr: np.ndarray) -> np.ndarray:
cv2 = self.cv2
hsv = cv2.cvtColor(color_bgr, cv2.COLOR_BGR2HSV)
if self.low[0] <= self.high[0]:
m = cv2.inRange(hsv, self.low, self.high)
else: # hue wraps through red
m = cv2.bitwise_or(
cv2.inRange(hsv, np.array([0, self.low[1], self.low[2]], np.uint8), self.high),
cv2.inRange(hsv, self.low, np.array([179, self.high[1], self.high[2]], np.uint8)),
)
m = cv2.morphologyEx(m, cv2.MORPH_OPEN, self.kernel)
return cv2.morphologyEx(m, cv2.MORPH_CLOSE, self.kernel)Optical-Z metres are rounded through float16, encoded as 255 near / 1 far / 0 invalid, then resized to 224 × 224. The image graph repeats grayscale across RGB channels.
def depth_to_theia_image_np(
depth_m: np.ndarray, near: float = DEPTH_NEAR_M, far: float = DEPTH_FAR_M
) -> np.ndarray:
"""numpy/cv2 twin of ``depth_dodge.depth_to_theia_image`` (255 near, 1 far, 0 invalid; bilinear to
224x224). Differs from the torch version by at most 1 grey level on ~1% of pixels (resize rounding)."""
import cv2 # noqa: PLC0415
depth = np.asarray(depth_m, np.float32).astype(np.float16).astype(np.float32)
valid = np.isfinite(depth) & (depth >= near) & (depth <= far)
value = np.round(np.clip(1.0 + 254.0 * (far - depth) / (far - near), 1.0, 255.0))
gray = np.where(valid, value, 0.0).astype(np.float32)
x = cv2.resize(gray, (THEIA_SIZE, THEIA_SIZE), interpolation=cv2.INTER_LINEAR)
return np.clip(np.round(x), 0, 255).astype(np.uint8)The full controller step shows the stock-SONIC branch, reference and proprioception packing, inference inputs, and residual-to-target conversion.
def _run_step(self, lowstate) -> dict:
q = np.array([lowstate.motor_state[m.value].q for m in G1_29_JointIndex], np.float32)
dq = np.array([lowstate.motor_state[m.value].dq for m in G1_29_JointIndex], np.float32)
quat = np.array(lowstate.imu_state.quaternion, np.float32)
quat = quat / (np.linalg.norm(quat) + 1e-8)
ang = np.array(lowstate.imu_state.gyroscope, np.float32)
frames = (ang, q - self.default_angles, dq, self.last_action.copy(), get_gravity_orientation(quat))
for hist, frame in zip(
(self.h_ang, self.h_q, self.h_dq, self.h_act, self.h_grav), frames, strict=True
):
if not hist: # training env fills the whole history from the first frame on reset
hist.extend([frame] * HISTORY_LEN)
else:
hist.append(frame)
if self._mode == "sonic":
# Stock SONIC on its neutral token; no token keys in ``action`` -> it holds the idle
# latent. Its residual comes back in its own joint order; keep ours (SDK order) fed so
# switching to dodge starts from a warm, consistent action history.
targets = self.sonic.run_step({}, lowstate)
self.last_action = self.sonic.last_action_mj[ISAACLAB_TO_MUJOCO].astype(np.float32)
return targets
policy = np.concatenate(
[np.concatenate(list(h)) for h in (self.h_ang, self.h_q, self.h_dq, self.h_act, self.h_grav)]
)
tokenizer = self._tokenizer(quat)
conditioning = np.concatenate([tokenizer, self._percept()[0]])
residual = self.actor.run(
None,
{
"tokenizer": tokenizer[None].astype(np.float32),
"policy": policy[None].astype(np.float32),
"conditioning": conditioning[None].astype(np.float32),
},
)[0][0].astype(np.float32)
if not np.isfinite(residual).all():
raise FloatingPointError("Non-finite Dodge action")
self.last_action = residual
target = self.default_angles + residual * self.action_scale
return {f"{m.name}.q": float(target[m.value]) for m in G1_29_JointIndex}Mechanisms verified in source. The recording’s timings are observations, not a controlled before/after benchmark.
The writer selects CUDA when available and warms up Theia with five calls. The small perception and actor sessions stay in the controller process, using CPU inference.
Inspect the writer ↘Store one Theia result per frame. The controller recomputes the 16-frame percept only when the history version changes, then reuses it between camera arrivals.
Inspect the cache ↘RealSense queues hold one frame. Shared memory publishes the newest tokens with a sequence counter; the reader accepts a copy only if the counter stayed unchanged.
Inspect the reader ↘The 60 Hz capture request is sampled toward 25 Hz for the policy. Missed cadence slots are skipped after stalls. Sixteen samples span 15 intervals: 0.60 s.
Inspect cadence ↘If a frame arrives while perception runs, the cache records the version it actually processed. The newer frame remains pending for the next computation.
Inspect the race fix ↘Use calibrated crop bounds and optical-Z depth. Round cached tokens through float16, then pass float32 to the exported adapter to match training features.
Inspect preprocessing ↘MJPEG rendering and JPEG encoding run on a separate thread. The writer submits at most every fourth frame while a viewer is connected; a pending preview frame is dropped.
Inspect the preview queue ↘Deployment uses the folded actor and a separate perception graph. The auxiliary head and critic stay in training. The training environment also caches frozen Theia features.
Inspect the learning path ↘These consecutive source lines crop and resize depth, call Theia once, match the cached precision, and publish tokens with a capture timestamp.
writer.write(depth_mm, t_cap)
policy_img = None
if theia is not None:
t0 = time.perf_counter()
roi = depth_mm[y0:y1, x0:x1]
small = cv2.resize(
roi, (POLICY_DEPTH_SHAPE[1], POLICY_DEPTH_SHAPE[0]), interpolation=cv2.INTER_NEAREST
)
depth_m = small.astype(np.float32) * 1e-3
policy_img = depth_to_theia_image_np(depth_m)
tokens = theia.run(None, {"image": policy_img[None]})[0][0]
tokens = tokens.astype(np.float16).astype(np.float32) # training cached fp16 features
dt_ms = (time.perf_counter() - t0) * 1e3
theia_ms = 0.9 * theia_ms + 0.1 * dt_ms
pwriter.write(tokens, depth_m, t_cap, dt_ms)The model runs outside the ring lock. The processed snapshot version is retained even if a new frame arrives during inference.
def _percept(self) -> np.ndarray:
with self._ring_lock:
version = self._ring_version
ring = np.stack(self.depth_ring)[None] if version != self._percept_version else None
if ring is not None:
# A new camera frame may arrive during inference. Record which
# snapshot was processed; never clear a newer frame's dirty flag.
percept = self.perception.run(None, {"depth_tokens": ring})[0]
self._cached_percept = percept
self._percept_version = version
return self._cached_perceptThe next deadline advances over missed slots instead of replaying a backlog of delayed frames.
class Cadence:
"""Pick frames at a fixed rate from a faster stream, by capture time (not arrival time), so the
ring contains 16 frames at 25 Hz like in training even though the writer runs at 60 Hz."""
def __init__(self, hz: float):
self.period = 1.0 / hz
self.next: float | None = None
def take(self, t_capture: float) -> bool:
if self.next is None:
self.next = t_capture + self.period
return True
if t_capture + 1e-4 < self.next:
return False
# Advance to the slot containing this frame; skip slots if we fell behind (stall).
self.next += self.period * max(1, math.floor((t_capture - self.next) / self.period) + 1)
return TrueOdd sequence values mean a write is in progress. The reader verifies the sequence before and after copying.
def try_read(self) -> tuple[np.ndarray, np.ndarray, dict] | None:
"""Return (tokens f32[197,192], depth_m f32[64,96], meta) for a new complete frame, else None."""
seq0, meta = self._header()
if seq0 & 1 or seq0 == self._last_seq or seq0 == 0:
return None
tokens, depth = self._tokens.copy(), self._depth.copy()
seq1, _ = self._header()
if seq1 != seq0:
return None
self._last_seq = seq0
meta["seq"] = seq0
return tokens, depth, metaThe worker performs rendering and JPEG encoding; the acquisition loop copies only what the preview needs.
def submit(self, render, *arrays) -> None:
"""Queue ``render(*arrays)`` for the worker; drops the frame if the previous one is still rendering."""
with self._pending_cond:
if self._pending is not None:
return
self._pending = (render, tuple(a.copy() if isinstance(a, np.ndarray) else a for a in arrays))
self._pending_cond.notify()If the camera times out, the runner refills the history with black-frame tokens and resets cadence until capture resumes. This is an input-recovery behavior; the dry-run and hardware validation record describe what has actually been checked.
PPO in simulation · 512 parallel worlds · one H100
PPO updates the temporal adapter, seven LoRAs, and critic while Theia and SONIC stay frozen. The critic reads privileged simulator state. A separate auxiliary update supervises the adapter and head using true ball state as its target; those auxiliary gradients do not update the decoder LoRAs.
Root-to-ball separation and low planar speed are combined with station holding, action-rate and joint-limit penalties. This task wrapper also adds hit and fall event penalties to the ORCS reward.
Reference, proprioception, base linear velocity, root state, previous actions, ball position and velocity, root-twist command, reward vector, and episode phase. The critic is never exported.
Robot randomization from ORCS plus a depth corruption model: pixel dropout, missing patches, repeated frames, range-dependent noise and 0–40 ms latency — so real camera artefacts are not a surprise.
The pure-PPO variants dodged 1.5–2.6% of resolved training throws. A head predicts normalized ball position and velocity from the percept, using its own Adam optimizer after PPO. It updates the adapter and head only, and is discarded at export.
512 worlds × 4,000 iterations, about 49 M transitions on one H100. The full lineage totals 14.05 h of logged iteration time; its 1001–4000 continuation totals 10.69 h. Frozen features are cached per frame rather than recomputed for every PPO minibatch.
Exact source excerpts; learned tensor maps above come from the released export.
The learned time code is added after token LayerNorm. A linear score pools the 197 tokens within each frame; the MLP combines the 16 pooled frame vectors.
def forward(self, depth=None, *, tokens=None):
if tokens is None:
tokens = self.raw_tokens(depth)
tokens = tokens.float()
tokens = self.token_norm(tokens)
tokens = tokens + self.time_embedding[None, :, None, :]
weights = torch.softmax(self.patch_score(tokens).squeeze(-1), dim=-1)
self.last_attention = weights.detach() # [B,T,197], display only
frame = (weights[..., None] * tokens).sum(dim=2)
return self.temporal(frame.reshape(frame.shape[0], -1))A separate optimizer runs after each PPO update. Actor.auxiliary_parameters includes the temporal adapter and auxiliary head, not the decoder LoRAs.
def _auxiliary_step(actor, algo, cfg, aux_optimizer):
if aux_optimizer is None or cfg.auxiliary_weight <= 0:
return None
stored = algo.storage.observations.flatten(0, 1)
total = 0.0
for _ in range(cfg.auxiliary_steps):
count = min(cfg.auxiliary_batch_size, stored.batch_size[0])
idx = torch.randperm(stored.batch_size[0], device=stored.device)[:count]
batch = stored[idx]
aux_optimizer.zero_grad(set_to_none=True)
raw = actor.auxiliary_loss(batch)
(cfg.auxiliary_weight * raw).backward()
torch.nn.utils.clip_grad_norm_(actor.auxiliary_parameters(), cfg.max_grad_norm)
aux_optimizer.step()
total += float(raw.detach())
return total / cfg.auxiliary_stepsTheia is frozen. This parameter list explicitly selects perception and auxiliary-head parameters only.
def auxiliary_parameters(self):
return [p for p in list(self.perception.trainable_parameters()) +
list(self.aux_head.parameters()) if p.requires_grad]Original training task · simulation, not hardware
Same task and frozen SONIC base, different visual learning recipes. Reward alone never taught the visual policy to dodge; the auxiliary ball-state loss on frozen Theia features did.
| Experiment | Worlds × iterations | Training¹ | Evaluation² |
|---|---|---|---|
| Theia / PPO only No token LayerNorm | 512 × 500 | 1.5% | 0.0% |
| Theia / PPO only With token LayerNorm | 512 × 1,000 | 2.6% | — |
| Theia / PPO only Smaller batch, with token LayerNorm | 256 × 2,000 | 1.6% | — |
| Depth CNN / auxiliary loss Vision trained from scratch | 512 × 2,000 | 2.1% | 0.0% |
| Theia / auxiliary loss Released policy · auxiliary weight 0.05 | 512 × 4,000 | 91.9% | 79.1% |
| Privileged oracle True ball state · simulation upper bound | 512 × 1,000 | 88.3% | 97.4% |
¹ Last 200 training iterations, with exploration and randomization. ² Deterministic actions, 64 worlds × 3,000 steps, three evaluation seeds, corruption on. “—” means not evaluated separately. Aborted throws and falls are separate outcomes.
The G1 scene, depth thumbnails, and attention overlay come from theia_aux_4000_60s (1).mp4. This page crops its displayed panels; it does not contain the original metric-depth arrays. The patch grid represents the 224² input after square resizing. The attention overlay is the recorded temporal adapter's patch weighting, not a newly computed transformer attention map.
Token highlights, pulses, and the percept grid are schematic. They do not claim to show measured activations, predicted ball state, or the policy's exact reaction latency. The 16-slot thumbnails show the frames represented by the history; each actual slot contains 197 × 192 tokens, not pixels.
The decoder heatmaps are sampled from real ONNX tensors. The plain and adapted exports have identical encoder tensors; the first decoder's 994-column base portion is identical. Later deltas are adapted weights minus matching plain weights. LoRA₀ is the 768-column conditioning portion. Original A/B factors were folded away; their displayed dimensions are exact, but their individual values are not reconstructed.
Weight colors use asinh(4096 × clip(w,−1,1)) / asinh(4096). Positive weights are red, negative weights blue, zero black. Matrix block proportions share a linear dimension scale. Thin rank-16 factors are outlined for visibility; the product is the actual scaled ΔW=(α/r)BA, rendered at 50% opacity.
Runtime inputs to actor.onnx: reference tokenizer[640], proprioception policy[930], and conditioning[768] = concat(tokenizer, percept). The encoder consumes the reference. Its FSQ output has 64 scalars; these join 930 proprioceptive values at the decoder. The exported first layer concatenates all three streams into 1762 inputs. LoRA₀ is folded into those extra 768 columns; later LoRAs are added to their base linear weights.
The final 29 outputs are normalized PD residuals. Per-joint action scales and the default pose convert them to joint position targets. These feed the G1; joint state, IMU and previous actions supply the next proprioceptive history.
Published depth-policy exports · Pinned Theia checkpoint. Theia is a DeiT-Tiny ViT, pretrained on RGB by distilling CLIP, Depth Anything, DINOv2, SAM and ViT. This experiment feeds it monochrome depth.
Reported training: 4000 PPO iterations and approximately 49 M transitions on one H100. The full iteration 1–4000 lineage totals 14.05 h of logged iteration time; the 1001–4000 continuation alone totals 10.69 h. Setup, compilation, and evaluation are additional. Original simulation evaluation: 79% dodged with vision, 0% with blank depth, 97% for the separately trained privileged oracle. These are the supplied report's results, not recomputed by this page.