Skip to the interactive architecture
Unitree G1 · perceptive whole-body control

Depth Dodge

A humanoid learns to dodge thrown balls from head-camera depth and its own body state.

nepyope · building on ViBe and ORCS

Scroll ↓ Real G1 ↓
G1 / SONIC / depth policy / iteration 4000

Depth → motion.

A frozen vision transformer, a learned temporal percept, and seven decoder adapters. Follow a recorded ball through the visual input and into the controller.

00.00 s

01 · See the ball

Recorded simulation · 25 camera frames / second

G1 in simulation

reference clip

The supplied iteration-4000 rollout. The animation follows its camera frame, not a synthetic ball trajectory.

The image entering Theia

224 × 224
196 patches · 16 × 16 pixels each

64 × 96 metric depth → near-bright uint8 → square resize → three identical channels.

Theia-Tiny

frozen · 5.5 M

197 × 192 tokens · CLS + 196 patches. Token marks show the layout, not measured activations. Select a patch to locate its token.

Frame 000

02 · Remember the approach

16 token frames · one every 40 ms · 0.64 s buffer, oldest → newest

Temporal adapter · trained · 625 k

The learned stages connect cached frozen tokens to the percept. Click a stage for details.

Per frame · ×16 · shared weights
Theia tokens197 × 192frozen · cached
Token LayerNorm192 per token384 params
+ time codeslot 1 … 1616 × 192
Patch scoreLinear 192 → 1193 params
Softmax pool197 tokens → 192attention weights
Across time · temporal MLP
Stack + flatten16 × 192 → 3072no params
LayerNorm30726,144 params
Linear3072 → 192590,016 params
SiLU192no params
Linear192 → 12824,704 params
LayerNorm → percept128256 params

Inside the learned adapter

13 actual parameter tensors · 624,769 parameters · checksum-verified iteration-4000 perception export

The matrices below use the same linear dimension scale as the SONIC decoder (0.082 SVG units per dimension). Thin vectors are outlined for visibility. Choose a matrix or a parameter to inspect its sampled values at a readable size.

Temporal mixing, grouped by history slot

The 3072 input columns are 16 consecutive blocks of 192. Each bar is the Frobenius norm of one column block, computed from the full tensor. Magnitude alone does not establish a slot’s causal importance.

Same color transform as SONIC: asinh(4096 × clip(w, −1, 1)) / asinh(4096). Positive red, negative blue, zero black. PNGs use nearest-neighbor samples at 256 pixels wide; statistics use every original tensor value. Tensor shapes, norms & provenance ↗

The auxiliary-head weights are absent from deployment ONNX and the released training checkpoint is unavailable. The percept animation remains labelled schematic: the replay contains rendered depth panels, not the original metric-depth arrays needed to reproduce measured activations.

percept · 128 values
schematic · no explicit ball-state output

03 · Turn perception into joint commands

Actual SONIC base and multiplied LoRA weights

Frozen baseTrained adapterTraining onlyRows = outputs · columns = inputs
−1+10 = black · asinh contrast · ΔW at 50% opacity

04 · On the real robot

Physical Unitree G1 · supplied deployment footage

From simulation
to a real G1.

The repository provides the LeRobot deployment path: head-camera depth and body state in, whole-body joint targets out. The supplied clip shows a physical G1 trial; its exact checkpoint and camera settings are not recorded alongside the file.

  1. ExportDecoder LoRAs are folded into the actor graph. The temporal adapter has its own perception graph; the auxiliary head and critic are omitted.
  2. SenseThe writer requests 60 Hz capture. The controller samples by capture time toward 25 Hz for its 16-frame Theia token history.
  3. ActThe LeRobot control loop targets 50 Hz, converting the actor’s residual actions into PD joint targets.

A hardware demonstration, not an evaluation: the simulation results below do not transfer automatically, and no hardware success rate has been measured yet.

Hardware setup and measurement protocol ↗

What the robot sees

Diagnostics screencast from a separate hardware trial

  1. Colour aligned to depth + detectionGreen indicates an accepted disc. Orange “hidden … m” indicates a detection rejected by the optional range gate.
  2. HSV maskColour thresholding and mask cleanup identify orange blobs. A white blob can still fail the downstream depth, size, or range checks.
  3. Depth published to the policyThe accepted disc keeps raw depth; the rest is far background. Uniform red in this excerpt is the 6 m background after the ball is hidden.
  4. Theia input · 224 × 224255 near, 1 far, 0 invalid. A uniform far background is almost black. The image graph repeats grayscale across three channels.
Control loop48–49 Hztrained at 50 Hz
Policy compute≈ 10 msbudget 20 ms per tick
Depth camera26 Hztrained at 25 Hz
Theia encode≈ 13 msonce per new frame
History span≈ 0.60 straining ring 0.64 s
Capture → history33–34 msnot capture-to-action latency

Approximate readouts from the separate diagnostics recording, including Dodge-mode log lines at elapsed runtime ≈ 54–57 s. The visible excerpt includes range-gated frames and nearly blank policy input. These readouts describe instrumentation, not dodge success, reaction latency, or the checkpoint in the newer deployment clip. The controller’s capture-to-history reading is inside the training corruption range of 0–40 ms; that alone does not establish transfer.

Earlier trial. Phone recording of a previous hardware test. Reference footage, not a measured success rate.

Inside the robot deployment

Source excerpts from the shipped controller and camera writer · pinned code revision ↗

The deployment uses two processes on the Jetson. A RealSense writer runs the frozen vision encoder on the GPU; the LeRobot process runs the temporal adapter and actor on the CPU. They exchange depth, tokens, and capture timestamps through shared memory.

CAMERA / GPU PROCESSRealSense → optional mask → Theia

480 × 270 at a requested 60 Hz. Crop using camera intrinsics, resize to 64 × 96, encode once per captured frame.

SHARED MEMORY197 × 192 tokens + depth + time

A sequence counter lets the reader reject incomplete or already-consumed snapshots.

LEROBOT / CPU PROCESS25 Hz history → 50 Hz joint targets

Sample by capture time. Reuse the percept between camera updates while the actor reads fresh body state.

Commands used by the integration

Run from the repository root in the matching Jetson environments

The writer and controller have separate Python environments because their RealSense, GPU, and DDS dependencies differ. Keep the working Jetson setup; the hardware guide records the known versions and remaining gaps.

1. Prepare the pinned overlay and model assets

Install the overlay into an existing controller environment without replacing its platform dependencies.

python3 scripts/prepare_lerobot.py --clone --apply --checkout external/lerobot
python3 scripts/fetch_assets.py --scope runtime
python -m pip install --no-deps -e external/lerobot
python robot/doctor.py --writer-python "$HOME/rs_pub_venv/bin/python" \
  --out results/robot/environment.json
python scripts/check_release.py --weights-dir weights/released
2. Check camera and inference without commanding joints

This uses synthetic standing body state and the real camera pipeline. It does not connect an actuator controller.

python robot/run_dodge.py --dry-run --camera shm --mode dodge --seconds 30 \
  --rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
  --log results/robot/dryrun.jsonl
3. Inspect orange-ball detection and masked policy input

The detector defaults below apply here. Range gating is off unless supplied through --writer_args.

python robot/run_dodge.py --dry-run --camera shm --mode dodge --seconds 30 \
  --rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
  --ball_only --view_port 8080 --writer_args "--trace" \
  --log results/robot/orange-mask-check.jsonl
4. Run the actuator path on the G1

Use the robot’s established low-level-control procedure, support arrangement, and physical emergency stop. Replace eth0 with the configured DDS interface. The runner starts in stock SONIC mode; Space toggles Dodge and q ends the session.

python robot/run_dodge.py --real --network_interface eth0 --seconds 120 \
  --rs_python "$HOME/rs_pub_venv/bin/python" --weights-dir weights/released \
  --log results/robot/run01.jsonl

Press h, m, or f to annotate a hit, miss, or fall. Outcomes are operator labels, not automatic contact measurements.

python robot/summarize_session.py results/robot/run01.jsonl \
  --out results/robot/run01-summary.json

Ball detection, depth masking, and temporal tracking

Raw full-scene depth is the default. --ball_only enables an additional RGB-assisted preprocessing path.

The camera helper makes a fresh spatial detection in each frame. HSV segmentation finds orange candidates; shape and depth checks choose a plausible ball. An accepted enclosing disc retains its measured depth. The learned adapter then uses depth history to infer motion — the detector’s centre, range, and score are never passed as a ball-state vector to the actor.

  1. Segment colour

    Align colour to the depth grid, threshold HSV, then apply a 5 × 5 elliptical opening and closing to the colour mask.

  2. Validate the blob

    Check area and circularity. Estimate centre distance from the median valid depth in the inner 70% radius, plus the configured ball radius. Check apparent metric size.

  3. Apply the optional range gate

    If an enabled maximum range rejects the detection, record it as “hidden.” The gate’s default is 0: disabled.

  4. Publish measured depth

    Keep the raw depth inside the selected disc. Replace other pixels with 6 m by default, or invalid zero when explicitly configured. No depth hole-filling is performed by this masker.

Camera-writer defaults in the pinned source
SettingValueMeaning
HSV lower / upper(3, 120, 70) / (22, 255, 255)Orange, in OpenCV’s HSV units
Ball radius / size tolerance0.12 m / 0.6Configured geometry; tolerance is relative radius error
Area / circularity25 px² / 0.5Minimum contour area and area ÷ enclosing-circle area
Depth supportAt least 10 nonzero samplesRequired inside the inner disc
Maximum ball range0 m: disabledOptional experiment setting; record any override
Non-ball background6 m: farEncodes as grayscale 1; invalid depth encodes as 0

Ball-only masking changes the input distribution from the full-scene depth used in training. The orange HSV thresholds are an experiment setting; they do not detect the black simulation ball by colour.

The code that selects the policy’s depth pixels

Exact BallMasker.apply implementation. Depth inside an accepted disc passes through unchanged.

def apply(self, color_bgr: np.ndarray, depth_mm: np.ndarray) -> np.ndarray:
    det = self.detect(color_bgr, depth_mm)
    self.last_far = None
    if det is not None and self.max_range > 0 and det["z"] > self.max_range:
        self.last_far, det = det, None
    self.last = det
    out = np.full_like(depth_mm, self.background_mm)
    if det is not None:
        disc = np.zeros(depth_mm.shape, np.uint8)
        self.cv2.circle(disc, (int(det["u"]), int(det["v"])), max(int(det["r_px"]), 1), 255, -1)
        keep = disc > 0
        out[keep] = depth_mm[keep]
    return out
How the detector rejects implausible orange blobs

Median-depth support, circularity, apparent radius, and a combined score select one candidate per frame.

def detect(self, color_bgr: np.ndarray, depth_mm: np.ndarray) -> dict | None:
    cv2 = self.cv2
    self.last_mask = self.mask(color_bgr)
    contours = cv2.findContours(self.last_mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)[0]
    best = None
    self.rejects: list[str] = []  # why the largest orange blobs were not accepted (for --trace)
    for contour in contours:
        area = cv2.contourArea(contour)
        if area < self.min_area:
            continue
        (u, v), r_px = cv2.minEnclosingCircle(contour)
        circ = area / max(np.pi * r_px**2, 1.0)
        if circ < self.min_circularity:
            self.rejects.append(f"{int(area)}px circ={circ:.2f}")
            continue
        inner = np.zeros(depth_mm.shape, np.uint8)
        cv2.circle(inner, (int(u), int(v)), max(int(0.7 * r_px), 2), 255, -1)
        inside = depth_mm[(inner > 0) & (depth_mm > 0)]
        if inside.size < 10:
            self.rejects.append(f"{int(area)}px no-depth")
            continue
        centre_m = float(np.median(inside)) / 1000.0 + self.radius
        metric_r = r_px * centre_m / self.fx
        size_err = abs(metric_r - self.radius) / self.radius if self.radius > 0 else 0.0
        if self.radius > 0 and size_err > self.size_tolerance:
            self.rejects.append(f"{int(area)}px r={metric_r * 100:.0f}cm@{centre_m:.1f}m")
            continue
        score = size_err + 0.5 * (1.0 - circ)
        if best is None or score < best["score"]:
            best = {"u": u, "v": v, "r_px": r_px, "z": centre_m, "circ": circ, "score": score}
    return best
HSV segmentation and morphological cleanup

The opening and closing operate on the colour mask. They do not fill holes in measured depth.

def mask(self, color_bgr: np.ndarray) -> np.ndarray:
    cv2 = self.cv2
    hsv = cv2.cvtColor(color_bgr, cv2.COLOR_BGR2HSV)
    if self.low[0] <= self.high[0]:
        m = cv2.inRange(hsv, self.low, self.high)
    else:  # hue wraps through red
        m = cv2.bitwise_or(
            cv2.inRange(hsv, np.array([0, self.low[1], self.low[2]], np.uint8), self.high),
            cv2.inRange(hsv, self.low, np.array([179, self.high[1], self.high[2]], np.uint8)),
        )
    m = cv2.morphologyEx(m, cv2.MORPH_OPEN, self.kernel)
    return cv2.morphologyEx(m, cv2.MORPH_CLOSE, self.kernel)
Depth → the image passed to Theia

Optical-Z metres are rounded through float16, encoded as 255 near / 1 far / 0 invalid, then resized to 224 × 224. The image graph repeats grayscale across RGB channels.

def depth_to_theia_image_np(
    depth_m: np.ndarray, near: float = DEPTH_NEAR_M, far: float = DEPTH_FAR_M
) -> np.ndarray:
    """numpy/cv2 twin of ``depth_dodge.depth_to_theia_image`` (255 near, 1 far, 0 invalid; bilinear to
    224x224). Differs from the torch version by at most 1 grey level on ~1% of pixels (resize rounding)."""
    import cv2  # noqa: PLC0415

    depth = np.asarray(depth_m, np.float32).astype(np.float16).astype(np.float32)
    valid = np.isfinite(depth) & (depth >= near) & (depth <= far)
    value = np.round(np.clip(1.0 + 254.0 * (far - depth) / (far - near), 1.0, 255.0))
    gray = np.where(valid, value, 0.0).astype(np.float32)
    x = cv2.resize(gray, (THEIA_SIZE, THEIA_SIZE), interpolation=cv2.INTER_LINEAR)
    return np.clip(np.round(x), 0, 255).astype(np.uint8)
Body state + percept → 29 joint targets

The full controller step shows the stock-SONIC branch, reference and proprioception packing, inference inputs, and residual-to-target conversion.

def _run_step(self, lowstate) -> dict:
    q = np.array([lowstate.motor_state[m.value].q for m in G1_29_JointIndex], np.float32)
    dq = np.array([lowstate.motor_state[m.value].dq for m in G1_29_JointIndex], np.float32)
    quat = np.array(lowstate.imu_state.quaternion, np.float32)
    quat = quat / (np.linalg.norm(quat) + 1e-8)
    ang = np.array(lowstate.imu_state.gyroscope, np.float32)

    frames = (ang, q - self.default_angles, dq, self.last_action.copy(), get_gravity_orientation(quat))
    for hist, frame in zip(
        (self.h_ang, self.h_q, self.h_dq, self.h_act, self.h_grav), frames, strict=True
    ):
        if not hist:  # training env fills the whole history from the first frame on reset
            hist.extend([frame] * HISTORY_LEN)
        else:
            hist.append(frame)

    if self._mode == "sonic":
        # Stock SONIC on its neutral token; no token keys in ``action`` -> it holds the idle
        # latent. Its residual comes back in its own joint order; keep ours (SDK order) fed so
        # switching to dodge starts from a warm, consistent action history.
        targets = self.sonic.run_step({}, lowstate)
        self.last_action = self.sonic.last_action_mj[ISAACLAB_TO_MUJOCO].astype(np.float32)
        return targets

    policy = np.concatenate(
        [np.concatenate(list(h)) for h in (self.h_ang, self.h_q, self.h_dq, self.h_act, self.h_grav)]
    )
    tokenizer = self._tokenizer(quat)
    conditioning = np.concatenate([tokenizer, self._percept()[0]])
    residual = self.actor.run(
        None,
        {
            "tokenizer": tokenizer[None].astype(np.float32),
            "policy": policy[None].astype(np.float32),
            "conditioning": conditioning[None].astype(np.float32),
        },
    )[0][0].astype(np.float32)
    if not np.isfinite(residual).all():
        raise FloatingPointError("Non-finite Dodge action")
    self.last_action = residual
    target = self.default_angles + residual * self.action_scale
    return {f"{m.name}.q": float(target[m.value]) for m in G1_29_JointIndex}

The engineering that keeps the loop moving

Mechanisms verified in source. The recording’s timings are observations, not a controlled before/after benchmark.

01 / SPLIT THE WORK

GPU vision, CPU control

The writer selects CUDA when available and warms up Theia with five calls. The small perception and actor sessions stay in the controller process, using CPU inference.

Inspect the writer ↘
02 / COMPUTE ONCE

Cache tokens and percepts

Store one Theia result per frame. The controller recomputes the 16-frame percept only when the history version changes, then reuses it between camera arrivals.

Inspect the cache ↘
03 / LIMIT BACKLOG

Latest complete snapshots

RealSense queues hold one frame. Shared memory publishes the newest tokens with a sequence counter; the reader accepts a copy only if the counter stayed unchanged.

Inspect the reader ↘
04 / MATCH TRAINING TIME

Sample by capture timestamp

The 60 Hz capture request is sampled toward 25 Hz for the policy. Missed cadence slots are skipped after stalls. Sixteen samples span 15 intervals: 0.60 s.

Inspect cadence ↘
05 / AVOID LOST UPDATES

Version the inference snapshot

If a frame arrives while perception runs, the cache records the version it actually processed. The newer frame remains pending for the next computation.

Inspect the race fix ↘
06 / PRESERVE THE CONTRACT

Match geometry and precision

Use calibrated crop bounds and optical-Z depth. Round cached tokens through float16, then pass float32 to the exported adapter to match training features.

Inspect preprocessing ↘
07 / KEEP MONITORING SEPARATE

Render diagnostics on a worker

MJPEG rendering and JPEG encoding run on a separate thread. The writer submits at most every fourth frame while a viewer is connected; a pending preview frame is dropped.

Inspect the preview queue ↘
08 / EXPORT THE RUNTIME PATH

Fold decoder LoRAs

Deployment uses the folded actor and a separate perception graph. The auxiliary head and critic stay in training. The training environment also caches frozen Theia features.

Inspect the learning path ↘
The per-frame GPU writer path

These consecutive source lines crop and resize depth, call Theia once, match the cached precision, and publish tokens with a capture timestamp.

writer.write(depth_mm, t_cap)
policy_img = None
if theia is not None:
    t0 = time.perf_counter()
    roi = depth_mm[y0:y1, x0:x1]
    small = cv2.resize(
        roi, (POLICY_DEPTH_SHAPE[1], POLICY_DEPTH_SHAPE[0]), interpolation=cv2.INTER_NEAREST
    )
    depth_m = small.astype(np.float32) * 1e-3
    policy_img = depth_to_theia_image_np(depth_m)
    tokens = theia.run(None, {"image": policy_img[None]})[0][0]
    tokens = tokens.astype(np.float16).astype(np.float32)  # training cached fp16 features
    dt_ms = (time.perf_counter() - t0) * 1e3
    theia_ms = 0.9 * theia_ms + 0.1 * dt_ms
    pwriter.write(tokens, depth_m, t_cap, dt_ms)
The versioned percept cache and race fix

The model runs outside the ring lock. The processed snapshot version is retained even if a new frame arrives during inference.

def _percept(self) -> np.ndarray:
    with self._ring_lock:
        version = self._ring_version
        ring = np.stack(self.depth_ring)[None] if version != self._percept_version else None
    if ring is not None:
        # A new camera frame may arrive during inference. Record which
        # snapshot was processed; never clear a newer frame's dirty flag.
        percept = self.perception.run(None, {"depth_tokens": ring})[0]
        self._cached_percept = percept
        self._percept_version = version
    return self._cached_percept
Capture-time cadence sampling

The next deadline advances over missed slots instead of replaying a backlog of delayed frames.

class Cadence:
    """Pick frames at a fixed rate from a faster stream, by capture time (not arrival time), so the
    ring contains 16 frames at 25 Hz like in training even though the writer runs at 60 Hz."""

    def __init__(self, hz: float):
        self.period = 1.0 / hz
        self.next: float | None = None

    def take(self, t_capture: float) -> bool:
        if self.next is None:
            self.next = t_capture + self.period
            return True
        if t_capture + 1e-4 < self.next:
            return False
        # Advance to the slot containing this frame; skip slots if we fell behind (stall).
        self.next += self.period * max(1, math.floor((t_capture - self.next) / self.period) + 1)
        return True
Accept only a new, complete shared-memory frame

Odd sequence values mean a write is in progress. The reader verifies the sequence before and after copying.

def try_read(self) -> tuple[np.ndarray, np.ndarray, dict] | None:
    """Return (tokens f32[197,192], depth_m f32[64,96], meta) for a new complete frame, else None."""
    seq0, meta = self._header()
    if seq0 & 1 or seq0 == self._last_seq or seq0 == 0:
        return None
    tokens, depth = self._tokens.copy(), self._depth.copy()
    seq1, _ = self._header()
    if seq1 != seq0:
        return None
    self._last_seq = seq0
    meta["seq"] = seq0
    return tokens, depth, meta
Drop preview work when a frame is already pending

The worker performs rendering and JPEG encoding; the acquisition loop copies only what the preview needs.

def submit(self, render, *arrays) -> None:
    """Queue ``render(*arrays)`` for the worker; drops the frame if the previous one is still rendering."""
    with self._pending_cond:
        if self._pending is not None:
            return
        self._pending = (render, tuple(a.copy() if isinstance(a, np.ndarray) else a for a in arrays))
        self._pending_cond.notify()

If the camera times out, the runner refills the history with black-frame tokens and resets cadence until capture resumes. This is an input-recovery behavior; the dry-run and hardware validation record describe what has actually been checked.

05 · Training with reinforcement learning

PPO in simulation · 512 parallel worlds · one H100

PPO updates the temporal adapter, seven LoRAs, and critic while Theia and SONIC stay frozen. The critic reads privileged simulator state. A separate auxiliary update supervises the adapter and head using true ball state as its target; those auxiliary gradients do not update the decoder LoRAs.

depth ×16 · proprio 930 · ref 640 29 joint targets privileged state + reward percept values aux → adapter + head only PPO → adapter + 7 LoRAs Simulator · 512 worldsMuJoCo Warp · frontal throwsrobot randomizationdepth corruption · 0–40 ms latency Actor · deployeddepth → Theia (frozen) → adapter → perceptSONIC (frozen) + 7 LoRAs → 29 actions Critic · training onlyprivileged state: true ball + robotasymmetric · never exported Aux head · training onlypercept → 128 → 64 → 6Smooth-L1 vs true ball · w 0.05 PPO updateclipped objective · GAEfixed exploration noise

Rewards

Root-to-ball separation and low planar speed are combined with station holding, action-rate and joint-limit penalties. This task wrapper also adds hit and fall event penalties to the ORCS reward.

What the critic sees

Reference, proprioception, base linear velocity, root state, previous actions, ball position and velocity, root-twist command, reward vector, and episode phase. The critic is never exported.

Randomization

Robot randomization from ORCS plus a depth corruption model: pixel dropout, missing patches, repeated frames, range-dependent noise and 0–40 ms latency — so real camera artefacts are not a surprise.

The auxiliary signal

The pure-PPO variants dodged 1.5–2.6% of resolved training throws. A head predicts normalized ball position and velocity from the percept, using its own Adam optimizer after PPO. It updates the adapter and head only, and is discarded at export.

Budget and optimization

512 worlds × 4,000 iterations, about 49 M transitions on one H100. The full lineage totals 14.05 h of logged iteration time; its 1001–4000 continuation totals 10.69 h. Frozen features are cached per frame rather than recomputed for every PPO minibatch.

Read the learning code

Exact source excerpts; learned tensor maps above come from the released export.

The actual temporal adapter forward pass

The learned time code is added after token LayerNorm. A linear score pools the 197 tokens within each frame; the MLP combines the 16 pooled frame vectors.

def forward(self, depth=None, *, tokens=None):
    if tokens is None:
        tokens = self.raw_tokens(depth)
    tokens = tokens.float()
    tokens = self.token_norm(tokens)
    tokens = tokens + self.time_embedding[None, :, None, :]
    weights = torch.softmax(self.patch_score(tokens).squeeze(-1), dim=-1)
    self.last_attention = weights.detach()  # [B,T,197], display only
    frame = (weights[..., None] * tokens).sum(dim=2)
    return self.temporal(frame.reshape(frame.shape[0], -1))
Auxiliary updates reach the adapter and head

A separate optimizer runs after each PPO update. Actor.auxiliary_parameters includes the temporal adapter and auxiliary head, not the decoder LoRAs.

def _auxiliary_step(actor, algo, cfg, aux_optimizer):
    if aux_optimizer is None or cfg.auxiliary_weight <= 0:
        return None
    stored = algo.storage.observations.flatten(0, 1)
    total = 0.0
    for _ in range(cfg.auxiliary_steps):
        count = min(cfg.auxiliary_batch_size, stored.batch_size[0])
        idx = torch.randperm(stored.batch_size[0], device=stored.device)[:count]
        batch = stored[idx]
        aux_optimizer.zero_grad(set_to_none=True)
        raw = actor.auxiliary_loss(batch)
        (cfg.auxiliary_weight * raw).backward()
        torch.nn.utils.clip_grad_norm_(actor.auxiliary_parameters(), cfg.max_grad_norm)
        aux_optimizer.step()
        total += float(raw.detach())
    return total / cfg.auxiliary_steps
Which parameters the auxiliary optimizer receives

Theia is frozen. This parameter list explicitly selects perception and auxiliary-head parameters only.

def auxiliary_parameters(self):
    return [p for p in list(self.perception.trainable_parameters()) +
            list(self.aux_head.parameters()) if p.requires_grad]

06 · Results

Original training task · simulation, not hardware

Same task and frozen SONIC base, different visual learning recipes. Reward alone never taught the visual policy to dodge; the auxiliary ball-state loss on frozen Theia features did.

Dodged fraction of resolved throws
ExperimentWorlds × iterationsTraining¹Evaluation²
Theia / PPO only No token LayerNorm512 × 5001.5%0.0%
Theia / PPO only With token LayerNorm512 × 1,0002.6%—
Theia / PPO only Smaller batch, with token LayerNorm256 × 2,0001.6%—
Depth CNN / auxiliary loss Vision trained from scratch512 × 2,0002.1%0.0%
Theia / auxiliary loss Released policy · auxiliary weight 0.05512 × 4,00091.9%79.1%
Privileged oracle True ball state · simulation upper bound512 × 1,00088.3%97.4%

¹ Last 200 training iterations, with exploration and randomization. ² Deterministic actions, 64 worlds × 3,000 steps, three evaluation seeds, corruption on. “—” means not evaluated separately. Aborted throws and falls are separate outcomes.

Released training lineage: hit fraction falls from nearly 1.0 to around 0.1 over approximately 49 million environment transitions, alongside auxiliary loss, perception gradient norm, and learning rate
Released lineage · 4,000 iterations · 14.05 h of logged iteration timeRaw logs ↗
Is it using vision? Dodged fraction, mean of three seeds. The oracle reads true ball state and is not the deployed actor.
Five moments in a simulated throw, showing the robot scene, corrupted depth input, and recorded patch weighting as the ball approaches
Scene, corrupted depth input, and recorded temporal-adapter patch weighting as the ball approaches.
What is recorded, measured, and schematic?

The G1 scene, depth thumbnails, and attention overlay come from theia_aux_4000_60s (1).mp4. This page crops its displayed panels; it does not contain the original metric-depth arrays. The patch grid represents the 224² input after square resizing. The attention overlay is the recorded temporal adapter's patch weighting, not a newly computed transformer attention map.

Token highlights, pulses, and the percept grid are schematic. They do not claim to show measured activations, predicted ball state, or the policy's exact reaction latency. The 16-slot thumbnails show the frames represented by the history; each actual slot contains 197 × 192 tokens, not pixels.

The decoder heatmaps are sampled from real ONNX tensors. The plain and adapted exports have identical encoder tensors; the first decoder's 994-column base portion is identical. Later deltas are adapted weights minus matching plain weights. LoRA₀ is the 768-column conditioning portion. Original A/B factors were folded away; their displayed dimensions are exact, but their individual values are not reconstructed.

Weight colors use asinh(4096 × clip(w,−1,1)) / asinh(4096). Positive weights are red, negative weights blue, zero black. Matrix block proportions share a linear dimension scale. Thin rank-16 factors are outlined for visibility; the product is the actual scaled ΔW=(α/r)BA, rendered at 50% opacity.

Export wiring and provenance

Runtime inputs to actor.onnx: reference tokenizer[640], proprioception policy[930], and conditioning[768] = concat(tokenizer, percept). The encoder consumes the reference. Its FSQ output has 64 scalars; these join 930 proprioceptive values at the decoder. The exported first layer concatenates all three streams into 1762 inputs. LoRA₀ is folded into those extra 768 columns; later LoRAs are added to their base linear weights.

The final 29 outputs are normalized PD residuals. Per-joint action scales and the default pose convert them to joint position targets. These feed the G1; joint state, IMU and previous actions supply the next proprioceptive history.

Published depth-policy exports · Pinned Theia checkpoint. Theia is a DeiT-Tiny ViT, pretrained on RGB by distilling CLIP, Depth Anything, DINOv2, SAM and ViT. This experiment feeds it monochrome depth.

Reported training: 4000 PPO iterations and approximately 49 M transitions on one H100. The full iteration 1–4000 lineage totals 14.05 h of logged iteration time; the 1001–4000 continuation alone totals 10.69 h. Setup, compilation, and evaluation are additional. Original simulation evaluation: 79% dodged with vision, 0% with blank depth, 97% for the separately trained privileged oracle. These are the supplied report's results, not recomputed by this page.