ODD-conditioned safety filters β trained policies
Trained policies for ODD-conditioned safety filters on a Unitree Go2 in MuJoCo (mjlab): safety filters whose guarantee survives runtime changes of the operating design domain β a payload loaded mid-mission, a motor that derates. The filter is an automaton over specification modes (stand/walk β rest) with certified transition funnels between them.
Code, documentation and the evaluation that reproduces every result:
SafeRoboticsLab/odd-conditioned (branch
cleanup). These weights are meant to be used with it.
Use
# in a clone of the code repository, after `source activate.sh`
bash scripts/fetch_weights.sh https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/c923466bd1c68132b643f6806b9e8a2c532e4f1f/odd-conditioned-weights-v1.tar.gz
pytest -q tests/ && bash scripts/reproduce.sh payload figures
fetch_weights.sh installs the policies under checkpoints/ and checks every file against the manifest
committed in the code repository (weights/MANIFEST.sha256, also included here).
Files
| path | contents |
|---|---|
odd-conditioned-weights-v1.tar.gz |
everything below, as one archive (what fetch_weights.sh installs) |
checkpoints/<name>/model.zip |
a two-player reach-avoid PPO twin (ReachAvoidPPO2P, safety-stable-baselines 0.4.0): control policy + state-value net, whose sign is the mode's certificate |
checkpoints/<name>/tensornormalize.pt |
its frozen observation-normalization statistics (48-d actor observation) |
checkpoints/<name>/config.yaml |
the exact training configuration of the run |
checkpoints/{rest,getup_stage1}/final/ |
final models used as warm starts when retraining descend and getup |
MANIFEST.sha256 |
sha256 of every file |
| policy | role |
|---|---|
stand |
STAND expert under a carried load W β [0, 120] N with a raised centre of mass; its value is V_stand |
rest |
REST expert: lie down and settle under any load β the anchor safe set |
getup |
certified get-up funnel REST β STAND; its value V_up gates the return and the abort |
descend |
certified descent funnel STAND β REST |
stand_wide |
STAND expert trained on W β [0, 150] N (a noisier certificate; the single-spec comparison) |
unified, unified_discounted |
single-specification baselines (one policy for stand-or-rest) |
leg_stand |
STAND expert for a derated front-right leg (the negative control) |
compound_stand, compound_rest |
STAND / REST experts for a leg that derates while carrying 80 N |
The nominal walking policy is not here: it ships with go2_atomic_skills.
Training
Every policy was trained with the two-player safety-PPO recipe of safety-stable-baselines on tasks of
robot-safety-sandbox
(branch project/odd-conditioned): 1024 environments, 50M steps, seed 0, a learned adversarial push of up to
25 N. The released checkpoint of each run is the one at 49,999,872 steps. scripts/train.sh in the code
repository retrains any of them.
Training variance. Each policy is a single training run, and the stance-type policies (stand,
compound_stand) vary a lot from run to run: retrained with other seeds, most stand experts fail the acceptance
test that the released one passes. The code repository documents an acceptance test (scripts/check_policies.py)
and a threshold-calibration script for retrained policies.
License
MIT, as the code.