ODD-conditioned safety filters β€” trained policies

Trained policies for ODD-conditioned safety filters on a Unitree Go2 in MuJoCo (mjlab): safety filters whose guarantee survives runtime changes of the operating design domain β€” a payload loaded mid-mission, a motor that derates. The filter is an automaton over specification modes (stand/walk ↔ rest) with certified transition funnels between them.

Code, documentation and the evaluation that reproduces every result: SafeRoboticsLab/odd-conditioned (branch cleanup). These weights are meant to be used with it.

Use

# in a clone of the code repository, after `source activate.sh`
bash scripts/fetch_weights.sh https://huggingface.co/buzinguyen/odd-conditioned-dev/resolve/c923466bd1c68132b643f6806b9e8a2c532e4f1f/odd-conditioned-weights-v1.tar.gz
pytest -q tests/ && bash scripts/reproduce.sh payload figures

fetch_weights.sh installs the policies under checkpoints/ and checks every file against the manifest committed in the code repository (weights/MANIFEST.sha256, also included here).

Files

path contents
odd-conditioned-weights-v1.tar.gz everything below, as one archive (what fetch_weights.sh installs)
checkpoints/<name>/model.zip a two-player reach-avoid PPO twin (ReachAvoidPPO2P, safety-stable-baselines 0.4.0): control policy + state-value net, whose sign is the mode's certificate
checkpoints/<name>/tensornormalize.pt its frozen observation-normalization statistics (48-d actor observation)
checkpoints/<name>/config.yaml the exact training configuration of the run
checkpoints/{rest,getup_stage1}/final/ final models used as warm starts when retraining descend and getup
MANIFEST.sha256 sha256 of every file
policy role
stand STAND expert under a carried load W ∈ [0, 120] N with a raised centre of mass; its value is V_stand
rest REST expert: lie down and settle under any load β€” the anchor safe set
getup certified get-up funnel REST β†’ STAND; its value V_up gates the return and the abort
descend certified descent funnel STAND β†’ REST
stand_wide STAND expert trained on W ∈ [0, 150] N (a noisier certificate; the single-spec comparison)
unified, unified_discounted single-specification baselines (one policy for stand-or-rest)
leg_stand STAND expert for a derated front-right leg (the negative control)
compound_stand, compound_rest STAND / REST experts for a leg that derates while carrying 80 N

The nominal walking policy is not here: it ships with go2_atomic_skills.

Training

Every policy was trained with the two-player safety-PPO recipe of safety-stable-baselines on tasks of robot-safety-sandbox (branch project/odd-conditioned): 1024 environments, 50M steps, seed 0, a learned adversarial push of up to 25 N. The released checkpoint of each run is the one at 49,999,872 steps. scripts/train.sh in the code repository retrains any of them.

Training variance. Each policy is a single training run, and the stance-type policies (stand, compound_stand) vary a lot from run to run: retrained with other seeds, most stand experts fail the acceptance test that the released one passes. The code repository documents an acceptance test (scripts/check_policies.py) and a threshold-calibration script for retrained policies.

License

MIT, as the code.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading