Explainer for NVIDIA-NeMo/RL #4266.
Permalinks pin the PR head 4db8f66; "before" links pin the base ed14402.
Ten pages on this stack โ this one: #4266: a training step that lands while Gym is paused.
step_10, and must describe the rollout data exactly as it was when those weights were saved.
With this PR, Gym's paused agents are saved in the same snapshot, so Gym has to stop first.
Dotted words have a short explanation: hover or tap them.An attempt records the current step (10) and prepares a folder under step_10
(single_controller.py:4626).
It then closes rollout admission and asks Gym to
drain and write its state
(:4648).
The training loop is not paused, so it keeps going. Only after Gym is done does the attempt take the data-plane barrier and compare the step again
(:4664).
If step 11 finished in between, the attempt deletes the folder, aborts Gym, and logs skipped
(:4911). A skipped attempt never counts as a failure.
Why it can't just finish into step_10 anyway. When step 11 commits, it deletes the rollout groups it trained on
(_cleanup_consumed_metas_unlocked).
A step_10 snapshot taken after that is missing those groups. But the step_10 weights never learned from them, so a restart from it would silently drop them.
So the check is right to throw the attempt away. The problem is that the throw-away now happens after a 70 s Gym pause, not before one.
How often. An attempt is wasted whenever a step finishes during the pause. In this example the pause (70 s) is longer than a step (40 s), so every attempt overlaps the end of some step, as long as the buffer holds enough finished rollouts to keep training going while new rollouts are stopped. The check before the pause only looks at whether a trainer checkpoint exists for the current step (:4598), not at whether a step is about to finish.
Nothing below is implemented.
Where the fix lands: RL only, one file.
single_controller.py:
the snapshot attempt opens a hold before Gym prepare and releases it after its cut; the training loop waits on that hold just before its optimizer step.
No Gym change: Gym only sees a normal prepare, commit and resume.
The two places the loop starts an optimizer step:
GRPO finish_train_step :3399
PPO critic epochs :3272
Holding before the optimizer step keeps the model and the data plane at step 10 until the cut is taken. The code already treats that point as safe: rows picked for the running step are kept in the snapshot and offered again after a restart (comment at :656). The cost moves from rollouts to training: step 11 waits 55 s, but only when a step happens to finish during a pause. Rollouts are stopped for the same 70 s either way, and now the snapshot lands. The hold must be released on every exit path, including a Gym abort, or training stops for good.
skipped with reason trainer_state_changed. The review asks the author to measure that rate first; the hold above is the fix if it is high.