Why a slow Gym snapshot can stop every later snapshot
NVIDIA-NeMo/RL #4266 (stack part 3 of 3).
RL links pin 98b3b6c (#4266) and ed14402 (#4265); Gym links pin f4fcf8c.
Reproduced with Gym's real checkpoint code and RL's real deadline code, in-process; not run on a cluster.
Before
RL's rollout deadline paused only while a colocated engine switched to training.
This PR
Adds Gym snapshots that park in-flight rollouts, but leaves RL's rollout deadline running while they are parked.
Status
Open — not fixed in this PR (#4266). Raised in the review comments.
Ten pages on this stack — this one: #4266: a snapshot that outlasts a rollout's deadline.
Background. A Gym snapshot pauses every in-flight rollout at a saved point (Gym calls this
parking)
until the snapshot is saved. RL gives each prompt group a time budget,
rollout_timeout_s.
Gym keeps every finished rollout until RL sends an
ACK
for it. Dotted words have a short explanation: hover or tap them.
The snapshot parks a rollout, but RL's clock for that rollout keeps running: the only place #4266
pauses RL's deadlines is the switch to training on a colocated engine.
If the snapshot lasts longer than the rollout's remaining budget, RL times the rollout out while it is
parked. RL then stops reading it but does not cancel it, so Gym finishes it after the snapshot, and nobody
ever ACKs it.
RL's deadline keeps running. One _Deadlinewraps the whole prompt-group stream, and its
registry is not paused for a snapshot.
At 180 s it fires while try 0 is parked, and RL gives up and starts try 1.
RL does not cancel try 0. It reads rows from a
Ray streaming call and never cancels it,
so the /run keeps going inside Gym.
Nobody ACKs it. RL gave up on try 0, so it never fetches its receipt. Gym's
retire() refuses a finished run, so nothing else can release it.
Every later snapshot fails. Gym's status()
is never ready while a finished run is un-ACKed, so snapshot 2 at 450 s and every one after it waits out
prepare_timeout_s (300 s by default) and fails.
This is a new way to reach the lost-reply stall: there, RL
never uses a finished reply; here, the snapshot itself is what makes RL give up on the reply. The numbers above
are the ones #4266's functional test uses (rollout_timeout_s=180) and #4266's default
prepare_timeout_s=300; any snapshot that outlasts a rollout's remaining budget does the same.
# scratch run with Gym's real AgentCheckpointParticipant and RL's real _Deadline (budget shortened to 0.05 s)
prepare: ready_to_commit=True parked=1 (RL deadline registry suspended=False)
RL outcome: {'timeout': 'NeMo-Gym prompt group exceeded 0.05s', 'state_when_rl_gave_up': 'parked'}
agent: bodies started=1 finished=1
next snapshot's prepare: ready_to_commit=False completed_unacknowledged=1
What would fix it — proposed, not in the PR
Where the fix lands: RL only, single_controller.py, plus a small change to
RequestDeadlineRegistry. No Gym change.
Nothing below is implemented. Pause RL's rollout deadlines for as long as a Gym snapshot holds admission,
the same way RL already pauses them for the switch to training:
# single_controller.py, _prepare_and_commit_gym_checkpoint (proposed)
self._gym_checkpoint_rollout_permitted.clear() # L1958: admission closes at 150 s
self._rollout_manager.suspend_request_deadlines() # try 0 has 30 s left; clock stops
...
# every place admission reopens resumes the clocks:# L2022 (prepare never attempted) and L2041 (after Gym resumes, commit or abort)
self._gym_checkpoint_rollout_permitted.set()
self._rollout_manager.resume_request_deadlines() # 210 s: try 0 still has 30 s left
One detail: RequestDeadlineRegistry keeps a single on/off flag. A second suspend() returns early, and resume() restarts every
live deadline. The switch to training uses the same flag, so without a count of who paused it, a snapshot that ends
while the engine is still training would restart the training switch's clocks early. The registry needs to count
holders and resume only when the last one lets go.
The same scratch run with the pause added: prepare: ready_to_commit=True parked=1 (RL deadline registry
suspended=True), then RL outcome: {'reply': {'reward': 1.0}} and finished=1.
Links: L2022,
L2041,
suspend_request_deadlines / resume_request_deadlines.
So what. With turn-level checkpointing on and a rollout deadline set, any snapshot that takes
longer than some rollout's remaining budget strands a finished result, and from then on no Gym-aware snapshot
can be saved for the rest of the process. Pausing the deadline during the snapshot removes the trigger; the
retire-then-adopt fix on the lost-reply page removes the stall itself.