Why a slow Gym snapshot can stop every later snapshot

NVIDIA-NeMo/RL #4266 (stack part 3 of 3). RL links pin 98b3b6c (#4266) and ed14402 (#4265); Gym links pin f4fcf8c. Reproduced with Gym's real checkpoint code and RL's real deadline code, in-process; not run on a cluster.

Before
RL's rollout deadline paused only while a colocated engine switched to training.
This PR
Adds Gym snapshots that park in-flight rollouts, but leaves RL's rollout deadline running while they are parked.
Status
Open — not fixed in this PR (#4266). Raised in the review comments.

Ten pages on this stack — this one: #4266: a snapshot that outlasts a rollout's deadline.

Background. A Gym snapshot pauses every in-flight rollout at a saved point (Gym calls this parking) until the snapshot is saved. RL gives each prompt group a time budget, rollout_timeout_s. Gym keeps every finished rollout until RL sends an ACK for it. Dotted words have a short explanation: hover or tap them.

The snapshot parks a rollout, but RL's clock for that rollout keeps running: the only place #4266 pauses RL's deadlines is the switch to training on a colocated engine. If the snapshot lasts longer than the rollout's remaining budget, RL times the rollout out while it is parked. RL then stops reading it but does not cancel it, so Gym finishes it after the snapshot, and nobody ever ACKs it.

One rollout, one snapshot

One rollout, budget rollout_timeout_s = 180 s; a snapshot starts at 150 s and takes 60 s 0 s 150 180 240 450 snapshot 1: rollout parked snapshot 2: 409 … RL, try 0 deadline fires while parked → RL gives up, starts try 1 Gym /run, try 0 finished, never ACKed Gym prepare ready → commits never ready running parked by the snapshot finished, not ACKed RL's rollout deadline RL stops reading try 0 at 180 s but never cancels it, so Gym resumes it after the snapshot and it finishes. Nobody holds its receipt, so nobody ACKs it, and Gym will not commit while it is un-ACKed. Snapshot 2 and every later one wait out prepare_timeout_s (300 s) and fail, for the rest of the process.

What happens, step by step

  1. A snapshot starts at 150 s. RL closes rollout admission (single_controller.py:1958) and Gym parks every running rollout at its next turn boundary.
  2. RL's deadline keeps running. One _Deadline wraps the whole prompt-group stream, and its registry is not paused for a snapshot. At 180 s it fires while try 0 is parked, and RL gives up and starts try 1.
  3. RL does not cancel try 0. It reads rows from a Ray streaming call and never cancels it, so the /run keeps going inside Gym.
  4. Snapshot 1 commits at 210 s and Gym resumes. Try 0 continues and finishes at 240 s. Gym keeps its reply until it is ACKed.
  5. Nobody ACKs it. RL gave up on try 0, so it never fetches its receipt. Gym's retire() refuses a finished run, so nothing else can release it.
  6. Every later snapshot fails. Gym's status() is never ready while a finished run is un-ACKed, so snapshot 2 at 450 s and every one after it waits out prepare_timeout_s (300 s by default) and fails.

This is a new way to reach the lost-reply stall: there, RL never uses a finished reply; here, the snapshot itself is what makes RL give up on the reply. The numbers above are the ones #4266's functional test uses (rollout_timeout_s=180) and #4266's default prepare_timeout_s=300; any snapshot that outlasts a rollout's remaining budget does the same.

# scratch run with Gym's real AgentCheckpointParticipant and RL's real _Deadline (budget shortened to 0.05 s) prepare: ready_to_commit=True parked=1 (RL deadline registry suspended=False) RL outcome: {'timeout': 'NeMo-Gym prompt group exceeded 0.05s', 'state_when_rl_gave_up': 'parked'} agent: bodies started=1 finished=1 next snapshot's prepare: ready_to_commit=False completed_unacknowledged=1

What would fix it — proposed, not in the PR

Where the fix lands: RL only, single_controller.py, plus a small change to RequestDeadlineRegistry. No Gym change.

Nothing below is implemented. Pause RL's rollout deadlines for as long as a Gym snapshot holds admission, the same way RL already pauses them for the switch to training:

# single_controller.py, _prepare_and_commit_gym_checkpoint (proposed) self._gym_checkpoint_rollout_permitted.clear() # L1958: admission closes at 150 s self._rollout_manager.suspend_request_deadlines() # try 0 has 30 s left; clock stops ... # every place admission reopens resumes the clocks: # L2022 (prepare never attempted) and L2041 (after Gym resumes, commit or abort) self._gym_checkpoint_rollout_permitted.set() self._rollout_manager.resume_request_deadlines() # 210 s: try 0 still has 30 s left

One detail: RequestDeadlineRegistry keeps a single on/off flag. A second suspend() returns early, and resume() restarts every live deadline. The switch to training uses the same flag, so without a count of who paused it, a snapshot that ends while the engine is still training would restart the training switch's clocks early. The registry needs to count holders and resume only when the last one lets go. The same scratch run with the pause added: prepare: ready_to_commit=True parked=1 (RL deadline registry suspended=True), then RL outcome: {'reply': {'reward': 1.0}} and finished=1. Links: L2022, L2041, suspend_request_deadlines / resume_request_deadlines.

🔴 review comment — single_controller.py:1958 (the review is still pending, so this link opens for others only after it is published)

So what. With turn-level checkpointing on and a rollout deadline set, any snapshot that takes longer than some rollout's remaining budget strands a finished result, and from then on no Gym-aware snapshot can be saved for the rest of the process. Pausing the deadline during the snapshot removes the trigger; the retire-then-adopt fix on the lost-reply page removes the stall itself.