How Megatron Inference router replay gets its expert routes into training, next to the vLLM path

Explainer for NVIDIA-NeMo/RL #4307. All NeMo-RL permalinks pin the PR head c7c37dc.

Before
Router replay worked only with vLLM generation. Megatron generation plus router replay was rejected at setup (setup.py:1260, router_replay.py:62, both before this PR).
This PR
Reshapes the expert routes from Megatron Inference into the same per-call rows vLLM already writes, so everything after staging reads them unchanged.
Status
Proposed — open PR, under review. The review raises a few issues next to this feature; this page only explains the feature.
Background. Router replay makes the trainer use the same experts, per token and per MoE layer, that the generation engine picked, so the two do not disagree on routing. Megatron Inference (MInf) is Megatron's own generation engine, the other choice next to vLLM. On the SingleController + Gym path, every model call is saved as it finishes, and only its new tokens are saved. Dotted words have a short explanation: hover or tap them.

Both engines give one route row per token except the last: shape [T-1, L, K] for T tokens, L MoE layers and top-K experts (MInf: inference_request.py:928). A later call in a multi-turn rollout stores only rows prev_len to T, where prev_len is what the earlier turns already stored. So the PR adds one all -1 row and drops the first prev_len rows (_delta_align_minf_routing_indices). NeMo-RL's vLLM worker already does the same two steps, but fills the added row with experts 0
K-1. The two paths differ only there, and in what happens when routes are bad.

Two engines, one staging row

Call 2 of the example below: 24 tokens, the first 7 already stored by call 1, 2 MoE layers, top-6 experts MInf engine routing_indices [23, 2, 6] int16 no row for the last token inference_request.py:928 NeMo-RL: MInf stager (new) add a -1 row → [24, 2, 6] keep rows [7:] → [17, 2, 6] missing or wrong-shape routes: call fails token_capture.py:96-104 _stage_admitted :347 Gym commit complete_call(extras=
) routes passed in by hand: Gym's Megatron reader returns no extras token_capture.py:361 Gym megatron.py:82 vLLM engine routed_experts [23, 2, 6] no row for the last token vllm/utils.py:258 NeMo-RL: vLLM worker add a 0
5 row → [24, 2, 6] keep rows [7:] → [17, 2, 6] bad routes: warn, drop, go on vllm/utils.py:325-339 vllm_worker_async.py:646 Gym commit complete_call_from_response Gym's vLLM reader pulls the trimmed routes out of the reply vllm_worker_async.py:686 Gym vllm.py:67 Shared from here TQTokenSink.stage rows must equal the new token count: 17 tq_token_sink.py:299 execute_route_plan joins both calls into one [24, 2, 6] row route_assembly.py:110 trainer -1 row → its own router router_replay.py:461 MInf only (new in this PR) vLLM only (already there) one shared code path for both

The same 2-turn rollout through both

Call 1: a 3-token prompt, 4 generated tokens, 7 in all. Call 2 continues it: those 7, then 8 new user tokens (a 15-token prompt), then 9 generated tokens, 24 in all. The model has 2 MoE layers and picks the top 6 experts.

# call 2 — prev_len = 7 (call 1's whole length), T = 24, so 17 new tokens MInf routing_indices [23, 2, 6] int16 # one row per token except the last add one all -1 row → [24, 2, 6] # token_capture.py:96-103 keep rows [7:] → [17, 2, 6] # :104 vLLM routed_experts [23, 2, 6] # also one row per token except the last add one 0
5 row → [24, 2, 6] # vllm/utils.py:325-339 keep rows [7:] → [17, 2, 6] # vllm_worker_async.py:646 both TQTokenSink.stage: 17 rows == 17 new tokens → staged # tq_token_sink.py:299
The full rollout the trainer sees: one column per token. Which route does each token get? call 1: prev_len 0, stores 7 rows call 2: prev_len 7, stores 17 rows token 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 MInf -1 -1 vLLM 0
5 0
5 prompt / user token generated token the engine's real route -1: trainer's router picks experts 0
5 forced Column 6 (dashed) is call 1's last generated token. In call 1 no engine had routed it yet, so each engine put its filler there. In call 2 both engines do route it, since it is part of call 2's prompt. But that row sits before prev_len = 7, so [7:] drops it. Column 23 is the rollout's last token. Its route only shapes its own output, which training does not score. documented in token-capture-ledger.md:132-139

What still differs after they meet

MInf (this PR)vLLM (already there)
Filler for a token with no route all -1: the trainer's own router picks (token_capture.py:96) experts 0
K-1, forced (utils.py:325)
Routes missing or the wrong shape the call fails, so Gym drops that rollout (token_capture.py:347, :313) routes dropped with a warning; the rollout is kept and those rows become -1 (vllm_worker_async.py:647, route_assembly.py:139)
Engine settings needs async_sched_mode: legacy and generation pipeline parallel size 1 (router_replay.py:76) no extra limits
Deferred route building not supported; setup rejects it (setup.py:1267) either setting works

Review comments on these rows: filler row — token-capture-ledger.md:133 engine settings — router_replay.py:112

Why keep both? vLLM stays the usual generation engine. MInf lets a run generate with Megatron itself, and before this PR such a run could not use router replay at all. In exchange, the MInf path is stricter — one bad route costs the whole rollout — and it needs the legacy scheduler and no pipeline split while generating.

What this means for you. Router replay now works with generation.backend: megatron on the SingleController + Gym token-capture path, and once the routes are staged the two engines share every line of code. Before you switch, set async_sched_mode: legacy and generation pipeline parallel size 1, keep deferred route building off, and expect missing routes to show up as dropped rollouts rather than warnings. The token at each turn boundary (column 6) gets no replayed route on either engine: MInf lets the trainer's router choose there, while vLLM forces experts 0
5.