How Megatron Inference router replay gets its expert routes into training, next to the vLLM path
Explainer for NVIDIA-NeMo/RL #4307.
All NeMo-RL permalinks pin the PR head c7c37dc.
Before
Router replay worked only with vLLM generation. Megatron generation plus router replay was rejected at setup
(setup.py:1260,
router_replay.py:62, both before this PR).
This PR
Reshapes the expert routes from Megatron Inference into the same per-call rows vLLM already writes, so everything after staging reads them unchanged.
Status
Proposed â open PR, under review. The review raises a few issues next to this feature; this page only explains the feature.
Background.Router replay makes the trainer use the same experts, per token and per MoE layer,
that the generation engine picked, so the two do not disagree on routing. Megatron Inference (MInf) is Megatron's own
generation engine, the other choice next to vLLM. On the SingleController + Gym path,
every model call is saved as it finishes, and only its new tokens are saved.
Dotted words have a short explanation: hover or tap them.
Both engines give one route row per token except the last: shape [T-1, L, K] for T tokens,
L MoE layers and top-K experts
(MInf: inference_request.py:928).
A later call in a multi-turn rollout stores only rows prev_len to T, where
prev_len
is what the earlier turns already stored. So the PR adds one all -1 row and drops the first prev_len rows
(_delta_align_minf_routing_indices).
NeMo-RL's vLLM worker already does the same two steps, but fills the added row with experts 0âŠK-1.
The two paths differ only there, and in what happens when routes are bad.
Two engines, one staging row
The same 2-turn rollout through both
Call 1: a 3-token prompt, 4 generated tokens, 7 in all. Call 2 continues it: those 7, then 8 new user tokens
(a 15-token prompt), then 9 generated tokens, 24 in all. The model has 2 MoE layers and picks the top 6 experts.
# call 2 â prev_len = 7 (call 1's whole length), T = 24, so 17 new tokens
MInf routing_indices [23, 2, 6] int16 # one row per token except the last
add one all -1 row â [24, 2, 6] # token_capture.py:96-103
keep rows [7:] â [17, 2, 6] # :104
vLLM routed_experts [23, 2, 6] # also one row per token except the last
add one 0âŠ5 row â [24, 2, 6] # vllm/utils.py:325-339
keep rows [7:] â [17, 2, 6] # vllm_worker_async.py:646
both TQTokenSink.stage: 17 rows == 17 new tokens â staged # tq_token_sink.py:299
Why keep both? vLLM stays the usual generation engine. MInf lets a run generate with
Megatron itself, and before this PR such a run could not use router replay at all. In exchange, the MInf path
is stricter â one bad route costs the whole rollout â and it needs the legacy scheduler and no pipeline split while generating.
What this means for you. Router replay now works with generation.backend: megatron on the
SingleController + Gym token-capture path, and once the routes are staged the two engines share every line of code.
Before you switch, set async_sched_mode: legacy and generation pipeline parallel size 1, keep deferred route building off,
and expect missing routes to show up as dropped rollouts rather than warnings. The token at each turn boundary (column 6)
gets no replayed route on either engine: MInf lets the trainer's router choose there, while vLLM forces experts 0âŠ5.