Putting the Inference Engine in an RSI Loop?
Over the last couple of months, I have been diving more into the RL post-training world and the systems around it. The goal is to leverage what I know about Inference Engines and Kernels already and understand what it would take to put the engine in an RSI loop. Getting the Kernels there in isolation is almost a solved and a more tractable problem. The complexity blows up exponentially when increasing that scope to the inference server + inference engine.
Inference engine itself can be broken down into numerous complex components. The inference stack holds the scheduler, KV cache manager, router, kernels, and so much more. So, when it comes to optimizing it, the obvious approach seems like to deploy a multi-agent swarm where each agent takes a part of the stack and optimizes it while coordinating with other agents.
So, at 50,000-feet level, how would one go about designing such a system? Some thoughts, drawing inspiration from the recent works from Kimi K2.5 and Prime Intellect's Multi-Agent Systems -
Environments
For kernels, the environments are reasonably tractable. You get a GPU, get nsys and ncu running, specify the metrics to look for and know how to design the reward function for the same. Put this in a loop and you have a kernel-generating machine. The last few model iterations have made this task more and more autonomous.
But when it comes to engines - the story is a bit more involved. The shape of the traffic directly influences the optimizations an engine should be running with, the distribution of the input length, output length, batch size, whether there is cache-reuse across conversations and requests, and so much more.
Multi-agent swarms only complicate this further. You need to account for prefix sharing across agents, some or all of the agents idling during tool calls and yet holding the KV Cache in the memory, one agent finishing much earlier and then waiting for others to finish.
The challenge here is designing an environment and the workload that provides all this experience to the agents. You need to have agents monitor, optimize, deploy, monitor, .... your stack. While there are some great multi-agent frameworks including Prime Intellect's PRIME-RL, MARTI, and AgentJet that allow training teams of agents and can be hooked up with environments like MultiAgentBench and TextArena, I didn't seem to find any that was providing engine environments and tasks, even less so multi-agent environments. There is certainly InferenceX AgentX that provides the ability to play agentic traffic through different inference engines (like vLLM, SGLang, TRT-LLM) but that, in my opinion, still falls short of an environment where agents can optimize, see the impact of the swarm of engines at work and go to work again.
Rewards
Now, once you have the environment available where you run real swarms tasks and are able to capture the metrics that the agent can act on, how do you go about designing the learning loop? The first question in that direction is what signals convey the impact of optimizations and design changes introduced by the agents.
In general, any optimization should ensure that at least the following are held true -
- Logprob drift D stays under
- Swarm's task success rate shouldn't drop (not more than epsilon)
- Requests not already terminating due to staleness now don't start dropping
- p90/p99 resume latency stays under the SLO
Considering all these, let's take a quick look into a potential Reward design.
Reward 1: Do agents (requests) from the same swarm go to the same engine or are distributed across
Whether a swarm's agents stay on one replica to take advantage of prefix hits or spread out and transfer KV between replicas is one of the many metrics the recipe needs to optimize for. The choice, at the end of the day, is to find a compromise between TTFT, ITL, memory consumed, KV transfer time and the reward function has to account for all of them.
So, designing the reward accounting for these in a multi-agent optimization system will look something like -
For the KV manager agent-
- - KV cache held by agents waiting on tool calls, in GB·s
- - tokens re-prefilled, either because their KV was evicted during a tool call or because the prefix existed somewhere in the fleet but not on the replica the request landed on
and for the router agent, again plus -
- - KV bytes moved between replicas
- - time requests spend queued on overloaded replicas
The reward function then becomes -
where , i.e. how much a metric improved over the stock engine.
This can be added on top of the end-to-end reward that every agent gets for the patch as a whole - how much faster the swarm finished, and perhaps how many more swarms can fit on the same deployment. So the KV manager agent gets
and the router agent
One thought here is whether should annealed to 0 over training, so that in the end the agents optimize only for the end-to-end and not the per-agent steps (to-be-decided). Each agent's policy then maximizes its expected reward -
Anyways, the exact shape of the reward function would very much depend on the cluster, application, engine design among other things. Moving on from the reward function and time to focus on the impact of the system performance on the policy design, i.e. Rollouts.
Trajectories and Rollouts
This might be a bit tricky at scale but let's look at it. In a trajectory, the agent would be expected to execute and profile, generate code, execute and profile again, update the code, execute and profile again and continue till it hits end of sequence or reaches the end of the trajectory. There are at least a couple of ways in which the rewards can be captured here - either at the end of the trajectory based on the metrics decided above or intermediate rewards.
One example of intermediate rewards for Kernel generation is that used by Cognition with Kevin - it scores every refinement turn and the reward for each turn is the discounted sum of the scores of the current kernel and all the ones after it.
with in their final run. For an engine, the intermediate would come from the local replays, and the last one from the judge on held-out tasks.
Now, scoping this out to the entire engine (and swarms of engine) execution immediately blows up the time each trajectory ends up taking. Kernels are a few seconds at most where every engine update requires restarting the engine - weight loading, warming up, capturing cuda graphs and so on. This way, a GRPO group of 8 rollouts for one task could potentially take hours each. I am still in the middle of designing this (and actually reading more about how to approach this) but so far it looks we need to approach this tiers. The first signals can come from speed-of-light models for kernels and op-level changes, and eventually move to execution environments.
These execution environments also need to be more tractable in the sense, doing a rack-scale optimization loop would absolutely suck and might not provided additional insights beyond what a Data Parallel group could offer. So obviously, smaller models and a smaller world size should certainly be more feasible and also potentially offer quicker iterations.
The trajectories get long enough that the trainer can't wait for all of them to finish - which brings us to the next part - Async RL.
Add some Asynchronous RL
With the trajectories potentially taking hours (and perhaps longer in some cases), a synchronous loop doesn't work. The trainer would need to wait for the slowest trajectory unless the trainer can go forward without waiting for all the trajectories to finish. I will avoid the discussion on what exactly this means for the objective function or how it works but in short, this divergence between the trainer and rollout can be corrected for with importance sampling. As for long trajectories, PipelineRL seems like a good candidate for updating weights mid-sequence and thus allow the generation to continue.
For engine tasks, continuing is the only practical option. However, how these sequences are used, the loss calculated, and the rewards associated are all policy design choices - relevant but not here. And last, truncation. Every trajectory runs against a budget, and there are multiple ways to tackle that. DeepSWE masks the loss for trajectories that hit the max context, the max steps or a 20-minute timeout, arguing that reward should only be assigned when the agent deliberately submits, and that rewarding lucky passes reinforces bad behavior.
So, at these trajectory lengths, async with mid-sequence weight updates is the only workable setup. The harder questions are less about the systems and more about what you do with the data - which tokens count, how stale is too stale, and what a truncated trajectory is worth - and those are the policy design choices I'll get into in a follow-up.
Conclusion
Wrapping up, I think putting the engine in an RSI loop is less about any individual goal and more about getting all the four things to line up: environments where a real swarms can run against the engine being changed, rewards that measure the swarm end to end while ensuring no reward hacking, trajectories that are cheap enough to run in large numbers, and an async loop that can work with trajectories that take hours.
References
- Kimi Team. Kimi K2.5: Visual Agentic Intelligence. arXiv:2602.02276, 2026.
- Prime Intellect. Multi-Agent Systems, and PRIME-RL on GitHub.
- MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference. ICLR 2026.
- AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning. arXiv:2606.04484, 2026.
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents. ACL 2025.
- TextArena. arXiv:2504.11442, 2025.
- SemiAnalysis. InferenceX AgentX: Agentic Benchmark for LLM Inference.
- Cognition. Kevin-32B: Multi-Turn RL for Writing CUDA Kernels.
- PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation. arXiv:2509.19128, 2025.
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning. arXiv:2505.24298, 2025.