Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
LoRA support has landed in TRL's AsyncGRPOTrainer via PR #7017 and ships with TRL v1.14: the asynchronous trainer can now train an adapter instead of the full model, and it syncs only the LoRA adapter to vLLM. This post covers a real-world project built on top of that feature, in which training and inference no longer share a machine. LoRA training is particularly well suited to RL — as Thinking Machines' "LoRA Without Regret" blog shows, LoRA can match full fine-tuning for policy-gradient RL even at rank 1, because the advantage function conveys only about O(1) bits of information per episode, leaving a rank-1 adapter with enough capacity to absorb it.
There is also a systems consequence to training with LoRA. A rank-1 adapter for a 1.5B model is only a few megabytes, while the full model is around 3 GB, so instead of sending the entire policy to the inference workers after every update, only the adapter has to be sent. vLLM can keep several adapters loaded at once, which means old rollouts finish with the policy they started with while new rollouts use the latest one.
TRL's AsyncGRPOTrainer already separates training from generation, letting the trainer and vLLM run on different machines at their own speed — straightforward on a single node or a cluster where both processes share a filesystem or can form an NCCL group. The goal was to reproduce that setup on Hugging Face Jobs, where each Job is one container on one VM and cannot spawn multiple nodes (with a cap of 8xH200 per node). With full-weight syncing this would go nowhere: every update would move gigabytes between machines, Jobs cannot communicate across nodes, and there is no shared local disk or shared localhost. With LoRA a sync is only a few megabytes, and HF Jobs provide volumes backed by Storage Buckets that can be mounted as a FUSE filesystem in every Job.
Architecturally, the trainer and the vLLM replicas run as separate HF Jobs on separate machines, and a small proxy in front of the replicas adds the auth header, routes each rollout to the replica that already holds its KV prefix, and broadcasts adapter loads to every replica. The AsyncGRPO metrics show where the bottleneck sits. Across five runs, the same recipe went from 3 h 27 min to 53 min for 500 steps.