Pre-training, Finetuning, RLHF and Preference Optimization
You can train any model on Transformer Lab. But we provide pre-written scripts for training advanced models from scratch or finetuning existing ones with production-ready implementations of DPO, ORPO, SIMPO, and GRPO that work out of the box. Complete RLHF pipeline with reward modeling handles the complex orchestration automatically, from data processing to final model outputs.