RL, DPO, GRPO, SFT, and execution-reward training for generating correct and solvable optimization models