Execution feedback, process reward models, self-reflection, verbal gradients, unit tests, and trajectory scoring for improving generated algorithms