no preview
DESCRIPTION
Second DPO training run results in a model that listens to your prompts even better than the last DPO experiment. I am training one final run to complete this and that one will be locked behind the top tier.
ClimbingTheMountain