Self trained zephyr-7b-dpo-qlora MT-bench score dropped to 1.88

huggingface / alignment-handbook

Robust recipes to align language models with human and AI preferences

Apache License 2.0

4.74k stars 412 forks source link

Hi, I just followed recipes/zephyr-7b-beta/dpo/config_qlora.yaml and hope to replicate the experiments. I was training on A10G, with 1 gpu, and the only modification I did was reducing the train_batch_size from 4 to 1 (due to memory constraint). However, my output models zephyr-7b-dpo-qlora only has mt-score of 1.88. I also did a mt-score benchmark with the downloaded zephyr-7b-sft-qlora and it had mt-bench score of 6.37 (which seems relatively normal). Does anyone else also have difficulties replicating this dpo experiments with qlora? Or is the batch size a critical difference for training?

huggingface / alignment-handbook

Self trained zephyr-7b-dpo-qlora MT-bench score dropped to 1.88 #188