OpenLLMAI / OpenRLHF

An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral)
https://openrlhf.readthedocs.io/
Apache License 2.0
1.71k stars 160 forks source link

action_log_probs重复计算 #301

Closed cdm114514 closed 1 month ago

cdm114514 commented 1 month ago

actor在做inference的时候不是可以直接返回logits吗,为什么experience maker里面还重算了一次

hijkzzz commented 1 month ago

to ensure accuracy. the accuracy of VLLM is different from HF

cdm114514 commented 1 month ago

Thanks!