TL;DL: this is a repo to align the large language models (LLMs) by online iterative RLHF. Also check out our technical report and Huggingface Repo! We present the workflow of Online Iterative ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results