👍RLHF
Reinforcement Learning from Human Feedback
Fine-tuning a model using human ratings of its answers, so it behaves more helpfully and safely.
AJ Learning HubReinforcement Learning from Human Feedback
Fine-tuning a model using human ratings of its answers, so it behaves more helpfully and safely.