#supervised-fine-tuning

1 episode

#121: Decoding RLHF: Why Your AI is So Annoyingly Nice

Ever wonder why AI is so polite? Herman and Corn dive into the mechanics of RLHF and how "niceness" gets baked into modern language models.