RLHF (Reinforcement Learning from Human Feedback)
Definition
A training step where people rate an AI's answers, and the model learns to produce more of what people preferred. After a model learns language from raw data, RLHF is part of how it's taught to be helpful, polite, and less likely to say something objectionable. Real humans sit and judge responses, and those judgements shape the tool's behaviour. It's a large part of why today's chatbots feel reasonable to talk to.
Why It Matters
It explains why two models trained on similar data can have very different manners. A lot of a chatbot's personality comes from this stage.
What does that look like in practice?
People rated thousands of a chatbot's replies as better or worse, and it learned to give more of the better ones. That's RLHF.
What people actually mean when they say this
When someone says a chatbot is 'polite' or 'well-behaved,' RLHF is a big part of why: humans shaped its manners after the initial training.
Last reviewed: June 2026
Knowing the word is just the start.
We help non-tech people go from looking things up to actually using AI confidently. Tutorials, prompts, plain English. All of it, for less than a cup of coffee a month.
Show Me How This Works →Start free. Cancel anytime. No judgment, ever.
