Skip to main content
7-Day Free Preview — No Credit Card Required
AI UnSpun
📖 AI Glossary

RLHF (Reinforcement Learning from Human Feedback)

Definition

A training step where people rate an AI's answers, and the model learns to produce more of what people preferred. After a model learns language from raw data, RLHF is part of how it's taught to be helpful, polite, and less likely to say something objectionable. Real humans sit and judge responses, and those judgements shape the tool's behaviour. It's a large part of why today's chatbots feel reasonable to talk to.

Why It Matters

It explains why two models trained on similar data can have very different manners. A lot of a chatbot's personality comes from this stage.

What does that look like in practice?

People rated thousands of a chatbot's replies as better or worse, and it learned to give more of the better ones. That's RLHF.

What people actually mean when they say this

When someone says a chatbot is 'polite' or 'well-behaved,' RLHF is a big part of why: humans shaped its manners after the initial training.

Last reviewed: June 2026

Share this

Knowing the word is just the start.

We help non-tech people go from looking things up to actually using AI confidently. Tutorials, prompts, plain English. All of it, for less than a cup of coffee a month.

Show Me How This Works →

Start free. Cancel anytime. No judgment, ever.