The setup

Reinforcement learning is the study of how an agent should act in an environment to maximize a reward signal it receives over time. That is the entire formalism. It is the simplest possible framing of intelligence-as-optimization, and it is also, perhaps surprisingly, the most general.

In supervised learning, the teacher tells the student the right answer. In unsupervised learning, there is no teacher at all. In reinforcement learning, the agent is given only a sparse, delayed signal — a number, occasionally, saying "that was good" or "that was bad" — and must figure out which of its many actions contributed to the outcome. This is called the credit assignment problem, and it is the central technical challenge of the field.

Rewards and the credit assignment problem

Imagine you play a long game of chess and lose. The reward — a single scalar, +1 for a win, −1 for a loss — arrives at the very end. Which of the forty moves you made was responsible for the loss? The opening where you weakened your king? The middlegame where you traded your bishop? The endgame where you let the pawn through? All played a role, but the reward function knows nothing of that.

The breakthrough of Q-learning, in the late 1980s and 90s, was to realize that an agent can learn an estimate of "how good is this state, on average" — a value function — and use the differences between successive value estimates as a teaching signal for which actions led to better outcomes. Combined with deep neural networks as function approximators (Deep Q-Networks, 2013), this idea scaled to settings that were previously unreachable.

Q(s, a) ← Q(s, a) + α [r + γ max Q(s', a') − Q(s, a)]

The line above is the heart of Q-learning. It says: update your estimate of "how good it is to take action a in state s" by moving it a little bit toward "the reward you actually got, plus your estimate of the best thing to do next". The math is small; the consequences are not.

The victories: games

RL earned its reputation by winning games that humans had thought were unreachable for machines. AlphaGo defeated the world champion at Go in 2016, combining deep networks with Monte Carlo tree search — a victory that, more than any other, announced the arrival of the deep RL era. AlphaZero generalized the approach, learning from scratch and surpassing all human knowledge in Go, chess, and shogi within hours of training. OpenAI Five won against the world champions of Dota 2; DeepMind's AlphaStar reached Grandmaster in StarCraft II.

Games are not the point. Games are the easiest environment we have — fast, repeatable, with reward functions that fit in a line of code. They are the field's training wheels. — RESEARCH NOTE, 2024

Beyond games: agents in the world

The harder question is what happens when you remove the training wheels. The real world does not reset after every episode. Reward signals are noisy, sparse, and often defined by humans who cannot fully articulate what they want. The agent cannot try the same episode a million times; it has to learn from each attempt as if it mattered.

This is where the field is now. Robotics uses RL to teach manipulation and locomotion, with sim-to-real transfer closing the gap between training and deployment. Recommendation systems frame long-term user satisfaction as a reward to be maximized. Industrial control uses RL to optimize HVAC, chip placement, and power grid dispatch. Autonomous driving treats the road as a multi-agent RL environment.

The dream — still distant, still real — is a general-purpose agent that can take a goal stated in natural language and figure out, through trial and error, how to achieve it in the physical world. We are not there yet. But the trajectory is clear.

RL + LLMs: the next frontier

The most consequential development of the last two years has been the merger of reinforcement learning with large language models. The basic idea: take a language model, give it access to tools, and train it with RL to produce trajectories of action rather than just tokens. The reward is set by either a human (RLHF) or another language model (RLAIF, constitutional AI).

The result is a new class of system — sometimes called an "agent" — that can plan, call tools, recover from errors, and pursue multi-step goals in digital environments. Today's agents can browse the web, write and execute code, query databases, and operate software for hours at a time. They are not yet reliable, but they are improving on a clear curve.

This convergence of RL and LLMs may be the most important technical development of the decade. It is also the source of the deepest alignment concerns: a system that can plan, act, and learn from feedback is a different kind of system than one that merely generates text. The choices we make about how to train it, what to reward, and what to constrain will shape the next fifty years.


Filed under: Reinforcement Learning · Agents · RLHF · Robotics
Cite as: Editorial Team (2026). Reinforcement Learning: from games to the real world. Signal, Vol. 1, Issue 8.