Glicko versus Elo: Which Rating Fits Play?

A pair of scales in the centre, blue chess pieces in the left pan and a dark knight in the orange right pan, a man with a laptop seated on the left, a woman with a clipboard on the right, and a line chart and dartboard behind them.

A 1,650 rating can mean very different things depending on how the system got there. Did the player earn it across hundreds of competitive matches? Did they just return from a long break? Are they a fast-improving newcomer who has played only a handful of games? That is the real question behind Glicko versus Elo. Both systems turn results into a rating, but Glicko adds context that matters when every ranked backgammon match should feel competitive.

Glicko versus Elo: the core difference

Elo is the rating system many players recognize first. It begins with a simple idea: before a match, each player has a rating, and the gap between those ratings produces an expected result. Beat a higher-rated opponent and gain more points. Lose to a lower-rated opponent and lose more. If the result matches expectations, the rating moves only a little.

That approach is elegant and easy to understand. A player rated 1,600 is expected to win more often than a player rated 1,400, so the 1,600 player has more to lose when they meet. Elo works especially well when players compete often, ratings have had time to settle, and the system has enough results to judge everyone consistently.

Its limitation is that a rating alone does not reveal how certain the system is. A 1,600 player with ten matches and a 1,600 player with 1,000 matches can look identical in a basic Elo table. They should not necessarily be treated the same by matchmaking.

Glicko addresses that gap by pairing a player rating with a rating deviation, often called RD. The rating estimates skill. The RD estimates how confident the system is in that estimate. A low RD means the player has a substantial, recent match history and their rating is relatively stable. A high RD means there is more uncertainty, perhaps because the player is new or has been inactive.

Glicko-2, the version used for skill-based ranked matchmaking on Online Backgammon, adds a third signal: rating volatility. The full walkthrough of how Glicko-2 rates a player covers all three numbers in order; this article is about what the comparison with Elo actually buys. Volatility helps the system recognize whether a player's results are changing more sharply than expected. It is not a judgment about whether someone is "streaky." It is a mathematical way to make rating updates respond appropriately when a player's competitive level may be moving.

Why uncertainty matters in backgammon

Backgammon is a strategy game with real variance. Strong decisions create an edge over time, but a single match can turn on a hit, a missed joker, a well-timed cube, or a difficult bear-off. A rating system cannot eliminate that variance, nor should it pretend that every short result perfectly measures skill.

Glicko-2 is useful here because it does not treat every player and every result as equally informative. Beating an established opponent near your level tells the system something meaningful. A brand-new player producing a surprising run tells it something too, but with more initial uncertainty around the estimate.

For a player who is improving through daily puzzles, match experience, and better cube decisions, this can feel more responsive than a fixed-update system. Early results can move a provisional or uncertain rating more quickly. As the player builds a record, the rating becomes more stable. That is the right trade-off: room to find your level, followed by a ranking that carries weight.

The same principle helps returning players. If someone steps away for months, their old rating is still useful history, but it should not be treated as perfect evidence of current form. Glicko raises uncertainty during inactivity. When that player returns, the system can learn faster from new results without automatically assuming they have become weaker.

Elo is not wrong - it solves a narrower problem

Calling Glicko better than Elo in every setting misses the point. Elo is a strong solution when simplicity, transparency, and long-term rating stability are the main priorities. It is also easier for players to calculate mentally. If you know the rating gap and the update factor, you can roughly predict what a win or loss will do.

Glicko is more nuanced, but that nuance makes individual rating changes harder to explain at a glance. Two players can win the same number of matches and see different movement because their opponents, rating deviations, and recent activity are different. That may surprise players who expect every win to be worth a fixed number of points.

Neither system is a shortcut around the fundamentals. A high rating is built by making stronger checker plays, understanding when to attack or anchor, using the doubling cube with discipline, and performing under pressure. The system's job is to place those results in useful context, not to create skill where it does not exist.

How Glicko-2 creates better ranked matches

The best ranked match is not simply a match against the closest visible rating. It is a match where both players have a credible chance to prove themselves and where the result improves the system's understanding of their skill.

With Glicko-2, the number a player carries into the queue has already been shaped by how certain the system is of it. An established player with a narrow rating deviation is a known quantity, and their rating moves slowly because it is close to a settled verdict. A newer player showing the same number but a wider deviation is not yet a known quantity, and their rating moves faster. Matchmaking pairs players by rating, so the deviation is what decides how quickly a rapidly improving player is carried to the level where their opponents belong.

It also makes the leaderboard more meaningful over time. A rating should reflect performance, but confidence matters when comparing players who have played very different amounts. Activity is not the only measure of quality, yet a rating earned through sustained competition deserves a different level of confidence than one built from a small sample.

For serious players, this is a feature rather than a complication. You are not only collecting points. You are building a competitive record against real opponents. Every ranked match adds evidence, and over time the system becomes harder to move with a brief hot streak or a short run of bad luck.

Rating fairness and dice fairness are different promises

A smart rating system cannot make unfair dice acceptable. Ratings measure match outcomes; they do not verify how the rolls were produced. In online backgammon, players need both competitive matchmaking and confidence that the game itself was not manipulated.

That is why fairness has two layers. Glicko-2 helps ensure that rankings react to results with appropriate confidence. A provably fair dice system addresses the integrity of the match record itself by allowing completed-match rolls to be re-derived from published records. One protects the quality of competition over time. The other gives players a way to verify the conditions of a specific game.

Keeping those ideas separate matters. A losing streak does not prove dice are unfair, and transparent dice do not guarantee that every rating update will feel satisfying. What players should expect is a system that can be checked, explained, and trusted even when the board does not go their way.

What players should watch instead of one rating change

A single post-match number is tempting to overread. After a hard loss, it can feel like the system has made a verdict on your ability. It has not. It has processed one result within a larger record that includes opponent strength, your rating certainty, and your recent activity.

Watch the longer pattern. Are you winning more often against players around your rating? Are your cube decisions improving? Do you recognize when a prime is stronger than a loose blot, or when a safe play gives up too much equity? Your rating will generally follow the quality of those decisions, though not in a straight line.

If you are new, play enough ranked matches to let the system find your level. If you are returning, expect your early games to carry more information than they did before your break. If you are established, do not chase every small rating fluctuation. Focus on match quality, learn from difficult positions, and show your strategy across a meaningful sample.

The useful question is not whether Glicko or Elo can predict every roll, every upset, or every brilliant comeback. It is whether the system gives honest results enough time to reveal the stronger player. Keep playing, keep learning, and let your record become the proof.

Put It Into Practice

The fastest way to learn backgammon is to play it. Match with real players right in your browser: free, no install, no sign-up needed.

Play Backgammon FreeIn your browser · no install