FIBS Rating versus Elo: Which Fits Backgammon?

A win is a win in chess. In backgammon it is not: a 1-point game decided by a late double six and a 25-point match ground out over two hours are both "a win", and no honest rating should score them the same. That is the real question behind FIBS rating versus Elo. Both systems turn results into a number, but only one of them was built with a backgammon match in mind, and Online Backgammon rates its ranked play with that one.
FIBS versus Elo: the core difference
Elo is the rating system many players recognize first. It begins with a simple idea: before a match, each player has a rating, and the gap between those ratings produces an expected result. Beat a higher-rated opponent and gain more points. Lose to a lower-rated opponent and lose more. If the result matches expectations, the rating moves only a little.
That approach is elegant and easy to understand. A player rated 1,600 is expected to win more often than a player rated 1,400, so the 1,600 player has more to lose when they meet. Elo works especially well in a game where every contest is the same size and players compete often enough for the numbers to settle.
Backgammon is not that game. The FIBS formula, introduced on the First Internet Backgammon Server in the 1990s and used by the backgammon community ever since, keeps Elo's good idea, the expected result from a rating gap, and adds the two facts Elo has no place for: how long the match was, and how much rated play the player already has behind them. The full walkthrough of how the rating works in ranked backgammon goes through the arithmetic step by step; this article is about what the comparison with Elo actually buys.
Match length is inside the formula
Elo has a single update factor, so a 1-point match and a 25-point match are worth the same, and the favourite is given the same chance in both. Anyone who has played a long match knows that is wrong. Luck decides a single game; over a long match the stronger player pulls away.
The FIBS formula scales the rating gap by the square root of the match length before it turns the gap into a probability, so the same 200-point gap gives the underdog about a 44 percent chance in a 1-pointer and about 24 percent in a 25-pointer. The stake grows with the square root of the length as well: a 25-point match can move a rating five times as far as a 1-point match. Longer matches say more about skill, so they count for more. Elo cannot express that without an operator inventing a table of factors on the side.
A newcomer's early results count for more
Elo treats a player's first game and their thousandth exactly alike. Systems built on top of it usually patch that with a provisional period or a separate uncertainty number that players then have to be taught about.
The FIBS formula uses something every backgammon player already understands: experience, measured as the points of rated matches you have played. While that total is small, every result is multiplied up, five-fold on the very first match, falling steadily to a multiplier of one at 400 points of experience, where it stays. Two brand-new accounts at 600 playing a 5-point match move 44 points each; two settled players in the same match move 9. A strong newcomer reaches the level where their opponents belong in a few evenings, and a settled rating stops jumping around on one result. There is no second number to interpret. How much you have played is the whole story.
Elo is not wrong - it solves a narrower problem
Calling the FIBS formula better than Elo in every setting misses the point. Elo is a strong solution when every contest is the same size and simplicity is the priority, which is exactly the situation in chess. It is also easy to calculate mentally: know the gap and the factor, and you can roughly predict what a win or loss will do.
The FIBS formula is a little more to hold in your head, because the same win is worth different amounts in a 3-point match and an 11-point match, and to a player with 50 points of experience and one with 500. But every one of those differences has a plain reason a player can state in a sentence, and the number of points at stake is known before the first roll.
Neither system is a shortcut around the fundamentals. A high rating is built by making stronger checker plays, understanding when to attack or anchor, using the doubling cube with discipline, and performing under pressure. The system's job is to score those results fairly, not to create skill where it does not exist.
How the formula creates better ranked matches
The best ranked match is not simply a match against the closest visible rating. It is a match where both players have a credible chance to prove themselves and where the result means something afterwards.
Because the FIBS formula carries newcomers quickly to their real level, the number a player brings into the queue is already close to the truth after a handful of matches, and matchmaking pairs players by that number. Because longer matches count for more, players who commit to real match play are the ones whose ratings settle fastest. And because no single match can move a rating by more than 200 points, a lucky evening cannot rewrite a ladder position that took months to earn.
It also makes the leaderboard more meaningful over time. Everyone starts from the same 600, so the distance between two players on the ladder is a record of rated results and nothing else. A rating earned through sustained match play is worth more than one built from a small sample, and the formula makes that difference show up in the number rather than in a footnote.
Rating fairness and dice fairness are different promises
A well-chosen rating formula cannot make unfair dice acceptable. Ratings measure match outcomes; they do not verify how the rolls were produced. In online backgammon, players need both competitive matchmaking and confidence that the game itself was not manipulated.
That is why fairness has two layers. The rating formula makes sure rankings react to results in proportion to what those results proved. A provably fair dice system addresses the integrity of the match record itself by allowing completed-match rolls to be re-derived from published records. One protects the quality of competition over time. The other gives players a way to verify the conditions of a specific game.
Keeping those ideas separate matters. A losing streak does not prove dice are unfair, and transparent dice do not guarantee that every rating update will feel satisfying. What players should expect is a system that can be checked, explained, and trusted even when the board does not go their way.
What players should watch instead of one rating change
A single post-match number is tempting to overread. After a hard loss, it can feel like the system has made a verdict on your ability. It has not. It has processed one result against one opponent over one match length, at your current level of experience.
Watch the longer pattern. Are you winning more often against players around your rating? Are your cube decisions improving? Do you recognize when a prime is stronger than a loose blot, or when a safe play gives up too much equity? Your rating will generally follow the quality of those decisions, though not in a straight line.
If you are new, play enough ranked matches to let the experience multiplier do its work. If you are established, do not chase every small rating fluctuation, and remember that a longer match is the fastest honest way to move. Focus on match quality, learn from difficult positions, and show your strategy across a meaningful sample.
The useful question is not whether FIBS or Elo can predict every roll, every upset, or every brilliant comeback. It is whether the system gives honest results enough time to reveal the stronger player. Keep playing, keep learning, and let your record become the proof.