THIS EXPLANATION
THE ROOM
TRV·39 Travel, Tourism & Hospitality 7 MIN · 7 STATIONS

Two-way ratings

A Socratic walk-through of two-way ratings — reasoned out one step at a time, not lectured.

04537916380462794538916280452791 abcdefgh
a

The question we started with

THE QUESTION #

Why do guest and host ratings both drift up to near-perfect once each side can rate the other?

On a platform where guests rate hosts and hosts rate guests, the distribution of scores collapses toward the top. Almost everything is five stars, or near enough; a four is read as a complaint and anything below it as a catastrophe. The scale nominally has five points and effectively has about one and a half.

That is a strange outcome for a measurement system, and it is stranger still because it is specific to mutuality. One-way rating systems inflate too, but not nearly as hard. Something about each side being able to rate the other compresses the scale further. Since the whole purpose of the ratings is to distinguish good from bad, a mechanism that reliably destroys the distinction is worth understanding.

b

Reasoning it through

REASONING #

Start with the reviewer's decision, and notice it is not costless. Writing an honest low rating has a benefit — it informs future users — but that benefit is diffuse, spread over strangers, and returns nothing to the writer. So even in a one-way system, the incentive to leave a negative review is weak relative to the effort, and only the strongly aggrieved bother. That alone produces some inflation.

Now make it two-way and add the crucial feature: the person you are rating can rate you back. Ask what a guest considering a three-star review is now weighing. The diffuse benefit to strangers is unchanged. But there is a new, concentrated, personal cost: the host may respond in kind, and the guest's own rating is an asset they need for future bookings. A single retaliatory low score does real damage in a distribution where everything sits at the top — precisely because everything sits at the top, a four stands out.

That is a self-reinforcing loop, and it is worth naming as such. The more compressed the scale becomes, the more damaging any individual low rating is, so the more reluctant everyone becomes to give one, so the more compressed the scale becomes.

There is a second, quieter mechanism: reciprocity. If someone has just given you five stars, giving them four feels like an insult you have to justify. If ratings are visible before you write yours, the first mover sets an anchor, and matching it is the path of least friction. Generosity is contagious in a way criticism is not.

And a third, which is pure selection. The worst interactions often end with nobody rating. A guest who had a miserable stay may want no further contact; a host who had a bad guest may prefer to move on. So the observations that would most inform the average are the ones most likely to be missing — the distribution is not merely shifted upward, it is missing its lower tail by construction.

Put the three together and you can predict something specific: the compression should be worst where ratings are sequentially visible and where both sides depend on their score. And that gives the fix. If neither party can see the other's rating until both have submitted — or until a deadline passes — then retaliation is impossible to target and the anchor is gone. Retaliation requires knowing what to retaliate against.

Platforms that moved to simultaneous reveal did see lower average ratings and more of them submitted, which is the observation the account predicts. Note what this tells you: the pre-change scores were not measuring quality, they were measuring a strategic equilibrium.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of two neighbours who each hold a reference letter for the other's landlord, written after they have both moved out but read by both.

Neither wants to write that the other played loud music, because the other still holds a blank sheet of paper and a pen. Both write warm letters. The letters are not lies exactly — each finds true generous things to say — but the landlord who reads them learns almost nothing, because the letters were shaped by the writers' exposure rather than by the tenants' conduct.

Now change one rule: both letters go into a sealed box and are opened together. The exposure vanishes, because neither can condition on what the other wrote.

WHERE IT BREAKS DOWN

Neighbours write once and part, whereas platform users face the same structure repeatedly with different counterparties, so their incentive is not about this one exchange but about protecting a score that persists across all of them — which makes the pressure steadier and less personal than the letters analogy suggests.

d

Clarifying the model

THE MODEL #

Inflation is not the same as uninformativeness, and this is the most important correction. It is tempting to conclude that a five-star average tells you nothing. But if almost everyone sits at 4.8 and a listing sits at 4.3, that gap is highly informative to a reader who knows the distribution — a compressed scale is still a scale. What is genuinely lost is resolution in the middle and the ability of a newcomer to read the numbers at face value. The practical consequence is that experienced users learn to read the text of reviews and the absence of enthusiasm rather than the number, which is a real cost: it makes the system harder to use for exactly the people it was meant to help.

Retaliation fear may be larger than actual retaliation. Evidence that the fear shapes behaviour is good; whether retaliation actually occurs at high rates is less established, and the two are easily conflated. A system can be distorted by an anticipated response that mostly does not occur.

Simultaneous reveal does not fix everything. It removes the targeting of retaliation and the anchoring, but not the diffuse-benefit problem, not the missing lower tail from silent non-raters, and not the fact that both parties may prefer a pleasant fiction. Ratings stay high after the change; they just stop being quite so uniformly high. The mechanism identified here is one contributor among several, and the change isolates it rather than eliminating inflation.

There is a legitimate reason for high scores that should not be dismissed. Platforms remove poor performers. If the bottom of the distribution is being deleted rather than rated, a high average is partly a real fact about who remains. That is a competing explanation for the same observation, and it predicts high ratings even under simultaneous reveal — which is roughly what is seen.

The falsification test. If retaliation fear and anchoring drive the compression, then moving from sequential to simultaneous reveal should lower average scores and raise submission rates, with the largest change among users whose own score matters most to them. If the change had produced no shift in either, the mechanism would be wrong and the inflation would have to be attributed entirely to selection, politeness and platform curation.

e

A picture of it

THE PICTURE #
Two-way ratings
Two-way ratings Follow the arrows down as time. The upper block is the loop that compresses the scale: a low rating is observed, answered, and the lesson generalises to every later exchange, which is why the effect persists long after any individual grievance. The lower block is the same exchange with one rule changed -- the platform withholds both until both are in -- and the point is what is missing from it: there is no arrow carrying information from one rater to the other before they commit. That absent arrow is the entire intervention. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/two-way-ratings.md","sourceIndex":1,"sourceLine":4,"sourceHash":"3d5fea81dfa24c5b11d04d3ccaba9bc23a5672a8b56addc3b46a7520a4c4196c","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":838,"height":1000},"qa":{"passed":true,"findings":[]}} Host 01 Platform 02 Guest 03 Sequential reveal Guest learns the lesson Simultaneous reveal submits 3 stars 1 shows the rating 2 submits 3 stars in return 3 guest score falls 4 next time, submits 5 stars 5 matches with 5 stars 6 submits honest rating 7 submits honest rating 8 both released together 9 neither could condition on the other 10
KINDSlifelineparticipantmessage

How to readFollow the arrows down as time. The upper block is the loop that compresses the scale: a low rating is observed, answered, and the lesson generalises to every later exchange, which is why the effect persists long after any individual grievance. The lower block is the same exchange with one rule changed — the platform withholds both until both are in — and the point is what is missing from it: there is no arrow carrying information from one rater to the other before they commit. That absent arrow is the entire intervention.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Two-way rating systems do not inflate because people are nice. They inflate because rating someone who can rate you back converts an honest assessment into a personal risk, because a compressed scale makes each low score more damaging and so tightens the compression further, and because the worst experiences go unrated altogether. The scores end up measuring an equilibrium rather than a quality — and the fix is not to exhort people to be honest but to remove the information that makes dishonesty rational.

g

Where to go next

ONWARD #
  • Why review text carries information the star rating has lost, and how readers learn to parse it.
  • How platform removal of poor performers produces high averages by a route that has nothing to do with rating behaviour.
  • Why one-way systems inflate too, and what remains when retaliation is impossible.

Nearby on the shelf

4