Nobody reads a store rating the way a dashboard does. A user deciding whether to install does not weigh 3.9 against 4.1 as adjacent points on a continuum. They register, in well under a second, which side of a line the number sits on — and the line, for most consumer categories, sits at four stars. Above it, the rating recedes into the background and the listing gets judged on its own merits. Below it, the rating becomes the merit. It is the first objection, raised before the product has said a word.
This is why the gap between 3.8 and 4.2 is not four tenths of a point. It is the difference between a storefront that converts visits and one that filters them.
A filter, not a score
The threshold behaviour has mechanical causes as well as psychological ones. Store interfaces round and badge ratings in ways that compress everything below 4.0 into a single impression. Comparison is ambient: the user rarely sees one app's rating in isolation — they see it next to three competitors, two of which sit above the line. And an increasing share of discovery now runs through ranked and recommended surfaces where the rating participates in whether the listing is shown at all, not merely how it looks once found.
The effect is that a sub-4.0 rating taxes every visit, from every channel, around the clock. Organic discovery, paid campaigns, press coverage, word of mouth — all of it arrives at the same checkpoint, and the checkpoint takes its share before the product gets to make its case.
What it does to paid spend
Paid acquisition makes the threshold sharper, not softer. Campaign spend buys arrivals at the listing; the rating then decides how many of them become installs. When conversion at the storefront sags, the cost of each install rises in direct proportion — the campaign did its job, and the storefront undid part of it. The arithmetic is unforgiving: a meaningful drop in listing conversion can erase the efficiency gains of an entire quarter of campaign optimisation.
What makes this expensive in practice is that it rarely appears as a line item. Blended metrics absorb it. Install cost creeps upward, the team tunes targeting and creative, results improve at the margin, and the underlying drag — the number at the top of the listing — sits outside the reporting loop entirely. The most common discovery path is a founder noticing that acquisition used to be cheaper and nobody being able to say precisely why.
Slow number, fast consequences
A rating is a slow-moving mass. On a base of a hundred thousand reviews, sentiment that took a year to settle does not lift in a week — every new review is diluted by the hundred thousand that preceded it. The consequences, by contrast, move at the speed of traffic: each day below the threshold is conversion lost that day, spend taxed that day. Slow cause, fast effect. That asymmetry is why drift below 4.0 deserves more urgency than it usually gets — waiting for the number to recover on its own means paying the tax for as long as the wait lasts.
The work, properly ordered, runs in three movements: read the review signal closely enough to separate what is dragging the rating from what is merely loud; arrest the erosion so the floor stops falling; then recover the number past the threshold and hold it there while the deeper causes are worked through. The order matters. Recovery attempted before diagnosis treats symptoms at random; diagnosis without recovery leaves the tax running while the product team works through a backlog.
The threshold is indifferent to effort. It does not credit the roadmap, the rebuild, or the support queue cleared last month. It asks one question — which side of the line is the number on today? — and prices every install accordingly.