Skip to Content

What drives product competitiveness on Amazon? Re-weighting a product evaluation index

Our study compared two scoring systems for Amazon products. Category rank and review volume mattered more, and price less, than expected.
January 15, 2025 by
Wellxpring Research
AMeasured. Findings from a Wellxpring study of Amazon product data. Correlations describe association, not cause. Sample details are available on request.

Sellers and analysts routinely score products to judge competitiveness: should we enter this niche, which competitor is strongest, which of our products is underperforming? We tested two scoring systems against Amazon’s own product ratings to see which factors carried real signal.

Two scoring systems

ComponentOriginal systemRevised system
Market performance (overall)45%56.39%
User feedback (overall)55%43.61%
Best Seller Rank / category rank35.38%42.30%
Customer rating29.42%reduced
Review countabout 11–12%14.82%
Sentiment scoreabout 11–12%14.82%
Correlation with Amazon ratings0.56160.5889

Shifting weight towards category rank and review volume, and away from raw rating, produced a score that tracked Amazon’s ratings more closely.

What moved together

PairCorrelationReading
Rating and review sentimentPearson 0.499, Spearman 0.497Star ratings and the tone of written reviews agree; ratings appear to reflect genuine experience.
Review count and category rankPearson 0.364, Spearman 0.405Products with more reviews tend to rank better. The direction of cause is not established.
Sentiment and rating stability−0.425Products with warmer sentiment saw more rating volatility.
Review count and rating stability0.323More reviews, steadier ratings.
Price and category rank0.136 (strongest price link)Price was only weakly related to performance in this sample.

What it means for sellers

  • Weight rank and review volume when screening niches and competitors; raw star rating alone is a weaker guide.
  • Build review volume deliberately and within platform rules: it is associated with both better rank and steadier ratings.
  • Do not assume price is the lever. In this sample, price explained little; test it before cutting it.
  • Watch volatility on new, well-liked products. Early enthusiasm can swing ratings until review volume builds.

Limits. These are associations in one dataset at one point in time. They are useful for building a screening score, not proof of what causes ranking. We treat the scoring weights as a starting point, recalibrated for each category.

Reading advertising through margin: TACoS, contribution after ads and incrementality
ACoS tells you what ads were credited with. Three better measures show what advertising actually did for profit.