The OkCupid attractiveness rating distributions
The OkCupid attractiveness ratings, showing that 80% of men are rated ‘below average’1 are one of those meme statistics that have been seared into the internet’s collective consciousness, sparking years of vigorous debate. Christian Rudder illustrates these divergent distributions in Dataclysm:
Although Kreager et al. (2014) didn’t explicitly name the dating site from which their data were drawn, it was almost certainly OkCupid. Men received a mean rating of 2.13 (SD = 0.70), compared to 2.84 (SD = 0.61) for women, with women’s evaluations of men showing a greater skew. This corresponds to a Cohen’s d of −1.08, a large effect.
In most sexually dimorphic, ornamented species, males are the more conspicuous and ‘beautiful’ ones, as females typically act as the sexual selectors due to their greater reproductive investment. A classic example is the peafowl, the inspiration for Darwin’s theory of sexual selection. To be slightly less generic, let’s instead show a picture of a male and female Andean cock-of-the-rock:


There are, however, edge cases like the polyandrous phalarope, where females compete for males and defend territory, making them the showier sex. Charles Darwin noted this reversal in The Descent of Man, and Selection in Relation to Sex:
Other characters proper to the males of the lower animals, such as bright colours and various ornaments, have been acquired by the more attractive males having been preferred by the females. There are, however, exceptional cases in which the males, instead of having been the selected, have been the selectors. We recognise such cases by the females having been rendered more highly ornamented than the males—their ornamental characters having been transmitted exclusively or chiefly to their female offspring. One such case has been described in the order to which man belongs, namely, with the Rhesus monkey.
Many interpret the ratings of men as objectively wrong because they don’t mirror women’s relatively neat, centred distribution. However, these results wouldn’t have surprised Darwin, who argued that owing to men’s historical dominance, women have been selectively bred to be more attractive than men, reversing the typical pattern:2
Man is more powerful in body and mind than woman, and in the savage state he keeps her in a far more abject state of bondage than does the male of any other animal; therefore it is not surprising that he should have gained the power of selection. Women are everywhere conscious of the value of their beauty; and when they have the means, they take more delight in decorating themselves with all sorts of ornaments than do men. They borrow the plumes of male birds, with which nature decked this sex in order to charm the females. As women have long been selected for beauty, it is not surprising that some of the successive variations should have been transmitted in a limited manner; and consequently that women should have transmitted their beauty in a somewhat higher degree to their female than to their male offspring. Hence women have become more beautiful, as most persons will admit, than men.
While the OkCupid dataset is the most widely cited, it’s far from the only source of attractiveness ratings. Several factors may have artificially inflated the attractiveness gap in the OkCupid data:
Women may be more skilled at selecting flattering photos, using better lighting and more strategic angles or crops.
Less attractive men may be overrepresented on online dating platforms compared to less attractive women.3
OkCupid had a feature that automatically notified users when they were rated 4–5 stars. This may have incentivized women to downrate men to avoid unwanted attention, while some men may have given higher ratings to signal interest.
Let’s see whether this disparity holds up in other datasets.
Does the attractiveness gender gap replicate in other datasets?
Marcus & Miller (2003) recruited 112 male and 112 female undergrads and sorted them into mixed-gender groups, wherein members rated each other’s attractiveness on a 7-point scale according to how ‘objectively attractive’ they perceived each target to be. Women were rated higher by both men and women, with women giving slightly higher ratings overall. Cross-sex ratings showed that men received a mean of 4.08 (SD = 0.42) and women 4.42 (SD = 0.33), yielding a Cohen’s d of −0.90—not far from the OkCupid gap. The difference between women’s ratings of each other (M = 4.54, SD = 0.49) and men’s ratings of other men (M = 3.96, SD = 0.49) was even larger. The average male rating was above the midpoint here, perhaps because seeing the full body and body language enhances perceived attractiveness or because in-person interactions tend to elicit more generous ratings.
Back et al. (2011) report data from a speed-dating event in which 30 opposite-sex raters (15 for younger and 15 for older participants) rated photos taken from pre-event video recordings of 190 men and 192 women on a 7-point scale. Mean ratings were 2.5 (SD = 0.75) for men and 3.16 (SD = 0.90) for women, a difference corresponding to d = −0.80. In this case, both men and women were rated significantly below the midpoint on average.
A LessWrong user analysed data from a speed-dating event where ratings were made after dates.4 Women were rated 0.5 standard deviations higher than men, again showing that in-person meetings tend to elicit more generous ratings. Also, both distributions appeared roughly normal.
Sidari et al. (2020) had 264 men and 275 women rate each other’s attractiveness after speed dates on a 7-point scale. Mean facial attractiveness was 4.18 (SD = 0.98) for men and 4.5 (SD = 0.88; d = −0.34) for women. Bodily attractiveness was rated 4.31 (SD = 1.01) for men and 4.66 (SD = 0.92; d = −0.36) for women. Overall attractiveness, including personality, was 4.58 (SD = 0.79) for men and 4.77 (SD = 0.76; d = −0.25) for women.
Hofer et al. (2021) had 10 independent observers (50% male) rate photos of male and female speed-daters before the event on a 0–5 scale. Men were rated 2.01 (SD = 1.01) and women 2.41 (SD = 1.03), corresponding to d = −0.40. After the dates, participants rated each other; men’s mean rose to 2.93 (SD = 0.63) and women’s to 3.48 (SD = 0.68), yielding d = −0.85.
Eastwick & Smith (2018) present norming data for 593 faces in the Chicago Face Database. Participants posed with a neutral expression, looked directly at the camera, and wore a plain grey t-shirt. Men had a mean rating of 3 (SD = 0.63) and women 3.45 (SD = 0.83), corresponding to d = −0.61. The midpoint was 4, so women were also rated below it. It also found that, while participants’ own ratings did not differentially affect romantic desire for the photos by gender, the norming attractiveness ratings showed a larger association for men (βdif = −.13). Additionally, finding a target uniquely attractive predicted romantic desire more for women.
Costa & Maestripieri (2023) had 87 male and 149 female undergrads rate photos of 86 male and 145 female undergrads. Participants uploaded digital photographs of their faces upright to the shoulder axis, with a neutral expression, a homogeneous background, and no hat, sunglasses, or makeup. On a 0–100 scale, men were rated 33.7 by women, while women were rated 54.2 by men—an even larger difference than the OkCupid data. Men were rated 42.8 by other men, and women 47.8 by other women. Self-perceived attractiveness was similar for both genders.
Ebner et al. (2018) had 154 participants—one third young, one third middle-aged, and one third older—rate photos of faces from the FACES Lifespan Database on a 0–100 scale. Young men rated young men at 30 (SD = 23) and women at 45 (SD = 27), while young women rated young men at 40 (SD = 27) and women at 44 (SD = 24). There was an interaction effect whereby age had a stronger negative impact on ratings of female faces than male faces, with the gap just about closing for older faces.
This gap isn’t a recent phenomenon brought about by the internet, either. Cross & Cross (1971) had eighty 7-year-olds, eighty 12-year-olds, eighty 17-year-olds, and sixty adults (M age = 36), rate 72 faces spanning different ages, sexes, and racial groups. As shown in this figure, both male and female judges rated female faces of all ages higher than male faces on average. Female judges generally gave more generous ratings, except in the case of adult male faces.
This pattern has also been observed in a Himba pastoralist community in Namibia, where members rated headshots of opposite-sex individuals for romantic desirability on a four-point scale. Men received 1.44, and women 1.88. As can be seen below, there were no himbo Himba to be found.
Meta-analysis on the ‘gender attractiveness gap’
Conveniently, a recent meta-analysis provides robust evidence for this pattern, examining 28 facial attractiveness datasets released over the past decade across more than 50 countries, with over 12,500 raters evaluating more than 11,000 facial images. Ratings were standardized prior to aggregation.
Female faces (N = 5,659) were rated 0.4 standard deviations higher than males faces (N = 5,532). Additionally, female stimuli approximated a normal distribution, whereas male ratings exhibited a modest right skew, with a higher concentration of lower scores and a tail of higher-rated faces.

Using linear multilevel modelling to account for face gender, rater gender, and their interaction, the researchers found:
Male faces received lower attractiveness ratings than female faces (d = −0.59).
Male raters gave lower ratings than female raters (d = −0.16).
Female raters showed a larger gender gap in their ratings than male raters (d = 0.21), driven by greater favourability towards other women.
In the subset of five studies with ratings aggregated across rater gender, the difference was d = −0.28.











