Where Strength Standards Actually Come From (and Which Ones to Trust)
Every strength standards table traces back to a source. Most of them will not tell you what it is. How the trustworthy ones are built, and the questions that expose the rest.
Every table traces back to somewhere
Strength standards feel like settled knowledge - the tables have been on the internet for twenty years, they mostly agree with each other, and they carry confident labels like Intermediate and Elite. But a standard is only as good as the dataset behind it, and if you follow the citations, nearly every table on the internet leads back to one of three places: somebody's expert opinion, numbers users typed into an app about themselves, or competition results relabelled.
Agreement between tables is not corroboration when they share an ancestor. Most of them do.
Competition data: the only verified source
There is exactly one kind of large-scale strength data where every number was performed in public, to a defined standard, and judged: powerlifting meet results. The OpenPowerlifting project collects them - millions of lifts across federations, released into the public domain - and it is the closest thing strength training has to ground truth.
It has real limitations, and honest tools state them. The population is self-selected: everyone in it chose to compete. Depth and pause standards are competition-grade, so your touch-and-go gym bench reads slightly generous against it. And it contains no overhead press at all, because that is not a contested lift - which is why a standards tool offering overhead press numbers is, by definition, using something weaker than meet data for them.
The gym-population illusion
The number everyone wants is "how do I compare to normal people who lift" - and it is precisely the number that cannot currently be produced, because no verified dataset of general-population one-rep maxes exists. Gyms do not test or record maxes. Surveys self-report, and self-reported maxes are systematically inflated.
The most widely used "gym standards" table is derived from competition data and disclaimed by its own author as not being population norms. That does not make it useless - as expert opinion it is informed - but it does make any percentile computed from it a fabrication. When an app shows you a competitive percentile of 20 and a gym percentile of 70, ask where the second number came from. There is no dataset it could have come from.
How honest percentiles are actually built
Raw meet results are not directly usable as standards - a 100 kg bench means something different at 74 kg bodyweight than at 105 kg, and something different for male and female lifters. The honest method fits statistical distributions per lift, per sex, per bodyweight class, interpolates between classes when a lifter sits between them, and evaluates where a given weight falls on the fitted curve.
Two things separate rigorous versions of this from hand-waving. First, validation: fitted curves can be checked against independently published, peer-reviewed strength norms and should reproduce them closely. Second, disclosed uncertainty: at bodyweight extremes the data thins out, and a trustworthy tool shows a sample size and a confidence band there instead of a single crisp number that implies precision the data does not contain.
The questions that expose a weak standard
You can audit any strength standards tool with four questions. What dataset is this, by name, and can I look at it? Who is in the comparison population, and does the tool say so next to the number? What happens at the edges of the data - does it admit uncertainty or does every input get a confident answer? And are there lifts on offer that no verified dataset covers?
A tool that names its source, labels its population, discloses its thin regions, and declines to invent numbers for uncovered lifts is telling you the truth. A tool that fails those checks may still be fun, but its percentiles are decoration.
Where Benchmark fits
Benchmark was built to pass its own audit. The source is named on a screen inside the app - OpenPowerlifting, CC0 licence, dataset version, derivation date, and the exact number of lifts compared. Curves are checked against peer-reviewed published norms to within 3%. Thin regions of the data show a sample size and a wider confidence band, overhead press is absent because no honest source for it exists, and there is no invented gym-population axis - one number we can stand behind, rather than two where the second is a guess.
Your lifts never leave your phone, the percentile is computed offline, bench press is free forever, and there is no account. See the Benchmark product page for screenshots and the full privacy policy.