Methodology

Every figure on this site is reproducible from published data plus the rules below.

Popularity and rank

Counts come from the Social Security Administration's national and state files, which record the name on every Social Security card application for a US birth. National files start in 1880, state files in 1910. Rank is the position of a name within its sex for a given year. The SSA withholds any name with fewer than five occurrences in a state-year, so state totals are always slightly lower than national ones.

Living population estimate

We take the number of births for a name in each year and apply a survival probability for that birth cohort, then sum across years. The survival curve is interpolated from period life-table figures and is applied separately by sex. This is an estimate: it does not account for immigration, emigration or name changes, so treat it as an order of magnitude rather than a census.

Median age

The age at which half of the estimated living bearers are older and half younger, calculated from the same survival-weighted birth series.

Phonetic similarity

Each name is reduced to a phoneme string by a rule set that collapses English spelling conventions, then compared to candidates on four measures: edit distance between phoneme strings, edit distance between spellings, edit distance between vowel skeletons, and whether the endings rhyme. The measures are weighted 0.55, 0.20, 0.15 and 0.10 and the result is reported from 0 to 1. Candidates are drawn from names sharing an opening sound, an ending or a syllable count, which is why the list is fast and why a name with an unusual shape returns fewer matches. The ranking then applies the same familiarity nudge described below, so a common name outranks an obscure one at equal phonetic distance.

Middle-name flow score

A first and middle name are scored on how the pair reads aloud. The score rewards a difference in syllable count, alternating stress, and a clean consonant-to-vowel junction between the two words. It penalises a shared initial, a rhyme between the two endings, a repeated sound where one name ends and the next begins, and two long names in a row. The result is reported from 0 to 1. Where two candidates score the same on sound, the one more widely used as a name in the birth record is placed first, by a margin small enough that it never overturns a genuine difference in how the pair reads. It is a readability heuristic, not a verdict.

Sibling fit

A sibling suggestion answers a narrower question than "is this a nice name": does it sound like it was chosen by the same people, at the same time, as the name you already have? Two measures do most of the work. The first is era similarity, the cosine between the two names' birth-year distributions in five-year buckets, which is high when two names rose and fell together. The second is comparable use, which falls away as the two names diverge in how many people carry them, because a top-ten name beside one almost nobody has chosen reads as an accident. Those are weighted 0.62 and 0.38, and then the collisions are subtracted: a shared initial, a shared ending, and above all names that sound too much alike, since siblings called Ellie and Ella get muddled for life. Candidates are drawn from names still in the top 800 today, and the list is balanced so it does not return six sisters.

One caveat worth stating plainly. Era similarity is only as sharp as the record it reads. The American files run from 1880, so a name's rise and fall is a distinctive shape and the matching is confident. The British files begin in 1996, which leaves a thirty-year window in which most names move in broadly the same direction; sibling suggestions on a British-only name page are correspondingly weaker, and the card there says so rather than pretending otherwise.

Two countries

American figures come from the Social Security Administration and cover 1880 onwards. British figures come from the Office for National Statistics and cover England and Wales from 1996. The two are counted differently and the populations differ by roughly a factor of six, so where a name page draws both curves each is scaled to its own peak. What that chart compares is the shape of a name's rise and fall, not the number of babies.

Pronunciation

Pronunciations come from the Carnegie Mellon Pronouncing Dictionary, a public-domain phonetic lexicon of North American English, converted from its ARPAbet notation into IPA. Where the dictionary has no entry we show nothing rather than generating a guess, which is why some name pages carry no pronunciation at all. These are American readings: a name said differently in Britain or Ireland will show the American form, and we would rather say that than pretend to a neutrality we do not have.

Surnames

The surname list is the 2010 US Census file of every surname borne by a hundred people or more. A first name is scored against a surname with the same rules as the middle-name score, since the question is the same one: how do these two words meet when spoken together.

Meanings and origins

Etymologies are researched and written by our editors from standard onomastic references. Where scholars disagree, the page says the origin is debated. Names without a defensible etymology carry no meaning field rather than a guess.

Corrections

If a meaning or a figure here is wrong, tell us and we will fix it and say so.