Start with a single number. ERIE is the third most common answer in American crosswords. It is also the 33,389th most common word in English. There are more than thirty thousand words you are more likely to meet in ordinary life than the name of a mid-sized lake, and yet inside a grid it outranks almost all of them.

That gap is the whole subject of this article. It has a name among solvers — crosswordese — but it is usually described rather than measured. It can be measured, and once you measure it two things become visible that a plain frequency list hides: crosswordese lives almost entirely in the shortest slots, and it is being replaced.

What we counted

The data is the xd corpus, an open archive of American crossword clues and answers. After cleaning it gives 8,006,563 answer usages and 368,549 distinct answers, spanning 1913 to 2026 and 46 publications. The New York Times is the largest single source at 2.46 million usages, followed by the Los Angeles Times (834K), Newsday (819K), Universal (714K) and USA Today (684K). So this is not a portrait of one newspaper's taste; it is American crosswords broadly.

For how common a word is in ordinary English we used LexiLab's own 46,832-word English frequency list, which is built from everyday written and spoken language and has nothing to do with puzzles. It matters that these two measurements are independent: one counts grids, the other counts language, and neither is derived from the other. Crosswordese is what you get when you subtract the second from the first.

One consequence of the first measurement is worth flagging, since it changed a tool while this article was being written. The Crossword Helper now ranks its suggestions by how often each word has actually been used as a crossword answer, rather than by how common it is in English — which is why it puts ERIE ahead of BRIE. The English list above is still involved, but only as one half of the ranking, never as the whole of it.

The 25 most common answers

Ranked by raw appearances, with each answer's rank in ordinary English alongside.

Most common crossword answers
8,006,563 answer usages, 46 American publications, 1913–2026.
#AnswerUsesRank in English
1ERA7,1874,679
2AREA7,084854
3ERIENAME5,98933,389
4ORE5,62213,261
5ALOE5,59834,157
6ARIA5,48012,012
7ONE5,37650
8ALE5,3659,058
9ERE5,35110,312
10ATE5,1641,889
11ELSE4,939232
12ARE4,90125
13ETA4,75712,257
14ELINAME4,7214,473
15ERR4,68511,072
16ANTE4,67716,207
17ALINAME4,6083,097
18EDENNAME4,5807,866
19ALA4,55929,517
20IDEA4,519283
21SPA4,4796,970
22IRE4,34339,304
23EAR4,3342,200
24ODE4,23326,652
25ORAL4,2179,843

NAME marks a proper noun. Every answer is 3 or 4 letters. Twenty-four of the twenty-five begin with a vowel — SPA is the only exception. This covers 1913–2026; restricting it to puzzles published since 2000 returns twenty-four of the same twenty-five words, with ORAL dropping out and OREO coming in at 13th. The list is not an artifact of old puzzles.

Two things stand out immediately. Every single entry is three or four letters long. And the list mixes two completely different kinds of word: ONE, ARE and IDEA are among the most common words in the language, while ERIE, ALOE, ALA and IRE are nowhere near it. Raw frequency cannot tell them apart. That is what the next section is for.

Measuring crosswordese

The intuition is simple: divide how often a word turns up in grids by how often it turns up in English. A word that is common in both, like AREA, scores low. A word that is common in grids and rare in English scores high. That ratio is crosswordese, expressed as a number.

One detail is worth flagging, because it changes the answer. The grid side is counted; the English side has to be estimated from a rank list. What the measure does not do is compare the two ranks directly, which is the obvious shortcut and gives a badly wrong result — ranks compress at the head of the distribution, so ERA and OREO differ by 1.7× in appearances but by 28× in rank. The exact formula is in the method note.

The most crosswordese answers
Ranked by grid frequency relative to ordinary English frequency.
AnswerUsesGrid rankRank in English
ERIENAME5,989333,389
ALOE5,598534,157
ENEABBR3,77338unranked
OREOBRAND4,1572841,763
IRE4,3432239,304
ETALPHRASE3,22266unranked
EPEE3,21469unranked
ETNANAME3,14272unranked
ASEA3,11973unranked
ESS3,8113637,054
ERAS3,4655038,914
ALAPHRASE4,5591929,517
OLEO2,831102unranked
ISEEPHRASE2,799107unranked

The top fourteen, with nothing filtered out — phrases and abbreviations are tagged and kept, on the same principle as the table above. "Unranked" means the word does not appear anywhere in a 46,832-word list of the most common English words. Those answers all share one floor value, so their order relative to each other is set by grid frequency alone: read this as a top tier rather than a strict 1-to-14 ranking.

This behaves the way a solver's instinct says it should. ERA, the single most repeated answer of all time, falls to 557th on this measure, because ERA is a perfectly ordinary English word that happens to be three letters of convenient shape. AREA falls to 1,441st. Meanwhile OREO rises into the top five.

The most extreme tier is the one where the right-hand column gives up entirely. 158 of the 2,000 most common answers are real dictionary words that fall outside the frequency list altogether — rare enough in everyday use that a list of the 46,832 most common English words does not reach them. EPEE, ASEA, OLEO, ALEE, APSE, ALIT, STET, ECRU, OLIO. Every one of these has a life outside puzzles — a fencer really does pick up an épée, an editor really does write stet — but for most solvers the crossword is where they meet them.

Crosswordese is a disease of short entries

Sort every answer ever printed by length, and the distribution is startlingly lopsided.

Share of all answer usages, by length
Three to five letters accounts for 73.9% of every answer ever printed.
18.2
32.1
23.5
10.7
5.8
3.2
1.7
1.5
0.9
0.5
0.4
0.3
0.9

The highlighted bar at 15 is the width of a standard American grid. Note the jump back up there. The curve decays smoothly to 0.29% at fourteen letters and then triples. Fifteen is the width of a standard American grid, so those are answers that span the puzzle from edge to edge — that spike is architecture, not vocabulary.

Now walk up the lengths and watch what happens to the answers themselves.

The inversion
The eight most repeated answers at each length, taken straight off the frequency count with nothing filtered out.

Three & four letters

ERA
AREA
ERIE
ORE
ALOE
ARIA
ONE
ALE

Ten letters

TURKEYTROT
OPENSESAME
SOURGRAPES
ROUNDROBIN
EASYSTREET
PAPERTIGER
HORSESENSE
RABBITEARS

Fifteen letters

STARSANDSTRIPES
MIDDLEOFTHEROAD
THROWINTHETOWEL
GONEWITHTHEWIND
SNAKEINTHEGRASS
BERMUDATRIANGLE
SINGININTHERAIN
CHINESECHECKERS

The left column is worth a second look, because it is not cherry-picked: it is simply the top eight, and it already contains both kinds of word. ERA, AREA and ONE are ordinary English; ERIE, ALOE and ARIA are not. Nothing comparable happens on the right. The transition between them is gradual — six-letter answers are still mostly single words (ESTATE, STEREO, ORIOLE, TSETSE), and by thirteen letters the change is complete: FLASHINTHEPAN, SMALLPOTATOES, COTTAGECHEESE, WHITEELEPHANT.

The most constrained slots in a crossword hold the least familiar language. The least constrained slots hold the most familiar.

This is not a coincidence, and it is not a matter of taste. It falls out of the geometry. A three-letter answer in a densely interlocked American grid is crossed on every single letter: all three of its letters must simultaneously work as part of three other answers running the other way. The pool of strings that can satisfy three constraints at once is small, and it skews heavily toward vowel-rich, consonant-light shapes. That is exactly the shape of ALOE, ASEA, OLEO, EPEE. Constructors do not use these words because they like them. They use them because at that length there is very little else that fits.

A fifteen-letter answer sits under the opposite pressure. It is usually a theme entry, chosen first and built around, with the rest of the grid arranged to accommodate it. It is selected to be lively — a phrase a solver will enjoy uncovering. So the entries with the most freedom get the most familiar language, and the entries with the least freedom get the strangest.

Why so many answers are not words at all

Look again at the fifteen-letter column and you will notice that none of those entries is a word. ONCEINABLUEMOON is five words with the spaces removed, and above about nine letters this is the normal case rather than the exception. It is the same pressure as crosswordese, running the other way.

A grid needs a letter run of an exact length that crosses cleanly in both directions. Ordinary single words are simply not numerous enough to fill every slot at every length, so the convention allows three ways of widening the pool, all of which appear in the data:

  • Multi-word phrases, run together: ETAL (et al.), ISEE (I see), AONE (A-one), INRE, ALOT. These dominate the long end.
  • Proper nouns: ERIE, ELI, ALI, EDEN, NERO, ETNA. Four of the twenty-five most repeated answers of all time are names.
  • Abbreviations and initialisms: SSE, RNA, ETS, NRA.

238 of the 2,000 most repeated answers fall into one of these three categories. We have marked them where they appear above rather than quietly dropping them, because removing ERIE from a list of the most common crossword answers would misrepresent the data — it really is the third most repeated answer there is.

The turnover

Everything so far describes crosswords as a single static object. They are not. And once you look at the same answers decade by decade, the most interesting result in the whole dataset appears.

This part needs care. Before 1994 the corpus is 99.7% New York Times, and broadly multi-publication afterwards, so a naive comparison across eras would mostly measure the archive changing shape rather than crosswords changing. To avoid that we restricted this section to the New York Times alone, which happens to be almost perfectly steady: between 308,000 and 333,000 answer usages in every decade from the 1960s to the 2010s. We then compared each answer's share of its decade, not its raw count.

Rising and falling in the New York Times
Share of all answers, 1990s–2010s compared with 1960s–1980s.
Rising
EPA7.6×
ATM5.8×
ECO5.1×
NERD4.7×
PSST4.0×
EURO3.8×
OREO3.7×
NCAA3.6×
Fading
ANTA0.17×
ESNE0.21×
ADIT0.27×
ORLE0.29×
ANIL0.34×
STOA0.38×
ITER0.48×
SNEE0.61×

New York Times only. Bar length is proportional to the size of the change, not to the number of appearances.

Read the two lists side by side and the pattern is hard to miss. The fading words are Latin, heraldry and classical architecture: ITER is a Roman road, STOA a Greek colonnade, ORLE a heraldic border, ANTA a pilaster, ESNE an Anglo-Saxon labourer, ADIT a mine entrance, ANIL an indigo shrub. The rising words are agencies, currencies, sports bodies, consumer brands and internet-era interjections.

The letter-shape pressure never changed — a constructor in 1965 and a constructor in 2015 both need vowel-heavy three-letter fill. What changed is the body of knowledge a puzzle assumes its solver has. Mid-century American crosswords were built for a reader with a classical education. Contemporary ones are built for a reader fluent in brands and acronyms. Crosswordese did not go away; it swapped registers.

When it happened

Decade buckets imply a slow tide. Plotted year by year, the change is not gradual at all: almost all of it lands inside a single year.

The chart below tracks thirteen classical crosswordese answers as a share of all New York Times answers: ITER, STOA, ESNE, ANTA, ORLE, ANIL, ADIT, SNEE, SMEE, ANOA, ETUI, SPAE and EDILE — the Latin, heraldry and dead-vocabulary tier.

Classical crosswordese in the New York Times, by year
Appearances per 10,000 answers, 1985–2025. Thirteen answers combined.

Dashed line: Will Shortz's first puzzle as New York Times crossword editor, 21 November 1993.

Maleska era and before Shortz era

The rate sits between 23 and 39 per 10,000 for every one of the nine years to 1993, then lands at 9.0 in 1994 and never returns. Across the whole Maleska-era window it averages 30.3 per 10,000; across 1994–1999, 9.8. Sixty-eight per cent of a category of vocabulary disappeared, essentially at a stroke.

The timing points somewhere specific. Eugene T. Maleska, who had edited the puzzle since 1977, died in August 1993. Will Shortz's first puzzle ran on 21 November 1993 — so 1993 is almost entirely his predecessor's work and 1994 is the first full year of his. He arrived with an explicit view of what was wrong: as he put it later, "the clues weren't playful, the answers weren't modern."

A change in what the archive contains would produce exactly the same-looking cliff, so that has to be ruled out first. It doesn't survive the check:

1993 compared with 1994
New York Times only. If the break were an artifact of the archive, every row would move together.
Measure
1993
1994
Change
Classical crosswordese, per 10k
23.6
9.0
−62%
ERIE / AREA / ALOE / ERA, per 10k
25.2
32.1
+27%
Distinct answers used
16,277
16,733
+2.8%
Total answers printed
31,305
31,200
−0.3%
Mean answer length
5.04
5.12
+1.6%

The puzzle kept printing the same number of answers, and the vowel-heavy workhorses went up. Only the obscure tier moved, which is what an editorial change looks like and not what a missing-data problem looks like.

That last row is quietly the most telling. Mean answer length rose and the number of distinct answers rose in the same year — a puzzle with longer entries and a wider vocabulary. And the geometric pressure this whole article is about did not go anywhere: the vowel-heavy pillars were used more in 1994, not less. The grid still demanded its ERIEs. What stopped was reaching for ESNE to get them.

The geometry stayed exactly as demanding. What changed was which words counted as an acceptable way to satisfy it.

One caution, and it is a real one. We cannot fully separate "the New York Times changed editors" from "American crosswords changed in the mid-1990s," because no other publication in this corpus has enough pre-1994 material to act as an external control — before 1994 the archive is 99.7% New York Times. Digital grid-filling software also became widely available in this period, and would independently have made obscure fill easier to avoid. What the data does establish is the shape of the change: a step, not a slope, at a moment when one publication changed editors. Gradual technology adoption does not produce a 62% drop between one December and the next January.

Note also what the chart shows after the break. The cliff did not finish the job — it started a long decline. The same thirteen words run at 7.9 per 10,000 through the 2000s, 6.4 through the 2010s, and 1.7 across 2020–2025. Against a Maleska-era baseline of 30.3, the classical tier is now running at about one-eighteenth of its former rate. ITER, STOA, ESNE, ANTA, ORLE and ADIT have not appeared in a New York Times puzzle since 2019.

The pillars do not move

One more result deserves a mention, because it cuts against the story. The very top of the list is remarkably stable. ERIE shifted by a factor of 0.79 between those eras, AREA by 0.88, ALOE by 0.94, ERA by 1.24. Essentially flat. The turnover is happening in the second tier, among the specialists. The handful of answers that solve the hardest geometric problems in the grid are apparently irreplaceable, and sixty years of changing taste has not dislodged them.

Run those same four answers through the periods used in the section above and the contrast is stark. Combined, they appear 25.8 times per 10,000 before 1994, 27.0 in 1994–99, 23.4 in the 2000s, 21.1 in the 2010s and 25.9 in 2020–25 — no trend at all, across the exact window in which the classical tier fell by a factor of eighteen. Whatever the 1994 break was, it was aimed at a specific register of vocabulary, not at short fill in general.


What this is useful for

The 25 above are the head of a longer list. The 200 most common crossword answers are collected on their own page, ranked on the post-2000 window and shown with the clue each one is given most often.

If you solve: the payoff of learning crosswordese is unusually concentrated, but it now has a shelf life. A few dozen short, vowel-heavy answers still cover a startling proportion of the fill you will meet, and the stable tier really is stable — ERIE and ALOE are as safe a bet today as in 1970. The classical tier is a different matter: if you learned ESNE and ANTA from a puzzle book, you learned words that a modern grid has essentially stopped using. And when you have a pattern and no idea, that is a search problem rather than a memory problem — which is what a pattern-based word finder is for.

If you build: the same pressure acts on any grid you make, including a small one built from a vocabulary list. The shorter and more interlocked your grid, the harder it pulls toward strange fill. The practical lever is geometry rather than word choice — a slightly more open grid, or one or two fewer crossings, does more for the quality of your fill than hunting for better three-letter words. That is worth knowing before you build a puzzle of your own, because it is much easier to loosen a grid than to rescue one that is already too tight.

More generally: crosswordese is worth measuring rather than just complaining about, because the measurement says something a word list cannot. It locates the problem in the geometry of the grid rather than in the taste of the people filling it, which is why the same handful of vowel-heavy answers survives every change in fashion. And it shows that the part which is taste — the classical tier, the Latin and the heraldry — turns out to be the part that can move, and has.


Method and caveats

Source. Answer counts come from the xd corpus of American crossword clues and answers. We counted 8,006,563 usages of 368,549 distinct answers across 46 publications, 1913–2026. Rankings in this article are restricted to the 22,504 answers appearing at least 50 times.

English frequency. Ranks come from LexiLab's 46,832-word English frequency list. The crosswordese score is log₂( (count / 8,006,563) × englishRank ). The first term is the answer's measured rate in grids; multiplying by rank stands in for dividing by an English rate, since Zipf's law makes frequency roughly proportional to 1/rank. The Zipf constant is dropped because it shifts every score equally and cannot change the ordering. We deliberately do not compare the two ranks, which compresses the head of the distribution and manufactures gaps the counts do not support.

The rank floor, and what it costs. An answer absent from the frequency list is assigned rank 46,833 — treated as rarer than everything on the list. This is the weakest part of the measure, and it has a specific consequence: the 317 unranked answers in our top 2,000 all share one denominator, so their order relative to each other reflects grid frequency only. EPEE is plainly more common in written English than ASEA, and the score cannot see that. We tested how much this matters by re-running the ranking with floors of 60,000, 100,000, 200,000 and 500,000. The membership of the top tier is robust — every answer in the top 25 sits beyond English rank 26,000 under any floor — but the ordering within it is not, and at higher floors the unranked words overtake ERIE and ALOE. Read the crosswordese table as a set, not a sequence. A continuous frequency source covering rare words would fix this properly and is the obvious next improvement.

What we excluded. Answers shorter than three letters, and answers containing digits or rebus symbols — these were rejected outright rather than stripped of their non-letters, since stripping would silently merge 1IRON into IRON. We also excluded runs of X, which the corpus uses as a placeholder where a solution was withheld rather than as an answer; left in, that placeholder would have ranked 58th overall.

The 1994 comparison. The thirteen tracked answers were chosen before the year-by-year rates were computed, on the basis of subject matter — Latin, heraldry, archaic trades and obsolete natural history — not on the basis of how they moved. Rates are per 10,000 New York Times answers in that year, so changes in how many puzzles the archive holds cannot produce the effect. The control rows in the 1993/1994 table exist to test the alternative explanation that the archive, rather than the puzzle, changed at that point.

Limits. This is American crosswords, weighted toward the New York Times and toward recent decades, and it says nothing about British cryptics, which work on entirely different principles. The trend and 1994 sections are New York Times only and should not be assumed to hold elsewhere. The 1994 break in particular has no external control: no other publication in the corpus has enough pre-1994 material to say whether the industry moved at the same time, so the association with the change of editor is a strong coincidence in timing rather than a demonstrated cause. Word validity was checked against Collins CSW24 and LexiLab's main English corpus; answers passing neither are marked as names, phrases or abbreviations rather than presented as words.