Transparent data methods

Five-Letter Word Methodology

How one immutable source becomes the independently validated 2–15-letter lists, counts, scores, and position statistics used across this website.

1 · Versioning

One pinned, reproducible input

The word-data input is pinned to an immutable release before extraction. Normal production builds use committed, verified artifacts and never fetch a mutable latest or main branch.

Dialect
US English
Generated
Five-letter baseline
6,880 verified records

2 · Extraction

What enters each length-specific list

  • US-English and neutral spellings only.
  • Exactly the selected length, from 2 through 15 lowercase ASCII letters a to z.
  • Maximum source coverage threshold 70 and variant level 1.
  • Abbreviations, spaces, hyphens, apostrophes, and rewritten source forms are excluded.
  • Duplicates resolve deterministically using the narrower source coverage level, then variant and spelling metadata.

3 · Bands

CORE and EXTENDED are coarse groups

5,164

CORE

The narrower default coverage tier.

1,716

EXTENDED

Broader optional coverage beyond CORE.

These groups help place broader or less familiar entries after CORE results. They are not precise popularity, difficulty, or validity scores.

4 · Computation

How the published statistics are calculated

Occurrences
Every tile is counted, so a repeated letter may contribute more than once per word.
Words containing a letter
Each word contributes at most once for that letter.
Position frequency
Each word contributes exactly one letter to every position in its length. Heatmap percentages use all verified records for that length as the denominator.
Repeated letters
A word is marked repeated when its unique-letter count is smaller than its word length.
Scrabble points
Standard English tile face values are summed without board multipliers. A score does not establish official Scrabble legality.

Explore the complete frequency and position report for five-letter words; equivalent reports are linked from every 2–15-letter hub.

5 · Quality controls

Reproducible validation

Tests regenerate the statistics and compare them exactly with the committed artifact. They also verify spelling shape, uniqueness, input metadata, band limits, computed word features, manifest counts, and artifact hashes.

Programmatic pages have a separate data-depth gate. A pattern count alone never creates an indexable route.

The same checks support the Word Finder and Word Generator; both tools read committed client projections rather than changing the canonical word data in the browser.

6 · Boundaries

What the data does not claim

  • It is not an official Wordle answer or accepted-guess list.
  • It is not an official Scrabble or NASPA dictionary.
  • The internal coverage tier is not presented as an exact commonness score.
  • Every length 2–15 has its own independently generated records, statistics, and manifest; counts are never inferred from the five-letter dataset.
  • Aggregate statistics are available as JSON/CSV for every length; a separate downloadable word-list product remains deferred.

Read the preserved data copyright and license notices.