How to analyse word frequency
- Paste your text into the box above, an article, an essay, a transcript or a page of copy.
- Tick Ignore common English words to hide the, and, of and their relatives.
- Raise the minimum length to 4 or 5 to filter out short words another way.
- Read the ranked list, which shows each word's count and its share of the text.
- Copy or download the list if you want to keep it.
What the analysis shows
- Every word ranked by frequency, with its count and percentage share.
- A stop-word filter that hides the common connecting words.
- A minimum length filter as an alternative way to cut the noise.
- Vocabulary variety, unique words as a share of the total.
- A count of words used exactly once, which indicates range.
- Case-insensitive matching with punctuation stripped, so Dog, dog and dog. are one word.
What frequency tells you about writing
Every writer has crutches. Words that felt right the first time and then kept arriving, "actually", "essentially", "leverage", "really", a favourite adjective applied to everything. They are invisible while you write and obvious to a reader, and the fastest way to find yours is to count.
Frequency analysis also exposes structural problems. If the top content word in a 2,000-word article appears forty times, the piece is probably circling rather than progressing. If a report about three topics mentions one of them five times as often as the others, the balance is off regardless of what the outline said.
Turn the stop-word filter on first
Without filtering, the top of any English text is always the same: the, of, and, to, a, in. These are function words. They carry grammar rather than meaning and they tell you nothing about your subject.
The filter removes about 200 of the most common English words, which pushes the words that actually carry your content to the top. That is the list worth reading. The minimum length control does a cruder version of the same job, set it to five and most function words disappear, and is the better option for text in other languages, since the stop-word list is English only.
Vocabulary variety and what it means
The variety figure is unique words divided by total words. It is a rough measure of lexical range, and it needs interpreting carefully because it falls naturally as a text gets longer. A 200-word paragraph might show 60 per cent, while a 5,000-word article showing 25 per cent is not necessarily repetitive.
Compare like with like. Two articles of similar length by the same author are comparable; a tweet and a thesis are not. The "used only once" figure is a useful companion: a high proportion suggests wide vocabulary, while a low one suggests the same terms recycling.
Practical uses beyond writing
- SEO: checking keyword density without guessing. Around one to two per cent for a primary term reads naturally; five per cent reads as stuffing.
- Research: finding recurring themes across interview transcripts or survey free-text responses.
- Editing: spotting an author's tics before a line edit.
- Teaching: comparing vocabulary range across student submissions.
- Translation: identifying the terms that recur often enough to need a consistent rendering.
How words are counted
Comparison is case-insensitive with punctuation stripped, so "Dog", "dog" and "dog." are one word. Words are not stemmed, which means "run", "runs" and "running" count separately. That is deliberate, since collapsing them would hide the difference between varied and repetitive phrasing.
Contractions keep their apostrophes, so "don't" is one word rather than two. Numbers are counted as words. If the text came from a PDF, extract it with the PDF to text tool first, and clean up any formatting debris with the text cleaner so stray characters do not skew the counts. Everything runs locally, so unpublished manuscripts and client copy stay private.