Keyword clustering: how to group keywords by topic
·4 min read
Keyword clustering is grouping a keyword list so that each group can be served by one page. You cluster either by SERP overlap, where two keywords belong together if Google ranks many of the same URLs for both, or by similarity, where they belong together if their words or meanings are close. Then each cluster becomes one page, targeted at its primary keyword, with the rest of the group as supporting terms.
The point is to stop writing five thin pages for five phrasings of the same question. Clustering is a planning step that sits between keyword research and writing. It is narrower than building topic clusters, which is about how a set of pages links together.
Two keyword clustering methods: SERP overlap vs similarity
SERP overlap asks whether Google already treats two queries as the same need. Similarity asks whether the queries look alike.
| SERP overlap | Similarity (word or semantic) | |
|---|---|---|
| Signal | Shared URLs in the top results for each keyword | Shared words, or close meaning in an embedding model |
| What it catches | Same intent, even with different wording | Phrasing variants of the same head term |
| What it misses | Angles nobody ranks a page for yet | Different intent behind similar words |
| Cost | One live results fetch per keyword | Runs on the list alone |
| Stability | Drifts as results change | Same input, same output |
SERP overlap. You pull the top 10 results for every keyword and group any pair that shares at least a set number of URLs. The threshold is yours to set. Three or four shared URLs out of ten is a sensible place to start. This is the closest proxy for intent you can get, because it reflects what the engine actually ranks. The downsides are cost and noise. Results differ by country and device, and thin niches show weak overlap even for queries that plainly belong together.
Similarity. Lexical clustering compares the words keywords share. Semantic clustering compares meaning through embeddings, so "cheap flights" and "budget airfare" can land together. Both are fast and deterministic enough to run on a few hundred terms in seconds. Both share one blind spot. "Apple pie recipe" and "apple pie calories" share most of their words, and Google may still want a recipe page for one and a nutrition answer for the other.
My take is to use similarity to get a first draft of groups, then check the clusters you plan to build pages for against the live results. Spot-checking 20 parent keywords by hand beats paying for overlap on 2,000 long-tail variants.
How to map one keyword cluster to one page
One cluster becomes one URL. Everything in the group is a phrasing the page should answer, and anything that needs a different page belongs in a different cluster.
- Clean the list first. Merge duplicates and plural variants so they do not create fake clusters.
- Cluster, then read every group with search intent in mind. Ask whether one page could satisfy every query in it. If one member wants a comparison and the rest want a how-to, move it out.
- Check your existing pages. If a cluster matches a page you already have, update that page instead of writing a new one. Two pages chasing the same cluster is how keyword cannibalization starts.
- Assign one URL per cluster and write it down.
- Keep singletons. A keyword that joined no group is still a candidate page, or a section on a related one.
Supporting keywords become H2s, FAQ answers or sentences in the body. If one needs more than a paragraph to answer, it probably wants its own page.
How to pick the primary keyword for each cluster
The primary keyword is the term the page is named for. It goes in the title, the H1 and the URL. Pick it by these tests, in order.
- Intent fit. It should describe what the whole page delivers, not one section of it.
- Breadth. The shortest phrasing that every member narrows is usually the right head term. "Running shoes" covers "trail running shoes" and "running shoes for beginners".
- Demand. When two candidates fit equally, take the one with more searches. Check volume with our bulk keyword volume checker rather than guessing.
- Winnability. If the broad head term is out of reach, a more specific member can make a better primary for now.
Group keywords with our keyword clustering tool
Our keyword clustering tool groups by word overlap only. It uses no embeddings, no language model, no search volume and no live search results. It drops stopwords such as "how", "for" and "the", ignores words shorter than three letters, and treats singular and plural as the same word. It then visits keywords broadest first. Each one joins an existing group when it contains every word of that group's parent, or shares at least two words with it at a Jaccard similarity of 0.5 or more. Otherwise it starts a new group. A keyword joins a parent directly, never through a chain, so "email" and "web hosting" cannot meet through "email hosting".
running shoes -> parent, /running-shoes
best running shoes
trail running shoes
running shoes for beginners
dress shoes -> singleton
email marketing -> singletonEach cluster comes back with its parent, which is the keyword with the fewest words, a suggested page path, the shared words and the members. The parent is the broadest phrasing, not the highest-volume one, so apply the demand test above before you commit. Paste one keyword per line or a Search Console export with a Query column. It takes up to 80 unique keywords per run, drops exact duplicates and skips URLs. Each run costs 3 credits. For hyphen and plural cleanup first, run our Keyword Deduplicator.