Scope and above

Cluster Analysis — Group Documents by Thematic Similarity

A corpus of 200 papers is not one conversation — it is many conversations happening simultaneously. Cluster Analysis finds those conversations and draws the borders between them.

▶ Cluster Analysis grouping a corpus of 60 papers into thematic groups · no upload

What is Cluster Analysis?

Cluster Analysis groups your documents by the similarity of their term frequency profiles. Documents that share the same pattern of terms end up in the same cluster. Documents with divergent term profiles end up in separate clusters. The result is a set of thematic groups — each cluster is a sub-field, a school of thought, or a methodological tradition within your corpus.

How it works

Gaptrium™ builds a term frequency vector for each document. Cluster Analysis compares those vectors using a deterministic similarity function and groups documents whose vectors are closest. The number of clusters, the grouping threshold, and the labelling of each cluster are all configurable. Every cluster membership is traceable — you can see exactly why each document was assigned to each group.

Available on Scope and above. Trial includes Heat Map and Gap Analysis. Start free →

Who uses Cluster Analysis?

🎓

Systematic Literature Review

Before writing your synthesis, understand the sub-fields within your inclusion set. Cluster Analysis shows you which papers are talking about the same thing and which are approaching the field from a different direction.

🔬

Research Landscape Mapping

Show funders and policy-makers the structure of a research field. Which clusters are growing? Which are contracting? Where is the field fragmenting into sub-disciplines?

📚

Curriculum Development

Cluster a body of educational literature to identify the distinct pedagogical schools present in the field — and which perspectives are absent from your reading list.

🏭

Competitive Analysis

Cluster competitor filings, product announcements, and press releases to identify the distinct strategic narratives across an industry.

Frequently asked questions

Is Cluster Analysis based on AI or machine learning?

No. Gaptrium™ Cluster Analysis uses deterministic frequency-based similarity. There is no trained model, no probabilistic inference, and no risk of hallucination. The same corpus produces the same clusters every time.

Can I edit cluster names and membership?

Yes. After the initial automatic clustering, you can rename each cluster, move documents between clusters, and add or remove terms from the cluster definition.

How many clusters does Gaptrium™ create?

You control this. Gaptrium™ suggests an optimal number based on the corpus structure, but you can set any number from 2 to 10.

Does Cluster Analysis work on small corpora?

Yes. Useful clustering is possible from as few as 10 documents. The algorithm adapts to corpus size.

Can I export cluster assignments?

Yes — as CSV (all plans), HTML and PNG (Scope and above), PDF (Education, Signal, Core).