Document Clustering and Visualization for Large Text Collections

dc.contributor.advisorBilal, Lounnas
dc.contributor.authorHadjer, Meghiche
dc.contributor.authorKhadidja Nour Elhouda, Gana
dc.date.accessioned2026-06-22T08:39:46Z
dc.date.issued2026-06-10
dc.description.abstractThe exponential growth of digital textual data across scientific, journalistic, and social me￾dia domains has created a critical need for automated document organisation and exploration tools. This study presents a comprehensive end-to-end pipeline for document clustering and visualisation applied to large text collections. Three complementary approaches are system￾atically evaluated: K-Means clustering with TF-IDF representations, Latent Dirichlet Alloca￾tion (LDA), and BERTopic, a state-of-the-art neural framework combining transformer-based embeddings with UMAP dimensionality reduction and HDBSCAN density-based clustering. Experiments are conducted on three benchmark corpora 20 Newsgroups, Reuters-21578, and a curated arXiv dataset representing diverse levels of semantic complexity and class distribu￾tion. Performance is assessed using six complementary metrics covering internal cluster qual￾ity, external label agreement, topic coherence and diversity, and computational efficiency. Results demonstrate that BERTopic achieves superior semantic coherence and topic diver￾sity, while K-Means delivers the highest structural accuracy and computational efficiency when paired with dense sentence embeddings. LDA, although interpretable, consistently un￾derperforms in capturing deep semantic structure. These findings provide practical guidelines for selecting document clustering methods in real-world information retrieval applications.
dc.identifier.urihttps://depot.univ-msila.dz/handle/123456789/48711
dc.language.isoen
dc.publisherUniversity of M'sila
dc.subjectDocument Clustering
dc.subjectBERTopic
dc.subjectK-Means
dc.subjectLDA
dc.subjectTF-IDF
dc.subjectSentence-BERT
dc.subjectUMAP
dc.subjectDimensionality Reduction
dc.subjectTopic Modelling
dc.subjectText Mining
dc.titleDocument Clustering and Visualization for Large Text Collections
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
document clustering final version.pdf
Size:
6.58 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections