Skip to main navigation Skip to search Skip to main content

Automatic Unsupervised Extraction of Unigrams of Terms and Named Entities Using the K-Means Clustering Algorithm

  • Aliya Kalykulova
  • , Bilal Saoud*
  • , Ibraheem Shayea
  • , Dauren Sagidullauly
  • *Corresponding author for this work
  • Astana IT University
  • Akli Mohand Oulhadj University of Bouira

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Effective extraction of domain-specific terms and named entities is a key challenge in text mining. This paper investigates the use of the k-means clustering algorithm for unsupervised extraction of unigrams and named entities from text data. The approach groups terms based on their vector representations, enabling the identification of semantically similar words without labeled data. Experiments conducted on the ACTER (Annotated Corpora for Term Extraction Research) corpus evaluate the method using precision, recall, and F1-score. Results show average scores of 25.79% precision, 40.05% recall and 30.47% F1-score, with optimal performance achieved using 40 to 60 clusters. Future work will explore algorithm optimization and comparisons with alternative extraction techniques.

Original languageEnglish
Title of host publicationSelected Papers from the International Conference on Artificial Intelligence - FICAILY2025 - Current Research, Industry Trends, and Innovations
EditorsAli Othman Albaji
PublisherSpringer Science and Business Media Deutschland GmbH
Pages375-386
Number of pages12
ISBN (Print)9783032002310
DOIs
Publication statusPublished - 2026
EventInternational Conference on AI: Current Research, Industry Trends, and Innovations, FICAILY 2025 - Tripoli, Libya
Duration: 9 Jul 202510 Jul 2025

Publication series

NameStudies in Computational Intelligence
Volume1229 SCI
ISSN (Print)1860-949X
ISSN (Electronic)1860-9503

Conference

ConferenceInternational Conference on AI: Current Research, Industry Trends, and Innovations, FICAILY 2025
Country/TerritoryLibya
CityTripoli
Period9/07/2510/07/25

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

Keywords

  • Automated term extraction
  • Clustering
  • NLP
  • Unigrams
  • Unsupervised annotator
  • k-means

Fingerprint

Dive into the research topics of 'Automatic Unsupervised Extraction of Unigrams of Terms and Named Entities Using the K-Means Clustering Algorithm'. Together they form a unique fingerprint.

Cite this