Abstract
Effective extraction of domain-specific terms and named entities is a key challenge in text mining. This paper investigates the use of the k-means clustering algorithm for unsupervised extraction of unigrams and named entities from text data. The approach groups terms based on their vector representations, enabling the identification of semantically similar words without labeled data. Experiments conducted on the ACTER (Annotated Corpora for Term Extraction Research) corpus evaluate the method using precision, recall, and F1-score. Results show average scores of 25.79% precision, 40.05% recall and 30.47% F1-score, with optimal performance achieved using 40 to 60 clusters. Future work will explore algorithm optimization and comparisons with alternative extraction techniques.
| Original language | English |
|---|---|
| Title of host publication | Selected Papers from the International Conference on Artificial Intelligence - FICAILY2025 - Current Research, Industry Trends, and Innovations |
| Editors | Ali Othman Albaji |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 375-386 |
| Number of pages | 12 |
| ISBN (Print) | 9783032002310 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | International Conference on AI: Current Research, Industry Trends, and Innovations, FICAILY 2025 - Tripoli, Libya Duration: 9 Jul 2025 → 10 Jul 2025 |
Publication series
| Name | Studies in Computational Intelligence |
|---|---|
| Volume | 1229 SCI |
| ISSN (Print) | 1860-949X |
| ISSN (Electronic) | 1860-9503 |
Conference
| Conference | International Conference on AI: Current Research, Industry Trends, and Innovations, FICAILY 2025 |
|---|---|
| Country/Territory | Libya |
| City | Tripoli |
| Period | 9/07/25 → 10/07/25 |
Bibliographical note
Publisher Copyright:© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
Keywords
- Automated term extraction
- Clustering
- NLP
- Unigrams
- Unsupervised annotator
- k-means
Fingerprint
Dive into the research topics of 'Automatic Unsupervised Extraction of Unigrams of Terms and Named Entities Using the K-Means Clustering Algorithm'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver