Skip to main content
94
Views
279
Downloads
Reviewed article

A Clustering Method for Information Summarization and Modelling a Subject Domain

How to cite:
Dmytro Lande, Ihor Subach, Olexander Puchkov, Artem Soboliev
"A Clustering Method for Information Summarization and Modelling a Subject Domain"
Information & Security: An International Journal,
50
no. 1
(2021):
79-86 .
https://doi.org/10.11610/isij.5013

A Clustering Method for Information Summarization and Modelling a Subject Domain

Source:

Information & Security: An International Journal,
Volume: 50,
Issue1,
p.79-86
(2021)

Abstract:

The article presents a discriminant cluster analysis method used to form real-time models of subject areas and digests based on automatic analysis of a large number of messages from social networks. It is based on estimating the discriminant value of terms. Cluster analysis, like the well-known LSA algorithm, provides a matrix representation of the data. The novelty is in using the most significant discriminant values as centroids to define clusters.

The algorithm is simplified; it does not involve referencing to the adjacency matrix, definition of eigenvectors. Its complexity is O(N2), where K is the number of clusters and N – the number of reference terms. If it is necessary to improve the quality of the proposed approach, the defined centroids can be transferred as input data for other known algorithms. Based on the above algorithm, toolkits for the formation of a language network and digests were developed and embedded in the “CyberAggregator” system, which provides accumulation, processing, summarization of data from social networks on cybersecurity issues.

94
Views
279
Downloads