TY - JOUR
KW - clustering method
KW - CyberAggregator
KW - information summarization
KW - social media monitoring
KW - subject domain
KW - visualization
KW - words network
AU - Dmytro Lande
AU - Ihor Subach
AU - Olexander Puchkov
AU - Artem Soboliev
AB - <p style="margin-left:19.85pt;">The article presents a discriminant cluster analysis method used to form real-time models of subject areas and digests based on automatic analysis of a large number of messages from social networks. It is based on estimating the discriminant value of terms. Cluster analysis, like the well-known LSA algorithm, provides a matrix representation of the data. The novelty is in using the most significant discriminant values as centroids to define clusters.</p><p style="margin-left:19.85pt;">The algorithm is simplified; it does not involve referencing to the adjacency matrix, definition of eigenvectors. Its complexity is O(N2), where K is the number of clusters and N – the number of reference terms. If it is necessary to improve the quality of the proposed approach, the defined centroids can be transferred as input data for other known algorithms. Based on the above algorithm, toolkits for the formation of a language network and digests were developed and embedded in the “CyberAggregator” system, which provides accumulation, processing, summarization of data from social networks on cybersecurity issues.</p>
BT - Information & Security: An International Journal
DO - https://doi.org/10.11610/isij.5013
IS - 1
N2 - <p style="margin-left:19.85pt;">The article presents a discriminant cluster analysis method used to form real-time models of subject areas and digests based on automatic analysis of a large number of messages from social networks. It is based on estimating the discriminant value of terms. Cluster analysis, like the well-known LSA algorithm, provides a matrix representation of the data. The novelty is in using the most significant discriminant values as centroids to define clusters.</p><p style="margin-left:19.85pt;">The algorithm is simplified; it does not involve referencing to the adjacency matrix, definition of eigenvectors. Its complexity is O(N2), where K is the number of clusters and N – the number of reference terms. If it is necessary to improve the quality of the proposed approach, the defined centroids can be transferred as input data for other known algorithms. Based on the above algorithm, toolkits for the formation of a language network and digests were developed and embedded in the “CyberAggregator” system, which provides accumulation, processing, summarization of data from social networks on cybersecurity issues.</p>
PY - 2021
SP - 79
EP - 86
T2 - Information & Security: An International Journal
TI - A Clustering Method for Information Summarization and Modelling a Subject Domain
VL - 50
ER -