Grasinančio turinio žinučių aptikimo naujienų komentaruose metodų tyrimas
Mikoliūnas, Aušrys |
Rastenis, Justinas | Recenzentas / Rewiewer |
Darbo gynimo komisijos pirmininkas / Thesis Defence Board Chairman | |
Darbo gynimo komisijos narys / Thesis Defence Board Member | |
Darbo gynimo komisijos narys / Thesis Defence Board Member | |
Darbo gynimo komisijos narys / Thesis Defence Board Member | |
Darbo gynimo komisijos narys / Thesis Defence Board Member | |
Darbo gynimo komisijos narys / Thesis Defence Board Member |
Magistrinio darbo tikslas - pritaikyti mašininio mokymo algoritmus grasinančio turinio tekstui atpažinti. Tikslui įgyvendinti buvo iškelti šie uždaviniai: atlikti teksto atpažinimui naudojamų metodų analizę, surinkti duomenų rinkinį, kuriuo bus apmokami minėtieji algoritmai, pasiūlyti metodą ar metodų kombinaciją kuri padėtų įgyvendinti darbo tikslą bei įgyvendinti pasiūlytus metodus ir pateikti rezultatus. Temos naujumas - toks tyrimas lietuvių kalbai dar nėra atliktas. Temos aktualumas - prisidėta prie lietuvių kalbos nagrinėjimo natūralios kalbos apdorojime. Įgyvendintas tikslas taip pat aktualus kaip filtras, filtruojantis grasinančius sakinius. Atlikus metodų analizę buvo išsiaiškinta, jog tekstui klasifikuoti egzistuoja daug klasifikatorių, kurie yra skirstomi į klasikinius ir giliojo mokymo algoritmus. Kiekvienas klasifikatorius turi privalumų ir trūkumų, todėl modelį reikia rinktis pagal esamą problemą ir siekiamą tikslą. Duomenų rinkinys buvo surinktas iš įvairių medijos portalų bei kitų kalbų egzistuojančių šaltinių, kurie buvo išversti į lietuvių kalbą. Rinkinys buvo sudarytas iš 500 grasinančių ir 500 neutralių sakinių. Tyrimams buvo pasiūlyta naudoti TF ir TF-IDF metodus, išmėginti dimensijų sumažinimą bei išbandyti skirtingus klasifikatorių parametrus. Galutiniams bandymas buvo išmėginti NB, SVM, BERT, „Decision Trees", „Gradinet Boosted", „Random Forest" ir MLP klasifikatoriai, TF ir TF-IDF bei LSA dimensijų mažinimas. Rezultatai parodė, jog dimensijų mažinimas pagerino tik DT ir GB klasifikatorių bendrą tikslumą, visų kitų klasifikatorių tikslumas nukrito. TF-IDF metodas buvo pranašesnis visuose atvejuose, išskyrus su NB klasifikatoriumi, kur TF metodas turėjo mažą pranašumą. Geriausius tikslumus parodė NB - 90 %, SVM - 89% ir MLP - 86%. Daugiausiai grasinančių sakinių atspėjo NB klasifikatorius, antroje vietoje MLP, todėl šie klasifikatoriai buvo pasirinkti integruoti į sakinio klasifikavimo API.
The aim of the Master's thesis is to apply machine learning algorithms to recognize text with threatening content. To achieve the goal, the following tasks were set: to perform an analysis of the methods used for text recognition, to collect a data set that will be used to train the aforementioned algorithms, to propose a method or a combination of methods that would help to implement the work's purpose, and to implement the proposed methods and present the results. The novelty of the topic is that such a study has not yet been carried out for the Lithuanian language. Relevance of the topic - contributed to the study of the Lithuanian language in natural language processing. A realized goal is also relevant as a filter that filters out threatening sentences. After analyzing the methods, it was found that there are many classifiers for text classification, which are divided into classic and deep learning algorithms. Each classifier has advantages and disadvantages, so the model should be chosen according to the existing problem and the desired goal. The data set was collected from various media portals and sources existing in other languages, which were translated into Lithuanian. The set consisted of 500 threatening and 500 neutral sentences. For research, it was proposed to use TF and TF-IDF methods, to try dimensionality reduction and to test different parameters of classifiers. For the final test, NB, SVM, BERT, Decision Trees, Gradinet Boosted, Random Forest and MLP classifiers, TF and TF-IDF and LSA dimensionality reduction were tested. The results showed that dimensionality reduction improved only DT and GB classifiers overall accuracy, all other classifiers dropped. The TF-IDF method was superior in all cases except for the NB classifier, where the TF method had a small advantage. The best accuracies were NB at 90%, SVM at 89%, and MLP at 86%. The most threatening sentences were guessed by the NB classifier, second to MLP, so these classifiers were chosen to be integrated into the sentence classification API.