An Empirical Comparative Study of Machine Learning, Deep Learning, and Transformer Models for Text Classification
DOI:
https://doi.org/10.19153/cleiej.29.2.3Keywords:
Natural Language Processing, Text Classification, Deep Learning, Transformer Models, Machine LearningAbstract
Text classification as a sentiment analysis, spam and document classification is an important task in Natural Language Processing (NLP). Recent developments of deep learning, and transformer-based models have contributed significantly to classifying accuracy. These models have difficulty meeting computational efficiency, inference latency, and interpretability criteria, and cannot be used in resource-constrained and real-time settings. The models are critically compared based on accuracy, computational efficiency and trade-offs on benchmark datasets in this paper. The results indicate that transformer models are much more accurate than other conventional concepts, but have a significant disadvantage of being computationally expensive, which severely limits their actual implementation. Also, XAI methods and fairness-conscious training approaches are required particularly when issues of transparency and fairness of the existing model emerge. New methods like zero-shot learning, few-shot learning, and self-supervised learning are promising directions to further improve generalization and flexibility of NLP models. Future studies must focus on identifying light, energy efficient NLP architecture that can achieve good accuracy, interpretability and scalability. This work fills the gaps between theory and practice, which is why it can be of interest to both academic researchers and practitioners in the industry.
References
J. Jia, W. Liang, and Y. Liang, “A Review of Hybrid and Ensemble in Deep Learning for Natural Language Processing,” 2023, arXiv (Cornell University). Accessed: Aug. 27, 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2312.05589
D. Tiwari, B. Nagpal, B. S. Bhati, A. Mishra, and M. Kumar, “A systematic review of social network sentiment analysis with comparative study of ensemble-based techniques,” Artif Intell Rev, vol. 56, no. 11, pp. 13407–13461, Nov. 2023, doi: 10.1007/s10462-023-10472-w.
A. Jerfy, O. Selden, and R. Balkrishnan, “The Growing Impact of Natural Language Processing in Healthcare and Public Health,” INQUIRY: The Journal of Health Care Organization, Provision, and Financing, vol. 61, Jan. 2024, doi: 10.1177/00469580241290095.
T. Bao, N. Ren, R. Luo, B. Wang, G. Shen, and T. Guo, “A BERT-Based Hybrid Short Text Classification Model Incorporating CNN and Attention-Based BiGRU,” Journal of Organizational and End User Computing, vol. 33, no. 6, pp. 1–21, Dec. 2021, doi: 10.4018/JOEUC.294580.
Supriyono, A. P. Wibawa, Suyono, and F. Kurniawan, “Advancements in natural language processing: Implications, challenges, and future directions,” Telematics and Informatics Reports, vol. 16, p. 100173, Dec. 2024, doi: 10.1016/j.teler.2024.100173.
K. Y. Thakkar and N. Jagdishbhai, “Exploring the capabilities and limitations of GPT and Chat GPT in natural language processing,” Journal of Management Research and Analysis, vol. 10, no. 1, pp. 18–20, Apr. 2023, doi: 10.18231/j.jmra.2023.004.
R. M. Samant, M. R. Bachute, S. Gite, and K. Kotecha, “Framework for Deep Learning-Based Language Models Using Multi-Task Learning in Natural Language Understanding: A Systematic Literature Review and Future Directions,” IEEE Access, vol. 10, pp. 17078–17097, 2022, doi: 10.1109/ACCESS.2022.3149798.
A. Özçift, K. Akarsu, F. Yumuk, and C. Söylemez, “Advancing natural language processing (NLP) applications of morphologically rich languages with bidirectional encoder representations from transformers (BERT): an empirical case study for Turkish,” Automatika, vol. 62, no. 2, pp. 226–238, Apr. 2021, doi: 10.1080/00051144.2021.1922150.
E. Kotei and R. Thirunavukarasu, “A Systematic Review of Transformer-Based Pre-Trained Language Models through Self-Supervised Learning,” Information, vol. 14, no. 3, p. 187, Mar. 2023, doi: 10.3390/info14030187.
W. Khan, A. Daud, K. Khan, S. Muhammad, and R. Haq, “Exploring the frontiers of deep learning and natural language processing: A comprehensive overview of key challenges and emerging trends,” Natural Language Processing Journal, vol. 4, p. 100026, Sep. 2023, doi: 10.1016/j.nlp.2023.100026.
A. Anuragi, D. S. Sisodia, and R. B. Pachori, “Mitigating the curse of dimensionality using feature projection techniques on electroencephalography datasets: an empirical review,” Artif Intell Rev, vol. 57, no. 3, p. 75, Feb. 2024, doi: 10.1007/s10462-024-10711-8.
S. Vora and H. Yang, “A comprehensive study of eleven feature selection algorithms and their impact on text classification,” in 2017 Computing Conference, IEEE, Jul. 2017, pp. 440–449. doi: 10.1109/SAI.2017.8252136.
Y. Huang, B. Giledereli, A. Köksal, A. Özgür, and E. Ozkirimli, “Balancing Methods for Multi-label Text Classification with Long-Tailed Class Distribution,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 8153–8161. doi: 10.18653/v1/2021.emnlp-main.643.
A. Benayas, S. Miguel-Ángel, and M. Mora-Cantallops, “Enhancing Intent Classifier Training with Large Language Model-generated Data,” Applied Artificial Intelligence, vol. 38, no. 1, Dec. 2024, doi: 10.1080/08839514.2024.2414483.
Z. Yang, Z. Dai, Y. Yan, J. Carbonell, R. Salakhutdinov, and Q. V. Le, “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” arXiv (Cornell University), 2019.
J. Chen, D. Tam, C. Raffel, M. Bansal, and D. Yang, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” Trans Assoc Comput Linguist, vol. 11, pp. 191–211, Mar. 2023, doi: 10.1162/tacl_a_00542.
C. A. Libbi, J. Trienes, D. Trieschnigg, and C. Seifert, “Generating Synthetic Training Data for Supervised De-Identification of Electronic Health Records,” Future Internet, vol. 13, no. 5, p. 136, May 2021, doi: 10.3390/fi13050136.
S. F. Sabbeh and H. A. Fasihuddin, “A Comparative Analysis of Word Embedding and Deep Learning for Arabic Sentiment Classification,” Electronics (Basel), vol. 12, no. 6, p. 1425, Mar. 2023, doi: 10.3390/electronics12061425.
J. Shi et al., “Two End-to-End Quantum-Inspired Deep Neural Networks for Text Classification,” IEEE Trans Knowl Data Eng, vol. 35, no. 4, pp. 4335–4345, Apr. 2023, doi: 10.1109/TKDE.2021.3130598.
A. Ramponi and B. Plank, “Neural Unsupervised Domain Adaptation in NLP—A Survey,” in Proceedings of the 28th International Conference on Computational Linguistics, Stroudsburg, PA, USA: International Committee on Computational Linguistics, 2020, pp. 6838–6855. doi: 10.18653/v1/2020.coling-main.603.
E. Ben-David, N. Oved, and R. Reichart, “PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains,” Trans Assoc Comput Linguist, vol. 10, pp. 414–433, Apr. 2022, doi: 10.1162/tacl_a_00468.
J. Chen, Y. Hu, J. Liu, Y. Xiao, and H. Jiang, “Deep Short Text Classification with Knowledge Powered Attention,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, pp. 6252–6259, Jul. 2019, doi: 10.1609/aaai.v33i01.33016252.
H. Wang, J. He, X. Zhang, and S. Liu, “A Short Text Classification Method Based on N ?Gram and CNN,” Chinese Journal of Electronics, vol. 29, no. 2, pp. 248–254, Mar. 2020, doi: 10.1049/cje.2020.01.001.
Ö. Çetino?lu, S. Schulz, and N. T. Vu, “Challenges of Computational Processing of Code-Switching,” in Proceedings of the Second Workshop on Computational Approaches to Code Switching, Stroudsburg, PA, USA: Association for Computational Linguistics, 2016, pp. 1–11. doi: 10.18653/v1/W16-5801.
Y. Sharma, R. Bhargava, and B. V. Tadikonda, “Named Entity Recognition for Code Mixed Social Media Sentences,” International Journal of Software Science and Computational Intelligence, vol. 13, no. 2, pp. 23–36, Apr. 2021, doi: 10.4018/IJSSCI.2021040102.
A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Corpus creation and language identification for code-mixed Indonesian-Javanese-English Tweets,” PeerJ Comput Sci, vol. 9, p. e1312, Jun. 2023, doi: 10.7717/peerj-cs.1312.
M. Ennab and H. Mcheick, “Designing an Interpretability-Based Model to Explain the Artificial Intelligence Algorithms in Healthcare,” Diagnostics, vol. 12, no. 7, p. 1557, Jun. 2022, doi: 10.3390/diagnostics12071557.
J. El Zini and M. Awad, “On the Explainability of Natural Language Processing Deep Models,” ACM Comput Surv, vol. 55, no. 5, pp. 1–31, May 2023, doi: 10.1145/3529755.
S. Häffner, M. Hofer, M. Nagl, and J. Walterskirchen, “Introducing an Interpretable Deep Learning Approach to Domain-Specific Dictionary Creation: A Use Case for Conflict Prediction,” Political Analysis, vol. 31, no. 4, pp. 481–499, Oct. 2023, doi: 10.1017/pan.2023.7.
N. V. D. S. S. V. Prasad Raju and P. N. Devi, “A Comparative Analysis of Machine Learning Algorithms for Big Data Applications in Predictive Analytics,” International Journal of Scientific Research and Management (IJSRM), vol. 12, no. 10, pp. 1608–1630, Oct. 2024, doi: 10.18535/ijsrm/v12i10.ec09.
S. Singh and A. Mahmood, “The NLP Cookbook: Modern Recipes for Transformer Based Deep Learning Architectures,” IEEE Access, vol. 9, pp. 68675–68702, 2021, doi: 10.1109/ACCESS.2021.3077350.
V. Perrone, M. Palma, S. Hengchen, A. Vatri, J. Q. Smith, and B. McGillivray, “GASC: Genre-Aware Semantic Change for Ancient Greek,” in Proceedings of the 1st International Workshop on Computational Approaches to Historical Language Change, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 56–66. doi: 10.18653/v1/W19-4707.
G. D. Rosin, I. Guy, and K. Radinsky, “Time Masking for Temporal Language Models,” in Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, New York, NY, USA: ACM, Feb. 2022, pp. 833–841. doi: 10.1145/3488560.3498529.
S. B A and R. P. K N, “ORDSAENet: Outlier Resilient Semantic Featured Deep Driven Sentiment Analysis Model for Education Domain,” Journal of Machine and Computing, pp. 408–430, Oct. 2023, doi: 10.53759/7669/jmc202303034.
K. Juluru, H.-H. Shih, K. N. Keshava Murthy, and P. Elnajjar, “Bag-of-Words Technique in Natural Language Processing: A Primer for Radiologists,” RadioGraphics, vol. 41, no. 5, pp. 1420–1426, Sep. 2021, doi: 10.1148/rg.2021210025.
A. Moreo, A. Esuli, and F. Sebastiani, “Word-class embeddings for multiclass text classification,” Data Min Knowl Discov, vol. 35, no. 3, pp. 911–963, May 2021, doi: 10.1007/s10618-020-00735-3.
M. Van Nguyen, V. D. Lai, A. Pouran Ben Veyseh, and T. H. Nguyen, “Trankit: A Light-Weight Transformer-based Toolkit for Multilingual Natural Language Processing,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 80–90. doi: 10.18653/v1/2021.eacl-demos.10.
P. Qi, Y. Zhang, Y. Zhang, J. Bolton, and C. D. Manning, “Stanza: A Python Natural Language Processing Toolkit for Many Human Languages,” arXiv (Cornell University), 2020.
E. Dias Canedo and B. Cordeiro Mendes, “Software Requirements Classification Using Machine Learning Algorithms,” Entropy, vol. 22, no. 9, p. 1057, Sep. 2020, doi: 10.3390/e22091057.
M. Errami, M. A. Ouassil, R. Rachidi, B. Cherradi, S. Hamida, and A. Raihani, “Sentiment Analysis on Moroccan Dialect based on ML and Social Media Content Detection,” International Journal of Advanced Computer Science and Applications, vol. 14, no. 3, 2023, doi: 10.14569/IJACSA.2023.0140347.
M. Veziro?lu, E. Veziro?lu, and ?. Ö. Bucak, “Performance Comparison between Naive Bayes and Machine Learning Algorithms for News Classification,” in Bayesian Inference - Recent Trends, IntechOpen, 2024. doi: 10.5772/intechopen.1002778.
H. Zhou, “Research of Text Classification Based on TF-IDF and CNN-LSTM,” J Phys Conf Ser, vol. 2171, no. 1, p. 012021, Jan. 2022, doi: 10.1088/1742-6596/2171/1/012021.
M. Zulqarnain, R. Ghazali, Y. M. M. Hassim, and M. Rehan, “A comparative review on deep learning models for text classification,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 19, no. 1, p. 325, Jul. 2020, doi: 10.11591/ijeecs.v19.i1.pp325-335.
N. Mohan, K. P. Soman, and R. Vinayakumar, “Deep power: Deep learning architectures for power quality disturbances classification,” in 2017 International Conference on Technological Advancements in Power and Energy ( TAP Energy), IEEE, Dec. 2017, pp. 1–6. doi: 10.1109/TAPENERGY.2017.8397249.
Y. Liu, M. He, M. Shi, and S. Jeon, “A Novel Model Combining Transformer and Bi-LSTM for News Categorization,” IEEE Trans Comput Soc Syst, vol. 11, no. 4, pp. 4862–4869, Aug. 2024, doi: 10.1109/TCSS.2022.3223621.
A. Aravinthan and C. Eugene, “Exploring Recent NLP Advances for Tamil: Word Vectors and Hybrid Deep Learning Architectures,” International Journal on Advances in ICT for Emerging Regions (ICTer), vol. 17, no. 2, pp. 85–101, Oct. 2024, doi: 10.4038/icter.v17i2.7279.
M. Kamyab, G. Liu, and M. Adjeisah, “Attention-Based CNN and Bi-LSTM Model Based on TF-IDF and GloVe Word Embedding for Sentiment Analysis,” Applied Sciences, vol. 11, no. 23, p. 11255, Nov. 2021, doi: 10.3390/app112311255.
A. Vaswani et al., “Attention Is All You Need,” arXiv (Cornell University), 2023.
S. Islam et al., “A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks,” arXiv (Cornell University), 2023.
A. Gillioz, J. Casas, E. Mugellini, and O. A. Khaled, “Overview of the Transformer-based Models for NLP Tasks,” in Proceedings of the 2020 Federated Conference on Computer Science and Information Systems, Sep. 2020, pp. 179–183. doi: 10.15439/2020F20.
P. Cheng and K. Erk, “Attending to Entities for Better Text Understanding,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 05, pp. 7554–7561, Apr. 2020, doi: 10.1609/aaai.v34i05.6254.
P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with Disentangled Attention,” arXiv (Cornell University), 2020.
O. Rohanian et al., “Lightweight transformers for clinical natural language processing,” Nat Lang Eng, vol. 30, no. 5, pp. 887–914, Sep. 2024, doi: 10.1017/S1351324923000542.
Z. Gao, A. Feng, X. Song, and X. Wu, “Target-Dependent Sentiment Classification With BERT,” IEEE Access, vol. 7, pp. 154290–154299, 2019, doi: 10.1109/ACCESS.2019.2946594.
C. Sun, X. Qiu, Y. Xu, and X. Huang, “https://doi.org/10.48550/arxiv.1905.05583,” arXiv (Cornell University), 2019.
A. Mariji? and M. Bagi? Babac, “Predicting song genre with deep learning,” Global Knowledge, Memory and Communication, vol. 74, no. 1/2, pp. 93–110, Jan. 2025, doi: 10.1108/GKMC-08-2022-0187.
P. Delobelle, T. Winters, and B. Berendt, “RobBERT: a Dutch RoBERTa-based Language Model,” in Findings of the Association for Computational Linguistics: EMNLP 2020, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 3255–3265. doi: 10.18653/v1/2020.findings-emnlp.292.
H. Rehana et al., “Evaluating GPT and BERT models for protein–protein interaction identification in biomedical text,” Bioinformatics Advances, vol. 4, no. 1, Jan. 2024, doi: 10.1093/bioadv/vbae133.
H. Sohn and H. Lee, “MC-BERT4HATE: Hate Speech Detection using Multi-channel BERT for Different Languages and Translations,” in 2019 International Conference on Data Mining Workshops (ICDMW), IEEE, Nov. 2019, pp. 551–559. doi: 10.1109/ICDMW.2019.00084.
X. Zheng, C. Zhang, and P. C. Woodland, “Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, Dec. 2021, pp. 162–168. doi: 10.1109/ASRU51503.2021.9688232.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient Estimation of Word Representations in Vector Space,” arXiv (Cornell University), 2013.
J. Pennington, R. Socher, and C. Manning, “Glove: Global Vectors for Word Representation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2014, pp. 1532–1543. doi: 10.3115/v1/D14-1162.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Comput, vol. 9, no. 8, pp. 1735–1780, Nov. 1997, doi: 10.1162/neco.1997.9.8.1735.
K. Cho et al., “Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2014, pp. 1724–1734. doi: 10.3115/v1/D14-1179.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
Tom Brown et al., “Language Models are Few-Shot Learners,” Neural Information Processing Systems, vol. 33, pp. 1877–1901, 2020.
C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer ,” Journal of Machine Learning Research, vol. 21, 2020.
OpenAI et al., “GPT-4 Technical Report,” arXiv (Cornell University), 2023.
A. Y. Muaad et al., “Arabic Document Classification: Performance Investigation of Preprocessing and Representation Techniques,” Math Probl Eng, vol. 2022, pp. 1–16, Apr. 2022, doi: 10.1155/2022/3720358.
K. M. Karaoglan and O. Findik, “Enhancing Aspect Category Detection Through Hybridised Contextualised Neural Language Models: A Case Study In Multi-Label Text Classification,” Comput J, vol. 67, no. 6, pp. 2257–2269, Jun. 2024, doi: 10.1093/comjnl/bxae004.
C. Zeng, S. Li, Q. Li, J. Hu, and J. Hu, “A Survey on Machine Reading Comprehension—Tasks, Evaluation Metrics and Benchmark Datasets,” Applied Sciences, vol. 10, no. 21, p. 7640, Oct. 2020, doi: 10.3390/app10217640.
B. Behera, G. Kumaravelan, and P. Kumar B., “Performance Evaluation of Deep Learning Algorithms in Biomedical Document Classification,” in 2019 11th International Conference on Advanced Computing (ICoAC), IEEE, Dec. 2019, pp. 220–224. doi: 10.1109/ICoAC48765.2019.246843.
“IMDB Dataset of 50K Movie Reviews,” www.kaggle.com. https://www.kaggle.com/datasets/lakshmi25npathi/imdb-dataset-of-50k-movie-reviews
“Sentiment Analysis on the Yelp Reviews Dataset,” kaggle.com. https://www.kaggle.com/code/omkarsabnis/sentiment-analysis-on-the-yelp-reviews-dataset
“Amazon reviews,” www.kaggle.com. https://www.kaggle.com/datasets/kritanjalijain/amazon-reviews
A. A. Jha, “Stanford Sentiment Treebank v2 (SST2),” Kaggle.com, 2020. https://www.kaggle.com/datasets/atulanandjha/stanford-sentiment-treebank-v2-sst2
SetFit, “sst5,” Huggingface.co, Sep. 18, 2023. https://huggingface.co/datasets/SetFit/sst5
A. Anand, “AG News Classification Dataset,” Kaggle.com, 2020. https://www.kaggle.com/datasets/amananandrai/ag-news-classification-dataset/data
“20 Newsgroups,” www.kaggle.com. https://www.kaggle.com/datasets/crawford/20-newsgroups
The Devastator, “Reuters-21578 (Text Categorization),” Kaggle.com, 2022. https://www.kaggle.com/datasets/thedevastator/uncovering-financial-insights-with-the-reuters-2
nasruddinaz, “BBCategorize - BBC News Category Classification,” Kaggle.com, Jul. 13, 2023. https://www.kaggle.com/code/nasruddinaz/bbcategorize-bbc-news-category-classification
sbhatti, “News Summarization,” Kaggle.com, 2018. https://www.kaggle.com/datasets/sbhatti/news-summarization
“Fake News Detection Datasets,” www.kaggle.com. https://www.kaggle.com/datasets/emineyetm/fake-news-detection-datasets
“Fake News Challenge,” Fakenewschallenge.org, 2016. http://www.fakenewschallenge.org/
Möbius, “COVID-19 Fake News Dataset,” Kaggle.com, 2020. https://www.kaggle.com/datasets/arashnic/covid19-fake-news
Narendra Prasath, “MultiLabel Classification - Reuters News Dataset,” Kaggle.com, 2020. https://www.kaggle.com/datasets/narendrageek/reuters21578-multilabel-classification-news
eshwarkoka, “GitHub - eshwarkoka/Medical-document-classification: Text classification on the medical abstracts in OHSUMED dataset,” GitHub, 2025. https://github.com/eshwarkoka/Medical-document-classification
S. University, “Stanford Natural Language Inference Corpus,” Kaggle.com, 2015. https://www.kaggle.com/datasets/stanfordu/stanford-natural-language-inference-corpus
“Welcome To Zscaler Directory Authentication,” Kaggle.com, 2026. https://www.kaggle.com/datasets/thedevastator/unlocking-language-understanding-with-the-multin
“Question Pairs Dataset,” www.kaggle.com. https://www.kaggle.com/datasets/quora/question-pairs-dataset
“The Stanford Question Answering Dataset,” rajpurkar.github.io. https://rajpurkar.github.io/SQuAD-explorer/
“google/boolq · Datasets at Hugging Face,” huggingface.co, Jan. 02, 2024. https://huggingface.co/datasets/google/boolq
facebookresearch, “GitHub - facebookresearch/XNLI: Evaluating Cross-lingual Sentence Representations,” GitHub, 2025. https://github.com/facebookresearch/XNLI
facebookresearch, “GitHub - facebookresearch/MLDoc: A Corpus for Multilingual Document Classification in Eight Languages.,” GitHub, 2025. https://github.com/facebookresearch/MLDoc
Y. Yang, Y. Zhang, C. Tar, and J. Baldridge, “PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification,” ACLWeb, Nov. 01, 2019. https://aclanthology.org/D19-1382/
Z. Ji, Q. Wei, and H. Xu, “BERT-based Ranking for Biomedical Entity Normalization,” AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science, vol. 2020, pp. 269–277, 2020, Available: https://pubmed.ncbi.nlm.nih.gov/32477646/
A. Mohan. kumar, “Legal Text Classification Dataset,” Kaggle.com, 2023. https://www.kaggle.com/datasets/amohankumar/legal-text-classification-dataset
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos, “MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Jan. 2021, doi: https://doi.org/10.18653/v1/2021.emnlp-main.559.
J. Yao and B. Yuan, “Optimization Strategies for Deep Learning Models in Natural Language Processing,” Journal of Theory and Practice of Engineering Science, vol. 4, no. 05, pp. 80–87, May 2024, doi: 10.53469/jtpes.2024.04(05).11.
S. Huang, Y. Liu, C. Fung, H. Wang, H. Yang, and Z. Luan, “Improving Log-Based Anomaly Detection by Pre-Training Hierarchical Transformers,” IEEE Transactions on Computers, vol. 72, no. 9, pp. 2656–2667, Sep. 2023, doi: 10.1109/TC.2023.3257518.
H. Ali et al., “Con-Detect: Detecting adversarially perturbed natural language inputs to deep classifiers through holistic analysis,” Comput Secur, vol. 132, p. 103367, Sep. 2023, doi: 10.1016/j.cose.2023.103367.
Warto et al., “Systematic Literature Review on Named Entity Recognition: Approach, Method, and Application,” Statistics, Optimization & Information Computing, vol. 12, no. 4, pp. 907–942, Feb. 2024, doi: 10.19139/soic-2310-5070-1631.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Rituraj Jain, Kamal Upreti, Ashish Sharma, Nilesh Kumar Sen, Ashish Kumar Mathur, Ramesh Babu P

This work is licensed under a Creative Commons Attribution 4.0 International License.
CLEIej is supported by its home institution, CLEI, and by the contribution of the Latin American and international researchers community, and it does not apply any author charges whatsoever for submitting and publishing. Since its creation in 1998, all contents are made publicly accesibly. The current license being applied is a (CC)-BY license (effective October 2015; between 2011 and 2015 a (CC)-BY-NC license was used).