Development and comparative evaluation of large language models for automated test case generation in banking software testing

Main Article Content

Hathairat Janwittaya
Ratthaslip Ranokphanuwat

Abstract

1) Software testing and test case generation are critical steps that require signifcant time and expertise, especially in complex and high-risk environments such as banking transaction systems. Since creating test cases is time-consuming and prone to errors, it is essential to apply artifcial intelligence technologies to accelerate the process and enhance test coverage. 2) This research develops and compares the performance of four Large Language Models (LLMs): LLaMA3.2-3B, Typhoon2-8b, Gemma-3-4b, and Qwen3-8B, for automated test case generation in software testing. 3) The research methodology consists of four main steps: 3.1) Utilizing a dataset of 20,000 test cases, which covers 15,763 positive and 4,237 negative cases within the banking transaction domain. 3.2) Fine-tuning the models using the LoRA technique and the LangChain framework, integrated with the ChromaDB vector database through a
Retrieval-Augmented Generation (RAG) architecture. 3.3) Evaluating performance using vector-based metrics, including Cosine Distance, Euclidean Distance, and Manhattan Distance. 3.4) Assessing test case coverage using Test Coverage, Functional Coverage, and Requirement Coverage, combined with evaluations from software testing experts. 4) The results indicate that Qwen3-8B achieved the lowest loss (0.1558), demonstrating the highest learning accuracy, while Gemma-3-4b obtained the lowest Euclidean Distance (0.5494) and Manhattan Distance (11.8533), indicating the closest similarity to the ground-truth data. In terms of coverage, Gemma-3-4b achieved the highest Test Coverage (93%) and Functional Coverage (91%), whereas Qwen3-8B achieved the highest Requirement Coverage (92%). Furthermore, expert evaluations revealed that Gemma-3-4b received the highest average score of 43.2, equivalent to 86.4%. In conclusion, the study demonstrates that employing LLMs for test case generation can signifcantly improve the coverage, accuracy, and effciency of software testing. Additionally, these models can be seamlessly integrated into chatbot systems to automatically generate test cases based on software requirements, highlighting their high practical potential for real-world applications in the software development industry.

Article Details

Section
Original Articles

References

Abbad, A., Abbad, K., & Tairi, H. (2016). Face recognition based on city-block and Mahalanobis cosine distance. In 2016 13th International Conference on Computer Graphics, Imaging and Visualization (CGiV) (pp. 112–114). IEEE.

Ammann, P., & Offutt, J. (2017). Introduction to software testing (2nd ed.). Cambridge University Press.

Bag, S., Gupta, A., Kaushik, R., & Jain, C. (2024). RAG beyond text: Enhancing image retrieval in RAG systems. In Proceedings of the 2024 International Conference on Electrical, Computer and Energy Technologies (ICECET) (pp. 310–315).

Bhatia, S., Gandhi, T., Kumar, D., & Jalote, P. (2024). Unit test generation using generative AI: A comparative performance analysis of autogeneration tools. In Proceedings of the 2024 International Workshop on Large Language Models for Code (LLM4Code ’24) (pp. 1–8). ACM.

Cheng, F. (2024). A comparative study of the performance of Spark-based k-means algorithm based on Euclidean distance and Manhattan distance. In Proceedings of the 3rd International Conference on Computer Communication and Artificial Intelligence (CCAI) (pp. 224–229).

Cheng, Y., Wang, M., Xiong, Y., Hao, D., & Zhang, L. (2016). Empirical evaluation of test coverage for functional programs. In 2016 IEEE International Conference on Software Testing, Verification and Validation (ICST) (pp. 255–265). IEEE.

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) (pp. 4171–4186).

Dhillon, I. S., Guan, Y., & Kogan, J. (2002). Refining clusters in high-dimensional text data. In Proceedings of the Workshop on Clustering High Dimensional Data and its Applications at the Second SIAM International Conference on Data Mining (pp. 71–82). SIAM.

Gao, X., & Li, G. (2016). A KNN model based on Manhattan distance to identify the SNARE proteins. Computational and Mathematical Methods in Medicine, 2016, Article 6480195. https://doi.org/10.1155/2016/6480195

Guan, W., & Fang, Y. (2025). Optimizing web-based AI query retrieval with GPT integration in LangChain: A CoT-enhanced prompt engineering approach. arXiv. https://arxiv.org/abs/2503.12345

Hoffmann, J., & Frister, D. (2024). Generating software tests for mobile applications using fine-tuned large language models. In Proceedings of the 45th International Conference on Software Engineering (ICSE) (pp. 1234–1245).

Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv. https://arxiv.org/abs/2106.09685

Hu, J., Liao, X., Gao, J., Qi, Z., Zheng, H., & Wang, C. (2024). Optimizing large language models with an enhanced LoRA fine-tuning algorithm for efficiency and robustness in NLP tasks. arXiv. https://arxiv.org/abs/2403.10567

Jacob, T. P., Bizotto, B. L. S., & Sathiyanarayanan, M. (2024). Constructing the ChatGPT for PDF files with LangChain – AI. In Proceedings of the 2024 International Conference on Inventive Computation Technologies (ICICT) (pp. 78–85).

Laplante, P. A., & Kassab, M. (2022). Requirements engineering for software and systems. Auerbach Publications.

Masuda, S., Nishi, Y., & Suzuki, K. (2020). Complex software testing analysis using international standards. In 2020 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW) (pp. 241–246). IEEE.

Nazi, A., Huang, Q., Shojaei, H., Esfeden, H. A., Mirhosseini, A., & Ho, R. (2022). Adaptive test generation for fast functional coverage closure. In DVCON USA.

Omar, S. F. (2013). A software traceability approach to support requirement-based test coverage analysis [Doctoral dissertation, Universiti Teknologi Malaysia].

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. arXiv. https://arxiv.org/abs/1905.00546

Schütze, H., Manning, C. D., & Raghavan, P. (2008). Introduction to information retrieval. Cambridge University Press.

Srikaewsiew, T., Khianchainat, K., Tharatipyakul, A., Pongnumkul, S., & Kanjanawattana, S. (2022). A comparison of the instructor-trainee dance dataset using cosine similarity, Euclidean distance, and angular difference. In 2022 26th International Computer Science and Engineering Conference (ICSEC) (pp. 235–240). IEEE.

Tiwari, D., Zhang, L., Monperrus, M., & Baudry, B. (2022). Production monitoring to improve test suites. IEEE Transactions on Reliability, 71(3), 1381–1397. https://doi.org/10.1109/TR.2022.3149480

Wang, Z., Guo, X., & Tsuchiya, T. (2025). Graph-centric approaches for coverage optimization in software requirement testing. In 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1270–1280). IEEE.

Wang, Z., Liu, K., Li, G., & Jin, Z. (2024). HITS: High-coverage LLM-based unit test generation via method slicing. In Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) (pp. 45–56).

Zhang, J., Wang, F., Ma, F., & Song, G. (2022). Text similarity calculation method based on optimized cosine distance. In 2022 International Conference on Electronics and Devices, Computational Science (ICEDCS) (pp. 37–39). IEEE.