การพัฒนาและเปรียบเทียบประสิทธิภาพของโมเดลภาษาขนาดใหญ่สำหรับการสร้างกรณีทดสอบอัตโนมัติในงานทดสอบซอฟต์แวร์ระบบธนาคาร

Main Article Content

หทัยรัตน์ เจนวิทยา
รัฐศิลป์ รานอกภานุวัชร์

บทคัดย่อ

1) การทดสอบซอฟต์แวร์หรือการสร้างกรณีทดสอบ (Test Case) เป็นขั้นตอนสำคัญที่ใช้เวลาและต้องอาศัยความเชี่ยวชาญ โดยเฉพาะในระบบที่ซับซ้อนและมีความเสี่ยงสูง เช่น ระบบธุรกรรมทางการเงินของธนาคาร เนื่องจากการสร้างกรณีทดสอบ ใช้เวลานานและอาจเกิดข้อผิดพลาด จึงมีความจำเป็นต้องประยุกต์ใช้เทคโนโลยีปัญญาประดิษฐ์เพื่อเพิ่มความรวดเร็วและครอบคลุมในการสร้างกรณีทดสอบ 2) งานวิจัยพัฒนาและเปรียบเทียบประสิทธิภาพของ Large Language Models (LLMs) จำนวน 4 โมเดล ได้แก่ LLaMA3.2-3B, Typhoon2-8b, Gemma-3-4b และ Qwen3-8B สำหรับการสร้างกรณีทดสอบอัตโนมัติในงานทดสอบซอฟต์แวร์ 3) วิธีการดำเนินงานวิจัยประกอบด้วย 4 ขั้นตอนหลัก ได้แก่ 3.1) ชุดข้อมูลกรณีทดสอบ 20,000 รายการ ครอบคลุมกรณี Positive 15,763 รายการ และ Negative 4,237 รายการ ในบริบทธุรกรรมธนาคาร 3.2) ปรับแต่งโมเดลด้วยเทคนิค LoRA และเฟรมเวิร์ก LangChain เชื่อมต่อกับฐานข้อมูลเวกเตอร์ ChromaDB ผ่านสถาปัตยกรรม Retrieval-Augmented Generation (RAG) 3.3) ประเมินผลใช้ตัวชี้วัดเชิงเวกเตอร์ ได้แก่ Cosine Distance, Euclidean Distance และ Manhattan Distance 3.4) ประเมินความครอบคลุมของ Test Case ใช้ตัวชี้วัด Test Coverage, Functional Coverage และ Requirement Coverage ร่วมกับความพึงพอใจจากผู้เชี่ยวชาญด้านการทดสอบซอฟต์แวร์ 4) ผลการวิจัยพบว่า Qwen3-8B มีค่า Loss ต่ำสุด (0.1558) มีความแม่นยำในการเรียนรู้สูงสุด ขณะที่ Gemma-3-4b มีค่า Euclidean Distance (0.5494) และ Manhattan Distance (11.8533) ต่ำสุด แสดงถึงความใกล้เคียงกับข้อมูลจริงมากที่สุด ในด้านความครอบคลุมพบว่า Gemma-3-4b มีค่า Test Coverage สูงสุด (93%) และ Functional Coverage (91%) ส่วน Qwen3-8B มี Requirement Coverage สูงสุด (92%) และจากการประเมินผู้เชี่ยวชาญพบว่า Gemma-3-4b ได้คะแนนเฉลี่ยสูงสุด (43.2) คิดเป็นร้อยละ (86.4%) สรุปผลการศึกษาการใช้ LLM ในการสร้าง Test Case สามารถเพิ่มความครอบคลุม ความแม่นยำ และประสิทธิภาพในงานทดสอบซอฟต์แวร์ได้อย่างมีนัยสำคัญ อีกทั้งยังสามารถประยุกต์ใช้งานผ่านระบบแชทบอทเพื่อสร้างกรณีทดสอบ จากข้อกำหนดความต้องการซอฟต์แวร์ได้โดยอัตโนมัติ ซึ่งมีศักยภาพสูงสำหรับการใช้งานจริงในกระบวนการพัฒนาซอฟต์แวร์ในอุตสาหกรรม

Article Details

ประเภทบทความ
Original Articles

เอกสารอ้างอิง

Abbad, A., Abbad, K., & Tairi, H. (2016). Face recognition based on city-block and Mahalanobis cosine distance. In 2016 13th International Conference on Computer Graphics, Imaging and Visualization (CGiV) (pp. 112–114). IEEE.

Ammann, P., & Offutt, J. (2017). Introduction to software testing (2nd ed.). Cambridge University Press.

Bag, S., Gupta, A., Kaushik, R., & Jain, C. (2024). RAG beyond text: Enhancing image retrieval in RAG systems. In Proceedings of the 2024 International Conference on Electrical, Computer and Energy Technologies (ICECET) (pp. 310–315).

Bhatia, S., Gandhi, T., Kumar, D., & Jalote, P. (2024). Unit test generation using generative AI: A comparative performance analysis of autogeneration tools. In Proceedings of the 2024 International Workshop on Large Language Models for Code (LLM4Code ’24) (pp. 1–8). ACM.

Cheng, F. (2024). A comparative study of the performance of Spark-based k-means algorithm based on Euclidean distance and Manhattan distance. In Proceedings of the 3rd International Conference on Computer Communication and Artificial Intelligence (CCAI) (pp. 224–229).

Cheng, Y., Wang, M., Xiong, Y., Hao, D., & Zhang, L. (2016). Empirical evaluation of test coverage for functional programs. In 2016 IEEE International Conference on Software Testing, Verification and Validation (ICST) (pp. 255–265). IEEE.

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) (pp. 4171–4186).

Dhillon, I. S., Guan, Y., & Kogan, J. (2002). Refining clusters in high-dimensional text data. In Proceedings of the Workshop on Clustering High Dimensional Data and its Applications at the Second SIAM International Conference on Data Mining (pp. 71–82). SIAM.

Gao, X., & Li, G. (2016). A KNN model based on Manhattan distance to identify the SNARE proteins. Computational and Mathematical Methods in Medicine, 2016, Article 6480195. https://doi.org/10.1155/2016/6480195

Guan, W., & Fang, Y. (2025). Optimizing web-based AI query retrieval with GPT integration in LangChain: A CoT-enhanced prompt engineering approach. arXiv. https://arxiv.org/abs/2503.12345

Hoffmann, J., & Frister, D. (2024). Generating software tests for mobile applications using fine-tuned large language models. In Proceedings of the 45th International Conference on Software Engineering (ICSE) (pp. 1234–1245).

Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv. https://arxiv.org/abs/2106.09685

Hu, J., Liao, X., Gao, J., Qi, Z., Zheng, H., & Wang, C. (2024). Optimizing large language models with an enhanced LoRA fine-tuning algorithm for efficiency and robustness in NLP tasks. arXiv. https://arxiv.org/abs/2403.10567

Jacob, T. P., Bizotto, B. L. S., & Sathiyanarayanan, M. (2024). Constructing the ChatGPT for PDF files with LangChain – AI. In Proceedings of the 2024 International Conference on Inventive Computation Technologies (ICICT) (pp. 78–85).

Laplante, P. A., & Kassab, M. (2022). Requirements engineering for software and systems. Auerbach Publications.

Masuda, S., Nishi, Y., & Suzuki, K. (2020). Complex software testing analysis using international standards. In 2020 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW) (pp. 241–246). IEEE.

Nazi, A., Huang, Q., Shojaei, H., Esfeden, H. A., Mirhosseini, A., & Ho, R. (2022). Adaptive test generation for fast functional coverage closure. In DVCON USA.

Omar, S. F. (2013). A software traceability approach to support requirement-based test coverage analysis [Doctoral dissertation, Universiti Teknologi Malaysia].

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. arXiv. https://arxiv.org/abs/1905.00546

Schütze, H., Manning, C. D., & Raghavan, P. (2008). Introduction to information retrieval. Cambridge University Press.

Srikaewsiew, T., Khianchainat, K., Tharatipyakul, A., Pongnumkul, S., & Kanjanawattana, S. (2022). A comparison of the instructor-trainee dance dataset using cosine similarity, Euclidean distance, and angular difference. In 2022 26th International Computer Science and Engineering Conference (ICSEC) (pp. 235–240). IEEE.

Tiwari, D., Zhang, L., Monperrus, M., & Baudry, B. (2022). Production monitoring to improve test suites. IEEE Transactions on Reliability, 71(3), 1381–1397. https://doi.org/10.1109/TR.2022.3149480

Wang, Z., Guo, X., & Tsuchiya, T. (2025). Graph-centric approaches for coverage optimization in software requirement testing. In 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1270–1280). IEEE.

Wang, Z., Liu, K., Li, G., & Jin, Z. (2024). HITS: High-coverage LLM-based unit test generation via method slicing. In Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) (pp. 45–56).

Zhang, J., Wang, F., Ma, F., & Song, G. (2022). Text similarity calculation method based on optimized cosine distance. In 2022 International Conference on Electronics and Devices, Computational Science (ICEDCS) (pp. 37–39). IEEE.