Open Access

Automated Testing Techniques for Enterprise Software Systems with GenAI Integration

4 Software Engineer, Cognizant, USA

Abstract

The rapid adoption of Generative Artificial Intelligence (GenAI) in enterprise workflows, from customer support to document processing, procurement and knowledge retrieval, has revealed a fundamental gap in conventional quality assurance practice. GenAI components are probabilistic, context-sensitive, dependent on retrieval corpora, tool integrations and evolving model versions, unlike deterministic software. Automated testing frameworks based on the assumption of stable input-output mappings are no longer sufficient to ensure the reliability, safety and compliance of these systems. In regulated and high-stakes enterprise settings, the lack of a scalable, CI/CD-friendly testing approach creates an unacceptable risk.

In this paper, we develop an automated testing framework that combines metamorphic testing and property-based testing to validate enterprise GenAI applications at scale. Metamorphic testing identifies inconsistencies and emergent faults between related input transformations without needing to define the expected outputs a priori, directly addressing the oracle problem faced by generative systems. Property-based testing is an alternative approach that creates a range of test cases from business rules, domain invariants and constraint specifications. Simultaneously, the framework evaluates hallucination rates, accuracy of retrieval grounding, correctness of tool calls, policy compliance, and prompt regression for RAG pipelines and agentic workflows.

To demonstrate how the proposed framework's metrics would be applied and evaluated in practice, this paper additionally presents an illustrative case scenario across four representative enterprise workflows customer support triage, procurement approvals, code review assistance, and knowledge-base Q&A conceptually modeled on published, evidence-driven quality-gate approaches for LLM applications. In this constructed 24-week quasi-experimental scenario, automated testing gates are illustrated as producing a

6.7 percentage point increase in task success rate (Cohen's d = 1.52), a 27.5% decrease in escalation rate, and a 31.0% decrease in rework rate. Business error cost per 1,000 workflow instances is illustrated as decreasing by 33.4%, and incorrect tool actions by 38.5%. User satisfaction is illustrated as increasing by 0.34 points on a 5-point scale. These illustrative results are shown to hold up under difference-in-differences and segmented-regression analyses, demonstrating how a framework's effectiveness could be assessed beyond simple pre/post comparison rather than reporting outcomes of a completed deployment. Taken together, the proposed framework and illustrative demonstration position metamorphic and property-based testing as scalable, CI/CD-compatible quality assurance techniques for enterprise GenAI systems, providing a practical path to deployment readiness, regulatory traceability, and sustained operational trust.

Keywords

References

Howard M, Lipner S: The Security Development Lifecycle. Microsoft Press, Redmond,; 2006.
OWASP FoundationOWASP: Top 10 for Large Language Model Applications. Open Worldwide Application Security Project. 2024,
National Institute of Standards and Technology: Artificial Intelligence Risk Management Framework (AI RMF 1.0. NIST, Gaithersburg, MD; 2023. 10.6028/NIST.AI.100-1
International Organization for Standardization: ISO/IEC 42001:2023 — Information Technology — Artificial Intelligence — Management System. ISO, Geneva 2023.
Bender EM, Gebru T, McMillan-Major A, Shmitchell S: On the dangers of stochastic parrots: can language models be too big?. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21).. 2021, 610-623. 10.1145/3442188.3445922
Zhang JM, Harman M, Ma L, Liu Y: Machine learning testing: survey, landscapes and horizons. IEEE Transactions on Software Engineering.. 2022, 48:1-36. 10.1109/TSE.2019.2962027
International Organization for Standardization: ISO/IEC 23894:2023 — Information Technology — Artificial Intelligence — Risk Management. ISO, Geneva. 2023,
European Union: Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act. Official Journal of the European Union 2024..
National Institute of Standards and Technology: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1. NIST, Gaithersburg, MD; 2024. 10.6028/NIST.AI.600-1
Ouyang L, Wu J, Jiang X, et al.: Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (NeurIPS 2022).. 2022, 35:27730-27744. 10.48550/arXiv.2203.02155
Shostack A: Threat Modeling: Designing for Security. Wiley, Indianapolis;
Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M: Not what you've signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec '23).. 2023, 79-90. 10.1145/3605764.3623985
Carlini N, Tramer F, Wallace E, et al.: Extracting training data from large language models. Proceedings of the 30th USENIX Security Symposium.. 2021, 2633-2650.
Raji ID, Bender EM, Paullada A, Denton E, Hanna: A: AI and the everything in the whole wide world benchmark. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.. 2021, 10.48550/arXiv.2111.15366
Huber T, Niklaus CLLMs meet: Bloom's taxonomy: a cognitive view on large language model evaluations. Proceedings of the 31st International Conference on Computational Linguistics (COLING 2025).. 2025, 5211-5246. 10.18653/v1/2025.coling-main.350
Basili VR, Caldiera G, Rombach HD: The Goal Question Metric approach. Encyclopedia of Software Engineering.. 1994, 528-532. 10.1002/0471028959.sof094
Garousi V, Felderer M, Kuhrmann M, Herkiloğlu K: Smells in software test code: a survey of knowledge in industry and academia. Journal of Systems and Software.. 2018, 138:1-10.1016/j.jss.2017.12.013
Shamshiri S, Rojas JM, Galeotti JP, Walkinshaw N, Fraser G: How do automatically generated unit tests influence software maintenance?. 2018 IEEE 11th International Conference on Software Testing, Verification and Validation (ICST).. 2018, 250-261. 10.1109/ICST.2018.00033
Luo Q, Hariri F, Eloussi L,: Marinov D: An empirical analysis of flaky tests. Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2014).. 2014, 643-653. 10.1145/2635868.2635920
Kochhar PS, Xia X, Lo D, Li S: Practitioners' expectations on automated fault localization. Proceedings of the 25th International Symposium on Software Testing and Analysis (ISSTA 2016).. 2016, 165-176. 10.1145/2931037.2931051
Liu J, Xia CS, Wang Y, Zhang: L: Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems (NeurIPS 2023).. 2023, 36:21558-21572.
Pearce H, Ahmad B, Tan B, Dolan-Gavitt B, Karri RAsleep at the keyboard?: Assessing the security of GitHub Copilot's code contributions. 2022 IEEE Symposium on Security and Privacy (SP).. 2022, 754-768. 10.1109/SP46214.2022.9833571
Fakhoury S, Naik A, Sakkas G, Chakraborty S, Lahiri SK: LLM-based test-driven interactive code generation: user study and empirical evaluation. IEEE Transactions on Software Engineering.. 2024, 50:2254-2268. 10.1109/TSE.2024.3428972
Just R, Jalali D,: Ernst MD: Defects4J: a database of existing faults to enable controlled testing studies for Java programs. Proceedings of the 2014 International Symposium on Software Testing and Analysis (ISSTA 2014).. 2014, 437-440. 10.1145/2610384.2628055
Christakis M, Bird: C: What developers want and need from program analysis: an empirical study. Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering (ASE 2016).. 2016, 332-343. 10.1145/2970276.2970347
Wing C, Simon K, Bello-Gomez RA: Designing difference in difference studies: best practices for public health policy research. Annual Review of Public Health.. 2018, 39:453-469. 10.1146/annurev-publhealth-040617-013507
Saltelli A, Bammer G, Bruno I, et al.: Five ways to ensure that models serve society: a manifesto. Nature.. 2020, 582:482-484. 10.1038/d41586-020-01812-9
Goodman-Bacon A: Difference-in-differences with variation in treatment timing. Journal of Econometrics. 2021, 225:254-277. 10.1016/j.jeconom.2021.03.014
Madaio MA, Stark L, Vaughan JW, Wallach H: Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. (CHI '20).. 2020, 1-14. 10.1145/3313831.3376445
Chang Y, Wang X, Wang J, et al.: A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology. 2024, 15:1-45. 10.1145/3641289
Mitchell M, Wu S, Zaldivar A, et al.: Model cards for model reporting. Proceedings of the 2019 Conference on Fairness, Accountability, and Transparency (FAT*'19)..2019,220-229. 10.1145/3287560.3287596
Weidinger L, Uesato J, Rauh M, et al.: Taxonomy of risks posed by language models. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT '22).. 2022, 214-229. 10.1145/3531146.3533088
Perez E, Huang S, Song F, et al.: Red teaming language models with language models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022). 2022, 3419-3448. 10.18653/v1/2022.emnlp-main.225
Maiorano AC: Automated self-testing as a quality gate: evidence-driven release management for LLM applications. arXiv.. 2026, 10.48550/arXiv.2603.15676

Most read articles by the same author(s)

<< < 1 2 3 4 5 6 7 8 > >> 

Similar Articles

1-10 of 35

You may also start an advanced similarity search for this article.