From Recon to Report: A Multi-Agent System to Harden Software Systems
2025 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Background. As software complexity continues to grow, traditional strategies to find vulnerabilities in the system and assessment of these become more labour-intensive and usually occur late in the development stage. To circumvent such limitations, automated tools that are capable of scanning and securing systems continuously are required.
Objectives. This thesis aims to (1) design and implement a multi-agent system (MAS) that autonomously searches for vulnerabilities in an unknown target, tries to exploit these and for each successful exploit, writes a report about it. (2) Evaluate the system's performance in doing this. And (3), validate its practical utility through feedback from software security professionals at Ericsson.
Methods. The MAS was developed using the CrewAI framework and given the name MASploit. A Design Science approach guided iterative problem identification, solution design, and evaluation. MASploit's performance was empirically tested on 12 targets, comparing three LLM models (4o-mini, Mistral-24b and o1-mini). Expert feedback was collected via a survey, focusing on report quality and system usefulness.
Result. The 4o-mini model achieved a 51.7\% success rate in exploiting the containers, Mistral-24b (5\%) and o1-mini (6.7\%). Failures primarily stemmed from the reconnaissance stage as well as model-specific issues like prompt refusals or incorrect output formatting. 86\% of survey respondents expressed willingness to adopt the system, praising its structured reports and potential. However, improvements were suggested for CVSS accuracy, remediation specificity, and execution transparency.
Conclusions. This research demonstrates the viability of MAS for automating cybersecurity tasks, reducing reliance on manual work while democratizing access to robust security practices and enabling developers to continuously and earlier test the software for vulnerabilities. The system's modular design allows for future integration of more tools. Further work should refine agent collaboration and expand the system's toolbox.
Place, publisher, year, edition, pages
2025.
Keywords [en]
Multi-Agent System, Penetration Testing, Vulnerability Assessment, Cybersecurity, Large Language Models.
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:bth-27959OAI: oai:DiVA.org:bth-27959DiVA, id: diva2:1962620
External cooperation
Ericsson
Subject / course
Degree Project in Master of Science in Engineering 30,0 hp
Educational program
DVAMI Master of Science in Engineering: AI and Machine Learning 300 hp
Supervisors
Examiners
2025-06-122025-06-012025-09-30Bibliographically approved