Large Language Models for Requirement Compliance Checking: A framework for road construction
2025 (English)Independent thesis Advanced level (professional degree), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Background: Requirement Engineering process is time consuming and very me-ticulous. Engineers may spend days trying to make sure requirements are struc-tured and worded correctly to avoid misinterpretation. To streamline the processthe Swedish Transport Administration in collaboration with IBM have developed aproof-of-concept system (Dokument-AI) for requirement compliance checking whichleverages an LLM to modify the input requirement to align with a pre-defined set oflinguistical and structural rules. However, being a proof of concept system which alsoleverages an LLM, it has to undergo an evaluation in order to assess its performance.Objectives: The objectives of this thesis are to create a framework for evaluatinghow Dokument-AI performs its given task, specifically we aim to measure precision.Method: The method consists of using LLM-as-a-Judge to judge whether the newproposed requirement follows the grammatical and linguistical requirements set byDokument-AI, in addition to that we use BLEU and chrF to measure how similarthe new requirement is to the original requirement.Results: The results show that Dokument-AI managed to generate requirementsthat passed in 75.5% of the total cases and from the manual review we can see thatLLM-as-a-Judge should be accurate in 95.2% of the cases.Conclusions: In conclusion, while Dokument-AI shows considerable promise for fastevaluation of potential new requirements, with re-formulations passing in 75.5% ofcases, its accuracy is not sufficient for critical applications. Although it can providevaluable reasoning and insights, even when its judgments are disputed, the risk ofinaccuracies is too high for the requirements engineering process, where mistakes canhave long-term negative impacts on a project. However, if used in conjunction withhuman supervision, the system’s performance can be drastically improved, makingthe data far more reliable and usable for practical applications.
Abstract [sv]
Bakgrund: Kravhanteringsprocessen är en mycket krävande process där ingenjörerkan lägga flera dagar på att säkerställa att krav är korrekt strukturerade och for-mulerade för att undvika feltolkningar. För att effektivisera processen har Trafik-verket, i samarbete med IBM, utvecklat ett proof-of-concept-system (Dokument-AI)för validering. Systemet använder LLM för att omformulera krav i enlighet med ettfördefinierat regelverk för språk och struktur. Eftersom det är ett proof-of-concept-system som bygger på en LLM krävs en utvärdering för att bedöma dess prestandaoch pålitlighet.Syfte: Syftet med detta examensarbete är att skapa ett ramverk för att utvärderaDokument-AI utifrån hur väl systemet utför sin uppgift med avseende på precision.Metod: Metoden består i att använda LLM-as-a-Judge för att bedöma huruvidadet omformulerade kravet följer de grammatiska och språkliga regler som fastställtsav Dokument-AI. Därutöver används BLEU och chrF för att mäta hur likt den nyaformuleringen är den ursprungliga.Resultat: Resultaten visar att Dokument-AI lyckades generera krav som godkändesi 75.5% av fallen. Den manuella granskningen indikerar att LLM-as-a-Judge är kor-rekt i 95.2% av fallen.Slutsats: Sammanfattningsvis visar Dokument-AI stor potential för snabb utvär-dering av nya krav, där omformuleringarna godkändes i 75.5% av fallen. Dock ärträffsäkerheten inte tillräcklig för kritiska tillämpningar. Även om systemet kan till-handahålla värdefulla resonemang och insikter, även vid ifrågasatta bedömningar, ärrisken för felaktigheter för hög inom kravhantering, där misstag kan få långsiktigtnegativa konsekvenser för ett projekt. Om systemet däremot används tillsammansmed mänsklig granskning kan dess prestanda förbättras avsevärt, vilket gör res-ultaten betydligt mer tillförlitliga och användbara i praktiken.
Place, publisher, year, edition, pages
2025. , p. 33
Keywords [en]
Large Language Model, Evaluation, Requirement Engineering, Automated Compliance Checking, Requirement Verification
Keywords [sv]
Stor språkmodell, Utvärdering, Kravhantering, Validering
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:bth-27911OAI: oai:DiVA.org:bth-27911DiVA, id: diva2:1966475
External cooperation
Trafikverket
Subject / course
Degree Project in Master of Science in Engineering 30,0 hp
Educational program
PAAMJ Master of Science in Engineering: Software Engineering 300,0 hp
Supervisors
Examiners
2025-08-222025-06-102025-09-30Bibliographically approved