Benchmarking Approaches for QnA LLMs with SPARQL on EnterpriseData: Integrating RAG, ReAct, Multiagent for query generation with Knowledge Graphs
2025 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE credits
Student thesis
Abstract [en]
Background: Large Language Models (LLMs) offer a promising avenue for Natural Language Question Answering (NLQA) on KGs, enabling users to "chat with their data" by converting natural language to SPARQL queries. However, effectively applying LLMs in complex enterprise environments with proprietary KGs presents challenges, particularly regarding context management and the lack of comparative evaluations of different advanced methodologies.
Objectives: This thesis aims to systematically evaluate and compare Text-to-SPARQL generation methodologies Direct LLM Prompting, Retrieval-Augmented Generation (RAG), Reasoning and Action (ReAct) frameworks, Multi-Agent Architectures, and a pre-built Langchain integration applied to an enterprise KG. The primary objective is to assess their performance regarding query accuracy, response latency, token usage, and robustness, providing empirically-grounded insights into their practical suitability for enterprise data.
Methods: A quantitative, experimental methodology was employed using theGPT4.1-mini-2025-04-14 LLM. Five distinct Text-to-SPARQL approaches were implemented and benchmarked against a proprietary enterprise KG using 30 curated question-query pairs, with 7 runs per question. Performance was measured by Execution Accuracy (EA), Overall Execution Accuracy (OEA), Average OEA (AOEA), response latency (Overall Average Latency OAL), token usage (Overall Average Total Tokens - OATT), and query execution success rate (robustness).
Results: The RAG approach achieved the highest accuracy (0.424 AOEA) and lowest latency (3.1s OAL). The Multi-Agent system was the most token-efficient (1689 average total tokens) and demonstrated high robustness (95% success rate). Direct Prompting was the most token-expensive (22668 average total tokens) with moderate accuracy (0.343 AOEA). The ReAct agent, in its current configuration, generally underperformed across multiple metrics. No method excelled universally.
Conclusions: The study concludes that the optimal Text-to-SPARQL method for enterprise KGs depends on specific operational priorities. RAG offers a strong balance of accuracy and speed. Multi-Agent systems are promising for cost-efficiency and robustness. Pre-built integrations can offer high execution reliability but may sacrifice accuracy and cost-effectiveness on specific KGs. The findings provide empirical evidence on the trade-offs of these LLM-based techniques in a realistic, data-scarce enterprise context.
Place, publisher, year, edition, pages
2025. , p. 63
Keywords [en]
Chat with your data, Natural Language Querying, Text-to-SPARQL, Knowledge Graphs, GraphRAG, Large Language Models
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:bth-28489OAI: oai:DiVA.org:bth-28489DiVA, id: diva2:1988908
External cooperation
In-Parallel
Subject / course
DV1478 Bachelor Thesis in Computer Science
Supervisors
Examiners
2025-08-212025-08-132025-09-30Bibliographically approved