Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Identifying LLM‑Powered Cyber Attacks with Timing Analysis and Honeytoken‑Based Deception
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
2026 (English)Independent thesis Advanced level (professional degree), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

The rapid emergence of large language models capable of autonomous offensive action has introduced a qualitatively new category of attacker to internet-exposed systems. Existing SSH honeypot-based intrusion detection systems are designed around abinary classification problem distinguishing automated scripts from human operators and are fundamentally ill-equipped to handle the distinct behavioral and temporal footprint produced by an LLM-powered agent. As such tools become commercially accessible, this classification gap represents an unaddressed blind spot in current defensive practice. This thesis investigates whether LLM-powered attackers can be reliably distinguished from automated bots and human attackers in SSH honeypot environments, and proposes an integrated detection framework that combines three complementary signal modalities: network-level timing characteristics derived from packet captures, session- and command-level behavioral patterns extracted from honeypot interaction logs, and Unicode-based honeytoken deception that exploits the tokenization properties of large language models. A quantitative controlled experiment was conducted using two virtual machines deployed on IBM Cloud. A Cowrie SSH honeypot received attacks from fourteen commercial LLM models accessed through an autonomous offensive agent framework, as well as from automated scripted bots. A literature-derived behavioral profile served as the human attacker reference. Honeytoken artefacts embedding invisible Unicode characters were planted in the fake filesystem to assess whether LLM agents would reproduce these characters in their output a behavior neither human operators nor conventional scripts are likely to exhibit. Features were extracted from network packet captures and structured honeypot log files. The results demonstrate complete distributional separation between LLM-powered agents and automated bots across all evaluated models: inter-session timing gaps ranged from 3 to 15 seconds for LLM agents compared to 6 to 8 milliseconds for scripted attacks, corresponding to an approximately 375×–2500× difference. Honeytoken trigger rates varied across models, with URL-based tokens achieving the highest detection rates; the honeytokens where designed to detect LLMs not the automated bots. The findings establish that LLM-powered attackers produce a consistent and detectable operational footprint using only signals observable at the defended endpoint. Timing-based and deception-based signals are orthogonal and mutually reinforcing, providing the empirical foundation for a four-layer detection framework for three-class attacker classification in SSH honeypot environments.

Abstract [sv]

Den snabba framväxten av stora språkmodeller med förmåga till autonom offensiv verksamhet har introducerat en kvalitativt ny kategori av angripare mot internetexponerade system. Befintliga SSH-honeypot-baserade intrångsdetekteringssystem är utformade för ett binärt klassificeringsproblem att särskilja automatiserade skript från mänskliga operatörer och saknar förutsättningar att hantera det distinkta beteendemässiga och tidsmässiga avtryck som produceras av en LLM-driven agent. I takt med att sådana verktyg blir kommersiellt tillgängliga utgör detta klassificeringsgap en oadresserad blind fläck i nuvarande defensiv praxis. Denna uppsats undersöker om LLM-drivna angripare på ett tillförlitligt sätt kan särskiljas från automatiserade botar och mänskliga angripare i SSH-honeypot-miljöer, och föreslår ett integrerat detekteringsramverk som kombinerar tre komplementära signalmodaliteter: tidsmässiga egenskaper på nätverksnivå härledda från paketfångster, sessions- och kommandonnivå-beteendemönster extraherade från honeypot-interaktionsloggar samt Unicode-baserad honeytoken-vilseledning som utnyttjar tokeniseringsegenskaper hos stora språkmodeller. Ett kvantitativt kontrollerat experiment genomfördes med två virtuella maskiner driftsatta på IBM Cloud. En Cowrie SSH-honeypot tog emot angrepp från fjorton kommersiella LLM-modeller åtkomliga via ett autonomt offensivt agentramverk, samt från automatiserade skriptade botar. En litteraturbaserad beteendeprofil användes som referens för mänskliga angripare. Honeytoken-artefakter med inbäddade osynliga Unicode-tecken planterades i det falska filsystemet för att utvärdera om LLM-agenter skulle reproducera dessa tecken i sin utdata ett beteende som varken mänskliga operatörer eller konventionella skript sannolikt uppvisar. Egenskaper extraherades från nätverkspaketfångster och strukturerade honeypot-loggfiler. Resultaten visar fullständig distributionell separation mellan LLM-drivna agenter och automatiserade botar för samtliga utvärderade modeller: tidsgap mellan sessioner varierade från 3 till 15 sekunder för LLM-agenter jämfört med 6 till 8 millisekunder för skriptade angrepp en skillnad på tre storleksordningar. Honeytokens triggerfrekvenser varierade mellan modellerna, med URL-baserade tokens som uppnådde de högsta detekteringsfrekvenserna; ingen automatiserad bot triggererade någon honeytokentyp under några omständigheter. Resultaten fastslår att LLM-drivna angripare producerar ett konsekvent och detekterbart operativt avtryck med enbart signaler som är observerbara vid den försvarade slutpunkten. Tidbaserade och vilseledningsbaserade signaler är ortogonala och ömsesidigt förstärkande, och utgör den empiriska grunden för ett detekteringsramverk med fyra lager för treklass-angriparklassificering i SSH-honeypot-miljöer.

Place, publisher, year, edition, pages
2026. , p. 81
Keywords [en]
SSH honeypot, LLM-powered attacker, behavioral detection, honeytoken, attacker classification
Keywords [sv]
SSH-honeypot, LLM-driven angripare, beteendedetektering, honeytoken, angriparklassificering
National Category
Security, Privacy and Cryptography Artificial Intelligence Networked, Parallel and Distributed Computing
Identifiers
URN: urn:nbn:se:bth-29714OAI: oai:DiVA.org:bth-29714DiVA, id: diva2:2080153
External cooperation
AI Sweden
Subject / course
Degree Project in Master of Science in Engineering 30,0 hp
Educational program
DVADS Master of Science in Engineering: Computer Security
Supervisors
Examiners
Available from: 2026-06-26 Created: 2026-06-26 Last updated: 2026-06-26Bibliographically approved

Open Access in DiVA

fulltext(13995 kB)56 downloads
File information
File name FULLTEXT01.pdfFile size 13995 kBChecksum SHA-512
3eee49cfa4adf6ee0a085352e57865c392a7e21e6f555d75fd461adbb8302284eae4eccd647b3fa6e7ef19204254e9dbb0fb0d2424412cd5977aeede8d3e5bda
Type fulltextMimetype application/pdf

By organisation
Department of Computer Science
Security, Privacy and CryptographyArtificial IntelligenceNetworked, Parallel and Distributed Computing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 106 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf