ElZemity, Adel, Li, Shujun, Arief, Budi (2026) Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis. In: 29th International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2026). (In press) (KAR id:116438)
|
PDF
Author's Accepted Manuscript
Language: English |
|
|
Download this file (PDF/829kB) |
Preview |
| Request a format suitable for use with assistive technology e.g. a screenreader | |
| Additional URLs: |
|
Abstract
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation, the opacity and escalating API costs of closed-weight frontier models motivate exploration of open-weight alternatives. However, many open-weight models are large, demanding significant compute resources and incurring non-trivial hosting costs that place them beyond reach for resource-constrained deployments. This paper investigates whether orchestrated ensembles of small language models (SLMs) can match or exceed single LLM performance on structured questions about malware detonation reports. We established baselines by testing eleven open-weight SLMs, three cyber security pre-trained models, and six frontier LLMs on Meta’s CyberSecEval Malware Analysis benchmark. We then designed and evaluated four orchestration architectures: (i) a multi-agent pipeline that decomposes analysis into structured evidence-collection and reasoning stages, (ii) an adversarial debate framework in which two agents iteratively critique each other’s reasoning, (iii) a hierarchical consultation system that pairs a general-purpose SLM with a cyber-specialised expert model, and (iv) a hybrid architecture that combines evidence-grounded pipelines with adversarial debate reasoning. The hybrid system (Qwen3-4B with Foundation-Sec-8B) achieved 35.30% overall accuracy, exceeding the strongest cyber-specialised baseline (22.54%) and the strongest ungrounded frontier baseline (34.77%); when given the same evidence pipeline, grounded Gemini remained the strongest configuration at 38.22%. Case studies on malware from the wild (UNC5142 and Lumma Stealer) illustrated the hybrid system’s ability to correct reasoning errors on novel evasion techniques such as EtherHiding and ClickFix. These findings show that evidence-grounded orchestration can substantially improve the performance of collaborative SLMs for supporting interpretation of malware detonation reports.
| Item Type: | Conference proceeding |
|---|---|
| Uncontrolled keywords: | Small language models · Malware analysis · Multi-agent systems · Orchestration · Large language models · Cyber security |
| Subjects: | Q Science > QA Mathematics (inc Computing science) |
| Institutional Unit: |
Schools > School of Computing Institutes > Institute of Cyber Security for Society |
| Former Institutional Unit: |
There are no former institutional units.
|
| Funders: | Engineering and Physical Sciences Research Council (https://ror.org/0439y7842) |
| Depositing User: | Budi Arief |
| Date Deposited: | 25 Sep 2026 15:26 UTC |
| Last Modified: | 25 Sep 2026 15:27 UTC |
| Resource URI: | https://kar.kent.ac.uk/id/eprint/116438 (The current URI for this page, for reference purposes) |
- Link to SensusAccess
- Export to:
- RefWorks
- EPrints3 XML
- BibTeX
- CSV
- Depositors only (login required):

https://orcid.org/0000-0002-5402-7837
Total Views
Total Views