Abstract
Obfuscation modifies code structure to impede reverse engineering and is widely used to protect intellectual property and evade malware detection. However, existing deobfuscation techniques largely rely on analyst-driven heuristics or are tailored to specific tools, suffering from low adaptability, scalability, and automation when handling complex or nested obfuscation methods. The core objective of this study is to develop an automated and scalable deobfuscation methodology that can generalize across diverse and complex nested obfuscation techniques by leveraging the capabilities of Large Language Models (LLMs). To achieve this, we present LLM4DOBF, a novel methodology that combines LLM fine-tuning to learn various obfuscation patterns with optimized prompt engineering. This approach guides the LLM to effectively analyze obfuscated code and restore it to its original form. Experimental results demonstrate that LLM4DOBF achieves superior code deobfuscation quality, as validated by high SacreBLEU scores, and shows notable code optimization improvements in the ‘Obfuscation Quality Quantification Framework’ evaluation. Furthermore, it consistently outperformed other LLMs in complex scenarios using major obfuscation tools like Tigress and OLLVM, proving its stability with a high response rate and a low compilation error rate. This study confirms that the proposed LLM4DOBF overcomes the limitations of existing heuristic-based methods, providing a robust, reliable, and generalizable solution for automated deobfuscation. This constitutes a significant contribution by presenting a scalable approach for modern obfuscated code analysis.
| Original language | English |
|---|---|
| Pages (from-to) | 30844-30859 |
| Number of pages | 16 |
| Journal | IEEE Access |
| Volume | 14 |
| DOIs | |
| Publication status | Published - 2026 |
Bibliographical note
Publisher Copyright:© 2013 IEEE.
Keywords
- Source code obfuscation
- large language models
- malware
- prompt engineering
- source code deobfuscation
ASJC Scopus subject areas
- General Computer Science
- General Materials Science
- General Engineering
Fingerprint
Dive into the research topics of 'Toward Efficient Deobfuscation via Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS