Skip to main navigation Skip to search Skip to main content

Toward Efficient Deobfuscation via Large Language Models

  • Byunggeon Choi
  • , Hongjoo Jin
  • , Dong Hoon Lee
  • , Wonsuk Choi*
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

Obfuscation modifies code structure to impede reverse engineering and is widely used to protect intellectual property and evade malware detection. However, existing deobfuscation techniques largely rely on analyst-driven heuristics or are tailored to specific tools, suffering from low adaptability, scalability, and automation when handling complex or nested obfuscation methods. The core objective of this study is to develop an automated and scalable deobfuscation methodology that can generalize across diverse and complex nested obfuscation techniques by leveraging the capabilities of Large Language Models (LLMs). To achieve this, we present LLM4DOBF, a novel methodology that combines LLM fine-tuning to learn various obfuscation patterns with optimized prompt engineering. This approach guides the LLM to effectively analyze obfuscated code and restore it to its original form. Experimental results demonstrate that LLM4DOBF achieves superior code deobfuscation quality, as validated by high SacreBLEU scores, and shows notable code optimization improvements in the ‘Obfuscation Quality Quantification Framework’ evaluation. Furthermore, it consistently outperformed other LLMs in complex scenarios using major obfuscation tools like Tigress and OLLVM, proving its stability with a high response rate and a low compilation error rate. This study confirms that the proposed LLM4DOBF overcomes the limitations of existing heuristic-based methods, providing a robust, reliable, and generalizable solution for automated deobfuscation. This constitutes a significant contribution by presenting a scalable approach for modern obfuscated code analysis.

Original languageEnglish
Pages (from-to)30844-30859
Number of pages16
JournalIEEE Access
Volume14
DOIs
Publication statusPublished - 2026

Bibliographical note

Publisher Copyright:
© 2013 IEEE.

Keywords

  • Source code obfuscation
  • large language models
  • malware
  • prompt engineering
  • source code deobfuscation

ASJC Scopus subject areas

  • General Computer Science
  • General Materials Science
  • General Engineering

Fingerprint

Dive into the research topics of 'Toward Efficient Deobfuscation via Large Language Models'. Together they form a unique fingerprint.

Cite this