CoDe-R: Refining Decompiler Output with LLMs via Rationale Guidance and Adaptive Inference
AI 摘要
CoDe-R利用LLM改进反编译器输出,通过认知增强和动态回退机制提高代码可执行性。
主要贡献
- 提出了语义认知增强(SCE)策略,引导模型恢复高级算法意图。
- 引入了动态双路径回退(DDPF)机制,平衡语义恢复和语法稳定性。
- 在HumanEval-Decompile基准测试上取得了SOTA结果
方法论
CoDe-R是一个两阶段框架,首先通过SCE进行语义注入,然后通过DDPF进行自适应推理和验证。
原文摘要
Binary decompilation is a critical reverse engineering task aimed at reconstructing high-level source code from stripped executables. Although Large Language Models (LLMs) have recently shown promise, they often suffer from "logical hallucinations" and "semantic misalignment" due to the irreversible semantic loss during compilation, resulting in generated code that fails to re-execute. In this study, we propose Cognitive Decompiler Refinement with Robustness (CoDe-R), a lightweight two-stage code refinement framework. The first stage introduces Semantic Cognitive Enhancement (SCE), a Rationale-Guided Semantic Injection strategy that trains the model to recover high-level algorithmic intent alongside code. The second stage introduces a Dynamic Dual-Path Fallback (DDPF) mechanism during inference, which adaptively balances semantic recovery and syntactic stability via a hybrid verification strategy. Evaluation on the HumanEval-Decompile benchmark demonstrates that CoDe-R (using a 1.3B backbone) establishes a new State-of-the-Art (SOTA) in the lightweight regime. Notably, it is the first 1.3B model to exceed an Average Re-executability Rate of 50.00%, significantly outperforming the baseline and effectively bridging the gap between efficient models and expert-level performance. Our code is available at https://github.com/Theaoi/CoDe-R.