LLM Reasoning 相关度: 8/10

CoDe-R: Refining Decompiler Output with LLMs via Rationale Guidance and Adaptive Inference

Qiang Zhang, Zhongnian Li
arXiv: 2604.12913v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

CoDe-R利用LLM改进反编译器输出,通过认知增强和动态回退机制提高代码可执行性。

主要贡献

  • 提出了语义认知增强(SCE)策略,引导模型恢复高级算法意图。
  • 引入了动态双路径回退(DDPF)机制,平衡语义恢复和语法稳定性。
  • 在HumanEval-Decompile基准测试上取得了SOTA结果

方法论

CoDe-R是一个两阶段框架,首先通过SCE进行语义注入,然后通过DDPF进行自适应推理和验证。

原文摘要

Binary decompilation is a critical reverse engineering task aimed at reconstructing high-level source code from stripped executables. Although Large Language Models (LLMs) have recently shown promise, they often suffer from "logical hallucinations" and "semantic misalignment" due to the irreversible semantic loss during compilation, resulting in generated code that fails to re-execute. In this study, we propose Cognitive Decompiler Refinement with Robustness (CoDe-R), a lightweight two-stage code refinement framework. The first stage introduces Semantic Cognitive Enhancement (SCE), a Rationale-Guided Semantic Injection strategy that trains the model to recover high-level algorithmic intent alongside code. The second stage introduces a Dynamic Dual-Path Fallback (DDPF) mechanism during inference, which adaptively balances semantic recovery and syntactic stability via a hybrid verification strategy. Evaluation on the HumanEval-Decompile benchmark demonstrates that CoDe-R (using a 1.3B backbone) establishes a new State-of-the-Art (SOTA) in the lightweight regime. Notably, it is the first 1.3B model to exceed an Average Re-executability Rate of 50.00%, significantly outperforming the baseline and effectively bridging the gap between efficient models and expert-level performance. Our code is available at https://github.com/Theaoi/CoDe-R.

标签

反编译 大语言模型 代码生成 逆向工程

arXiv 分类

cs.SE cs.AI cs.CR