AI Agents 相关度: 9/10

From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python

Jinhua Wang, Biswa Sengupta
arXiv: 2604.11518v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出一种基于基准的LLM辅助代码跨语言迁移方法,实现Rust到Python的迁移和功能扩展。

主要贡献

  • LLM辅助的跨语言代码迁移方法
  • 基于基准的迭代优化和调试
  • 实现功能扩展的Python端口
  • 验证Python性能接近Rust,代码量显著减少

方法论

使用LLM将Rust代码翻译为Python,通过基准测试驱动迭代改进,并增加新功能。

原文摘要

Cross-language migration of large software systems is a persistent engineering challenge, particularly when the source codebase evolves rapidly. We present a methodology for LLM-assisted continuous code translation in which a large language model translates a production Rust codebase (648K LOC, 65 crates) into Python (41K LOC, 28 modules), with public agent benchmarks as the objective function driving iterative refinement. Our subject system is Codex CLI, a production AI coding agent. We demonstrate that: (1) the Python port resolves 59/80 SWE-bench Verified tasks (73.8%) versus Rust's 56/80 (70.0%), and achieves 42.5% on Terminal-Bench versus Rust's 47.5%, confirming near-parity on real-world agentic tasks; (2) benchmark-driven debugging, revealing API protocol mismatches, environment pollution, a silent WebSocket failure mode, and an API 400 crash, is more effective than static testing alone; (3) the architecture supports continuous upstream synchronisation via an LLM-assisted diff-translate-test loop; and (4) the Python port has evolved into a capability superset with 30 feature-flagged extensions (multi-agent orchestration, semantic memory, guardian safety, cost tracking) absent from Rust, while preserving strict parity mode for comparison. Our evaluation shows that for LLM-based agents where API latency dominates, Python's expressiveness yields a 15.9x code reduction with negligible performance cost, while the benchmark-as-objective-function methodology provides a principled framework for growing a cross-language port from parity into an extended platform.

标签

LLM 代码迁移 跨语言 AI Agent 基准测试

arXiv 分类

cs.SE cs.AI