英文标题:DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis
作者:Abhinav Rajeev Kumar, Harshit Arora, Varun Singh, Manikandan Nanjappan
arXiv ID:2609.11504 | 分类:cs.LG | 发表:2026-09-10
许可:CC-BY
摘要 一个结构上有效的DeFi工作流仍可能授权一笔代价高昂的交易。{{NL}}我们提出DeFiFlowBench,一个包含207条由团队撰写的提示词的基准,用于自然语言DeFi工作流合成。它衡量图覆盖率、配置完整性和声明的安全谓词,{{NL}}然后在本地EVM上测试受支持的交易配置。直接、{{NL}}约束和少样本提示在固定的{{PT_MATH_3}}价格影响上限下,每个配置产生{{PT_MATH_1}}–{{PT_MATH_2}}次不安全的留出执行。{{NL}}由报价推导出的滑点界限并不能阻止订单本身的价格影响。我们提出Koan-Safe,它结合了仅提示的意图解析器、可替换的生成器,以及带有默
A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。