英文标题:Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search
作者:Rui Liu, Tao Zhe, Yanyong Huang, Sankha Narayan Guria, Xiao Luo, Wei Fan, Yanjie Fu, Dongjie Wang
arXiv ID:2609.10225 | 分类:cs.LG | 发表:2026-09-09
许可:CC-BY
摘要。 特征变换通过从原始特征中构建信息丰富的抽象表示,提升表格数据的预测性能。{{NL}}近期的生成式方法将变换知识编码到连续嵌入空间中,以便高效探索候选策略,但面临三个关键局限:{{NL}}1)忽视了低层特征、操作与高层抽象之间的层次关系;{{NL}}2)对本质上具有置换不变性的变换序列施加了顺序敏感的嵌入,引入了系统性偏差;{{NL}}3)依赖基于梯度的搜索,难以适用于非凸的变换空间。{{NL}}我们提出了一个包含两个互补组件的框架。{{NL}}首先,一个置换不变的层次化模块捕捉特征、操作与抽象层级之间的交互,并采用自注意力池化机制,将语义等价的结构映射为与下游性能一致的嵌入。{{NL}
Feature transformation improves predictive performance on tabular data by constructing informative abstractions from raw features. Recent generative approaches encode transformation knowledge into continuous embedding spaces for efficient exploration of candidate strategies, but face three key limitations: (1) overlooking hierarchical relationships between low-level features, operations, and high-level abstractions; (2) enforcing order-sensitive embeddings on inherently permutation-invariant transformation sequences, thereby introducing systematic bias; and (3) relying on gradient-based search, which is ill-suited to non-convex transformation spaces. We propose a framework with two complementary components. First, a permutation-invariant hierarchical module captures interactions across features, operations, and abstraction levels, with a self-attention pooling mechanism that maps semantically equivalent structures to consistent embeddings aligned with downstream performance. Second, a policy-guided multi-objective reinforcement learning strategy initializes the search from empirically strong seeds and jointly optimizes predictive accuracy and transformation efficiency. Extensive experiments on diverse tabular benchmarks demonstrate the effectiveness and robustness of our framework against strong baselines. Our code and data are publicly available at: https://github.com/RayLiu1103/PHER.
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。