Back to Search View Original Cite This Article

Abstract

<p>Are random forests, the workhorse of supervised machine learning methods in the social sciences, still “good enough” versus new methods that tout dramatic performance benefits? In this article we present a large, diverse Monte Carlo study and existing real-world data benchmarks to compare tree-based methods with TabPFN, a new pretrained transformer foundation model that performs in-context learning over millions of synthetic datasets designed for tabular data. We compare tree-based methods with TabPFN in the popular R-learner framework for conditional average treatment effect estimation. This allows us to assess both predictive model performance and resulting gains for downstream inference. First, our simulations suggest that TabPFN does outperform random forest, achieving near-oracle results for conditional effect estimation. TabPFN’s improvements manifest in the most difficult simulation setups, where the data generating process is complex and there are fewer observations. Second, in real-world data analyses TabPFN performs well, outperforming random forest in some cases, especially as sample dwindles and the number of predictive covariates increases. Our results suggest that tree-based methods are still well suited for social science data, but TabPFN specifically, and the prior data fitted network approach generally, is a strong competitor worthy of consideration.</p>

Show More

Keywords

data methods tabpfn random treebased

Related Articles

PORE

About

Connect