Abstract
<jats:p>Oxidative coupling of methane (OCM) performance depends on both operating conditions and catalyst formulation. This study evaluates whether explicit catalyst identity improves prediction of C2 yield in a small experimental dataset containing 291 OCM runs. A Random Forest regressor was trained using four operational variables (temperature, CH4/O2 ratio, total flow, and PAr), first without and then with one-hot-encoded catalyst descriptors (M1, M2, M3, and support). Using a fixed 80/20 split (random_state = 42), the operational-only model produced R2 = -0.161, RMSE = 3.77 percentage points, and MAE = 2.51 percentage points. Adding catalyst identity increased performance to R2 = 0.152, RMSE = 3.22, and MAE = 2.36. Five-fold random cross-validation gave mean R2 values of 0.256 ± 0.231 for the operational model and 0.398 ± 0.172 for the catalyst-aware model. Group-aware validation of the catalyst-aware workflow yielded mean R2 = 0.382 ± 0.151 when folds were grouped by primary catalyst component (M1) and 0.212 ± 0.262 when grouped by support. These results show that catalyst descriptors add reproducible predictive information, while the moderate error and fold-to-fold variability indicate that the model should be interpreted as a screening baseline rather than a validated optimization system. The analysis provides a transparent benchmark for subsequent work on physicochemical descriptors, uncertainty calibration, and prospective experimental validation.</jats:p>