Abstract
<title>Abstract</title> <p>Motivation: Causal discovery and gene regulatory network (GRN) inference methods re- cover edges well on small synthetic graphs, which has motivated applying them to Perturb-seq. Whether they recover direct regulatory edges from real interventional data is usually assessed by comparing >1 million candidate gene pairs against a few hundred curated reference edges. That comparison is dominated by prevalence dilution, so we restrict evaluation to a regulator-centric eligible universe: edges originating from perturbed genes that are annotated regulators, ranked within each regulator’s own candidate target list. Our question is not which method wins, but whether this benchmark can answer the question at all. Results: We benchmark seven distinct causal-discovery and GRN-inference methods (an eighth pipeline label, “GRNBoost2”, is excluded from the primary aggregate because no code in the project produces its stored output, so it cannot be described as a method) plus nine abla- tions on three Replogle et al. [2022] Perturb-seq datasets (K562 genome-scale, K562 essential, RPE1 essential; 20,000 cells × 1,024 genes), against three curated references (TRRUST v2, DoRothEA A+B, CollecTRI). The eligible universe is small and the evaluable subset far smaller: per dataset × reference cell, between 0 and 28 regulators have a non-self curated target in the measured gene set, and the whole benchmark rests on 39 distinct regulators (63 regula- tor × reference cells; 894 method rows are not 894 independent observations, MixedLM ICC = 0.255). Against 2,000 matched random rankings per regulator and a three-level hierarchical bootstrap that resamples regulator identities, the mean average precision above matched ran- dom, pooled over the seven distinct methods, is +0.0085 AP, 95 % CI [−0.0200, +0.0467]: the interval includes zero. A seeded random-control arm carried end-to-end lands at −0.0013 AP, [−0.0109, +0.0076], confirming the machinery is calibrated. Only 3 of the 7 methods have intervals excluding zero; all three are run on a single dataset, so their intervals price no between- dataset variance, and one of the three—the vanilla autoencoder—sits significantly below random (−0.0095, [−0.0184, −0.0039]). Both methods evaluated on all three datasets cover zero. No per-method mean survives Bonferroni correction. The published synthetic-to-real comparison is not measurable as stated, and once it is made measurable the transfer failure is roughly one order of magnitude, not total. Matching metric, universe and prevalence across five synthetic scales cuts a ∼59× raw discrepancy to 3.9× in macro-AP at 1,000 nodes; a genuine residual then remains and is large. We report it as an additive contrast in chance-corrected excess AP, +0.1118 [+0.0746, +0.1489] at 1,000 nodes, rather than as a ratio: the denominator of any such ratio is the real excess AP, whose own interval covers zero, so the ratio is not bounded away from arbitrarily large values. The additive gap shrinks monotonically with graph scale (+0.3129 at 50 nodes to +0.1118 at 1,000). The conventional 15-node benchmark cannot support the claim at all: its structural prevalence floor is 8.22 %, a random ranker scores 4.8× the best real method there, and NOTEARS—the method whose synthetic F1 anchors the transfer argument—scores below chance under the real metric. Methods do score curated direct targets above unrelated genes (score ratios 1.22×–21.64×; mean-rank ratios 1.06×–2.01×), but so do shuffled-label and no-label ablations, so part of that gradient is expression correlation rather than perturbation information. Conclusion: On the eligible universe reachable with these three curated references and a 1,024-gene panel, aggregate alignment between method rankings and curated targets is not dis- tinguishable from matched random in either direction. This is a finding about this benchmark’s resolving power, not a general claim about the methods or about Perturb-seq: with an effective n of 39 regulators, this benchmark can neither establish direct-edge recovery nor rule it out. The binding constraint we can demonstrate is the number of evaluable regulators, which is set by the overlap between the curated references and the measured gene panel; adding cells does not change it. Whether a larger evaluable set would resolve the question is untested here. Availability: Code, processed result tables, figure-value sidecars and plotting scripts are avail- able at https://github.com/ChimdiWalter/Perturb-seq.</p>