Abstract
<title>Abstract</title> <p>Distinguishing genuine cancer drivers from the much larger set of passenger mutations remains one of the persistent bottlenecks in translating tumor sequencing data into biological insight. We present CCNIF (Cancer Causal Network Inference Framework), a reproducible computational pipeline that prioritizes candidate driver genes from a cohort's mutation and expression data, then builds a multi-domain evidence profile differential expression, functional enrichment, and, for a fully characterized case study, clinical correlation, protein interaction network position, and survival association around each candidate. Applied to 505 TCGA lung adenocarcinoma (LUAD) patients, CCNIF ranked 12,947 genes by combined mutation-frequency and expression evidence and retained the top 50 as a candidate driver panel. Framework-derived confidence scores stratified this panel into four tiers (five Tier A, nine Tier B, five Tier C, thirty-one Tier D genes; mean confidence 66.17, mean quality 67.51), with SFTPB, ZFHX4, EGFR, MUC16, and IGHA1 receiving the highest confidence. Benchmarking the panel against four independent, non-overlapping external resources IntOGen, OncoKB, CancerMine, and the Network of Cancer Genes (NCG) showed substantial, resource-dependent overlap (76.0%, 26.0%, 24.0%, and 70.0% coverage respectively), including correct recovery of canonical LUAD drivers such as TP53, EGFR, KRAS, and KEAP1. TP53, the cohort's most frequently mutated gene (48.5%), was carried through the complete evidence pipeline as a case study: differential expression against 23,814 background genes, gene set enrichment dominated by E2F target and G2M checkpoint programs, a 431-node, 1,788-edge STRING interaction network resolving into 108 communities, and survival analysis that found no significant log-rank association with mutation status in this cohort (p = 0.228) a result we report and discuss rather than treat as a negative outcome to be omitted. We report CCNIF's design, its validation against external driver databases, and its current scope honestly: the 49 non-TP53 drivers carry framework-level confidence and quality scores and database-verified external support, but have not yet been carried through the same clinical, network, and survival sub-analyses as TP53, and we describe them explicitly as computationally prioritized candidates for downstream experimental follow-up rather than fully characterized drivers. CCNIF's source code, configuration, and full per-driver evidence trail are structured for direct reuse in other cohorts and tumor types.</p>