Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<title>Abstract</title> <p>Purpose: Sparse alternatives to softmax can reduce attention support, but their value in compact classifiers may depend on the task and implementation. This study tests whether fixed and adaptive sparse normalizers improve predictive quality under a shared compact-classifier protocol. Methods: Dense softmax, sparsemax, entmax-1.5, three fixed-ratio top-k softmax settings, and head-wise adaptive entmax (HAE) were evaluated over ten matched seeds. Evidence covers CIFAR-10, Fashion-MNIST, 20 Newsgroups, a synthetic marker task, and a separate KMNIST follow-up. Paired tests, within-family Holm correction, attention-support summaries, and implementation-specific timing diagnostics were used. No validation split was used; top-k ratio selection and adjusted p values are therefore descriptive within exploratory comparison families rather than preregistered confirmatory evidence. Results: Fixed top-k improved over softmax on CIFAR-10 (Δ accuracy = 0.0263, Holm p = 0.0051), Fashion-MNIST (Δ = 0.0260, Holm p = 0.0055), and KMNIST (Δ = 0.0186, Holm p = 0.0381). The corresponding top-k support fractions were 0.294, 0.176, and 0.294; these fractions are imposed by the retained-ratio design rather than learned from the data. In the unmasked fixed-length 20 Newsgroups implementation, which included position-bearing padded slots and used an all-slot density denominator, all six sparse alternatives improved over softmax after Holm correction, while HAE and entmax-1.5 had the largest mean accuracy differences. HAE did not show evidence of improvement over fixed entmax-1.5 under this protocol; equivalence was not tested. The synthetic task did not separate methods, and timing results did not establish deployment efficiency. Conclusion: The evidence supports task- and implementation-specific normalizer assessment, not universal sparse-attention superiority or a general selection rule.</p>