Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>We propose the Mamba-based framework with Reusable Transformer (MRT) for scene text recognition (STR) Although Mamba provides linear-time global modeling, its sequential scanning and causal formulation are misaligned with the bidirectional and spatially localized nature of STR. To address these limitations, MRT introduces Mamba with embedded Neighborhood Attention (NA-Mamba) that replaces causal convolutions with standard convolutions and augments the State Space Model (SSM) pathway with a parallel Neighborhood Attention branch for local spatial aggregation. To inject bidirectional global context, we further incorporate Reusable Transformer that is applied at multiple stages with shared weights, avoiding excessive parameters and empirically accelerating convergence. Experiments on six common benchmarks and the Union-14M benchmark show that MRT achieves a strong accuracy–efficiency trade-off. MRT-B (Base) reaches 96.96% average accuracy, outperforming comparably sized or larger models, while the lightweight MRT-T (Tiny) attains 95.06% with only 4.72M parameters, avoiding over-parameterization and facilitating more efficient optimization.</p>

Show More

Keywords

reusable transformer mamba global causal

Related Articles

PORE

About

Connect