Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Designing circuits that are algorithmically correct, low cost under a target fault-tolerant quantum computing (FTQC) architecture, and executable on near-term hardware remains a major bottleneck in the transition to scalable quantum computing. We introduce RubriQ, a scalable framework that formulates circuit synthesis as a large language model (LLM) code-generation task, optimized via group relative policy optimization (GRPO). Unlike conventional black-box neural critics, RubriQ employs a domain-grounded programmatic rubric as the reinforcement learning reward function, separating surface-code Clifford+T cost metrics from near-term hardware validation and unitary fidelity. To support high-throughput training, RubriQ integrates GPU-accelerated CUDA-Q simulation directly into the reinforcement learning (RL) loop and is deployed on NERSC Perlmutter using DeepSpeed ZeRO2 across multinode NVIDIA A100 clusters. On benchmark tasks, RubriQ attains a per-dataset correctness pass rate of at least 96% and a mean T-gate compression of 3.31× over correct-only circuits , significantly outperforming sparse-reward RL baselines (2.05×), converging 2–3× faster, and maintaining less than 1% near-term hardware-constraint violations. Validated on IBM and IonQ quantum processors, RubriQ establishes an automated, high-performance computing (HPC)-driven pipeline for generating surface-code-cost-aware circuits with separate near-term hardware validation.</p>

Show More

Keywords

rubriq nearterm circuits quantum computing

Related Articles

PORE

About

Connect