Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>In the research and development of large models at present, Scaling Laws have become a widely adopted empirical guiding framework across the industry. Summarized based on massive standardized training experiments, they clearly depict the power-law correlation between model performance and model parameters, training data volume, as well as hardware computing capacity, laying a crucial practical foundation for the large-scale research and development of general large models over the past five years. As industrial implementation advances, the industry has long taken the expansion of computing hardware as its core development path. Nevertheless, the model of continuously stacking hardware and enlarging model scales has gradually brought about the phenomenon of diminishing marginal returns. Based on observations from frontline engineering practices, this paper proposes the CAI law of tripartite balance between computing power, algorithms and information, with the quantitative relational expression: 𝑪 = 𝑨/𝑰(𝑰 ≠ 𝟎). In this formula: denotes the effective computing power output of the system, representing the computational capacity convertible into real business value; stands for the driving potential of a complete algorithm system, covering full-chain algorithm capabilities including underlying model architectures, domain-specific Tokenizers, and fine-tuning optimization schemes; refers to the system operation loss brought by generalized information resources, consisting of various computing-consuming elements such as redundant data, noise samples, conflicting texts, invalid instructions and scheduling overhead. As a complementary perspective to existing mature theories, this paper completes qualitative derivation of the formula and multi-gradient dimensionless numerical simulation, and conducts comparative analysis on relevant scenarios to systematically elaborate the mutually restrictive and collaboratively promotable internal logic of computing power, algorithms and information. Simulation results intuitively demonstrate that optimizing algorithms and purifying information resources to reduce information loss can significantly boost the effective computing power of AI systems on the premise of fixed hardware infrastructure. This paper provides reference ideas for resource allocation to R&amp;D enterprises and industrial application entities of large models, reminding industry practitioners not to solely focus on the expansion of computing hardware. Instead, they should dynamically balance the proportions of computing power, algorithms and data information according to business development stages, so as to achieve wider application coverage, more accurate reasoning and more economical resource consumption for large models. It also serves as an auxiliary analytical tool with practical inspiration for medium- and long-term intelligent resource planning.</p>

Show More

Keywords

computing information model hardware power

Related Articles

PORE

About

Connect