Provisioning for Intelligence: A Reference Architecture for Enterprise AI/ML Infrastructure at Scale
Keywords:
AI infrastructure, MLOps, GPU scheduling, machine learning platform, model serving, data engineering, cost governance, reference architectureAbstract
Enterprises adopting machine learning at scale frequently discover that their existing cloud platform, however mature, is the wrong shape for the workload. General-purpose cloud architecture is optimized around abundant, cheap, elastic commodity compute. AI workloads run on scarce, expensive accelerators, are dominated by the movement and lineage of large datasets, and split cleanly into a training world and a serving world with almost opposite operational characteristics. Treating AI infrastructure as ordinary cloud with graphics processors bolted on produces stranded capacity, runaway cost, and models that cannot be reproduced or safely retired. This article presents a five-layer reference architecture for enterprise AI/ML infrastructure (accelerated compute, data foundation, ML platform, serving, and a cross-cutting governance plane) and argues that each layer inverts an assumption that general-purpose cloud architecture takes for granted. I describe how accelerators should be pooled and scheduled rather than allocated per team, why data gravity constrains where the lower layers can live, how the model lifecycle is properly understood as a closed loop rather than a pipeline with an end, and why serving must be split into latency-bound and throughput-bound paths with opposing cost profiles. The reference architecture is provider-neutral and intended as a durable design vocabulary rather than a product recipe.
References
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, "Hidden Technical Debt in Machine Learning Systems," in Advances in Neural Information Processing Systems 28 (NeurIPS), 2015.
Google Cloud, "MLOps: Continuous Delivery and Automation Pipelines in Machine Learning," Google Cloud Architecture documentation, 2020.
M. Kleppmann, Designing Data-Intensive Applications. Sebastopol, CA: O'Reilly Media, 2017.
B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, "Borg, Omega, and Kubernetes," Communications of the ACM, vol. 59, no. 5, pp. 50-57, 2016.
Amazon Web Services, "AWS Well-Architected Framework: Machine Learning Lens," AWS Whitepaper, 2020.
B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Eds., Site Reliability Engineering: How Google Runs Production Systems. Sebastopol, CA: O'Reilly Media, 2016.
N. Forsgren, J. Humble, and G. Kim, Accelerate: The Science of Lean Software and DevOps. Portland, OR: IT Revolution Press, 2018.
Downloads
Issue
Section
License
Copyright (c) 2023 Siri Chandana Chava Journal (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




