In brief
The challenge. Public-cloud platforms make experimentation easy, but production workloads can expose constraints in portability, data location and cost control. These constraints differ by workload, provider, contract and jurisdiction.
The response. A sovereign AI architecture gives an organisation explicit control over where data and workloads run, who can access them, and when an external provider may be used.
What this paper covers. This paper describes private multi-node execution, workload-led orchestration, policy-based routing, and a maturity model for adopting them.
The strategic drivers of digital sovereignty
Portability, data control and infrastructure economics shape the case for sovereign AI. Their importance depends on the workload, jurisdiction and supplier relationship.
Gartner forecasts worldwide generative AI spending of $644 billion in 2025, 76.4% above 2024, with roughly 80% allocated to hardware. It also forecasts sovereign-cloud IaaS spending of $80 billion in 2026, a year-on-year increase of 35.6%, and estimates that geopolitical pressures will move 20% of current workloads from global to local cloud providers. These forecasts indicate demand for greater infrastructure control; they do not determine the right architecture for an individual organisation.
Vendor lock-in and margin compression
Provider-specific APIs and managed services can make workloads harder or more costly to move. The effect depends on the services used, data volumes, egress terms and contractual commitments.
Regulatory and data custody demands
Regulated environments may require evidence of provenance, access, logging and data location for training data, fine-tuning artefacts and inference requests. The EU AI Act (Regulation (EU) 2024/1689) entered into force in August 2024, with obligations phasing in through December 2027. Its application depends on the system, role and risk category; it should not be reduced to a universal data-locality requirement.
Infrastructure economics
For stable, well-utilised workloads, private infrastructure may offer a lower total cost of ownership than on-demand cloud capacity. The comparison must include utilisation, financing, power, support, refresh cycles and residual value.
The Setloop GPU Maturity Framework
A phased maturity model allows organisations to test workload assumptions before committing capital. Each stage has an architectural deliverable and a decision gate.
| Stage | Focus area | Architectural deliverable | Governance gate |
|---|---|---|---|
| Stage 1: Assessment | Workload profiling | Memory footprint analysis, latency SLAs, and token profiling. | Independent evaluation of platform claims and procurement plans. |
| Stage 2: Foundation | Secure tenancy | Bare-metal or virtualised GPU fabric, high-throughput storage and base interconnect topology (NVLink, NVSwitch, InfiniBand, RoCEv2). | Multi-tenant isolation and data boundary verification. |
| Stage 3: Optimisation | Compute efficiency | KV-cache allocation, dynamic batching and kernel optimisation (vLLM, TensorRT-LLM). | Measured improvement against an agreed utilisation, latency and cost baseline. |
| Stage 4: Federation | Hybrid orchestration | Dynamic cross-plane routing across private clusters and hyperscalers based on latency and cost. | End-to-end token attribution and cost reconciliation. |
Reference architecture: the sovereign hybrid AI fabric
The reference architecture separates compute, workload orchestration, and policy and observability. Clear interfaces between these layers can reduce the cost of changing a provider or hardware platform, although they do not remove migration work.
3.1Compute & physical infrastructure layer
Multi-node private fabrics
Deployment of non-blocking InfiniBand or RoCEv2 networks linking private accelerator clusters for distributed training and low-latency inference. Within the node and across the rack, fifth-generation NVLink and NVSwitch supply up to 1.8 TB/s of bidirectional GPU-to-GPU bandwidth per GPU on NVIDIA Blackwell platforms, complementing the inter-node fabric.
High-throughput recoverable storage
Clustered file systems engineered for sustained checkpointing and high-concurrency parameter weights access.
3.2Workload orchestration & abstraction layer
Dynamic workload routing
Policy-driven gateways classify supported requests, route suitable batch work to private compute, and send excess demand to external providers when policy permits.
Workload-first decisions
Infrastructure parameters are selected from model topology, precision requirements, memory demand, latency objectives and interconnect constraints.
3.3Policy, observability & security layer
AI agent gateways
Inline controls, including LLMTrace patterns, can detect configured PII, known prompt attacks and anomalous payload structures before requests reach the execution cluster. They reduce exposure but cannot eliminate every failure mode.
Autonomous reliability
Infrastructure telemetry feeds governed operational workflows, such as AutoOps, for SLO monitoring, proposed remediation and an auditable record of actions.
Engineering implementation roadmap
The transition to a sovereign fabric is executed in three phases, sequenced so that governance and profiling evidence precede capital expenditure.
Discovery & Profiling
- Workload benchmarking
- GPU infrastructure strategy
- Build-vs-buy definition
Platform Construction
- Bare-metal orchestration
- Private storage fabric
- Secure tenancy configuration
Workload Migration & Hardening
- Distributed inference engine
- AI agent gateway integration
- Autonomous SRE
- Workload-driven sizing. Profile token generation, memory demand and latency constraints before selecting an accelerator class or committing capital.
- Platform integration. Construct the sovereign substrate using governed multi-tenancy and recoverable storage systems.
- Production hardening. Connect telemetry to incident diagnosis and record operational decisions and actions for later review.
How Setloop can help
Moving selected workloads from centralised cloud services to private or hybrid infrastructure requires systems engineering, performance measurement and operational governance. Setloop supports architecture, benchmarking and implementation within the customer's environment.
Bespoke architecture delivery
Setloop engineers design, benchmark and deploy private GPU infrastructure against agreed security, data-location, performance and operational requirements.
Proprietary accelerators
Engagements can use LLMTrace for agent-gateway controls, AutoOps for governed SRE automation, and GPU Cloud Platform designs as implementation components.
Structured advisory
Independent review of hardware plans, data-centre contracts and supplier architectures gives decision-makers evidence before capital is committed.
What infrastructure leaders should remember
- Sovereignty is an operating property, not a vendor label. Define control over data, execution, access and portability across compute, orchestration and policy layers.
- Profile before you procure. Workload-driven sizing (token curves, memory footprints, latency SLAs) must precede any accelerator capital expenditure.
- Gate each stage. Progress only when the agreed security, isolation, performance and cost criteria have been tested.
- Keep the hybrid option open. Policy-based routing can use external capacity when cost, security and data rules permit. Sovereignty does not require isolation.
References
- Gartner. “Gartner Forecasts Worldwide GenAI Spending to Reach $644 Billion in 2025.” Gartner newsroom, March 2025.
- Gartner. “Gartner Says Worldwide Sovereign Cloud IaaS Spending Will Total $80 Billion in 2026.” Gartner newsroom, February 2026.
- UNCTAD. “Data Protection and Privacy Legislation Worldwide.” Global Cyberlaw Tracker, 2024.
- European Union. Regulation (EU) 2024/1689 (EU AI Act). In force August 2024; phased application through December 2027.
- Flexera. “2025 State of the Cloud Report.” March 2025.
- NVIDIA. “NVLink and NVSwitch: fifth-generation specifications” and “GB200 NVL72.” NVIDIA product documentation.
About Setloop
Setloop is an engineering consultancy and product studio for organisations building GPU workloads, cloud GPU platforms, private AI infrastructure and AI factory architectures. Its engineers design, benchmark and deploy operational infrastructure in customers' private and hybrid-cloud environments across the UK and EU.
The Setloop Technical White Paper Series distils reference architectures from client work and the Setloop product portfolio, including LLMTrace, AutoOps, GPU Cloud Platform, AI FinOps and Automatic RL Research.
- Web: setloop.io
- Email: [email protected]
- LinkedIn: /company/setloop
- Book a GPU architecture review: setloop.io/contact