StrategyWorking Paper

Acting on Two Intersecting Axes: Canada's AI Opportunity

Canada's AI opportunity hinges on two intersecting axes — workload types (training, inference, conversational, agentic) and the full stack from energy to applications. This briefing argues the right national capacity metric is GPU/accelerator slots and AI-ready megawatts, not data-centre floor space, and maps Canada's narrow, Québec-concentrated installed footprint against gigawatt-scale global benchmarks like OpenAI's Stargate.

Richard St-Pierre·October 2, 2025·9 min read
ai-infrastructuregpu-computesovereign-aicanadadata-centersacceleratorsai-policyliquid-cooling

Key finding: Modern AI is accelerator-bound. The relevant national capacity metric is GPU (and other accelerator) slots tied to high-density, liquid-cooling-capable data halls — not data-centre square footage or CPU colocation megawatts. Count accelerators, not buildings.

Canada's AI opportunity — and risk — now hinges on two intersecting axes:

  • Axis 1 — Workload types: frontier training, scaled inference, conversational/generative assistants, and increasingly agentic systems, alongside perception/analytics and edge deployments. These workloads have very different compute, latency, and safety profiles, and they pull distinctively on infrastructure.

  • Axis 2 — The full stack: energy → data centres/cooling → networking/interconnect → accelerators (GPUs/other) → system software → models/data → middleware/orchestration → applications. Value creation (and capture) now requires capacity and competence across this stack, not at one layer alone.

What changed since 2022: AI shifted from narrow pilots to an industrial-scale platform. Frontier training has become a semi-industrial activity; inference has become a global utility delivered from accelerator-dense campuses; and conversational AI has driven mass adoption. Hyperscalers and model labs are now building AI-optimized facilities, not generic cloud farms (e.g., OpenAI's Stargate program — 1 GW clusters, 200 MW first phase in 2026 — explicitly optimized for inference at scale). (OpenAI)

The policy hinge: Modern AI is accelerator-bound. The relevant national capacity metric is GPU (and other accelerator) slots tied to high-density, liquid-cooling-capable data halls — not data-centre square footage or CPU colocation megawatts. Today, direct liquid cooling is still used by a minority of operators (about one-fifth report some DLC use), while AI rack densities push past 25–100 kW/rack and rising. Canada's analysis and permitting should therefore pivot to AI-ready MW and GPU slots, not total IT MW. (Uptime Institute)

Canada's footing (installed, operator-anchored): GPU-ready capacity is concentrated in Québec, led by QScale Q01 (142 MW secured power; ~96 MW protected for IT), plus Vantage/Cologix sites; Ontario's and Alberta's large metros are still mostly CPU-weighted; British Columbia's Bell AI Fabric has commissioned non-GPU inference capacity (Groq LPU) and announced additional phases. This validates the core message: CPU-only colocation does not equal AI capacity. (qscale.com)

Canadian IP and platforms: Cohere remains a strategic asset and signal of ambition (federal $240M support; private buildout with partners), but one model company — without scale accelerator capacity and the surrounding stack — cannot deliver national AI readiness by itself. (Government of Canada)

Bottom line for the centre: To keep value from being diverted abroad, Canada must (i) measure what matters (GPU/accelerator slots; AI-ready MW), (ii) permit and power AI-first campuses, (iii) secure accelerator supply, (iv) fund sovereign models and middleware in bilingual/regulated domains, and (v) adopt AI at scale across public services to create domestic demand signals. The dossier below distills the threads into actionable key findings.

Key Findings at a Glance

The AI Economy Is Now Organized by Workload Type and by the Full Stack

  • Training concentrates capital and risk; it demands long-running, tightly coupled GPU superclusters with very high interconnect bandwidth and liquid cooling. Inference is permanent, elastic, latency-sensitive utility compute that drives regional replication and cost discipline. Conversational/generative workloads determine the mix and surge profile on the inference layer and force safety/guardrail middleware to mature. (Global infrastructure programs such as OpenAI's Stargate reflect this specialization — first 200 MW phase online in 2026.) (OpenAI)

  • Applications (Gov/health/finance/mining/retail) monetize the stack; middleware (RAG, vector DBs, safety/guardrails, agent orchestration) is now a control point for adoption and compliance; system software (CUDA/ROCm, compilers, schedulers) keeps developers tethered to accelerator ecosystems (notably NVIDIA). NVIDIA's CUDA-Q/QODA strategy aligns future quantum-classical acceleration with the same developer surface, reinforcing ecosystem lock-in. (NVIDIA Newsroom)

Implication: Policy that funds only "research" or only "data centres" misses the flywheel. Every layer — energy to apps — must be present domestically to capture value.

The Relevant Capacity Metric Is Accelerators, Not CPUs or Floor Space

  • AI servers are rapidly outgrowing legacy envelopes: common AI racks run at 25–100 kW today; state-of-the-art designs exceed that, pushing facility designs to direct-to-chip liquid cooling and high-pressure water loops. DLC adoption is rising but remains far from universal (~22% of operators report some DLC today), creating a near-term ceiling on how much of any generic colo can host GPUs without retrofit. (Uptime Institute)

  • Canonical training-class accelerators (e.g., NVIDIA H100 SXM) draw ~700 W each; AI clusters saturate facility-level power and cooling before floor area. Counting "buildings" or "MVA at the fence" is a poor proxy for AI readiness; GPU slots and AI-ready MW (liquid-cooling-capable, high-density halls, 400G/800G fabrics) are the correct measures. (NVIDIA)

Implication: Federal reporting and incentives should shift to GPU slot creation (with associated interconnect) and AI-ready MW, not aggregate DC MW.

Canada's Installed, GPU-Ready Footprint Is Real — but Narrow and Uneven

  • Québec leads: QScale Q01 (Lévis) is the flagship (142 MW secured; ~96 MW protected IT), purpose-built for HPC/AI with OCP-Ready design. Vantage QC2 (Québec City, 86 MW) and Montréal campuses (50 MW, 30 MW, 11 MW) add modern hyperscale capacity. Cologix MTL10 (35 MW) provides additional high-density potential. (qscale.com)

  • Ontario hosts large metros (e.g., Digital Realty TOR1, Equinix TR2, STACK TOR01A), but most installed power is still CPU-weighted without AI-first retrofits (liquid loops, 50–100 kW/rack rooms). The TOR1 engineering case cites 64 MW capacity; Equinix TR2 publishes 16 MW. (Stantec)

  • British Columbia is emerging for inference: Bell AI Fabric has two 7 MW sites (Kamloops, Merritt) online with Groq LPU accelerators; additional 26 MW + 26 MW phases are announced for 2026/27 (pipeline). GPU training anchor capacity remains limited. (Bell)

  • Alberta has modern footprint (e.g., eStruxture CAL-2 ~20 MW), and the AWS Canada West region, but public GPU-specific disclosures remain thin.

  • National counts remain misleading: Canada has ~239 operating data centres, but installed AI-ready capacity is a subset of total IT MW. (Research notes also cite full-build colo capacity >800 MW — again, not equivalent to GPU-ready megawatts.) (CER)

Implication: Today's GPU-ready capacity is concentrated; without accelerating AI-first buildouts in ON/AB/BC, model training and high-throughput inference workloads — and related spend — will continue to flow abroad.

Cohere Is a Critical — but Insufficient — Pillar

  • The federal commitment of $240M to Cohere (as part of the Sovereign AI Compute Strategy) is the right signal: anchor demand for domestic compute, attract private capital, and keep Canadian model IP onshore. But one model company cannot substitute for a national footprint of accelerators, liquid-ready capacity, and middleware integrators. (Government of Canada)

Implication: Pair sovereign model investments (Cohere, institutes, open bilingual models) with sovereign compute (GPU slots) and AI-first data-centre incentives, or the value chain will remain fractured.

The Global Benchmark Is Moving Very Fast

  • Frontier deployments now measure in gigawatts. OpenAI's Stargate is the clearest indicator: 1 GW clusters, 200 MW first phase online in 2026; Microsoft, AWS and others are fielding GB300/Blackwell-class clusters at supercomputer scale with liquid cooling and 800G fabrics. Canada's capacity, even with pipeline, is two orders of magnitude smaller. (Converge Digest)

Implication: Canada must pick where to lead (e.g., sovereign bilingual models; public-sector AI; climate/resource AI; quantum-classical pilots) and ensure hard capacity exists domestically for those bets.

What to Do Next (Policy-Shaped Actions)

  1. Measure and publish: Adopt an official GPU-slot and AI-ready MW inventory (by province and campus). Tie permits, grid allocations, and tax credits to the creation of AI-ready capacity (liquid cooling, high-density rooms, interconnect).

  2. Accelerate "AI-first" builds/retrofits: Fast-track liquid-ready retrofits in Toronto/Montreal/Calgary/Vancouver; prioritize 50–100 kW/rack rooms and DLC water loops in new permits.

  3. Secure accelerators: Pursue bulk-allocation agreements with suppliers; coordinate with allies for assured delivery to Canadian research/industry; support domestic chip-design efforts that align with CUDA/QODA-style ecosystems to preserve developer leverage. (NVIDIA Newsroom)

  4. Anchor demand: Scale public-sector adoption (challenge-based procurement; safe sandboxes) and sector pilots (health, resources, finance).

  5. Sovereign models & middleware: Fund bilingual, compliant models and guardrail/middleware stacks usable across departments and regulated industries — deployed on Canadian accelerators.

One-liner for decision-makers: Count accelerators, not buildings. If a megawatt cannot cool and interconnect GPUs, it is not AI capacity.

Sources (selected, load-bearing): QScale Q01 (142 MW secured; OCP-Ready; protected IT load) — QScale OCP page, OCP site assessment. Vantage Québec capacities: QC2 (86 MW), Montréal QC4 (50 MW), QC6 (30 MW), QC1 (11 MW) — renx.ca. Cologix MTL10: 35 MW critical power (Longueuil). Digital Realty TOR1 (Toronto): 64 MW engineering case; Equinix TR2: 16 MW spec — Stantec. Bell AI Fabric: Kamloops & Merritt 7 MW sites; additional phases; Groq LPU inference. National counts: ~239 operating data centres (CER Market Snapshot). Installed capacity anchors (Toronto share): CleanBridge GDC2025 (Toronto ≈ 335 MW ≈ 45% of installed). DLC adoption & densification: Uptime Institute Cooling Survey 2024 (~22% report some DLC use); AI racks 25–100 kW+ (Vertiv). CUDA-Q/QODA (quantum-classical under CUDA umbrella). OpenAI Stargate: 1 GW cluster; 200 MW first phase 2026. Cohere support: Government of Canada investment of up to $240M; CoreWeave partnership reports.

← Back to all essays

Stay informed

New essays on digital sovereignty, AI governance, and national strategy — delivered when published.