Data Sources
This calculator uses a combination of official pricing, published benchmarks, and modeled reference assumptions. Each source is tagged with its provenance class throughout the product.
API Pricing
Anthropic
Model pricing from anthropic.com/pricing. Cached as of 2026-03-11. Provenance: Official
OpenAI
Model pricing from openai.com/api/pricing. Cached as of 2026-03-11. Provenance: Official
Throughput Benchmarks
InferenceX (SemiAnalysis)
Benchmark cards are normalized from the InferenceX public benchmark program and scaled to Vera Rubin VR200 NVL72-class hardware as a modeled projection (VR200 delivered-throughput data is not yet public). Run conditions: FP8/FP4 precision, varying input/output lengths and concurrency. Provenance: Benchmark + Modeled (VR200 scaling)
Infrastructure Assumptions
NVIDIA Vera Rubin VR200 NVL72 Architecture
Component composition (72 Rubin packages / 144 dies, Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs, NVLink 6 fabric) from public NVIDIA materials. Provenance: Official
VR200 rack cost and power
Rack hardware cost of ~$7.8M is a reported analyst estimate (Morgan Stanley, via Tom's Hardware, 2026); our $10.2M/rack buy-equivalent adds a modeled data-center share (~$12M per IT MW of shell, power, and cooling). Reported VR200 rack power spans ~120–230 kW depending on configuration; we model ~190 kW per compute rack. A 12 MW IT deployment is therefore 60 compute racks (~11.4 MW) plus a lower-draw head row of network, storage, and management racks inside the same gross envelope. Provenance: Modeled (from reported estimates)
Beyond-rack infrastructure multipliers
Networking (15%), storage/context-memory (12%), facilities/cooling (10%), software/ops (8%) are reference assumptions derived from typical enterprise DGX SuperPOD deployments. These are NOT official NVIDIA BOMs and require real vendor quotes for procurement decisions. Provenance: Modeled
Rack lease rates
Default $100,000/rack/month is a reference estimate consistent with the ~$10.2M/rack buy-equivalent on a multi-year master lease. Actual quotes vary significantly by term, geography, and commitment size. Provenance: Quote-driven
Power pricing
Default 7.5 cents/kWh derived from EIA industrial electricity averages. Provenance: Modeled
Reported Industry Signals (used in the landing story)
The landing film cites publicly reported events and figures. These are journalism and company disclosures — not Swarm measurements — and are labeled “reported” wherever they appear.
- Uber — 2026 AI budget exhausted by April; engineer caps. Uber's CTO told The Information the company burned its full-year 2026 AI budget in four months as Claude Code spread across ~5,000 engineers; it now caps agentic-coding spend at $1,500/month per engineer per tool. Reported by Fortune, Forbes, and AI Magazine (May 2026).
- Microsoft — internal Claude Code licenses cancelled on cost. Microsoft cancelled most internal Claude Code licenses across its Experiences & Devices division (deadline June 30, 2026) after per-engineer costs reached $500–$2,000/month, directing engineers to its in-house Copilot CLI. Reported by Fortune and TheNextWeb (May–June 2026).
- Coinbase — AI spend halved via cheapest-model defaults; usage still compounding. Coinbase cut AI spending ~50% by defaulting its ~1,200 agents and engineers to cheaper open-weight models, aggressive caching, and routing — while token usage kept rising. Reported by The New Stack and Fortune/Yahoo (June 2026).
- Platform token growth — Google. Google-disclosed monthly token processing: ~480T (I/O, May 2025), ~980T (July 2025), ~1.3Q (October 2025), ~3.2Q (I/O, May 2026) — roughly 7× in twelve months. See Google I/O 2026 keynote and coverage.
- Marketplace token growth — OpenRouter. OpenRouter grew from a ~100T tokens/year run rate to ~1.5Q tokens/year in roughly a year (~15×). Per Menlo Ventures (May 2026).
- Cloud exits — Dropbox, X. Dropbox disclosed $74.6M of infrastructure savings over two years in its S-1; X's engineering team reported cutting its cloud bill by roughly 60% after moving workloads on-prem.
- GPU-cloud pricing cross-check (margin context). The Swarm all-in reference rate (~$980/kW-month, compute included) is a quote-driven, margin-inclusive value. As a sanity check against published GPU-cloud rates: CoreWeave lists GB200 NVL72 capacity at roughly $10.50/GPU-hour on-demand (~$42/hr per 4-GPU node), with representative reserved discounts of ~35% (1-yr) to ~55% (3-yr): roughly $2,000+/kW-month equivalent even at 3-year reserved rates. Dedicated single-tenant buildings on long-term master leases are structurally cheaper than multi-tenant GPU-cloud retail; that spread is where both the customer's savings and Swarm's margin live. See CoreWeave pricing and GB200 provider comparisons (2026). Final pricing is always a Swarm proposal.
Data freshness: Seed data last updated 2026-03-11. Live connectors are optional and not required for the demo path. Contact Swarm Systems for current quotes and updated benchmark data.