Your AI budget just doubled, again.
Swarm Systems
Take control of your AI spend.
Own your infrastructure.
Leave the cloud.
The two-minute version of the decision every AI-heavy enterprise is hitting this year: why the API meter can't be your infrastructure, and the program that replaces it.
Act I · The curve everyone is on
The token era has a shape.
Reported, not modeled: one platform's disclosed token consumption. Watch the dates. Each point is only months apart.
It's an exponential. Project it forward at HALF that pace, and last year becomes a footnote on the chart.
Now price it. The amber line is that same curve as a monthly bill for an enterprise spending $13.9M/mo today, with prices modeled falling a third every year.
Strip the tokens away and look at spend alone. We model the price per token falling 33% a year, 70% cheaper by month 36, and the bill still compounds.
≈5× the plan in three years: $71.8M a month against today's $13.9M. Cheaper tokens aren't a plan; you'll spend them. What follows is the plan.
Act I · The math
Meet the deployment the board approved.
Headcount
200,000 employees
company-wide AI rollout
Model
Claude Opus 4.8
$5 in / $25 out per MTok, list
The budget
$0/mo
$100 per employee per month
Month-one usage
0 tokens
300.0K per employee per workday, modeled agent-era
Month-one bill
$0/mo
69% of budget. Everyone's happy.
Month one is comfortable. This story is about every month after that.
Act I · The curve
Success compounds. Budgets don't.
Zoomed in, month by month, it looks manageable. Tokens per employee climb ~5% a month, a fraction of the measured industry pace, and still compounding.
Month 9: you slip past the budget you promised the board. Quietly, barely over the line.
Zoom out. $102.5M a month, ≈5× the plan, the same multiple the industry curve just showed you. On the meter your options are throttle your best people, or explain this line every quarter. Forever.
Act I · This year, everywhere
2026 is the year of AI rationing.
The best tools worked. The meters exploded. And the biggest names in tech responded the only way a meter lets you:
Uber
burned its entire 2026 AI budget by April, four months in, as Claude Code spread across ~5,000 engineers. Now it caps engineers at $1,500/month per tool.
Fortune · Forbes · The Information, May 2026
Microsoft
pulled most internal Claude Code licenses across the Windows / M365 org after per-engineer costs hit $500–$2,000 a month, and moved engineers to its cheaper in-house tool.
Fortune · TNW, May–June 2026
Coinbase
cut AI spend ~50% by defaulting 1,200 agents and its engineers to the cheapest open-weight models. Token usage kept compounding anyway.
The New Stack · Fortune, June 2026
Caps, revoked licenses, defaults to the cheapest model: rationing the one tool that's actually working. The workload isn't shrinking. The meter is just unaffordable. If you want your people on the best models, you have to own the economics underneath them.
Act I · The answer
Run the models on machines you control.
Same work, same tokens, priced as infrastructure instead of retail. The meter stops being the business model.
Per token
68% less
at delivered capacity, full Rubin stack
Same budget
3.1× the tokens
or the headroom to stop rationing
Your first deployment
$0/mo
60 racks · 12 MW IT, less than this month's API bill
The engine's verdict on this workload: the economic threshold is already crossed. Dedicated capacity wins today. This isn't a someday decision.
Dropbox saved $74.6M in two years after moving off the public cloud
S-1 filingX cut its cloud bill by roughly 60% after moving on-prem
as reportedAct I · The clock
Compute arrives 12 months after you sign, not after you decide.
That's the lease-to-live lead time. Wait six months to sign and months 12–18 from today still run on the meter, at those months' grown prices. The math on waiting:
The threshold is already crossed, and every month of deliberation has a price tag. Deciding now is the cheapest thing in this story.
Act II · The catch
One detail: those racks need a home. The self-build list:
Or partner with Swarm, and check every box.
Every box, one address: your building.
Find land
Win entitlements
Join the interconnection queue
Post utility securities
Contract grid power
Order long-lead equipment
Track chip roadmaps
Design years ahead
Raise the capital
Manage the build
Commission it
Operate it 24/7
Every line is measured in months, some in years. ~36 months to first token if everything goes right, while the meter collects $1.3B. Keep scrolling.
Sites secured. Queues already joined. Power contracted. Operations staffed. One master lease covers all of it, and your compute lands in 12 months, not 36.
A dedicated, single-tenant Swarm building: 16–20+ MW IT, leased as 12 MW deployments. Land, power, cooling, fabric, and operations, behind one door with your name on it. Step inside.
Act II · The factory
Step inside. Meet your private AI factory.
Deployment 1 of your building: sixty Vera Rubin VR200 racks at ~190 kW each, plus the lower-draw head row the reference architecture actually requires: scale-out fabric, context-memory and shared storage, management. You don't build it, finance it, or staff it. You rent it, all-in.
Your private AI factory · Deployment 1
≈12 MW IT gross · single-tenant
Head-row racks draw far less power than compute racks. Everything fits inside the ≈12 MW gross envelope.
60× Vera Rubin VR200 racks
~190 kW per rack. Your racks, your models, your data
3.3T tokens/month delivered
benchmark-based, at modeled utilization
Networking · storage · cooling
the full pod stack, not just GPUs
Staffed 24/7 operations
Swarm runs it; you consume it
A dedicated single-tenant building
16–20+ MW IT buildings. No cage, no neighbors
Leased as 12 MW IT deployments
your next deployments land in the same building program
The rent
$11.8M/MO ALL-IN
≈ $981/kW-month, margin-inclusive: compute, facility, power, and ops
If you bought it instead
$612.0M
capex, on your balance sheet
Act II · Scale in place
Then it multiplies.
1 DEPLOYMENT · 12 MW IT
Deployment 1 · 12 MW
Deployment 2 · 12 MW · next
Deployment 3 · 12 MW · next
Deployment 1 lands with headroom. At this workload's pace, headroom has a shelf life.
A second deployment lands ~6 months after you order it. Same building, same lease, no new project.
A third: 36 MW IT, the scale this workload actually reaches ~29 months after first power (engine-computed). Your own building, filled on schedule.
Act II · The program
The first factory is the first step.
Programmatic scaling: 12 MW deployments across Swarm's secured buildings, each landing in months, ordered before you need them. On this workload the program reaches 36 MW on deployment 3 and keeps stepping toward ~72 MW inside four years. A self-build gets one step, three years from now.
The path to 36 MW
3 deployments
12 MW each · ~29 months after first power
Every step
68% less per token
at delivered capacity, with headroom to stop rationing
Or blend
Hybrid
keep an API burst lane for spikes and frontier work
Cards on the table: growth headroom costs money before you fill it. Even carrying it, this stepped ramp runs 47% under the meter across four years of 5%-a-month growth. No new land hunt. No new interconnection queue. And no retail margin on the fastest-growing line of your P&L.
Rent the intelligence.
Own the economics.
5-year savings
$0
equivalent service, engine-computed
Threshold
Crossed
dedicated capacity already wins
Cost of waiting a year
$0
savings forgone by deferring
Capacity online
12 mo
from reservation to first token
The story is a calculator
Every number you just saw is an input you can change.
The worked example (200,000 employees, 300.0K tokens per person per day, Opus 4.8 at list) is loaded in the calculator. Swap in your headcount, your usage, your growth. The engine recomputes the threshold, the timing, and the program in front of you.
01
Start simple
Pick the workload that looks like yours, or keep the worked example.
02
See your threshold
One minute to the commit / wait verdict, with the math visible.
03
Go as deep as you want
Engineer mode opens every assumption: racks, power, benchmarks, provenance.
Cards on the table
Nothing in this story is a quote. Pricing is official provider list pricing. Usage, growth, and infrastructure costs are modeled on a full Rubin-era stack including the networking, storage, cooling, and operations that rack-only math skips. Lease rates are quote-driven reference values. Company savings figures come from public filings and reports. Every figure carries its label, and the final check is always a Swarm proposal.