Your AI budget just doubled, again.

Swarm Systems

Take control of your AI spend.

Own your infrastructure.

Leave the cloud.

The two-minute version of the decision every AI-heavy enterprise is hitting this year: why the API meter can't be your infrastructure, and the program that replaces it.

Scroll · the story plays as you go

Act I · The curve everyone is on

The token era has a shape.

MAY '25480T TOKENS/MOREPORTED
1.4Q2.7Q4.2Q480TMAY '25980TJUL '251.3QOCT '253.2QMAY '26TOKENS/MO · ONE PLATFORM$25.0M$50.0M$75.0M$100.0MTHE BILL · $/MOTHE PLAN · $13.9M/MO TODAY$/TOKEN MODELED −33%/YR · −70% BY +36 MO5× THE PLAN$71.8M VS $13.9M/MOTODAY+12 MO+24 MO+36 MO

Reported, not modeled: one platform's disclosed token consumption. Watch the dates. Each point is only months apart.

It's an exponential. Project it forward at HALF that pace, and last year becomes a footnote on the chart.

Now price it. The amber line is that same curve as a monthly bill for an enterprise spending $13.9M/mo today, with prices modeled falling a third every year.

Strip the tokens away and look at spend alone. We model the price per token falling 33% a year, 70% cheaper by month 36, and the bill still compounds.

≈5× the plan in three years: $71.8M a month against today's $13.9M. Cheaper tokens aren't a plan; you'll spend them. What follows is the plan.

Reported figuresModeled projectionModeled price decline

Act I · The math

Meet the deployment the board approved.

Headcount

200,000 employees

company-wide AI rollout

Model

Claude Opus 4.8

$5 in / $25 out per MTok, list

The budget

$0/mo

$100 per employee per month

Month-one usage

0 tokens

300.0K per employee per workday, modeled agent-era

Month-one bill

$0/mo

69% of budget. Everyone's happy.

Month one is comfortable. This story is about every month after that.

Official pricingModeled usageWorked example

Act I · The curve

Success compounds. Budgets don't.

Month 6$17.7M/MO0.9× THE PLAN
$12.5M$25.1M$38.0MTHE PLAN$20.0M/MOYOUR API BILLOVER PLAN · MO 9MO 6TODAY

Zoomed in, month by month, it looks manageable. Tokens per employee climb ~5% a month, a fraction of the measured industry pace, and still compounding.

Month 9: you slip past the budget you promised the board. Quietly, barely over the line.

Zoom out. $102.5M a month, ≈5× the plan, the same multiple the industry curve just showed you. On the meter your options are throttle your best people, or explain this line every quarter. Forever.

Official pricingModeled growth

Act I · This year, everywhere

2026 is the year of AI rationing.

The best tools worked. The meters exploded. And the biggest names in tech responded the only way a meter lets you:

Uber

burned its entire 2026 AI budget by April, four months in, as Claude Code spread across ~5,000 engineers. Now it caps engineers at $1,500/month per tool.

Fortune · Forbes · The Information, May 2026

Microsoft

pulled most internal Claude Code licenses across the Windows / M365 org after per-engineer costs hit $500–$2,000 a month, and moved engineers to its cheaper in-house tool.

Fortune · TNW, May–June 2026

Coinbase

cut AI spend ~50% by defaulting 1,200 agents and its engineers to the cheapest open-weight models. Token usage kept compounding anyway.

The New Stack · Fortune, June 2026

Caps, revoked licenses, defaults to the cheapest model: rationing the one tool that's actually working. The workload isn't shrinking. The meter is just unaffordable. If you want your people on the best models, you have to own the economics underneath them.

Public reportsFull citations on the data sources page.

Act I · The answer

Run the models on machines you control.

Same work, same tokens, priced as infrastructure instead of retail. The meter stops being the business model.

RENTEDAPI · Opus 4.8 list, your mix
$11.00/MTOK
OWNEDYour cluster, at delivered capacity
$3.56/MTOK

Per token

68% less

at delivered capacity, full Rubin stack

Same budget

3.1× the tokens

or the headroom to stop rationing

Your first deployment

$0/mo

60 racks · 12 MW IT, less than this month's API bill

The engine's verdict on this workload: the economic threshold is already crossed. Dedicated capacity wins today. This isn't a someday decision.

Dropbox saved $74.6M in two years after moving off the public cloud

S-1 filing

X cut its cloud bill by roughly 60% after moving on-prem

as reported
Official pricingModeled infraQuote-driven lease

Act I · The clock

Compute arrives 12 months after you sign, not after you decide.

That's the lease-to-live lead time. Wait six months to sign and months 1218 from today still run on the meter, at those months' grown prices. The math on waiting:

Sign todaybaseline
signature todaycompute lands · month 12
Wait 6 months+$98.7M to the meter
signature month 6compute lands · month 18
Wait 12 months+$254.9M to the meter
signature month 12compute lands · month 24
Wait 18 months+$488.3M to the meter
signature month 18compute lands · month 30
12-mo lead: lease signed → compute live extra months on the meter: the cost of waiting

The threshold is already crossed, and every month of deliberation has a price tag. Deciding now is the cheapest thing in this story.

Modeled lead timeOfficial pricing

Act II · The catch

One detail: those racks need a home. The self-build list:

Or partner with Swarm, and check every box.

Every box, one address: your building.

1

Find land

2

Win entitlements

3

Join the interconnection queue

4

Post utility securities

5

Contract grid power

6

Order long-lead equipment

7

Track chip roadmaps

8

Design years ahead

9

Raise the capital

10

Manage the build

11

Commission it

12

Operate it 24/7

DEPLOY 1 · 12 MWDEPLOY 2DEPLOY 3SINGLE-TENANT · 16–20+ MW IT

Every line is measured in months, some in years. ~36 months to first token if everything goes right, while the meter collects $1.3B. Keep scrolling.

Sites secured. Queues already joined. Power contracted. Operations staffed. One master lease covers all of it, and your compute lands in 12 months, not 36.

A dedicated, single-tenant Swarm building: 16–20+ MW IT, leased as 12 MW deployments. Land, power, cooling, fabric, and operations, behind one door with your name on it. Step inside.

Act II · The factory

Step inside. Meet your private AI factory.

Deployment 1 of your building: sixty Vera Rubin VR200 racks at ~190 kW each, plus the lower-draw head row the reference architecture actually requires: scale-out fabric, context-memory and shared storage, management. You don't build it, finance it, or staff it. You rent it, all-in.

Your private AI factory · Deployment 1

12 MW IT gross · single-tenant

60× VR200 compute 5× scale-out fabric 3× context-memory / storage 1× management

Head-row racks draw far less power than compute racks. Everything fits inside the ≈12 MW gross envelope.

60× Vera Rubin VR200 racks

~190 kW per rack. Your racks, your models, your data

3.3T tokens/month delivered

benchmark-based, at modeled utilization

Networking · storage · cooling

the full pod stack, not just GPUs

Staffed 24/7 operations

Swarm runs it; you consume it

A dedicated single-tenant building

16–20+ MW IT buildings. No cage, no neighbors

Leased as 12 MW IT deployments

your next deployments land in the same building program

The rent

$11.8M/MO ALL-IN

$981/kW-month, margin-inclusive: compute, facility, power, and ops

If you bought it instead

$612.0M

capex, on your balance sheet

Quote-driven leaseModeled reference layout

Act II · Scale in place

Then it multiplies.

1 DEPLOYMENT · 12 MW IT

Deployment 1 · 12 MW

Deployment 2 · 12 MW · next

Deployment 3 · 12 MW · next

Deployment 1 lands with headroom. At this workload's pace, headroom has a shelf life.

A second deployment lands ~6 months after you order it. Same building, same lease, no new project.

A third: 36 MW IT, the scale this workload actually reaches ~29 months after first power (engine-computed). Your own building, filled on schedule.

Modeled expansion

Act II · The program

The first factory is the first step.

Programmatic scaling: 12 MW deployments across Swarm's secured buildings, each landing in months, ordered before you need them. On this workload the program reaches 36 MW on deployment 3 and keeps stepping toward ~72 MW inside four years. A self-build gets one step, three years from now.

6.3T12.5T18.8T25.0TSELF-BUILD: FIRST CAPACITY, MONTH 36DEMANDSWARMCAPACITYCOD+12 MO+24 MO+36 MO+47 MO

The path to 36 MW

3 deployments

12 MW each · ~29 months after first power

Every step

68% less per token

at delivered capacity, with headroom to stop rationing

Or blend

Hybrid

keep an API burst lane for spikes and frontier work

Cards on the table: growth headroom costs money before you fill it. Even carrying it, this stepped ramp runs 47% under the meter across four years of 5%-a-month growth. No new land hunt. No new interconnection queue. And no retail margin on the fastest-growing line of your P&L.

Modeled expansionQuote-driven lease

Rent the intelligence.

Own the economics.

5-year savings

$0

equivalent service, engine-computed

Threshold

Crossed

dedicated capacity already wins

Cost of waiting a year

$0

savings forgone by deferring

Capacity online

12 mo

from reservation to first token

The story is a calculator

Every number you just saw is an input you can change.

The worked example (200,000 employees, 300.0K tokens per person per day, Opus 4.8 at list) is loaded in the calculator. Swap in your headcount, your usage, your growth. The engine recomputes the threshold, the timing, and the program in front of you.

01

Start simple

Pick the workload that looks like yours, or keep the worked example.

02

See your threshold

One minute to the commit / wait verdict, with the math visible.

03

Go as deep as you want

Engineer mode opens every assumption: racks, power, benchmarks, provenance.

Cards on the table

Nothing in this story is a quote. Pricing is official provider list pricing. Usage, growth, and infrastructure costs are modeled on a full Rubin-era stack including the networking, storage, cooling, and operations that rack-only math skips. Lease rates are quote-driven reference values. Company savings figures come from public filings and reports. Every figure carries its label, and the final check is always a Swarm proposal.