How Many GPUs
Would It Take?
Token budgets for eight enormous jobs, from reading every U.S. law to examining every patient on Earth.
Research by grok 4.6 agents, method review by GPT-6 Pro, built with Claude · September 2026
Reading everything people have written, once, is cheap. Re-reading, watching and thinking about everything, all the time, is not. A single GPU could read every federal statute in 1.3 GPU-hours. A daily medical workup for everyone in a care episode takes about 26,900 GPUs running around the clock. And the one job here that rivals the entire projected 2030 chip supply sounds the most ordinary: a robot helper in every home, which needs 85 million GPUs.
What separates those answers isn’t the size of the pile. It’s how many things a model tends to at once, how often it has to process them again, and how much it writes each time. Change the design and the same job moves from a handful of GPUs to thousands.
Below, eleven large jobs are converted into tokens, the word-sized chunks models read and write, and then into GPUs. Each job has a one-time cost, such as reading every record once, and a recurring cost, such as the daily workups that follow. Every number updates when you change the serving speed or the one-time deadline, and every estimate carries a 95% range.
The Short Answer
Nine of the eleven jobs could run every day on a few thousand GPUs or fewer. Medicine needs tens of thousands, and a robot in every home needs tens of millions. One-time costs are larger but finite: reading every medical record on Earth takes 211,000 GPUs to do this in one day, and rewriting every app in both stores takes 6,220 GPUs.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more.
How to read this chart. Each job gets up to two bars. A solid bar is recurring work, counted as GPUs busy through its window: around the clock for most jobs, a workday for company messages, all year for courts. An outlined bar is one-time work, counted as the GPUs needed to finish it by the deadline you pick. The whisker on each bar is our 95% range on the job’s token size; click one to jump to its estimate table. Click a job’s name to hide it and the axis rescales to what remains.
Eleven Enormous Jobs
Each chapter shows several ways to do the same job, with one-time and recurring costs side by side. The underlined design is the one we would plan with. Hide designs to rescale the chart; hover for tokens; click a bar for the math.
Workload 01Read all U.S. federal law
A single GPU could read every federal statute and regulation in less than a day, and keeping up with new law costs almost nothing. The U.S. Code takes 1.3 GPU-hours to read; adding the Code of Federal Regulations brings the job to 6.9 GPU-hours.
The U.S. Code runs 24.4 million words, which a modern tokenizer turns into about 37 million tokens, and the regulations add more than 100 million words. GPUs read in parallel, and reading is the cheap half of their work. The recurring cost is the daily Federal Register and new public laws: a few hundred thousand tokens a day, or less than 0.01 GPUs.
Most legal text is case law. Every published American opinion comes to about 19 billion tokens, and the world’s published opinions to roughly 350 billion, a one-time job of 12,200 GPU-hours.
Reading isn’t legal analysis. Checking every provision against every other multiplies the bill many times over. A sensible system reads once, builds an index, and pulls the relevant sections for each question.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
U.S. Code
Code + regulations
+ State law
+ U.S. case law
World opinions
The Evidence
Measured,
the U.S. Code runs 24.4 million words, the largest it has been since at least 1991. Source
Measured,
legal text tokenizes at about 1.5 tokens per word, a bit worse than everyday English (1.33). Source
Reported,
the Code of Federal Regulations holds over 100 million words, not counting agency guidance. Source
Workload 02Watch a livestream all day
Watching one livestream all day takes a small slice of one GPU. Watching every surveillance camera on Earth takes millions. One stream at default resolution needs 0.02 GPUs around the clock, and catching up on its last 30 days takes 0.5 GPUs to do this in one day.
Video models don’t watch the way people do. Gemini samples one frame per second and turns each into 70 tokens by default, or 280 at high resolution, plus 32 tokens for every second of audio. A day of one stream is about 8.8 million tokens at default resolution and 27 million at high resolution.
At scale, the totals jump. Every Twitch, YouTube and Kick stream plus 4,000 satellite channels needs 2,400 GPUs. The world’s roughly one billion surveillance cameras need 13 million GPUs, and indexing the whole YouTube library once takes 1.3 million GPUs to do this in one day.
Two caveats apply. One frame per second misses fast motion, and naming strangers or reconstructing a day’s events is harder than describing a scene. Those are accuracy problems, not just more tokens. Video priced at text-serving rates is also an approximation: a real pipeline runs vision and audio encoders too.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Default res
High res
5 fps
All major streams
Every CCTV camera
YouTube library
The Evidence
Reported,
Gemini samples video at 1 frame per second and charges 70 tokens per frame by default, 258–280 at high resolution. Source
Reported,
audio adds 32 tokens for every second of video. Source
Measured,
NVIDIA runs 33 live captioning streams on one H100 with a small video model and 10-second chunks. Source
Workload 03Defend 1,000 servers
Defending 1,000 servers takes about 5.7 GPUs, as long as ordinary software condenses the logs before a model reads them. Re-checking the last 90 days once takes 405 GPUs to do this in one day.
A thousand busy servers produce around 10,000 security events a second. Send every one straight to a language model and it reads about 216 billion tokens a day, using 255 GPUs before it writes a single verdict.
Real defense tools condense first. The planning design has each server send one 2,000-token evidence packet every 100 seconds and runs 1,000 deep investigations a day, which is where the number above comes from. Elastic reports its triage agents settle at 7 to 9 model calls per investigation after tuning.
Compute isn’t the limit here. No amount of reading can find an attacker the logs never recorded, and a wrong containment decision is expensive at any GPU count.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Correlated only
+ Investigations
10× attack surge
Raw logs
Raw, verbose
The Evidence
Rule of thumb,
SIEM sizing tables budget about 10 events per second for a Linux server and 30 for Windows. Source
Measured,
a 1 KB JSON security event is about 350 tokens. Logs tokenize denser than prose. Source
Reported,
Microsoft Defender correlates millions of signals into incidents before it acts, rather than judging each event. Source
Workload 04Fly 20,000 drones
Flying 20,000 drones takes almost no AI. Asking a language model to fly each one takes hundreds to thousands of GPUs, and it still wouldn’t work.
Autopilots already fly drones with plain arithmetic. In February 2026, 22,580 drones flew at once from a single computer, on paths calculated in advance. Planning such a show is a one-time job: a model writes 31 formations and the show software fills in the paths, which takes 1.5 GPU-hours.
Put a model in charge of 200 squads, briefing each every 10 seconds, and it needs 8 GPUs. Give every drone its own decision each second and it needs 800 GPUs. Let each one reason for 1,000 tokens first and it needs 10,600 GPUs.
The GPU count is the smaller problem. In a published benchmark with 16,384 simultaneous requests, a large reasoning model took almost six seconds to start answering and then wrote about 27 tokens a second per request. A drone at 10 meters a second covers 60 meters on stale orders in that time. Robot models avoid this by thinking a few times a second and passing faster motor commands to a small controller.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Squad leaders
Every drone, 1 Hz
Watch all cameras
Every drone, 10 Hz
Every drone thinks
The Evidence
Reported,
PX4 autopilots take numeric setpoints and only need a 2 Hz heartbeat from an outside controller. Source
Reported,
inside the autopilot, position runs at 50 Hz, attitude at 250 Hz and rotation rate at 1,000 Hz. Source
Reported,
22,580 drones flew at once from a single computer in February 2026, on flight paths computed in advance. Source
Workload 05Read a company's messages
An AI can label everything a 100,000-person company writes in a day on about 6.3 GPUs. Thinking hard about each message takes 178 GPUs. Reading the kept archive once, about two years of de-duplicated email and chat, takes 1,040 GPUs to do this in one day.
The average Microsoft 365 worker receives 117 emails and 153 Teams messages a day, but most are copies of something sent to many people, and Teams chats are copied into every participant’s mailbox. Counted once, that is about 100 messages per person. With quoted replies removed, the median email reply is 43 words.
A short label per message, such as “payroll” or “customer complaint”, adds little. Writing 1,000 tokens of reasoning about every message multiplies the output a hundredfold, and writing is the expensive half. Scaled to all 450 million paid Microsoft 365 seats, even the light version needs 9,380 GPUs.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Light, workday
Light, spread out
Every inbox copy
Deep reasoning
Every M365 seat
The Evidence
Measured,
the average Microsoft 365 worker receives 117 emails a day and 153 Teams messages per weekday. Source
Measured,
with quoted text stripped, half of email replies are shorter than 43 words. Source
Reported,
Microsoft has more than 450 million paid commercial Microsoft 365 seats. Source
Workload 06Examine every patient on Earth
Reading every person’s medical record once is a bounded job; the daily bill depends on whether each workup re-reads it. Reading all 8.3 billion records and writing a summary of each takes 211,000 GPUs to do this in one day. After that, a daily workup for everyone in a care episode, using the summary plus new notes, takes 26,900 GPUs around the clock.
Who counts is a choice, not a measurement. On an average day about 1.5% of humanity sees a clinician. A classic study found that each month, about a third of people consider seeking care, which works out to about 5% a day if each episode lasts five days. We plan for 5% and show the other choices as separate bars.
Record length varies enormously. The median patient arriving at a U.S. emergency room in 2022 had 98,000 tokens of notes on file, and one in five had more text than Moby-Dick. Most of the world has far shorter or paper records, so a global average is closer to 12,000 tokens.
That is why the design matters more than the population. Reading a full U.S.-sized 200,000-token chart at every workup takes 117,000 GPUs. Re-sending it with each of eight questions takes 789,000 GPUs. A full assessment for every person, every day, takes 5.9 million GPUs, and the most wasteful version needs 150 million GPUs, more AI chips than exist.
One warning applies to every medical number here. The serving speeds were measured on 2,000-token prompts, and long records cost more per token as attention grows with length. A real system would read records in windows of 32,000 to 64,000 tokens and retrieve what each question needs, so these figures are reference work rather than a tested design.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Summarized charts
US ER-size charts
Clinic visits
US only
Re-read each turn
Symptomatic (15%)
Everyone, daily
Literal maximum
The Evidence
Reported,
the world had about 8.3 billion people in 2026. Source
Measured,
people worldwide make about 5.4 outpatient visits a year, roughly 1.5% of humanity on any given day. Source
Measured,
each month, 800 of every 1,000 Americans have symptoms and 327 consider seeking care. Source
Workload 07Argue every court case
Arguing every new court case in the world, each at the depth its file needs, takes about 202 GPUs running all year. Clearing the existing backlog once is as much reading and writing as a full year of new cases. Arguing the roughly 280 million pending cases takes 74,300 GPUs to do this in one day.
The world opens roughly 280 million court cases a year. U.S. state courts alone receive 70 million filings, most of them traffic tickets. Chinese courts accepted 37.5 million cases in 2025, about 29% of them enforcement stages of earlier disputes, and India has a backlog of 54 million.
Most files are short. A traffic case is a page or two, and even a federal appeal brief is capped at 13,000 words. The expensive American habit is discovery, the exchange of internal documents before trial: a large corporate production reviews about 100 GB, nearly half a billion tokens per case. Give that to every case and the bill becomes 466,000 GPUs.
Token budgets are the biggest unknown. A far more generous docket, with two million tokens for every contested case and ten billion for each of the largest discovery fights, needs 2,790 GPUs.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Case by case
Re-argue backlog
Generous budgets
All get discovery
All get big law
The Evidence
Reported,
U.S. state courts received 70 million filings in 2024; 57% of cases are traffic. Source
Reported,
Chinese courts accepted 37.5 million cases in 2025. Source
Reported,
Brazil opened 40.9 million new cases in 2025 and carries 75.5 million pending. Source
Workload 08Call every game on TV
Watching every televised game and updating win probability on every play takes about 16 GPUs at the busiest hour of a Saturday. Sports produce few decision points: a handful per second, worldwide.
Sportradar streamed over 525,000 matches in 2025. Games average under two hours, so roughly 135 are live at any moment and about 500 at a Saturday peak, when European football overlaps with American college football. A 2.5-hour game at high resolution is about 2.8 million tokens.
Production systems compute probabilities with small statistical models; NFL Next Gen Stats runs 75 of them per play in under a second. Having a language model re-reason every play instead raises the peak to 20 GPUs. Reading decades of play-by-play once to prime the captions takes 258 GPU-hours.
As a stress test, watching every data-covered match with 3,000 feeds live at once and a detailed update every 45 seconds would need 124 GPUs.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Stats only
Average day
Saturday peak
Peak + LLM
Every touch
Every match
The Evidence
Reported,
Sportradar distributes more than 650,000 live sports streams a year. Source
Reported,
Stats Perform covers 500,000+ matches a year across 3,900 competitions. Source
Reported,
NFL Next Gen Stats runs 75 machine learning models on every play in under a second. Source
Workload 09Rewrite every app
Rewriting every app in both stores is a one-time job of about 6,220 GPUs to do this in one day. Keeping them all current afterward takes about 5.6 GPUs.
The App Store lists about 2.6 million apps and Google Play about 2.6 million, but 37% of apps are on both, so there are roughly 3.4 million distinct products. Most are small, and a third haven’t been updated in two years. A few giants pull the average up: Uber’s Android codebase grew from 10,000 lines in 2008 to 10 million in 2021.
Code runs about 7 to 10 tokens per line. A coding agent reads far more than it writes, but most of those reads are cached: in 731 benchmark sessions, 97.6% of input tokens were cache hits, which cost almost nothing to serve again. We count new tokens only.
A heavy loop that explores, rewrites, tests and reviews every app takes 22,100 GPUs to do this in one day, and review is a large share: in one study of agent traces, code review used 59.4% of all tokens.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Port each app once
Agentic + review
Small apps only
The Evidence
Measured,
Claude Opus 4.7 averaged 3.46 million input tokens per SWE-bench Pro session, and 97.6% of them were cache reads. Source
Reported,
Apple reviewed 9,100,620 app submissions in 2025 and rejected 2,093,244 of them. Source
Reported,
Apple's App Store saw 557,000 new app submissions in 2025, up 24% from 2024. Source
Workload 10Advise every trader
A model that checks every order a person places, worldwide, needs about 2,360 GPUs. Reading every filing and a decade of financial news first takes 564 GPUs to do this in one day.
Owning stock isn’t trading. China has 251 million securities investors and India 238 million demat accounts, but only 8% of India’s retail traders traded on more than 50 days last year. Counting orders people place themselves gives about 120 million a day, led by China, where Shanghai and Shenzhen logged 68.8 billion two-sided transfers in 2025.
An always-on agent that briefs each of 40 million active traders every 15 minutes during their market hours needs 8,890 GPUs. Calling a model on every order and cancel, including algorithmic flow, needs 39,500 GPUs. Market ticks themselves aren’t read as text.
These are averages across the day. Load peaks at about three times the average when Asian markets open, so a real deployment would provision for the peak.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
Advice per trade
15-minute agent
LLM on every order
The Evidence
Reported,
the Shanghai and Shenzhen exchanges recorded 68.848 billion two-sided share transfers in 2025, about 142 million trades a day. Source
Reported,
just over 1.9 million French people bought or sold shares in 2025, making 56 million retail equity trades. Source
Reported,
investors in India held 237.7 million demat accounts in August 2026, but only about 46 million were active NSE clients. Source
Workload 11Give every home a robot helper
A robot helper in every home, with its planning done in a datacenter, needs about 85 million GPUs running around the clock. Even on OpenAI’s faster Jalapeño chips at 12,000 tokens a second, that is 49% of all the AI chips projected for 2030. It is the only job here on that scale, and a robot for every person needs 1.7 billion GPUs.
There are about 2.2 billion households and 8.3 billion people. A robot that listens all waking day hears 32 audio tokens a second. Robot models plan a few times a second: Gemini Robotics takes about 250 milliseconds from camera images to an action chunk, and each second of motion encodes to about 30 tokens per arm.
The one-time cost is scanning each home and learning the family: 3.4 million GPUs to do this in one day. The recurring cost is what matters, and it depends on where the robot thinks. If robots run their models on board and call the cloud only for hard questions, datacenter demand drops to 624,000 GPUs. Four 30-frames-per-second cameras with nonstop reasoning for every person would need 40 billion GPUs.
Bodies, not tokens, limit this decade. Goldman Sachs projects 890,000 humanoid robots shipped in 2030, a tiny fraction of 2.2 billion homes. If the robots arrive, their datacenter demand would exceed the planning estimates of every other job on this page combined, many times over.
Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.
One per home
Robot per person
On-device brain
30 fps, always on
The Evidence
Measured,
π0-FAST takes about 750 ms to generate one 1-second action chunk, which runs to roughly 30 tokens per robot arm. Source
Reported,
Gemini Robotics runs its main model in the cloud, and it takes about 250 ms to go from camera images to an action chunk. Source
Reported,
Figure’s Helix runs both its 7–9 Hz vision-language model and its 200 Hz motor policy on GPUs inside the robot. Source
What Makes a Job Big?
A job’s size depends less on how much text exists than on how often a model processes it again and how much it writes each time.
GPUs ∝ things handled at once × passes per second × tokens per pass
The map below plots every design’s recurring cost: tokens read per second across, tokens written per second up. Dashed curves mark equal GPU counts. Jobs that mostly read sit low; jobs that write, and especially jobs that reason, climb toward the costly corner, because at Blackwell planning speeds a GPU writes tokens five times more slowly than it reads them.
Serving speed moves every point at once. Pick a speed to compare how much the hardware matters against how much the design does. Each row shows the design we would plan with, its one-time and recurring costs, and a button that cycles through the other designs.
Read all U.S. federal law
Read the U.S. Code and federal regulations once, then keep up with each day's new rules
One-time
0.3 GPUs to do this in one day
95% range 0.2–0.4
■ = 1/100 of a GPU
Reads 200M, writes 10M tokens in total
Recurring
less than 0.01 GPUs around the clock
95% range less than 0.01–less than 0.01
■ = 1/100 of a GPU
Reads 370K, writes 18K tokens per day
Watch a livestream all day
One stream at Gemini's default resolution, 1 frame per second
One-time
0.5 GPUs to do this in one day
95% range 0.09–2
■ = 1/100 of a GPU
Reads 260M, writes 30M tokens in total
Recurring
0.02 GPUs around the clock
95% range 0.01–0.03
■ = 1/100 of a GPU
Reads 8.8M, writes 1M tokens per day
Defend 1,000 servers
Evidence packets plus a deep AI investigation of 1,000 alerts a day
One-time
405 GPUs to do this in one day
95% range 15–3,940
■ = 10 GPUs
Reads 160B, writes 39B tokens in total
Recurring
5.7 GPUs around the clock
95% range 0.2–41
■ = 1/10 of a GPU
Reads 2.2B, writes 530M tokens per day
Fly 20,000 drones
Autopilots fly the paths; a model briefs 200 squads every 10 seconds
One-time
0.06 GPUs to do this in one day
95% range 0.02–2.8
■ = 1/100 of a GPU
Reads 7.3M, writes 9.7M tokens in total
Recurring
8 GPUs around the clock
95% range 2.5–30
■ = 1/10 of a GPU
Reads 5.2B, writes 350M tokens per day
Read a company's messages
Label every unique message (about 100 per person) within the 8-hour workday
One-time
1,040 GPUs to do this in one day
95% range 149–5,840
■ = 100 GPUs
Reads 650B, writes 50B tokens in total
Recurring
6.3 GPUs through the workday
95% range 1.6–26
■ = 1/10 of a GPU
Reads 1.3B, writes 100M tokens per workday
Examine every patient on Earth
Daily workups for people in a care episode (5% a day), using a chart summary
One-time
211,000 GPUs to do this in one day
95% range 76,900–1.1 million
■ = 10,000 GPUs
Reads 100T, writes 17T tokens in total
Recurring
26,900 GPUs around the clock
95% range 4,020–71,900
■ = 1,000 GPUs
Reads 6.6T, writes 3.3T tokens per day
Argue every court case
Every new case worldwide, argued at the depth its file needs
One-time
74,300 GPUs to do this in one day
95% range 30,800–216,000
■ = 1,000 GPUs
Reads 27T, writes 7.4T tokens in total
Recurring
202 GPUs all year
95% range 79–586
■ = 10 GPUs
Reads 27T, writes 7.4T tokens per year
Call every game on TV
Saturday peak: watch 500 games at once and caption every play
One-time
11 GPUs to do this in one day
95% range 2.9–68
■ = 1 GPU
Reads 5.1B, writes 850M tokens in total
Recurring
16 GPUs around the clock
95% range 5.8–31
■ = 1 GPU
Reads 14B, writes 15M tokens per day
Rewrite every app
Port every unique App Store and Google Play app once, then keep up
One-time
6,220 GPUs to do this in one day
95% range 1,330–24,700
■ = 100 GPUs
Reads 1.5T, writes 770B tokens in total
Recurring
5.6 GPUs around the clock
95% range 1.7–22
■ = 1/10 of a GPU
Reads 1.5B, writes 660M tokens per day
Advise every trader
A model checks each order a person places, with no calls on algorithmic orders
One-time
564 GPUs to do this in one day
95% range 289–1,630
■ = 10 GPUs
Reads 390B, writes 20B tokens in total
Recurring
2,360 GPUs around the clock
95% range 475–11,900
■ = 100 GPUs
Reads 840B, writes 240B tokens per day
Give every home a robot helper
One robot per home, with its planning done in a datacenter
One-time
3.4 million GPUs to do this in one day
95% range 1.2 million–110 million
■ = 100,000 GPUs
Reads 2.5Q, writes 88T tokens in total
Recurring
85 million GPUs around the clock
95% range 17 million–300 million
■ = 1 million GPUs
Reads 45Q, writes 5.7Q tokens per day
So how many GPUs would it take? For most of these jobs, a few thousand at most. For jobs that re-read long records for billions of people, watch every camera, or give every home a robot, more than the world has.
The pattern holds across all eleven. A stock of text is finite: even the useful public web is a one-time read that a large cluster finishes in weeks to months. A flow never stops. Work that runs forever, for many people or devices, re-reading long records or writing at length, is what data centers are built for.
For scale: Google said in May 2026 that it processes 3.2 quadrillion tokens a month, which at these speeds is about 222,000 GPUs working around the clock. Epoch AI counts about 24 million AI accelerators shipped by mid-2026, and projections for 2030 center on about 100 million. Only the robot job, watching every camera, and the most wasteful version of medicine reach even a tenth of that.
Hardware will keep getting faster. If OpenAI’s Jalapeño chip delivers 12,000 blended tokens a second, it trims reading-heavy jobs and cuts writing-heavy ones sharply. It doesn’t change which jobs are big. Design does: cache the chart, condense the logs, let the autopilot fly, and let the robot think on board.
Data and Method
What a token is. Models read and write text in tokens, chunks of about three-quarters of an English word. We count tokens read (input, or “prefill”) separately from tokens written (output, or “decode”, which includes hidden reasoning), because a GPU writes tokens several times more slowly than it reads them.
The conversion. GPU-seconds = tokens read ÷ read speed + tokens written ÷ write speed. GPUs = GPU-seconds ÷ the seconds available: a day for ongoing work, eight hours for a workday, a year for court caseloads. One-time reads are shown as GPU-hours in their chapter and, marked “once” with outlined blocks, as GPUs needed to finish within a day on the ladder. Only recurring work describes GPUs running continuously. For the blended Jalapeño tier, GPU-seconds = all tokens ÷ 12,000.
Serving speeds. Dense model (4K / 800): A slower reference rate for dense 100B–400B-class models, long contexts, or fast per-user speeds. Llama 3.1 405B serves ~140–270 generated tokens/s per GPU in MLPerf. Blackwell planning (10K / 2K): A fixed reference rate for a large open mixture-of-experts model at ~50–100 tokens/s per user, discounted from short-prompt benchmarks: 38% of measured peak prefill and 57% of MLPerf interactive decode. Not a guarantee for long contexts. Blackwell optimized (20K / 6K): Disaggregated serving with 4-bit weights on NVL72 racks at relaxed latency. Near SGLang's 26K prefill and CoreWeave's 6.5K decode per GPU. Future: OpenAI Jalapeño (12K blended): OpenAI's first custom inference chip (with Broadcom), unveiled June 2026. Assumed 12,000 blended tokens/s per chip for input and output alike. OpenAI reports mixed tokens/s per kW on an 8K-in/1K-out test; ~12,700/s per 700 W chip is derived, not stated. The blend is input-heavy, so it flatters generation-heavy work.
Long records. The serving speeds come from benchmarks with 2,000-token prompts. Attention cost grows with prompt length, and memory for a model’s working cache grows with it too: about 35 GB for one 500,000-token sequence on DeepSeek’s compressed attention. There is no single honest correction factor, so the medical and legal numbers assume records are read in bounded windows of 32,000 to 64,000 tokens with extraction and retrieval, and should be read as reference work, not throughput predictions for giant prompts. Video is priced as text-token-equivalent work; a real pipeline also runs vision and audio encoders.
What the numbers are not. They are GPU-equivalents of work, not a purchase order. A large model must fit in memory, and the vLLM benchmark behind our planning rate used 16 GPUs for one serving setup, so a small job can still need a whole rack. They ignore latency: a stack that meets throughput can still miss a 100-millisecond control deadline. They ignore power, networking, storage, retries and failed runs. A stack ten times slower makes every number ten times larger.
How the estimates were checked. Ten research agents (grok 4.6 at extra-high reasoning) re-derived every figure in the original brief against primary sources in September 2026, marking each claim confirmed, adjusted, unverifiable or an assumption. Three new workloads (medicine, courts, sports) were built from scratch the same way. A GPT-6 Pro method review then pushed back. It flagged that medical “need” is a prevalence choice rather than a consultation rate, that short-prompt benchmark speeds do not transfer to 500,000-token records, that court token budgets matter more than case counts, and that one-time reads must stay visibly one-time; the medical ceiling, the generous court docket, the sports stress peak and the long-record caveats come from that review. Corrections from the research pass include Gemini’s default video rate (70 tokens per frame, not 258), 10,200 GPUs rather than 10,000 for reasoning drones, 1.5 rather than 1.6 tokens per legal word, and Sportradar’s 650,000 streams a year. Every figure lives in one typed data file; a script recomputes all of them at every speed and fails if any drifts from the research.
Why no confidence intervals. The dominant unknowns are choices, not sampling error: which model, how much reasoning, which deadline, whether a chart is cached. Each chapter shows those choices as separate bars instead.
95% ranges. Each cost line has a range on its token size: a low end we would expect to be undercut about one time in 40 and a high end exceeded about one time in 40. They are judgment-based, built from the sourced low and high values of each driver (people, frequency, tokens per item), not statistical confidence intervals. The GPU range follows from the token range at the chosen speed.
One-time and recurring. One-time work (reading every record once, rewriting an app, clearing a backlog) is shown with outlined bars as the GPUs needed to finish by a deadline: one day, 30 days or one year. Recurring work (daily workups, per-play updates, continuous perception) is shown with solid bars as GPUs running for its whole window.
Estimate tables
Every job, design and cost line, with token ranges and GPUs at the speed and deadline selected above.
Read all U.S. federal law
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Read the U.S. Code once with short notes, then keep up with new public laws One-time: Read the 24.4 million words of the U.S. Code once, with short notes (in total) | 37M 35M–48M | 1.8M 710K–4.8M | 0.05 0.05–0.08 to do this in one day | Low uses the operative word count at 1.45 tokens/word, the lowest measured on Code text, with 2% notes. High uses a broader 28.8M-word file-text count at 1.65 tokens/word and 10% notes.Math and sources
|
| Recurring: Keep up each day: new public laws in the Statutes at Large (per day) | 10K 6K–27K | 515 120–2.7K | less than 0.01 less than 0.01–less than 0.01 around the clock | Low is the recent pace: the 118th Congress enacted 4,350 pages over two years and 119-1 enacted 2,008, about 1.5–1.6M words a year. High treats a busy Congress's 6M words as one year (117th: 8,742 pages over two years).Math and sources
|
| Read the U.S. Code and federal regulations once, then keep up with each day's new rules (plan with) One-time: Read the U.S. Code and the Code of Federal Regulations once, with short notes (in total) | 200M 180M–250M | 10M 3.6M–25M | 0.3 0.2–0.4 to do this in one day | CFR words 100–122M: 100M is the White House floor; 194,395 pages at ~600 words/page gives about 117–122M. Tokens/word 1.45–1.65, as measured on CFR volumes. Code 35.4–47.5M tokens. Notes 2–10%.Math and sources
|
| Recurring: Keep up each day: new federal laws and the Federal Register (per day) | 370K 250K–540K | 18K 5K–54K | less than 0.01 less than 0.01–less than 0.01 around the clock | FR page counts: 60,917 (2025, lowest) to 107,262 (2024 gross, highest), 82,730 central. Tokens/page 1,450–1,750 spans five measured issues (1,504–1,704). Law band as in the U.S. Code design. Notes 2–10%.Math and sources
|
| Add all 50 states' statutes and regulations to federal law One-time: Read federal law plus all 50 states' statutes and administrative codes once (in total) | 1.5B 1.3B–1.8B | 73M 26M–180M | 2.1 1.7–3.2 to do this in one day | Low scales StateCodes' 268M words of regulations for 45 states to 50 states (298M) and uses 1.45 tokens/word. High uses 550M words of statutes, State RegData's 416M words of regulations and 1.65 tokens/word. Notes 2–10%.Math and sources
|
| Recurring: Keep up each day: new federal and state laws and regulations (per day) | 640K 330K–1.7M | 32K 6.5K–170K | less than 0.01 less than 0.01–less than 0.01 around the clock | Bill counts are sourced (Ballotpedia ~18,300 a year; Quorum 15,951 in the first half of 2025). Words per bill, 800–8,000, is assumed. State regulation flow is assumed at 3–12% a year of the stock. Federal lines as above.Math and sources
|
| Add every published American court opinion to all U.S. statutes and regulations One-time: Read all U.S. statutes, regulations and published court opinions once (in total) | 23B 18B–32B | 1.2B 370M–3.2B | 34 23–55 to do this in one day | Case law 17–30B tokens: low is near the 6.92M-document Common Pile set (78 GB ≈ 19.5B); high allows CourtListener's 'over nine million' decisions at longer length. Stacked on the all-codes band. Notes 2–10%.Math and sources
|
| Recurring: Keep up each day: new U.S. laws, regulations and court opinions (per day) | 1.8M 870K–3.9M | 89K 17K–390K | less than 0.01 less than 0.01–less than 0.01 around the clock | Opinion clusters filed per year were 109,276–117,797 in 2023–2025; band 100,000–160,000 allows for late additions. A cluster can hold several opinions, so 2,000–5,000 tokens. The homepage 10-day add rate (~255,000 a year) includes backfill and is not used.Math and sources
|
| Read every published court decision in the world (opinions only, no statutes) One-time: Read the world's published court decisions once (opinions only) (in total) | 350B 200B–1T | 18B 4B–100B | 506 255–1,740 to do this in one day | No global census exists. Low assumes only China, the U.S. and a small share of other countries are online in full text. High allows a decade of Brazil's ~40M yearly decisions and longer Chinese documents. Notes 2–10%.Math and sources
|
| Recurring: Keep up each day: new court decisions published worldwide (per day) | 140M 55M–330M | 6.8M 1.1M–33M | 0.2 0.07–0.6 around the clock | Low counts China, the U.S. and a fraction of Brazil's decisions at short length. High allows longer Brazilian decisions, full Indian and European output and Chinese documents near 1,700 characters.Math and sources
|
Watch a livestream all day
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| One stream at Gemini's default resolution, 1 frame per second (plan with) One-time: Index the last 30 days of this stream (default resolution, 1 fps) (in total) | 260M 59M–790M | 30M 3M–180M | 0.5 0.09–2 to do this in one day | Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.Math and sources
|
| Recurring: Watch one 24-hour livestream (default resolution, 1 fps) (per day) | 8.8M 8.5M–13M | 1M 430K–3M | 0.02 0.01–0.03 around the clock | Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.Math and sources
|
| One stream at high resolution for on-screen text, 1 frame per second One-time: Index the last 30 days of this stream (high resolution, 1 fps) (in total) | 810M 180M–2.4B | 30M 3M–180M | 1.1 0.2–3.8 to do this in one day | Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.Math and sources
|
| Recurring: Watch one 24-hour livestream (high resolution, 1 fps) (per day) | 27M 25M–40M | 1M 430K–3M | 0.04 0.03–0.06 around the clock | Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.Math and sources
|
| One stream at high resolution, 5 frames per second One-time: Index the last 30 days of this stream (high resolution, 5 fps) (in total) | 3.7B 800M–11B | 30M 3M–180M | 4.5 0.9–14 to do this in one day | Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.Math and sources
|
| Recurring: Watch one 24-hour livestream (high resolution, 5 fps) (per day) | 120M 110M–190M | 1M 430K–3M | 0.1 0.1–0.2 around the clock | Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.Math and sources
|
| Every Twitch, YouTube and Kick livestream plus 4,000 satellite TV channels, at once One-time: Index the last 30 days of every one of these channels' broadcasts (in total) | 40T 7.7T–160T | 4.5T 390B–36T | 71,900 11,200–392,000 to do this in one day | A chosen lookback of 7 / 30 / 60 days (Twitch default / between Affiliate and Partner / Partner maximum), not a measured archive. YouTube keeps many stream recordings for years; linear TV keeps none publicly. Channel band as in the daily watch.Math and sources
|
| Recurring: Watch every Twitch, YouTube and Kick stream and 4,000 satellite TV channels (per day) | 1.3T 1.1T–2.6T | 150B 56B–600B | 2,400 1,600–6,530 around the clock | Channel low 130,000 allows for YouTube Live and Gaming overlap; central 150,000 and high 200,000 allow for uncounted linear feeds and smaller platforms. Per-channel token and notes bands as in video-low.Math and sources
|
| Every surveillance camera on Earth, around the clock One-time: Index the footage every camera still keeps on its recorder (in total) | 190Q 26Q–1,200Q | 31Q 2Q–410Q | 400 million 42 million–3.8 billion to do this in one day | Retention low 7 days (EDPB: erase after a few days); central 31 days (ICO keeps its own footage one month); high 90 days (common for commercial systems). Banks and police can keep longer. Camera band as in the daily watch.Math and sources
|
| Recurring: Watch every surveillance camera, around the clock, video only (per day) | 6Q 3.8Q–14Q | 1Q 280T–4.5Q | 13 million 6 million–42 million around the clock | Camera low = Transforma Insights 658 million connected CCTV devices (2025). Central = IHS Markit forecast of about 1 billion installed from 2021, a forecast rather than a count. High 1.5 billion is an unsourced ceiling. Per-camera band as video-low without audio.Math and sources
|
| Index the whole YouTube library, then each day's new uploads One-time: Index every video on YouTube once, at default resolution (in total) | 730T 420T–2.2Q | 83T 22T–500T | 1.3 million 615,000–5.4 million to do this in one day | YouTube publishes no catalog-hours total. A 2022 random sample found a 615 s mean; 9.8–14.8 billion videos give 1.7–2.5 billion hours. Shorts pull the mean down, so central is 2.0 billion; low 1.2B, high 4.0B allows for growth since 2024.Math and sources
|
| Recurring: Index each day's new YouTube uploads, at default resolution (per day) | 260B 250B–1T | 30B 13B–230B | 480 369–2,540 around the clock | Low and central use YouTube's last official figure, more than 500 hours a minute (May 2019). High = about 4 billion videos uploaded in 2023 (UMass estimate) × the 615 s sampled mean ≈ 1.87 million hours a day; short Shorts make this an upper bound.Math and sources
|
Defend 1,000 servers
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Condense each server's logs into evidence packets and send only those to the model One-time: Re-check 90 days of logs as packets; inventory and config-scan every server (in total) | 160B 6.4B–1.5T | 39B 1.3B–380B | 405 15–3,940 to do this in one day | Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.Math and sources
|
| Recurring: A 2,000-token evidence packet per server every 100 seconds (10 per second) (per day) | 1.7B 86M–10B | 430M 17M–2.6B | 4.5 0.2–27 around the clock | Low 1/s = one packet per server every ~17 minutes × (1,000 in, 200 out); high 30/s = one per server every ~33 s × (4,000 in, 1,000 out). No vendor publishes this rate; it is a design choice.Math and sources
|
| Evidence packets plus a deep AI investigation of 1,000 alerts a day (plan with) One-time: Re-check 90 days of logs as packets; inventory and config-scan every server (in total) | 160B 6.4B–1.5T | 39B 1.3B–380B | 405 15–3,940 to do this in one day | Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.Math and sources
|
| Recurring: Evidence packets every 100 s per server, plus 1,000 deep investigations a day (per day) | 2.2B 94M–16B | 530M 19M–3.9B | 5.7 0.2–41 around the clock | Packets as in the packet-only design. Investigations low: 100 alerts/day (Prophet typical) × 8 calls (Elastic after tuning) × (10,000 in, 2,000 out). High: 3,000/day (large enterprise) × 150 calls × (12,000 in, 3,000 out); Elastic's 36,000-token calls with fewer calls land inside this.Math and sources
|
| During an active intrusion: ten times the evidence packets One-time: Re-check 90 days of normal-rate logs as packets; inventory and scan servers (in total) | 160B 6.4B–1.5T | 39B 1.3B–380B | 405 15–3,940 to do this in one day | Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.Math and sources
|
| Recurring: Each attack day: 100 evidence packets per second plus 1,000 deep investigations (per day) | 18B 2.6B–75B | 4.4B 520M–19B | 46 6–194 around the clock | Packets low 30/s × (1,000 in, 200 out); high 200/s × (4,000 in, 1,000 out); central is 10× the normal rate. Investigations use the same band as the headline design and are not multiplied.Math and sources
|
| Skip condensing: send every raw log event to the model One-time: Read 90 days of every past raw event, inventory every server, and scan configs (in total) | 19T 950B–120T | 78B 2.6B–4T | 23,000 1,110–162,000 to do this in one day | Days 30 / 90 / 365 combined in log quadrature with the daily raw band (1 EPS × 150 tokens to 20 EPS × 400 tokens): ÷20 and ×6.2 on input. Output low 0 (read only); high uses 5% notes. Inventory and scan add 36.5M low and 580M high.Math and sources
|
| Recurring: Every raw event from 1,000 servers, read by the model with a 1-token verdict (per day) | 220B 13B–690B | 860M 86M–35B | 255 16–1,000 around the clock | Events/s 1/10/20: Linux servers sit below Logmanager's Windows Server 30 EPS, which belongs to the verbose design; filtered EDR runs near 1–3. Tokens 150 syslog / 250 mixed / 400 JSON. Output 0 if the model only reads, up to 5% notes.Math and sources
|
| Every raw event from Windows-style servers logging 30 events per second One-time: Read 90 days of past verbose Windows-style events; inventory and scan servers (in total) | 82T 13T–570T | 230B 26B–20T | 95,900 15,700–768,000 to do this in one day | Days 30 / 90 / 365 combined in log quadrature with the daily verbose band (10 EPS × 250 to 100 EPS × 400): ÷6.1 and ×6.9 on input. Output low 0 (read only); high uses 5% notes. Inventory and scan add 36.5M low and 580M high.Math and sources
|
| Recurring: Every raw event from 1,000 verbose Windows-style servers, with a 1-token verdict (per day) | 910B 220B–3.5T | 2.6B 860M–170B | 1,070 255–5,000 around the clock | Logmanager Windows Server 30 EPS central, Active Directory 100 EPS high, Linux 10 EPS low. Elastic's unfiltered 48k events/host/hour is 13.3 EPS. Tokens 250 / 350 / 400 for JSON-heavy events. Output 0 if read only, up to 5% notes.Math and sources
|
Fly 20,000 drones
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Autopilots fly the paths; a model briefs 200 squads every 10 seconds (plan with) One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total) | 7.3M 1.7M–460M | 9.7M 2.4M–380M | 0.06 0.02–2.8 to do this in one day | Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.Math and sources
|
| Recurring: Brief 200 squads every 10 seconds while 20,000 drones fly all day (per day) | 5.2B 1.7B–17B | 350M 86M–1.7B | 8 2.5–30 around the clock | Call rate held at 20 per second. Input per call: low 1,000 (squad totals and exceptions only), high 10,000 (full state, neighbours and recent events per drone). Output per call 50 / 200 / 1,000; briefing length is an assumption. A 100-drone status list measured about 4,200 tokens (o200k).Math and sources
|
| A model makes a decision for every drone once per second One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total) | 7.3M 1.7M–460M | 9.7M 2.4M–380M | 0.06 0.02–2.8 to do this in one day | Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.Math and sources
|
| Recurring: One model decision per drone per second, all day (per day) | 520B 170B–1.7T | 35B 14B–140B | 800 280–2,800 around the clock | Rate held at 1 Hz. Input per call: low 100 (own state and goal, no neighbours), high 1,000 (full neighbour states and recent events). Output per call 8 / 20 / 80 tokens, an assumption for a short setpoint command. One compact drone status line measured 14 to 42 tokens.Math and sources
|
| Every drone streams its camera to a model all day One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total) | 7.3M 1.7M–460M | 9.7M 2.4M–380M | 0.06 0.02–2.8 to do this in one day | Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.Math and sources
|
| Recurring: Watch every drone's camera at 1 frame per second, high resolution, all day (per day) | 480B 180B–2.5T | 20B 4B–40B | 676 227–3,100 around the clock | Low: Gemini 3 default 70 tokens per frame + 32 audio = 102/s × 86,400 × 20,000 = 176.3B in, 200,000 notes per stream. High: 5 fps at 280 tokens + 32 audio = 1,432/s × 86,400 × 20,000 = 2,474.5B in, 2,000,000 notes per stream.Math and sources
|
| A model makes a decision for every drone ten times per second One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total) | 7.3M 1.7M–460M | 9.7M 2.4M–380M | 0.06 0.02–2.8 to do this in one day | Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.Math and sources
|
| Recurring: One model decision per drone, ten times per second, all day (per day) | 5.2T 1.7T–17T | 350B 140B–1.4T | 8,000 2,800–28,000 around the clock | 10× the 1 Hz case. Input per call 100 / 300 / 1,000 (no neighbours to full neighbour states); output per call 8 / 20 / 80. Rate held at 10 Hz.Math and sources
|
| Every drone reasons for 1,000 tokens once per second One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total) | 7.3M 1.7M–460M | 9.7M 2.4M–380M | 0.06 0.02–2.8 to do this in one day | Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.Math and sources
|
| Recurring: Every drone reasons for 1,000 tokens, once per second, all day (per day) | 520B 170B–1.7T | 1.7T 350B–6.9T | 10,600 2,200–42,000 around the clock | Input per call 100 / 300 / 1,000, as in the 1 Hz case. Generated reasoning per call 200 / 1,000 / 4,000; CoreWeave's benchmark responses averaged over 2,000 tokens, which sits inside the band.Math and sources
|
Read a company's messages
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Label every unique message (about 100 per person) within the 8-hour workday (plan with) One-time: Read the company's kept email and chat once (about 2 years, de-duplicated) (in total) | 650B 97B–3.3T | 50B 6.4B–350B | 1,040 149–5,840 to do this in one day | Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, new tokens 30–150, wrapper 1.2–2.5× (quoted threads not re-read; earlier messages already counted), 0.5–7 retained-equivalent years (heavy deleters to full journaling).Math and sources
|
| Recurring: Label each workday's new unique messages (about 100 per person) (per workday) | 1.3B 350M–5.1B | 100M 22M–450M | 6.3 1.6–26 through the workday | Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, mean new tokens 30–150 (Kooti email mean 153 words), prompt wrapper 1.2–5× (5× keeps quoted threads), label 3–40 tokens.Math and sources
|
| Label every unique message, spread over the following 24 hours One-time: Read the company's kept email and chat once (about 2 years, de-duplicated) (in total) | 650B 97B–3.3T | 50B 6.4B–350B | 1,040 149–5,840 to do this in one day | Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, new tokens 30–150, wrapper 1.2–2.5× (quoted threads not re-read; earlier messages already counted), 0.5–7 retained-equivalent years (heavy deleters to full journaling).Math and sources
|
| Recurring: Label one workday's unique messages over the next 24 hours (weekends idle) (per day) | 1.3B 350M–5.1B | 100M 22M–450M | 2.1 0.5–8.5 around the clock | Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, mean new tokens 30–150 (Kooti email mean 153 words), prompt wrapper 1.2–5× (5× keeps quoted threads), label 3–40 tokens.Math and sources
|
| Scan every received copy with no de-duplication One-time: Read every inbox copy of kept email and chat once (about 2 years) (in total) | 1.8T 270B–8T | 140B 18B–870B | 2,810 417–14,300 to do this in one day | Judgment 95% band: log-normal mix, not every extreme at once. Received items 117–275 (low drops Teams), new tokens 30–150, wrapper 1.2–2.5× (threads not re-read), 0.5–7 retained-equivalent years, label 3–40 tokens.Math and sources
|
| Recurring: Scan all 270 received items per person each workday, duplicates included (per workday) | 3.5B 1B–12B | 270M 62M–1.1B | 17 4.5–61 through the workday | Judgment 95% band: log-normal mix, not every extreme at once. Received items 117–275 (low drops Teams), new tokens 30–150, wrapper 1.2–5×, label 3–40 tokens. Same wrapper as the de-duplicated design.Math and sources
|
| Reason for 1,000 tokens about every unique message One-time: Write 1,000 tokens about every kept unique message (about 2 years) (in total) | 650B 97B–3.3T | 5T 490B–35T | 29,700 2,950–206,000 to do this in one day | Judgment 95% band: log-normal mix, not every extreme at once. Unique 40–180/person, 200–4,000 generated tokens per message, 0.5–7 retained-equivalent years; input band as the light read.Math and sources
|
| Recurring: Write 1,000 tokens about each workday's unique messages (per workday) | 1.3B 350M–5.1B | 10B 1.6B–45B | 178 29–799 through the workday | Judgment 95% band: log-normal mix, not every extreme at once. Unique 40–180/person, 200–4,000 generated tokens (short summary to long reasoning); input band as light.Math and sources
|
| Label the messages of all 450 million paid Microsoft 365 seats One-time: Read kept unique email and chat once for all 450 million M365 seats (in total) | 2.9Q 440T–15Q | 230T 29T–1.6Q | 4.7 million 677,000–27 million to do this in one day | Light one-time band scaled by seats. Seats 450–475 million (over 450 million reported in Jan 2026; +6% for FY26). Log-normal mix, not every extreme at once; the token recipe range dominates.Math and sources
|
| Recurring: Label one global workday's unique messages for every M365 seat (24 hours) (per day) | 5.8T 1.6T–23T | 450B 99B–2T | 9,380 2,420–38,200 around the clock | Light recurring band scaled by seats 450–475 million. Log-normal mix, not every extreme at once. The Americas/Europe overlap can raise the peak to about 1.5× the 24-hour average.Math and sources
|
Examine every patient on Earth
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Daily workups for people in a care episode (5% a day), using a chart summary (plan with) One-time: Read every person's record once (12,000-token average) and write a summary (in total) | 100T 33T–460T | 17T 6.6T–100T | 211,000 76,900–1.1 million to do this in one day | Band on the global mean record, not person-level tails. Low 4,000: thinner high-income files and few-hundred-token paper cards elsewhere. High 55,600: every person has a UF Health-sized 10-year note file (82B words / 2.48M patients x 1.68 tokens/word). Summaries 800–12,000 tokens; no measured length exists.Math and sources
|
| Recurring: Daily workup for 5% of people: summary, excerpts, notes, labs (16,000 tokens) (per day) | 6.6T 1.3T–17T | 3.3T 430B–9.1T | 26,900 4,020–71,900 around the clock | Monte Carlo over driver ranges (share of people 1.1–8% a day, context 10k–40k tokens, output 2k–20k), 2.5th–97.5th percentile, instead of multiplying every extreme together.Math and sources
|
| If every care-seeker had a US emergency-room-sized chart, read in full at each workup Recurring: Daily workup for 5% of people reading a full 200,000-token chart once each (per day) | 84T 4.7T–270T | 3.3T 180B–13T | 117,000 6,550–388,000 around the clock | Share 1.1–8%. Chart 50,000–400,000 + 2,000–5,000 new; the 400k high is about 2x Patterson's ED-conditional mean, allowing for excluded labs and imaging. Generated 2,000–20,000. A 1M chart is a person-level tail, not a band on the mean.Math and sources
|
| Only people who actually see a clinician on a given day (1.5%) Recurring: Daily workup for the 1.5% of people with a clinic visit that day (per day) | 4.1T 2.5T–15T | 250B 110B–750B | 6,200 3,500–22,100 around the clock | Share 1.3–1.8% (Moses UI 4.88–5.99 / 365 to OECD 6.5 / 365). Record 20,000–100,000 + 3,000: clinic attenders have larger files than the 12,000 global mean; 100k is a US-ED-sized chart. Generated 1,000–5,000.Math and sources
|
| US-only subset: read every resident's record once, then daily summary workups for 5% One-time: Read each US resident's record once (55,600 tokens), write a 4,000-token summary (in total) | 19T 6.5T–35T | 1.4T 680B–4.2T | 30,000 11,400–64,600 to do this in one day | Population 340M (under the Census clock) to 349M (CBO Social Security area). Record 19,000 (children, people rarely in care) to 100,000 (Patterson ED-median notes plus structure applied to all). Summary 2,000–12,000.Math and sources
|
| Recurring: Daily workup for 5% of US residents: summary, excerpts, new notes and labs (per day) | 310B 38B–1.1T | 140B 7.5B–550B | 1,150 87–4,450 around the clock | Share 1.1% (one workup per care-seeking episode) to 8% (Green with a 7-day episode). Context 10,000–40,000. Generated 2,000–20,000.Math and sources
|
| Same full charts, but re-send the whole chart on each of 8 turns Recurring: Re-send the full 200,000-token chart on each of 8 turns for 5% of people a day (per day) | 670T 23T–2.7Q | 3.3T 180B–13T | 789,000 27,700–3.2 million around the clock | Share 1.1–8%. Chart 50,000–400,000 (400k is about 2x Patterson's ED-conditional mean). Passes 5–10. Low = 1.1% x (5 x 50k + 2k). High = 8% x (10 x 400k + 4k). Generated 2,000–20,000.Math and sources
|
| Everyone with a symptom today (15%), with 1-million-token records Recurring: Daily full-record workup for the 15% with symptoms, 1-million-token charts (per day) | 1.2Q 190T–3.2Q | 25T 5.3T–63T | 1.6 million 248,000–4 million around the clock | Share 8–19% (Green 800/1,000/month with 3–7 symptomatic days). Record 280,000–2,000,000 (Patterson 75th percentile notes to beyond War and Peace, plus structure) + 3,000–5,000 new. Generated 8,000–40,000.Math and sources
|
| A full-record assessment for every person on Earth, every day Recurring: Every person, every day: 510,000 tokens read once, 20,000 generated (per day) | 4.2Q 500T–8.5Q | 170T 66T–420T | 5.9 million 961,000–12 million around the clock | Coverage fixed at 100%. Record plus new 60,000–1,020,000 (the site's text-record sensitivity range, not measured percentiles). Generated 8,000–50,000 (short workup to long reasoning plus checking).Math and sources
|
| Every chronic patient (40%), 5-million-token charts, re-read 8 times a day Recurring: Daily 8-pass re-read of 5-million-token charts for 40% of people (per day) | 130Q 15Q–240Q | 66T 23T–190T | 150 million 17 million–270 million around the clock | Share 35% (global adult chronic-disease guess) to 57% (US 76.4% of adults applied to all ages). Record 1M–5M. Passes 5–10. Low = 35% x (5 x 1M + 3k). High = 57% x (10 x 5M + 5k). Generated 8,000–40,000.Math and sources
|
Argue every court case
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Every new case worldwide, argued at the depth its file needs (plan with) One-time: Once: argue the ~280 million pending cases and read the world's case law (in total) | 27T 11T–81T | 7.4T 3.2T–21T | 74,300 30,800–216,000 to do this in one day | Judgment band from combined driver ranges, not all extremes at once: pending 200–400M; traffic 35–60%; simple 20–40%; eDiscovery share 0.02–0.4%; per-case ladders as on the recurring line. Case law 200–700B in, notes 2.5–10% (5–70B out). 2.5th ≈ 10.6T/3.2T, 97.5th ≈ 81T/21T. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.Math and sources
|
| Recurring: Every year: argue the ~280 million new cases, each at the depth its file needs (per year) | 27T 10T–80T | 7.4T 3T–21T | 202 79–586 all year | Judgment band from combined driver ranges: cases 180–400M; traffic 35–60%; simple 20–40%; eDiscovery share 0.02–0.4%; per case L0 1k–10k/0.5k–5k, L1 10k–80k/5k–40k, L2 80k–500k/25k–200k, L3 8M–80M/1M–15M. 2.5th ≈ 10T/3T, 97.5th ≈ 80T/21T. Every high at once is ~150T/40T.Math and sources
|
| Every live file re-argued from scratch each year, backlog included One-time: Once: read the world's published case law, with short notes (in total) | 350B 200B–1T | 18B 4B–100B | 506 255–1,740 to do this in one day | Low: China corpus half-purged and India orders counted as short (200B). High: China at ~2,500 tokens a document on 160M documents plus a large India order stock (700B). Notes 2.5–10% of input. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.Math and sources
|
| Recurring: Every year: re-argue every live file from scratch (new filings plus backlog) (per year) | 47T 15T–160T | 13T 4.5T–42T | 353 119–1,170 all year | Low: 1.5 × the case-by-case low (10T/3T) = 15T/4.5T, more of the stock stale or suspended. High: 2.0 × the case-by-case high (80T/21T) = 160T/42T, pending ≈ new with no haircut.Math and sources
|
| 280 million cases a year with far larger token budgets per tier One-time: Once: ~280 million pending cases at generous budgets, plus world case law (in total) | 680T 420T–3.2Q | 41T 16T–290T | 1 million 582,000–5.4 million to do this in one day | Same mega-share sensitivity as the yearly line, × 280/300: 0.001% mega → 422.8T/15.9T; 0.1% mega → 3,194.8T/293.1T. Case law 200–700B in, notes 2.5–10% (5–70B out). Volume and per-tier budgets held at central. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.Math and sources
|
| Recurring: Every year: argue 280 million new cases at generous per-tier budgets (per year) | 670T 420T–3.2Q | 41T 16T–290T | 2,790 1,590–14,800 all year | Only the mega share is varied, holding 300M cases: 0.001% → 453T/17.04T; 0.1% → 3,423T/314.04T. Volume (240–400M) and per-tier budgets stay at central, so this is a sensitivity range used as the band.Math and sources
|
| Every new case gets a typical hosted eDiscovery review (10 GB) One-time: Once: a 10 GB hosted review for each of ~280M pending cases, plus case law (in total) | 13Q 1.6Q–48Q | 1.4Q 200T–6Q | 23 million 3 million–90 million to do this in one day | Low: 200M pending × 8M in / 1M out. High: 400M × 120M in / 15M out (10 GB at 12M tok/GB, or ~27 GB at 4.5M). Case law 200–700B in, notes 2.5–10% (5–70B out). World case law sized as in the law chapter: 200B–1T input, 4–100B notes.Math and sources
|
| Recurring: Every year: a 10 GB hosted eDiscovery review for each of ~280 million new cases (per year) | 13Q 1.4Q–48Q | 1.4Q 180T–6Q | 62,100 7,420–247,000 all year | Low: 180M cases × 8M in / 1M out (a 1–2 GB matter). High: 400M × 120M in / 15M out. Tokens per GB 2.5–12M: DWR's 15,000 pages per GB at 200–600 words a page. DWR is the average hosted matter, not the median filing.Math and sources
|
| Every new case gets a large-company U.S. discovery production (100 GB) One-time: Once: a ~100 GB large-company review for each pending case, plus case law (in total) | 130Q 20Q–600Q | 4.2Q 1Q–16Q | 170 million 29 million–790 million to do this in one day | Low: 200M pending × 100M in / 5M out. High: 400M × 1.5B in / 40M out (~330 GB at 4.5M or ~125 GB at 12M tok/GB). Case law 200–700B in, notes 2.5–10% (5–70B out). World case law sized as in the law chapter: 200B–1T input, 4–100B notes.Math and sources
|
| Recurring: Every year: a ~100 GB large-company review for each of ~280 million new cases (per year) | 130Q 18Q–600Q | 4.2Q 900T–16Q | 466,000 71,300–2.2 million all year | Low: 180M cases × 100M in / 5M out (~22 GB, the cheap end of RAND's sample). High: 400M × 1.5B in / 40M out (~330 GB at 4.5M or ~125 GB at 12M tok/GB). RAND's sample is large-company productions, not the median case.Math and sources
|
Call every game on TV
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Numeric win-probability model with a one-line caption per play (average day) One-time: Read historical play-by-play once for the NFL, MLB, NBA and Premier League (in total) | 5.1B 1.7B–30B | 850M 170M–5.9B | 11 2.9–68 to do this in one day | Low: sum of per-league lows (33.8M events) at 50 in / 5 out tokens. High: sum of per-league highs (73.75M) at a full 400-token nflfastR-style row and 80 out.Math and sources
|
| Recurring: One-line caption per win-probability update, 135 games at a time, no video (per day) | 9.9M 910K–48M | 4M 300K–24M | 0.03 less than 0.01–0.2 around the clock | Low: 42 games live, 30 updates/h, 30 in / 10 out tokens. High: 208 games live, 120 updates/h, 80 in / 40 out tokens.Math and sources
|
| Watch every live game at high-res on an average day, after watching four league archives One-time: Read four leagues' play-by-play and watch their archived games once (in total) | 240B 92B–370B | 1B 180M–6.9B | 289 108–465 to do this in one day | Play-by-play low/high: 33.8M events × 50 tokens / 73.75M × 400. Low 86,890 h (MLB only 3 recent seasons, NFL 4,631+960, NBA 17,000, PL 7,600 games) at 290 tok/s. High 260,882 h (full vaults: NFL 6,000, NBA 20,000, MLB 57,500, all 13,166 PL games) at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Recurring: Watch about 135 live games at high-res 1 fps and caption every play (per day) | 3.6B 1.1B–5.5B | 4M 300K–24M | 4.2 1.2–6.5 around the clock | Video low: 42 games live (200,000 premium events a year) at 290 tok/s; high: 208 (1 million matches, all with video) at 300 tok/s. Captions as in the stats-only design: 30 to 120 updates/h, 30/10 to 80/40 tokens. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Saturday peak: watch 500 games at once and caption every play (plan with) One-time: Read historical play-by-play once to seed the per-play captions (in total) | 5.1B 1.7B–30B | 850M 170M–5.9B | 11 2.9–68 to do this in one day | Low: sum of per-league lows (33.8M events) at 50 in / 5 out tokens. High: sum of per-league highs (73.75M) at a full 400-token nflfastR-style row and 80 out.Math and sources
|
| Recurring: Watch 500 games at once at high-res 1 fps, held all day, and caption every play (per day) | 14B 5B–26B | 15M 1.4M–120M | 16 5.8–31 around the clock | Video low: 200 games at once at 290 tok/s; high: 1,000 games at 300 tok/s. Captions low: 200 games, 30 updates/h, 30 in / 10 out tokens; high: 1,000 games, 120 updates/h, 80 in / 40 out. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Saturday peak: watch 500 games and have a language model re-reason every play One-time: Read historical play-by-play once as full 400-token state rows (in total) | 17B 14B–30B | 3.4B 2.7B–5.9B | 39 31–68 to do this in one day | Tokens per row held at 400 in / 80 out. Low: sum of per-league lows (33.8M events). High: sum of per-league highs (73.75M events).Math and sources
|
| Recurring: Watch 500 games at once, held all day, and have an LLM re-reason every play (per day) | 15B 5.3B–32B | 480M 88M–1.5B | 20 6.6–45 around the clock | Play rate held at 61/h. Low: 200 games at 290 tok/s, 1,000 in / 300 out per play. High: 1,000 games at 300 tok/s, 4,000 in / 1,000 out per play. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Peak, plus an LLM call on every logged on-ball event One-time: Read all 14 billion Opta event data points and watch four league archives once (in total) | 1.4T 300B–3.1T | 17B 4.2B–35B | 1,670 373–3,770 to do this in one day | Opta low: 15 tokens per point (a point is one field), 500-token match summaries; high: 200 tokens per point, 4,000-token summaries. Footage low: 86,890 h at 290 tok/s; high: 260,882 h of full league vaults at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Recurring: Watch 500 games at once, held all day, with an LLM call on every on-ball event (per day) | 19B 6B–43B | 1.9B 290M–8.4B | 34 8.6–98 around the clock | Low: 200 games at 290 tok/s, 200 events/h, 1,000 in / 300 out per event. High: 1,000 games at 300 tok/s, 350 events/h (soccer-heavy mix at ~1,650 events a match), 2,000 in / 1,000 out. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Scale-up: every data-covered match, 3,000 feeds at once, an update every 45 seconds One-time: Read all Opta event data and watch a decade of betting streams (hypothetical) (in total) | 9.3T 3.1T–30T | 25B 8.2B–55B | 10,900 3,670–35,100 to do this in one day | Opta as in the every-touch design (15 to 200 tokens per point). Video low: 5 years × 400,000 streams × 1.4 h at 290 tok/s. High: 15 years × 700,000 × 2.4 h at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
| Recurring: Watch 3,000 feeds at once, held all day, with a detailed update every 45 seconds (per day) | 92B 29B–150B | 2.9B 580M–9.6B | 124 37–228 around the clock | Cadence held at one update per 45 seconds. Low: 1,000 feeds at 290 tok/s, 300 generated per update. High: 5,000 feeds at 300 tok/s, 1,000 generated per update. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).Math and sources
|
Rewrite every app
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| Port every unique App Store and Google Play app once, then keep up (plan with) One-time: Port every unique app once: read its code, write the new app and tests (in total) | 1.5T 300B–5.8T | 770B 170B–3.1T | 6,220 1,330–24,700 to do this in one day | Driver ranges: unique apps 3.0–4.5M, 3,000–40,000 lines per app, 6.5–11 tokens per line, input 2–7× and output 1.4–4× source, combined in log space rather than stacked. Writing the new code is about 70% of GPU time, so lines per app and the output multiple move the number most.Math and sources
|
| Recurring: Port each day's ~25,000 app updates and ~2,400 new apps (new tokens only) (per day) | 1.5B 420M–6.4B | 660M 220M–2.6B | 5.6 1.7–22 around the clock | Updates 12,000–45,000 a day, 15,000–150,000 new input and 8,000–50,000 output per update; new apps 1,500–4,000 a day at 2,000–20,000 lines. Each part's ranges are combined in log space, then the two lows and two highs are added.Math and sources
|
| A heavy agent loop with tests and repeated review for every app and update One-time: Rewrite every unique app with a heavy loop: explore, rewrite, test, review (in total) | 3.8T 680B–20T | 3.1T 570B–16T | 22,100 4,060–117,000 to do this in one day | Same app, line and token ranges as the one-app port, with input 4–30× and output 3.5–25× source, combined in log space. Epoch's MirrorCode rebuilds without source used 280M tokens for a 17,000-line program, well above this band per app.Math and sources
|
| Recurring: Heavy loop on each day's ~25,000 updates and ~2,400 new apps (per day) | 4.3B 1B–22B | 3B 820M–15B | 23 5.9–110 around the clock | Update and new-app ranges as in the one-app port, plus update multiples of 1.5–6× input and 2.5–10× output and new-app multiples of 4–30× and 3.5–25×. Each part is combined in log space; lows and highs are then added.Math and sources
|
| Only the small apps: 87% of apps but about a quarter of the code One-time: Translate the ~3M tiny and small apps once; skip medium to mega apps (in total) | 170B 40B–490B | 110B 26B–320B | 838 198–2,440 to do this in one day | Small apps 2.6–3.9M, 800–8,000 mean lines, 6.5–11 tokens per line, input 1.5–3× and output 1.1–2× source (near-direct translation), combined in log space.Math and sources
|
| Recurring: Port each day's ~12,500 small-app updates and ~2,100 new small apps (per day) | 370M 100M–1.3B | 150M 45M–490M | 1.3 0.4–4.3 around the clock | Updates 6,000–25,000 a day at 8,000–60,000 input and 3,000–15,000 output; new small apps 1,300–3,500 a day at 800–8,000 lines. Each part combined in log space; lows and highs added.Math and sources
|
Advise every trader
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| A model checks each order a person places, with no calls on algorithmic orders (plan with) One-time: Once: read SEC EDGAR, other countries' filings and a finance news archive (in total) | 390B 210B–1.1T | 20B 8B–67B | 564 289–1,630 to do this in one day | Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.Math and sources
|
| Recurring: Reason once before each of ~120 million orders people place per trading day (per day) | 840B 200B–4.5T | 240B 42B–1.1T | 2,360 475–11,900 around the clock | Decisions (60–250 million a day) and tokens per decision (2k–32k in, 400–8k out) combined as independent uncertainties, not the product of both extremes. Low is fewer, smaller orders with a short check; high counts every small A-share lot and a long reasoning pass.Math and sources
|
| An agent briefs each active trader every 15 minutes while their markets are open One-time: Once: the same filings and news archive as the advice design (in total) | 390B 210B–1.1T | 20B 8B–67B | 564 289–1,630 to do this in one day | Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.Math and sources
|
| Recurring: Brief each of 40 million active traders every 15 minutes in their market hours (per day) | 5.1T 1.5T–22T | 510B 110B–2.6T | 8,890 2,350–40,500 around the clock | Traders (20–100 million), digest size (1,500–12,000 in, 100–1,500 out) and weighted session (6–10 hours) combined as independent uncertainties at the fixed 15-minute cadence. Faster or slower cadences are different designs, not uncertainty.Math and sources
|
| A model is called on every order and cancel, including algorithmic orders One-time: Once: the same filings and news archive; market ticks are not read as text (in total) | 390B 210B–1.1T | 20B 8B–67B | 564 289–1,630 to do this in one day | Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.Math and sources
|
| Recurring: Call the model on each of ~9.75 billion orders and cancels per trading day (per day) | 24T 3.1T–160T | 1.9T 160B–15T | 39,500 4,510–272,000 around the clock | Fills (400 million to 1.5 billion a day; low = China 142M + NSE 85M + US 150M + crypto), orders per fill (6–50) and per-call tokens (400–8,000 in, 20–800 out) combined as independent uncertainties. Excludes raw quote updates, which outnumber orders.Math and sources
|
Give every home a robot helper
| Design · cost line | Tokens read | Tokens written | GPUs (10K / 2K) | Why this range |
|---|---|---|---|---|
| One robot per home, with its planning done in a datacenter (plan with) One-time: Scan each home once and take a guided tour with the family (in total) | 2.5Q 930T–93Q | 88T 16T–670T | 3.4 million 1.2 million–110 million to do this in one day | Low is a quick walkthrough of a small home and a short tour. Central is a ~45-minute 3D scan and a 2-hour tour. High is a three-camera high-res scan of a large home plus a four-hour two-camera tour. Tour length and note sizes are assumptions.Math and sources
|
| Recurring: Per day: listening, cloud planning while the robot moves, and conversation (per day) | 45Q 9.1Q–190Q | 5.7Q 1.2Q–14Q | 85 million 17 million–300 million around the clock | Monte Carlo over the stated driver ranges (homes, awake hours, planning cadence 2–10 Hz, tokens per tick and per second of motion), 2.5th–97.5th percentile, instead of multiplying every extreme together.Math and sources
|
| Every person gets a robot that follows them One-time: Scan each home and spend half a waking day learning each person (in total) | 25Q 8.5Q–360Q | 150T 35T–900T | 30 million 10 million–420 million to do this in one day | Low is a short four-hour camera session per person with a small home scan. Central is half a waking day per person. High is a full waking day with two high-res cameras. Everyone, children included, is in every band.Math and sources
|
| Recurring: Per day: two cameras and audio following each person, plus cloud reasoning (per day) | 730Q 100Q–8,300Q | 140Q 8.3Q–1,100Q | 1.7 billion 170 million–16 billion around the clock | Low is one 2 fps stream and a slow small planner for 14 h. Central is two cameras at 5 fps plus 2 Hz reasoning for 16 h. High is four cameras at 10 fps, high-res, 5 Hz reasoning for 18 h. Bands add each part’s extremes, so they are wider than 95%.Math and sources
|
| Robots run their models on board and call the cloud only for hard questions One-time: Upload a compact home map and a preference interview to the cloud (in total) | 290T 60T–5Q | 62T 14T–550T | 691,000 150,000–9 million to do this in one day | Low is a room list and a short form. Central is a detailed text map of rooms and objects plus a long interview. High is a near-complete 3D copy of the home stored as tokens. All sizes are assumptions.Math and sources
|
| Recurring: Per day: cloud calls when language or a stuck task needs a larger model (per day) | 97T 8.1T–2Q | 88T 6T–1.9Q | 624,000 44,100–13 million around the clock | Call rate is an assumption. Low is a quiet home where on-robot models handle nearly everything. Central is several real questions per waking hour. High is a cloud assistant consulted every few minutes with long reasoning.Math and sources
|
| Every person, four high-res cameras at 30 fps and nonstop 10 Hz reasoning One-time: Dense photogrammetry of every home plus a million-token file on every person (in total) | 190Q 5Q–890Q | 640T 140T–1.8Q | 220 million 6.5 million–1 billion to do this in one day | Low reuses the ordinary home scan and a half-million-token file per person. Central is a two-hour four-camera high-res scan. High is four hours with six cameras at 15 fps and a two-million-token file. File sizes are assumptions.Math and sources
|
| Recurring: Per day: four 30 fps high-res cameras, audio and 10 Hz reasoning per person (per day) | 28,000Q 870Q–51,000Q | 1,400Q 76Q–5,800Q | 40 billion 1.4 billion–92 billion around the clock | Low is two default-res cameras at 10 fps for 16 waking hours with slow reasoning. Central is four high-res cameras at 30 fps around the clock. High adds two cameras and 2,000-token prompts. Bands add each part’s extremes.Math and sources
|
Sources. Driving vLLM WideEP and Large-Scale Serving Toward Maturity on Blackwell (Part I) (vLLM Blog, 2026-02-03); CoreWeave Leads MLPerf 0.7 Endpoints Benchmark with DeepSeek-R1 (CoreWeave); Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP (Part II) (LMSYS, 2025-09-25); MLPerf Inference v6.0 results (MLCommons, 2026-04-01); InferenceX: Kimi K3 on B200 (SemiAnalysis, 2026-09-12); AI Chip Sales data explorer (Epoch AI, 2026-08-27); US Code growth 1991–2025 (arXiv:2511.13747, 2025-11-13); United States Code (GovInfo); Economic Report of the President, 2026 (Chapter 2) (Council of Economic Advisers, 2026); StateCodes: all statutory code from 50 U.S. states (Hariri & Ho, arXiv:2508.19365, 2025); Caselaw Access Project (Common Pile) (Hugging Face); Video understanding (Gemini API docs, Google, 2026-09-10); Media resolution (Gemini API docs, Google, 2026-09-02); RTVI-VLM Performance (NVIDIA VSS docs); The world's most surveilled cities (citing IHS Markit) (Comparitech); Streamlabs and Stream Hatchet Q2 2026 Live Streaming Report (Streamlabs, 2026); Automatic attack disruption in Microsoft Defender XDR (Microsoft Learn, 2026-06-11); EPS calculator (Logmanager); How Elastic InfoSec optimizes Elastic Defend (Elastic Security Labs, 2026-01-27); Dropzone AI (catalog entry) (Risky Business); Microsoft Digital Defense Report 2025 (Microsoft, 2025); Offboard Mode (PX4 Guide); Multicopter control architecture diagram (PX4 Autopilot); Most multirotors airborne simultaneously from a single computer (outdoors) (Guinness World Records, 2026-02-03); GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv:2503.14734); Breaking down the infinite workday (Microsoft Work Trend Index, 2025-06-17); Evolution of Conversations in the Age of Email Overload (Kooti et al., WWW 2015); Microsoft FY26 Q2 results (Office 365 for IT Pros, 2026-01-30); World Population Prospects 2024 (United Nations, 2024-07-11); The Ecology of Medical Care Revisited (Green et al., New England Journal of Medicine, 2001); Funding and services needed to achieve universal health coverage (Moses et al., Lancet Public Health, 2019); Health at a Glance 2025 (OECD, 2025); The size of the electronic health record at the point of care (Patterson et al., JAMIA Open, 2024); Health workforce levels and trends 2026 (World Health Organization, 2026); Court Statistics Project data (National Center for State Courts); State courts play a key role in American life (Pew Charitable Trusts, 2025-04); Supreme People's Court work report (Supreme People's Court of China, 2026-03-09); Pendency in Indian courts (Data for India); Justiça em Números 2026 (recap) (JUDIT, citing CNJ, 2026-06); Where the Money Goes: Understanding Litigant Expenditures for Producing Electronic Discovery (RAND Institute for Civil Justice, 2012); How many documents in a gigabyte? 2025 statistics for eDiscovery (Digital WarRoom, 2025); Federal Rule of Appellate Procedure 32 (Cornell LII); Live Streams (Sportradar); Coverage and data rights (Stats Perform); A decade of NFL Next Gen Stats innovation (Amazon Science); nflfastR: win probability models (nflverse); Why do you have a delay on placing bets on a market that is in-play? (Betfair Developer Support); Wikipedia:Size of Wikipedia (Wikipedia, 2026-09-13); Books of the world, stand up and be counted! (Inside Google Books, 2010-08-05); Will we run out of data? Limits of LLM scaling based on human-generated data (Epoch AI, 2024); Sundar Pichai at I/O 2026 (Google, 2026-05-19); Are women really more talkative than men? (Mehl et al., Science, 2007); A large language model for electronic health records (GatorTron) (Yang et al., npj Digital Medicine, 2022); Oral health fact sheet (World Health Organization, 2025-03-17); DeepSeek-R1 config.json (Hugging Face); Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures (arXiv:2505.09343, 2025); Jalapeño first results (OpenAI, 2026-08-25); OpenAI Jalapeño ASIC at Hot Chips 2026 (ServeTheHome, 2026-08-25).