TokenSizing

How Many GPUs
Would It Take?

Token budgets for eight enormous jobs, from reading every U.S. law to examining every patient on Earth.

Research by grok 4.6 agents, method review by GPT-6 Pro, built with Claude · September 2026

Reading everything people have written, once, is cheap. Re-reading, watching and thinking about everything, all the time, is not. A single GPU could read every federal statute in 1.3 GPU-hours. A daily medical workup for everyone in a care episode takes about 26,900 GPUs running around the clock. And the one job here that rivals the entire projected 2030 chip supply sounds the most ordinary: a robot helper in every home, which needs 85 million GPUs.

What separates those answers isn’t the size of the pile. It’s how many things a model tends to at once, how often it has to process them again, and how much it writes each time. Change the design and the same job moves from a handful of GPUs to thousands.

Below, eleven large jobs are converted into tokens, the word-sized chunks models read and write, and then into GPUs. Each job has a one-time cost, such as reading every record once, and a recurring cost, such as the daily workups that follow. Every number updates when you change the serving speed or the one-time deadline, and every estimate carries a 95% range.

The Short Answer

Nine of the eleven jobs could run every day on a few thousand GPUs or fewer. Medicine needs tens of thousands, and a robot in every home needs tens of millions. One-time costs are larger but finite: reading every medical record on Earth takes 211,000 GPUs to do this in one day, and rewriting every app in both stores takes 6,220 GPUs.

Show
One-time deadline
Scale
Click a job to hide or show it; the axis rescales to what is visible.

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more.

0.010.11101001K10K100K1M10M100M1B10BRead English Wikipedia (one-time)Read every book (one-time)Google's AI trafficBlackwell GPUs shipped (mid-2026)All AI chips (H100-equiv., mid-2026)Projected AI chips, 20300.3 onceless than 0.01 2.1 onceless than 0.01 34 onceless than 0.01 506 once0.2 0.5 once0.02 71,900 once2,400 400 million once13 million 1.3 million once480 405 once5.7 0.06 once8 1,040 once6.3 4.7 million once9,380 211,000 once26,900 74,300 once202 11 once16 10,900 once124 6,220 once5.6 564 once2,360 3.4 million once85 million

How to read this chart. Each job gets up to two bars. A solid bar is recurring work, counted as GPUs busy through its window: around the clock for most jobs, a workday for company messages, all year for courts. An outlined bar is one-time work, counted as the GPUs needed to finish it by the deadline you pick. The whisker on each bar is our 95% range on the job’s token size; click one to jump to its estimate table. Click a job’s name to hide it and the axis rescales to what remains.

Eleven Enormous Jobs

Each chapter shows several ways to do the same job, with one-time and recurring costs side by side. The underlined design is the one we would plan with. Hide designs to rescale the chart; hover for tokens; click a bar for the math.

Workload 01Read all U.S. federal law

A single GPU could read every federal statute and regulation in less than a day, and keeping up with new law costs almost nothing. The U.S. Code takes 1.3 GPU-hours to read; adding the Code of Federal Regulations brings the job to 6.9 GPU-hours.

The U.S. Code runs 24.4 million words, which a modern tokenizer turns into about 37 million tokens, and the regulations add more than 100 million words. GPUs read in parallel, and reading is the cheap half of their work. The recurring cost is the daily Federal Register and new public laws: a few hundred thousand tokens a day, or less than 0.01 GPUs.

Most legal text is case law. Every published American opinion comes to about 19 billion tokens, and the world’s published opinions to roughly 350 billion, a one-time job of 12,200 GPU-hours.

Reading isn’t legal analysis. Checking every provision against every other multiplies the bill many times over. A sensible system reads once, builds an index, and pulls the relevant sections for each question.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.010.11101001K10K0.05less than 0.010.3less than 0.012.1less than 0.0134less than 0.015060.2

U.S. Code

Code + regulations

+ State law

+ U.S. case law

World opinions

The Evidence

Measured,

the U.S. Code runs 24.4 million words, the largest it has been since at least 1991. Source

2025 - arXiv:2511.13747

Measured,

legal text tokenizes at about 1.5 tokens per word, a bit worse than everyday English (1.33). Source

2025 - arXiv:2511.13747

Reported,

the Code of Federal Regulations holds over 100 million words, not counting agency guidance. Source

2026 - Council of Economic Advisers

Workload 02Watch a livestream all day

Watching one livestream all day takes a small slice of one GPU. Watching every surveillance camera on Earth takes millions. One stream at default resolution needs 0.02 GPUs around the clock, and catching up on its last 30 days takes 0.5 GPUs to do this in one day.

Video models don’t watch the way people do. Gemini samples one frame per second and turns each into 70 tokens by default, or 280 at high resolution, plus 32 tokens for every second of audio. A day of one stream is about 8.8 million tokens at default resolution and 27 million at high resolution.

At scale, the totals jump. Every Twitch, YouTube and Kick stream plus 4,000 satellite channels needs 2,400 GPUs. The world’s roughly one billion surveillance cameras need 13 million GPUs, and indexing the whole YouTube library once takes 1.3 million GPUs to do this in one day.

Two caveats apply. One frame per second misses fast motion, and naming strangers or reconstructing a day’s events is harder than describing a scene. Those are accuracy problems, not just more tokens. Video priced at text-serving rates is also an approximation: a real pipeline runs vision and audio encoders too.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.010.11101001K10K100K1M10M100M1B10B0.50.021.10.044.50.171,9002,400400 million13 million1.3 million480

Default res

High res

5 fps

All major streams

Every CCTV camera

YouTube library

The Evidence

Reported,

Gemini samples video at 1 frame per second and charges 70 tokens per frame by default, 258–280 at high resolution. Source

2026 - Gemini API docs, Google

Reported,

audio adds 32 tokens for every second of video. Source

2026 - Gemini API docs, Google

Measured,

NVIDIA runs 33 live captioning streams on one H100 with a small video model and 10-second chunks. Source

NVIDIA VSS docs

Workload 03Defend 1,000 servers

Defending 1,000 servers takes about 5.7 GPUs, as long as ordinary software condenses the logs before a model reads them. Re-checking the last 90 days once takes 405 GPUs to do this in one day.

A thousand busy servers produce around 10,000 security events a second. Send every one straight to a language model and it reads about 216 billion tokens a day, using 255 GPUs before it writes a single verdict.

Real defense tools condense first. The planning design has each server send one 2,000-token evidence packet every 100 seconds and runs 1,000 deep investigations a day, which is where the number above comes from. Elastic reports its triage agents settle at 7 to 9 model calls per investigation after tuning.

Compute isn’t the limit here. No amount of reading can find an attacker the logs never recorded, and a wrong containment decision is expensive at any GPU count.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.11101001K10K100K1M4054.54055.74054623,00025595,9001,070

Correlated only

+ Investigations

10× attack surge

Raw logs

Raw, verbose

The Evidence

Rule of thumb,

SIEM sizing tables budget about 10 events per second for a Linux server and 30 for Windows. Source

Logmanager

Measured,

a 1 KB JSON security event is about 350 tokens. Logs tokenize denser than prose. Source

2026 - Elastic Security Labs

Reported,

Microsoft Defender correlates millions of signals into incidents before it acts, rather than judging each event. Source

2026 - Microsoft Learn

Workload 04Fly 20,000 drones

Flying 20,000 drones takes almost no AI. Asking a language model to fly each one takes hundreds to thousands of GPUs, and it still wouldn’t work.

Autopilots already fly drones with plain arithmetic. In February 2026, 22,580 drones flew at once from a single computer, on paths calculated in advance. Planning such a show is a one-time job: a model writes 31 formations and the show software fills in the paths, which takes 1.5 GPU-hours.

Put a model in charge of 200 squads, briefing each every 10 seconds, and it needs 8 GPUs. Give every drone its own decision each second and it needs 800 GPUs. Let each one reason for 1,000 tokens first and it needs 10,600 GPUs.

The GPU count is the smaller problem. In a published benchmark with 16,384 simultaneous requests, a large reasoning model took almost six seconds to start answering and then wrote about 27 tokens a second per request. A drone at 10 meters a second covers 60 meters on stale orders in that time. Robot models avoid this by thinking a few times a second and passing faster motor commands to a small controller.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.010.11101001K10K100K0.0680.068000.066760.068,0000.0610,600

Squad leaders

Every drone, 1 Hz

Watch all cameras

Every drone, 10 Hz

Every drone thinks

The Evidence

Reported,

PX4 autopilots take numeric setpoints and only need a 2 Hz heartbeat from an outside controller. Source

PX4 Guide

Reported,

inside the autopilot, position runs at 50 Hz, attitude at 250 Hz and rotation rate at 1,000 Hz. Source

PX4 Autopilot

Reported,

22,580 drones flew at once from a single computer in February 2026, on flight paths computed in advance. Source

2026 - Guinness World Records

Workload 05Read a company's messages

An AI can label everything a 100,000-person company writes in a day on about 6.3 GPUs. Thinking hard about each message takes 178 GPUs. Reading the kept archive once, about two years of de-duplicated email and chat, takes 1,040 GPUs to do this in one day.

The average Microsoft 365 worker receives 117 emails and 153 Teams messages a day, but most are copies of something sent to many people, and Teams chats are copied into every participant’s mailbox. Counted once, that is about 100 messages per person. With quoted replies removed, the median email reply is 43 words.

A short label per message, such as “payroll” or “customer complaint”, adds little. Writing 1,000 tokens of reasoning about every message multiplies the output a hundredfold, and writing is the expensive half. Scaled to all 450 million paid Microsoft 365 seats, even the light version needs 9,380 GPUs.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.11101001K10K100K1M10M100M1,0406.31,0402.12,8101729,7001784.7 million9,380

Light, workday

Light, spread out

Every inbox copy

Deep reasoning

Every M365 seat

The Evidence

Measured,

the average Microsoft 365 worker receives 117 emails a day and 153 Teams messages per weekday. Source

2025 - Microsoft Work Trend Index

Measured,

with quoted text stripped, half of email replies are shorter than 43 words. Source

Kooti et al., WWW 2015

Reported,

Microsoft has more than 450 million paid commercial Microsoft 365 seats. Source

2026 - Office 365 for IT Pros

Workload 06Examine every patient on Earth

Reading every person’s medical record once is a bounded job; the daily bill depends on whether each workup re-reads it. Reading all 8.3 billion records and writing a summary of each takes 211,000 GPUs to do this in one day. After that, a daily workup for everyone in a care episode, using the summary plus new notes, takes 26,900 GPUs around the clock.

Who counts is a choice, not a measurement. On an average day about 1.5% of humanity sees a clinician. A classic study found that each month, about a third of people consider seeking care, which works out to about 5% a day if each episode lasts five days. We plan for 5% and show the other choices as separate bars.

Record length varies enormously. The median patient arriving at a U.S. emergency room in 2022 had 98,000 tokens of notes on file, and one in five had more text than Moby-Dick. Most of the world has far shorter or paper records, so a global average is closer to 12,000 tokens.

That is why the design matters more than the population. Reading a full U.S.-sized 200,000-token chart at every workup takes 117,000 GPUs. Re-sending it with each of eight questions takes 789,000 GPUs. A full assessment for every person, every day, takes 5.9 million GPUs, and the most wasteful version needs 150 million GPUs, more AI chips than exist.

One warning applies to every medical number here. The serving speeds were measured on 2,000-token prompts, and long records cost more per token as attention grows with length. A real system would read records in windows of 32,000 to 64,000 tokens and retrieve what each question needs, so these figures are reference work rather than a tested design.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

101001K10K100K1M10M100M1B211,00026,900117,0006,20030,0001,150789,0001.6 million5.9 million150 million

Summarized charts

US ER-size charts

Clinic visits

US only

Re-read each turn

Symptomatic (15%)

Everyone, daily

Literal maximum

The Evidence

Reported,

the world had about 8.3 billion people in 2026. Source

2024 - United Nations

Measured,

people worldwide make about 5.4 outpatient visits a year, roughly 1.5% of humanity on any given day. Source

2019 - Moses et al., Lancet Public Health

Measured,

each month, 800 of every 1,000 Americans have symptoms and 327 consider seeking care. Source

2001 - Green et al., New England Journal of Medicine

Workload 07Argue every court case

Arguing every new court case in the world, each at the depth its file needs, takes about 202 GPUs running all year. Clearing the existing backlog once is as much reading and writing as a full year of new cases. Arguing the roughly 280 million pending cases takes 74,300 GPUs to do this in one day.

The world opens roughly 280 million court cases a year. U.S. state courts alone receive 70 million filings, most of them traffic tickets. Chinese courts accepted 37.5 million cases in 2025, about 29% of them enforcement stages of earlier disputes, and India has a backlog of 54 million.

Most files are short. A traffic case is a page or two, and even a federal appeal brief is capped at 13,000 words. The expensive American habit is discovery, the exchange of internal documents before trial: a large corporate production reviews about 100 GB, nearly half a billion tokens per case. Give that to every case and the bill becomes 466,000 GPUs.

Token budgets are the biggest unknown. A far more generous docket, with two million tokens for every contested case and ten billion for each of the largest discovery fights, needs 2,790 GPUs.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

101001K10K100K1M10M100M1B74,3002025063531 million2,79023 million62,100170 million466,000

Case by case

Re-argue backlog

Generous budgets

All get discovery

All get big law

The Evidence

Reported,

U.S. state courts received 70 million filings in 2024; 57% of cases are traffic. Source

2025 - Pew Charitable Trusts

Reported,

Chinese courts accepted 37.5 million cases in 2025. Source

2026 - Supreme People's Court of China

Reported,

Brazil opened 40.9 million new cases in 2025 and carries 75.5 million pending. Source

2026 - JUDIT, citing CNJ

Workload 08Call every game on TV

Watching every televised game and updating win probability on every play takes about 16 GPUs at the busiest hour of a Saturday. Sports produce few decision points: a handful per second, worldwide.

Sportradar streamed over 525,000 matches in 2025. Games average under two hours, so roughly 135 are live at any moment and about 500 at a Saturday peak, when European football overlaps with American college football. A 2.5-hour game at high resolution is about 2.8 million tokens.

Production systems compute probabilities with small statistical models; NFL Next Gen Stats runs 75 of them per play in under a second. Having a language model re-reason every play instead raises the peak to 20 GPUs. Reading decades of play-by-play once to prime the captions takes 258 GPU-hours.

As a stress test, watching every data-covered match with 3,000 feeds live at once and a detailed update every 45 seconds would need 124 GPUs.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.010.11101001K10K100K110.032894.2111639201,6703410,900124

Stats only

Average day

Saturday peak

Peak + LLM

Every touch

Every match

The Evidence

Reported,

Sportradar distributes more than 650,000 live sports streams a year. Source

Sportradar

Reported,

Stats Perform covers 500,000+ matches a year across 3,900 competitions. Source

Stats Perform

Reported,

NFL Next Gen Stats runs 75 machine learning models on every play in under a second. Source

Amazon Science

Workload 09Rewrite every app

Rewriting every app in both stores is a one-time job of about 6,220 GPUs to do this in one day. Keeping them all current afterward takes about 5.6 GPUs.

The App Store lists about 2.6 million apps and Google Play about 2.6 million, but 37% of apps are on both, so there are roughly 3.4 million distinct products. Most are small, and a third haven’t been updated in two years. A few giants pull the average up: Uber’s Android codebase grew from 10,000 lines in 2008 to 10 million in 2021.

Code runs about 7 to 10 tokens per line. A coding agent reads far more than it writes, but most of those reads are cached: in 731 benchmark sessions, 97.6% of input tokens were cache hits, which cost almost nothing to serve again. We count new tokens only.

A heavy loop that explores, rewrites, tests and reviews every app takes 22,100 GPUs to do this in one day, and review is a large share: in one study of agent traces, code review used 59.4% of all tokens.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

0.11101001K10K100K1M6,2205.622,100238381.3

Port each app once

Agentic + review

Small apps only

The Evidence

Measured,

Claude Opus 4.7 averaged 3.46 million input tokens per SWE-bench Pro session, and 97.6% of them were cache reads. Source

2026 - Netpreme on Hugging Face

Reported,

Apple reviewed 9,100,620 app submissions in 2025 and rejected 2,093,244 of them. Source

2026 - Apple

Reported,

Apple's App Store saw 557,000 new app submissions in 2025, up 24% from 2024. Source

2025 - Appfigures

Workload 10Advise every trader

A model that checks every order a person places, worldwide, needs about 2,360 GPUs. Reading every filing and a decade of financial news first takes 564 GPUs to do this in one day.

Owning stock isn’t trading. China has 251 million securities investors and India 238 million demat accounts, but only 8% of India’s retail traders traded on more than 50 days last year. Counting orders people place themselves gives about 120 million a day, led by China, where Shanghai and Shenzhen logged 68.8 billion two-sided transfers in 2025.

An always-on agent that briefs each of 40 million active traders every 15 minutes during their market hours needs 8,890 GPUs. Calling a model on every order and cancel, including algorithmic flow, needs 39,500 GPUs. Market ticks themselves aren’t read as text.

These are averages across the day. Load peaks at about three times the average when Asian markets open, so a real deployment would provision for the peak.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

1001K10K100K1M5642,3605648,89056439,500

Advice per trade

15-minute agent

LLM on every order

The Evidence

Reported,

the Shanghai and Shenzhen exchanges recorded 68.848 billion two-sided share transfers in 2025, about 142 million trades a day. Source

2026 - Sina Finance

Reported,

just over 1.9 million French people bought or sold shares in 2025, making 56 million retail equity trades. Source

2026 - Autorité des marchés financiers

Reported,

investors in India held 237.7 million demat accounts in August 2026, but only about 46 million were active NSE clients. Source

2026 - Business Today

Workload 11Give every home a robot helper

A robot helper in every home, with its planning done in a datacenter, needs about 85 million GPUs running around the clock. Even on OpenAI’s faster Jalapeño chips at 12,000 tokens a second, that is 49% of all the AI chips projected for 2030. It is the only job here on that scale, and a robot for every person needs 1.7 billion GPUs.

There are about 2.2 billion households and 8.3 billion people. A robot that listens all waking day hears 32 audio tokens a second. Robot models plan a few times a second: Gemini Robotics takes about 250 milliseconds from camera images to an action chunk, and each second of motion encodes to about 30 tokens per arm.

The one-time cost is scanning each home and learning the family: 3.4 million GPUs to do this in one day. The recurring cost is what matters, and it depends on where the robot thinks. If robots run their models on board and call the cloud only for hard questions, datacenter demand drops to 624,000 GPUs. Four 30-frames-per-second cameras with nonstop reasoning for every person would need 40 billion GPUs.

Bodies, not tokens, limit this decade. Goldman Sachs projects 890,000 humanoid robots shipped in 2030, a tiny fraction of 2.2 billion homes. If the robots arrive, their datacenter demand would exceed the planning estimates of every other job on this page combined, many times over.

Show
One-time deadline
Scale

Now showing: solid bars are GPUs busy through each recurring job's window (around the clock, a workday, or all year, as labeled); outlined bars are GPUs needed to do this in one day for one-time work; whiskers are the 95% range of token size. At Blackwell planning speed (10K / 2K tokens/s per GPU), log scale, chosen automatically because the visible values span 100× or more. Underlined: the design we plan with. Click a bar for the math, a whisker for the estimate table.

10K100K1M10M100M1B10B100B3.4 million85 million30 million1.7 billion691,000624,000220 million40 billion

One per home

Robot per person

On-device brain

30 fps, always on

The Evidence

Measured,

π0-FAST takes about 750 ms to generate one 1-second action chunk, which runs to roughly 30 tokens per robot arm. Source

2025 - Physical Intelligence (arXiv)

Reported,

Gemini Robotics runs its main model in the cloud, and it takes about 250 ms to go from camera images to an action chunk. Source

2025 - Google DeepMind (arXiv)

Reported,

Figure’s Helix runs both its 7–9 Hz vision-language model and its 200 Hz motor policy on GPUs inside the robot. Source

2025 - Figure AI

What Makes a Job Big?

A job’s size depends less on how much text exists than on how often a model processes it again and how much it writes each time.

GPUs ∝ things handled at once × passes per second × tokens per pass

The map below plots every design’s recurring cost: tokens read per second across, tokens written per second up. Dashed curves mark equal GPU counts. Jobs that mostly read sit low; jobs that write, and especially jobs that reason, climb toward the costly corner, because at Blackwell planning speeds a GPU writes tokens five times more slowly than it reads them.

110010K1M100M10B1T100T0.01110010K1M100M10B1T100T1 GPU100 GPUs10,000 GPUs1 million GPUs100 million GPUs10 billion GPUsTokens read per second →Tokens written per second →
Recurring cost only. Filled dots: the design each job plans with; open dots: other designs. Dashed curves: equal GPU counts at the chosen speed.

Serving speed moves every point at once. Pick a speed to compare how much the hardware matters against how much the design does. Each row shows the design we would plan with, its one-time and recurring costs, and a button that cycles through the other designs.

Read all U.S. federal law

Read the U.S. Code and federal regulations once, then keep up with each day's new rules

One-time

0.3 GPUs to do this in one day

95% range 0.20.4

■ = 1/100 of a GPU

Reads 200M, writes 10M tokens in total

Recurring

less than 0.01 GPUs around the clock

95% range less than 0.01less than 0.01

■ = 1/100 of a GPU

Reads 370K, writes 18K tokens per day

Watch a livestream all day

One stream at Gemini's default resolution, 1 frame per second

One-time

0.5 GPUs to do this in one day

95% range 0.092

■ = 1/100 of a GPU

Reads 260M, writes 30M tokens in total

Recurring

0.02 GPUs around the clock

95% range 0.010.03

■ = 1/100 of a GPU

Reads 8.8M, writes 1M tokens per day

Defend 1,000 servers

Evidence packets plus a deep AI investigation of 1,000 alerts a day

One-time

405 GPUs to do this in one day

95% range 153,940

■ = 10 GPUs

Reads 160B, writes 39B tokens in total

Recurring

5.7 GPUs around the clock

95% range 0.241

■ = 1/10 of a GPU

Reads 2.2B, writes 530M tokens per day

Fly 20,000 drones

Autopilots fly the paths; a model briefs 200 squads every 10 seconds

One-time

0.06 GPUs to do this in one day

95% range 0.022.8

■ = 1/100 of a GPU

Reads 7.3M, writes 9.7M tokens in total

Recurring

8 GPUs around the clock

95% range 2.530

■ = 1/10 of a GPU

Reads 5.2B, writes 350M tokens per day

Read a company's messages

Label every unique message (about 100 per person) within the 8-hour workday

One-time

1,040 GPUs to do this in one day

95% range 1495,840

■ = 100 GPUs

Reads 650B, writes 50B tokens in total

Recurring

6.3 GPUs through the workday

95% range 1.626

■ = 1/10 of a GPU

Reads 1.3B, writes 100M tokens per workday

Examine every patient on Earth

Daily workups for people in a care episode (5% a day), using a chart summary

One-time

211,000 GPUs to do this in one day

95% range 76,9001.1 million

■ = 10,000 GPUs

Reads 100T, writes 17T tokens in total

Recurring

26,900 GPUs around the clock

95% range 4,02071,900

■ = 1,000 GPUs

Reads 6.6T, writes 3.3T tokens per day

Argue every court case

Every new case worldwide, argued at the depth its file needs

One-time

74,300 GPUs to do this in one day

95% range 30,800216,000

■ = 1,000 GPUs

Reads 27T, writes 7.4T tokens in total

Recurring

202 GPUs all year

95% range 79586

■ = 10 GPUs

Reads 27T, writes 7.4T tokens per year

Call every game on TV

Saturday peak: watch 500 games at once and caption every play

One-time

11 GPUs to do this in one day

95% range 2.968

■ = 1 GPU

Reads 5.1B, writes 850M tokens in total

Recurring

16 GPUs around the clock

95% range 5.831

■ = 1 GPU

Reads 14B, writes 15M tokens per day

Rewrite every app

Port every unique App Store and Google Play app once, then keep up

One-time

6,220 GPUs to do this in one day

95% range 1,33024,700

■ = 100 GPUs

Reads 1.5T, writes 770B tokens in total

Recurring

5.6 GPUs around the clock

95% range 1.722

■ = 1/10 of a GPU

Reads 1.5B, writes 660M tokens per day

Advise every trader

A model checks each order a person places, with no calls on algorithmic orders

One-time

564 GPUs to do this in one day

95% range 2891,630

■ = 10 GPUs

Reads 390B, writes 20B tokens in total

Recurring

2,360 GPUs around the clock

95% range 47511,900

■ = 100 GPUs

Reads 840B, writes 240B tokens per day

Give every home a robot helper

One robot per home, with its planning done in a datacenter

One-time

3.4 million GPUs to do this in one day

95% range 1.2 million110 million

■ = 100,000 GPUs

Reads 2.5Q, writes 88T tokens in total

Recurring

85 million GPUs around the clock

95% range 17 million300 million

■ = 1 million GPUs

Reads 45Q, writes 5.7Q tokens per day

So how many GPUs would it take? For most of these jobs, a few thousand at most. For jobs that re-read long records for billions of people, watch every camera, or give every home a robot, more than the world has.

The pattern holds across all eleven. A stock of text is finite: even the useful public web is a one-time read that a large cluster finishes in weeks to months. A flow never stops. Work that runs forever, for many people or devices, re-reading long records or writing at length, is what data centers are built for.

For scale: Google said in May 2026 that it processes 3.2 quadrillion tokens a month, which at these speeds is about 222,000 GPUs working around the clock. Epoch AI counts about 24 million AI accelerators shipped by mid-2026, and projections for 2030 center on about 100 million. Only the robot job, watching every camera, and the most wasteful version of medicine reach even a tenth of that.

Hardware will keep getting faster. If OpenAI’s Jalapeño chip delivers 12,000 blended tokens a second, it trims reading-heavy jobs and cuts writing-heavy ones sharply. It doesn’t change which jobs are big. Design does: cache the chart, condense the logs, let the autopilot fly, and let the robot think on board.

Data and Method

What a token is. Models read and write text in tokens, chunks of about three-quarters of an English word. We count tokens read (input, or “prefill”) separately from tokens written (output, or “decode”, which includes hidden reasoning), because a GPU writes tokens several times more slowly than it reads them.

The conversion. GPU-seconds = tokens read ÷ read speed + tokens written ÷ write speed. GPUs = GPU-seconds ÷ the seconds available: a day for ongoing work, eight hours for a workday, a year for court caseloads. One-time reads are shown as GPU-hours in their chapter and, marked “once” with outlined blocks, as GPUs needed to finish within a day on the ladder. Only recurring work describes GPUs running continuously. For the blended Jalapeño tier, GPU-seconds = all tokens ÷ 12,000.

Serving speeds. Dense model (4K / 800): A slower reference rate for dense 100B–400B-class models, long contexts, or fast per-user speeds. Llama 3.1 405B serves ~140–270 generated tokens/s per GPU in MLPerf. Blackwell planning (10K / 2K): A fixed reference rate for a large open mixture-of-experts model at ~50–100 tokens/s per user, discounted from short-prompt benchmarks: 38% of measured peak prefill and 57% of MLPerf interactive decode. Not a guarantee for long contexts. Blackwell optimized (20K / 6K): Disaggregated serving with 4-bit weights on NVL72 racks at relaxed latency. Near SGLang's 26K prefill and CoreWeave's 6.5K decode per GPU. Future: OpenAI Jalapeño (12K blended): OpenAI's first custom inference chip (with Broadcom), unveiled June 2026. Assumed 12,000 blended tokens/s per chip for input and output alike. OpenAI reports mixed tokens/s per kW on an 8K-in/1K-out test; ~12,700/s per 700 W chip is derived, not stated. The blend is input-heavy, so it flatters generation-heavy work.

Long records. The serving speeds come from benchmarks with 2,000-token prompts. Attention cost grows with prompt length, and memory for a model’s working cache grows with it too: about 35 GB for one 500,000-token sequence on DeepSeek’s compressed attention. There is no single honest correction factor, so the medical and legal numbers assume records are read in bounded windows of 32,000 to 64,000 tokens with extraction and retrieval, and should be read as reference work, not throughput predictions for giant prompts. Video is priced as text-token-equivalent work; a real pipeline also runs vision and audio encoders.

What the numbers are not. They are GPU-equivalents of work, not a purchase order. A large model must fit in memory, and the vLLM benchmark behind our planning rate used 16 GPUs for one serving setup, so a small job can still need a whole rack. They ignore latency: a stack that meets throughput can still miss a 100-millisecond control deadline. They ignore power, networking, storage, retries and failed runs. A stack ten times slower makes every number ten times larger.

How the estimates were checked. Ten research agents (grok 4.6 at extra-high reasoning) re-derived every figure in the original brief against primary sources in September 2026, marking each claim confirmed, adjusted, unverifiable or an assumption. Three new workloads (medicine, courts, sports) were built from scratch the same way. A GPT-6 Pro method review then pushed back. It flagged that medical “need” is a prevalence choice rather than a consultation rate, that short-prompt benchmark speeds do not transfer to 500,000-token records, that court token budgets matter more than case counts, and that one-time reads must stay visibly one-time; the medical ceiling, the generous court docket, the sports stress peak and the long-record caveats come from that review. Corrections from the research pass include Gemini’s default video rate (70 tokens per frame, not 258), 10,200 GPUs rather than 10,000 for reasoning drones, 1.5 rather than 1.6 tokens per legal word, and Sportradar’s 650,000 streams a year. Every figure lives in one typed data file; a script recomputes all of them at every speed and fails if any drifts from the research.

Why no confidence intervals. The dominant unknowns are choices, not sampling error: which model, how much reasoning, which deadline, whether a chart is cached. Each chapter shows those choices as separate bars instead.

95% ranges. Each cost line has a range on its token size: a low end we would expect to be undercut about one time in 40 and a high end exceeded about one time in 40. They are judgment-based, built from the sourced low and high values of each driver (people, frequency, tokens per item), not statistical confidence intervals. The GPU range follows from the token range at the chosen speed.

One-time and recurring. One-time work (reading every record once, rewriting an app, clearing a backlog) is shown with outlined bars as the GPUs needed to finish by a deadline: one day, 30 days or one year. Recurring work (daily workups, per-play updates, continuous perception) is shown with solid bars as GPUs running for its whole window.

Estimate tables

Every job, design and cost line, with token ranges and GPUs at the speed and deadline selected above.

Read all U.S. federal law

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Read the U.S. Code once with short notes, then keep up with new public laws
One-time: Read the 24.4 million words of the U.S. Code once, with short notes (in total)
37M
35M48M
1.8M
710K4.8M
0.05
0.050.08
to do this in one day
Low uses the operative word count at 1.45 tokens/word, the lowest measured on Code text, with 2% notes. High uses a broader 28.8M-word file-text count at 1.65 tokens/word and 10% notes.
Math and sources
  1. 24,409,969 words of operative federal law (U.S. Code, published 2025)
  2. × 1.50 tokens/word (measured, o200k) = 36,614,954 input tokens
  3. × 5% notes = 1,830,748 output tokens
  4. Low: 24,409,969 × 1.45 = 35,394,455; notes 2%
  5. High: 28.8M words (file-text count, +18%) × 1.65 = 47,520,000; notes 10%
  • arxiv.org (“The latest version of the US Code —published in 2025— is the largest since 1991, encompassing over 24.4 million words.”)
  • govinfo.gov (“Measured with o200k: U.S. Code title 15 about 1.45–1.50 tokens/word; title 18 about 1.63.”)
  • govinfo.gov (“The United States Code, is the codification by subject matter of the general and permanent laws of the United States.”)
Recurring: Keep up each day: new public laws in the Statutes at Large (per day)10K
6K27K
515
1202.7K
less than 0.01
less than 0.01less than 0.01
around the clock
Low is the recent pace: the 118th Congress enacted 4,350 pages over two years and 119-1 enacted 2,008, about 1.5–1.6M words a year. High treats a busy Congress's 6M words as one year (117th: 8,742 pages over two years).
Math and sources
  1. Central: 2.5M new words/yr (midpoint of GovTrack's 4–6M per two-year Congress)
  2. × 1.50 tokens/word = 3.75M tokens/yr
  3. ÷ 365.25 = 10,300 input tokens/day; + 5% notes = 513
  4. Low: 1.5M words/yr (118th and 119-1 pace: ~2,000–2,200 pages/yr × 740) × 1.45 / 365.25 = 6,000
  5. High: 6.0M words/yr (a busy Congress's words in one year) × 1.65 / 365.25 = 27,100
  • govtrack.us (“Since World War II (the earliest we have data), Congress has typically enacted 4-6 million words of new law in each two-year Congress.”)
  • archives.gov (“Total number of public laws: 72. Total number of pages: 2,008.”)
  • archives.gov (“Total number of public laws: 240. Total number of pages: 3,236.”)
Read the U.S. Code and federal regulations once, then keep up with each day's new rules (plan with)
One-time: Read the U.S. Code and the Code of Federal Regulations once, with short notes (in total)
200M
180M250M
10M
3.6M25M
0.3
0.20.4
to do this in one day
CFR words 100–122M: 100M is the White House floor; 194,395 pages at ~600 words/page gives about 117–122M. Tokens/word 1.45–1.65, as measured on CFR volumes. Code 35.4–47.5M tokens. Notes 2–10%.
Math and sources
  1. U.S. Code: 36.6M tokens
  2. CFR: 107M words × 1.50 tokens/word ≈ 160M tokens
  3. 36.6M + 160M ≈ 197M, rounded to 200M input tokens
  4. × 5% notes = 10M output tokens
  5. Low: 100M CFR words × 1.45 + 35.4M Code = 180.4M
  6. High: 122M CFR words × 1.65 + 47.5M Code = 248.8M
  • whitehouse.gov (“At the latest count, the CFR contains over 100 million words of regulatory text, and this does not include agency guidance documents.”)
  • raw.githubusercontent.com (“2024,245,194395”)
  • quantgov.org (“it would take the average person over three years to read the entire US Code of Federal Regulations (CFR), which contains more than 100 mill…”)
  • arxiv.org (“The latest version of the US Code —published in 2025— is the largest since 1991, encompassing over 24.4 million words.”)
Recurring: Keep up each day: new federal laws and the Federal Register (per day)370K
250K540K
18K
5K54K
less than 0.01
less than 0.01less than 0.01
around the clock
FR page counts: 60,917 (2025, lowest) to 107,262 (2024 gross, highest), 82,730 central. Tokens/page 1,450–1,750 spans five measured issues (1,504–1,704). Law band as in the U.S. Code design. Notes 2–10%.
Math and sources
  1. Laws: 10,300 input tokens/day (as in U.S. Code design)
  2. FR: 82,730 pages/yr (2020s average) × 1,580 tokens/page (measured) = 130.7M tokens/yr
  3. FR: ÷ 365.25 = 357,900 input tokens/day
  4. Total 368,100 input tokens/day; + 5% notes = 18,400
  5. Low: laws 6,000 + FR 60,917 pages × 1,450 / 365.25 = 241,800 → 247,800
  6. High: laws 27,100 + FR 107,262 pages × 1,750 / 365.25 = 513,900 → 541,000
  • cei.org (“The 2025 Federal Register closed out at 60,917 pages. ... Six years into the 2020s ... the average is 82,730 annually.”)
  • brookings.edu (“2023 90,402 2024 107,262”)
  • federalregister.gov (“Measured on 5 issues (API raw text, o200k): 933–1,007 words/page, 1.52–1.69 tokens/word, 1,504–1,704 tokens/page.”)
  • govtrack.us (“Since World War II (the earliest we have data), Congress has typically enacted 4-6 million words of new law in each two-year Congress.”)
  • archives.gov (“Total number of public laws: 72. Total number of pages: 2,008.”)
Add all 50 states' statutes and regulations to federal law
One-time: Read federal law plus all 50 states' statutes and administrative codes once (in total)
1.5B
1.3B1.8B
73M
26M180M
2.1
1.73.2
to do this in one day
Low scales StateCodes' 268M words of regulations for 45 states to 50 states (298M) and uses 1.45 tokens/word. High uses 550M words of statutes, State RegData's 416M words of regulations and 1.65 tokens/word. Notes 2–10%.
Math and sources
  1. Federal Code + CFR: 200M tokens
  2. 50-state statutes: 488M words × 1.50 = 732M tokens
  3. 50-state administrative codes: ~350M words × 1.50 = 525M tokens
  4. Total 1.457B input tokens; + 5% notes = 72.85M output
  5. Low: 180.4M + 488M × 1.45 + 298M × 1.45 = 1.320B
  6. High: 248.8M + 550M × 1.65 + 416M × 1.65 = 1.843B
  • arxiv.org (“StateCodes, which includes all statutory code from 50 U.S. states, containing 488M words ... state regulations from 45 U.S. states, containi…”)
  • whitehouse.gov (“At the latest count, the CFR contains over 100 million words of regulatory text, and this does not include agency guidance documents.”)
Recurring: Keep up each day: new federal and state laws and regulations (per day)640K
330K1.7M
32K
6.5K170K
less than 0.01
less than 0.01less than 0.01
around the clock
Bill counts are sourced (Ballotpedia ~18,300 a year; Quorum 15,951 in the first half of 2025). Words per bill, 800–8,000, is assumed. State regulation flow is assumed at 3–12% a year of the stock. Federal lines as above.
Math and sources
  1. Laws 10,300 + FR 357,900 input tokens/day (as above)
  2. State bills: 18,000 enacted/yr × 2,500 words (ASSUMPTION) × 1.50 / 365.25 = 184,800
  3. State regs: 20M new words/yr (ASSUMPTION) × 1.50 / 365.25 = 82,100
  4. Total 635,100 input tokens/day; + 5% notes = 31,750
  5. Low: 6,000 + 241,800 + 12,000 × 800 × 1.45/365.25 + 10M × 1.45/365.25 = 325,600
  6. High: 27,100 + 513,900 + 25,000 × 8,000 × 1.65/365.25 + 50M × 1.65/365.25 = 1,670,400
  • ballotpedia.org (“The report is based on a dataset of 243,529 bills adopted from January 2011 to May 2024 ... The mean number of bills passed per year was 325…”)
  • quorum.us (“So far in 2025, state lawmakers have introduced 130,120 bills (excluding carryover) and enacted 15,951 of them.”)
  • arxiv.org (“state regulations from 45 U.S. states, containing 268M words across 2.2M sections”)
  • cei.org (“The 2025 Federal Register closed out at 60,917 pages. ... Six years into the 2020s ... the average is 82,730 annually.”)
  • federalregister.gov (“Measured on 5 issues (API raw text, o200k): 933–1,007 words/page, 1.52–1.69 tokens/word, 1,504–1,704 tokens/page.”)
  • govtrack.us (“Since World War II (the earliest we have data), Congress has typically enacted 4-6 million words of new law in each two-year Congress.”)
Add every published American court opinion to all U.S. statutes and regulations
One-time: Read all U.S. statutes, regulations and published court opinions once (in total)
23B
18B32B
1.2B
370M3.2B
34
2355
to do this in one day
Case law 17–30B tokens: low is near the 6.92M-document Common Pile set (78 GB ≈ 19.5B); high allows CourtListener's 'over nine million' decisions at longer length. Stacked on the all-codes band. Notes 2–10%.
Math and sources
  1. Common Pile: 78 GB / 6.92M documents ≈ 2,650 tokens per opinion
  2. CourtListener 8.3M precedential opinions × ~2,650 ≈ 22B tokens
  3. + all statutes and regulations: 1.457B + 22B = 23.457B input tokens
  4. + 5% notes = 1.173B output tokens
  5. Low: 1.320B + 17B = 18.32B
  6. High: 1.843B + 30B = 31.84B
  • huggingface.co (“Documents 6,919,240 | UTF-8 GB 78”)
  • courtlistener.com (“8,300,000 Number of precedential opinions in CourtListener.”)
  • wiki.free.law (“Our commitment to completeness is reflected in our database's size, which currently contains over nine million legal decisions drawn from ov…”)
  • arxiv.org (“StateCodes, which includes all statutory code from 50 U.S. states, containing 488M words across 1.8M sections”)
Recurring: Keep up each day: new U.S. laws, regulations and court opinions (per day)1.8M
870K3.9M
89K
17K390K
less than 0.01
less than 0.01less than 0.01
around the clock
Opinion clusters filed per year were 109,276–117,797 in 2023–2025; band 100,000–160,000 allows for late additions. A cluster can hold several opinions, so 2,000–5,000 tokens. The homepage 10-day add rate (~255,000 a year) includes backfill and is not used.
Math and sources
  1. Laws, FR, state bills and state regs: 635,100 input tokens/day (as above)
  2. Opinions: 120,000 clusters/yr (CourtListener, 117,797 filed in 2025) × 3,500 tokens
  3. Opinions: = 420M tokens/yr ÷ 365.25 = 1,149,900/day
  4. Total 1,785,000 input tokens/day; + 5% notes = 89,250
  5. Low: 325,600 + 100,000 × 2,000 / 365.25 = 873,200
  6. High: 1,670,400 + 160,000 × 5,000 / 365.25 = 3,860,700
  • courtlistener.com (“Opinion clusters by filing date, all statuses: 109,276 (2023), 110,002 (2024), 117,797 (2025).”)
  • courtlistener.com (“8,300,000 Number of precedential opinions in CourtListener.”)
  • huggingface.co (“Documents 6,919,240 | UTF-8 GB 78”)
  • ballotpedia.org (“The report is based on a dataset of 243,529 bills adopted from January 2011 to May 2024 ... The mean number of bills passed per year was 325…”)
  • cei.org (“The 2025 Federal Register closed out at 60,917 pages. ... Six years into the 2020s ... the average is 82,730 annually.”)
  • govtrack.us (“Since World War II (the earliest we have data), Congress has typically enacted 4-6 million words of new law in each two-year Congress.”)
Read every published court decision in the world (opinions only, no statutes)
One-time: Read the world's published court decisions once (opinions only) (in total)
350B
200B1T
18B
4B100B
506
2551,740
to do this in one day
No global census exists. Low assumes only China, the U.S. and a small share of other countries are online in full text. High allows a decade of Brazil's ~40M yearly decisions and longer Chinese documents. Notes 2–10%.
Math and sources
  1. China Judgments Online ~110–150M documents × ~1,000–1,300 tokens ≈ 150B
  2. U.S.: CourtListener 8.3M precedential × ~2,650 ≈ 22B
  3. Brazil, India, EU and rest of world ≈ 180B (ASSUMPTION)
  4. Total ≈ 350B input tokens; + 5% notes = 17.5B output
  5. Low 200B; high 1,000B
  • court.gov.cn (“上网公布裁判文书1097.9万份,同比增长13.3%。”)
  • arxiv.org (“China Judgments Online corpus: criminal 5,428,717 docs, avg length 962.84; civil 17,387,874 docs, avg length 1,353.03 (Table 2).”)
  • courtlistener.com (“8,300,000 Number of precedential opinions in CourtListener.”)
  • huggingface.co (“Documents 6,919,240 | UTF-8 GB 78”)
  • migalhas.com.br (“Houve pico histórico de produtividade, com 44,8 milhões de processos baixados (+19,9%)”)
Recurring: Keep up each day: new court decisions published worldwide (per day)140M
55M330M
6.8M
1.1M33M
0.2
0.070.6
around the clock
Low counts China, the U.S. and a fraction of Brazil's decisions at short length. High allows longer Brazilian decisions, full Indian and European output and Chinese documents near 1,700 characters.
Math and sources
  1. China: 10.979M judgments × ~1,300 chars × ~0.75–0.8 o200k tokens/char ≈ 11B tokens/yr
  2. Brazil: ~44M decisions (44.8M cases closed in 2024) × ~600 tokens ≈ 25B (length ASSUMPTION)
  3. U.S. ~0.4B + India, EU and rest of world ~10B (ASSUMPTION)
  4. Total ≈ 50B tokens/yr ÷ 365.25 = 136.9M input tokens/day
  5. + 5% notes = 6.84M output tokens/day
  6. Low: 20B/yr ÷ 365.25 = 54.8M; high: 120B/yr ÷ 365.25 = 328.5M
  • court.gov.cn (“上网公布裁判文书1097.9万份,同比增长13.3%。”)
  • arxiv.org (“China Judgments Online corpus: criminal 5,428,717 docs, avg length 962.84; civil 17,387,874 docs, avg length 1,353.03 (Table 2).”)
  • migalhas.com.br (“Houve pico histórico de produtividade, com 44,8 milhões de processos baixados (+19,9%)”)
  • courtlistener.com (“Opinion clusters by filing date, all statuses: 109,276 (2023), 110,002 (2024), 117,797 (2025).”)

Watch a livestream all day

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
One stream at Gemini's default resolution, 1 frame per second (plan with)
One-time: Index the last 30 days of this stream (default resolution, 1 fps) (in total)
260M
59M790M
30M
3M180M
0.5
0.092
to do this in one day
Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.
Math and sources
  1. Lookback days × tokens per stream-day
  2. Central: 30 × 8,812,800 = 264,384,000 input; 30 × 1,000,000 = 30,000,000 notes
  3. Low: 7 × 8,467,200 = 59,270,400 input; 7 × 432,000 = 3,024,000 notes
  4. High: 60 × 13,219,200 = 793,152,000 input; 60 × 3,000,000 = 180,000,000 notes
  • help.twitch.tv (“Twitch Partners, Prime and Twitch Turbo users will have past broadcasts saved for 60 days before being deleted. Twitch Affiliates have their…”)
  • support.google.com (“If your live stream is less than 12 hours, YouTube can automatically archive it for you. If your stream exceeds 12 hours, it may not be capt…”)
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
Recurring: Watch one 24-hour livestream (default resolution, 1 fps) (per day)8.8M
8.5M13M
1M
430K3M
0.02
0.010.03
around the clock
Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.
Math and sources
  1. Gemini 3 default: 70 tokens/frame + 32 audio tokens/s = 102 tokens/s
  2. Central: 102 × 86,400 s = 8,812,800 input tokens per day
  3. Low: older models' 66 + 32 = 98 tokens/s × 86,400 = 8,467,200
  4. High: 8,812,800 × 1.5 = 13,219,200 (chunk prompts and overlap)
  5. Notes: 1,000,000 per day; low 432,000 (100 tokens per 20 s); high 3,000,000
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
  • ai.google.dev (“If media_resolution is set to low, frames are tokenized at 66 tokens per frame. Otherwise, frames are tokenized at 258 tokens per frame. Aud…”)
  • docs.nvidia.com (“OSL = 100 (Captioning) — the VLM generates a descriptive caption of the video chunk. Captioning (OSL=100) | 33. Chunk duration | 10 seconds.”)
One stream at high resolution for on-screen text, 1 frame per second
One-time: Index the last 30 days of this stream (high resolution, 1 fps) (in total)
810M
180M2.4B
30M
3M180M
1.1
0.23.8
to do this in one day
Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.
Math and sources
  1. Lookback days × tokens per stream-day
  2. Central: 30 × 26,956,800 = 808,704,000 input; 30 × 1,000,000 = 30,000,000 notes
  3. Low: 7 × 25,056,000 = 175,392,000 input; 7 × 432,000 = 3,024,000 notes
  4. High: 60 × 40,435,200 = 2,426,112,000 input; 60 × 3,000,000 = 180,000,000 notes
  • help.twitch.tv (“Twitch Partners, Prime and Twitch Turbo users will have past broadcasts saved for 60 days before being deleted. Twitch Affiliates have their…”)
  • support.google.com (“If your live stream is less than 12 hours, YouTube can automatically archive it for you. If your stream exceeds 12 hours, it may not be capt…”)
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
Recurring: Watch one 24-hour livestream (high resolution, 1 fps) (per day)27M
25M40M
1M
430K3M
0.04
0.030.06
around the clock
Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.
Math and sources
  1. Gemini 3 high: 280 tokens/frame + 32 audio = 312 tokens/s
  2. Central: 312 × 86,400 = 26,956,800 input tokens per day
  3. Low: older models' 258 + 32 = 290 tokens/s × 86,400 = 25,056,000
  4. High: 26,956,800 × 1.5 = 40,435,200 (chunk prompts and overlap)
  5. Notes: 1,000,000 per day; low 432,000; high 3,000,000
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
  • ai.google.dev (“Approximately 300 tokens per second of video at high media resolution. Models with a 1M context window can process videos up to 3 hours long…”)
  • docs.cloud.google.com (“The default resolution for videos is 70 tokens per frame. MEDIA_RESOLUTION_HIGH: 280 tokens per frame. For models earlier than Gemini 3, eac…”)
  • docs.nvidia.com (“OSL = 100 (Captioning) — the VLM generates a descriptive caption of the video chunk. Captioning (OSL=100) | 33. Chunk duration | 10 seconds.”)
One stream at high resolution, 5 frames per second
One-time: Index the last 30 days of this stream (high resolution, 5 fps) (in total)
3.7B
800M11B
30M
3M180M
4.5
0.914
to do this in one day
Low 7 days = Twitch default past-broadcast retention; central 30 days sits between Affiliate 14 and Partner 60 days; high 60 days = Twitch Partner maximum. YouTube may not archive streams over 12 hours. Token rates as in the daily watch.
Math and sources
  1. Lookback days × tokens per stream-day
  2. Central: 30 × 123,724,800 = 3,711,744,000 input; 30 × 1,000,000 = 30,000,000 notes
  3. Low: 7 × 114,220,800 = 799,545,600 input; 7 × 432,000 = 3,024,000 notes
  4. High: 60 × 185,587,200 = 11,135,232,000 input; 60 × 3,000,000 = 180,000,000 notes
  • help.twitch.tv (“Twitch Partners, Prime and Twitch Turbo users will have past broadcasts saved for 60 days before being deleted. Twitch Affiliates have their…”)
  • support.google.com (“If your live stream is less than 12 hours, YouTube can automatically archive it for you. If your stream exceeds 12 hours, it may not be capt…”)
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
Recurring: Watch one 24-hour livestream (high resolution, 5 fps) (per day)120M
110M190M
1M
430K3M
0.1
0.10.2
around the clock
Input low uses the lowest documented token rate; central is the current Gemini 3 rate with no pad; high is 1.5× central for per-chunk prompts and 50% overlap when a day is split into 1-hour calls. Notes low = a caption every 20 s; high = 10-s captions plus a transcript.
Math and sources
  1. Tokens scale with frame rate: 5 frames/s
  2. Central: 280 × 5 + 32 audio = 1,432 tokens/s × 86,400 = 123,724,800
  3. Low: older 258 × 5 + 32 = 1,322 tokens/s × 86,400 = 114,220,800
  4. High: 123,724,800 × 1.5 = 185,587,200 (chunk prompts and overlap)
  5. Notes: 1,000,000 per day; low 432,000; high 3,000,000
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
  • docs.cloud.google.com (“Use a higher Frame Per Second (FPS) sampling rate for videos requiring granular temporal analysis, such as fast-action understanding or high…”)
  • docs.nvidia.com (“OSL = 100 (Captioning) — the VLM generates a descriptive caption of the video chunk. Captioning (OSL=100) | 33. Chunk duration | 10 seconds.”)
Every Twitch, YouTube and Kick livestream plus 4,000 satellite TV channels, at once
One-time: Index the last 30 days of every one of these channels' broadcasts (in total)
40T
7.7T160T
4.5T
390B36T
71,900
11,200392,000
to do this in one day
A chosen lookback of 7 / 30 / 60 days (Twitch default / between Affiliate and Partner / Partner maximum), not a measured archive. YouTube keeps many stream recordings for years; linear TV keeps none publicly. Channel band as in the daily watch.
Math and sources
  1. Channels × lookback days × default-res tokens per stream-day
  2. Central: 150,000 × 30 × 8,812,800 = 39,657,600,000,000 input
  3. Low: 130,000 × 7 × 8,467,200 = 7,705,152,000,000 input
  4. High: 200,000 × 60 × 13,219,200 = 158,630,400,000,000 input
  5. Notes: 150,000 × 30 × 1,000,000 = 4,500,000,000,000; low 130,000 × 7 × 432,000; high 200,000 × 60 × 3,000,000
  • help.twitch.tv (“Twitch Partners, Prime and Twitch Turbo users will have past broadcasts saved for 60 days before being deleted. Twitch Affiliates have their…”)
  • support.google.com (“If your live stream is less than 12 hours, YouTube can automatically archive it for you. If your stream exceeds 12 hours, it may not be capt…”)
  • streamlabs.com (“Twitch declined 8.3% to 2.07 million [ACV]. Average concurrent viewers per channel rose 7.2% to 22.6. YouTube Live ... 12.97 billion hours w…”)
Recurring: Watch every Twitch, YouTube and Kick stream and 4,000 satellite TV channels (per day)1.3T
1.1T2.6T
150B
56B600B
2,400
1,6006,530
around the clock
Channel low 130,000 allows for YouTube Live and Gaming overlap; central 150,000 and high 200,000 allow for uncounted linear feeds and smaller platforms. Per-channel token and notes bands as in video-low.
Math and sources
  1. Twitch: 2.07M concurrent viewers ÷ 22.6 per channel = 91,593 live channels
  2. YouTube Live: 12.97B hours ÷ (91 days × 24) ÷ 308 per channel = 19,281
  3. YouTube Gaming 10,508 + Kick 7,218 → internet total ≈ 128,600; + 4,001 satellite TV ≈ 132,600
  4. Central: 150,000 × 8,812,800 = 1,321,920,000,000 input per day
  5. Low: 130,000 × 8,467,200 = 1,100,736,000,000; high: 200,000 × 13,219,200 = 2,643,840,000,000
  6. Notes: 150,000 × 1,000,000 = 150,000,000,000; low 130,000 × 432,000; high 200,000 × 3,000,000
  • streamlabs.com (“Twitch declined 8.3% to 2.07 million [ACV]. Average concurrent viewers per channel rose 7.2% to 22.6. YouTube Live ... 12.97 billion hours w…”)
  • lyngsat.com (“LyngSat Stream has public links to 4001 linear satellite TV channels and 3326 linear satellite radio channels transmitting online.”)
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
Every surveillance camera on Earth, around the clock
One-time: Index the footage every camera still keeps on its recorder (in total)
190Q
26Q1,200Q
31Q
2Q410Q
400 million
42 million3.8 billion
to do this in one day
Retention low 7 days (EDPB: erase after a few days); central 31 days (ICO keeps its own footage one month); high 90 days (common for commercial systems). Banks and police can keep longer. Camera band as in the daily watch.
Math and sources
  1. Cameras × retention days × silent default-res tokens per camera-day
  2. Central: 1,000,000,000 × 31 × 6,048,000 = 187,488,000,000,000,000 input
  3. Low: 658,000,000 × 7 × 5,702,400 = 26,265,254,400,000,000 input
  4. High: 1,500,000,000 × 90 × 9,072,000 = 1,224,720,000,000,000,000 input
  5. Notes: 1B × 31 × 1,000,000 = 31,000,000,000,000,000; low 658M × 7 × 432,000; high 1.5B × 90 × 3,000,000
  • edpb.europa.eu (“the personal data should in most cases (e.g. for the purpose of detecting vandalism) be erased, ideally automatically, after a few days. The…”)
  • cy.ico.org.uk (“Images are routinely retained for one month, but may be retained longer in the event that they are required as part of an investigation.”)
  • comparitech.com (“At the end of 2021, over one billion surveillance cameras were estimated to have been installed worldwide, according to IHS Markit's latest …”)
  • transformainsights.com (“The report charts the growth in the number of devices, which will grow from 658 million in 2025 to 829 million in 2035.”)
Recurring: Watch every surveillance camera, around the clock, video only (per day)6Q
3.8Q14Q
1Q
280T4.5Q
13 million
6 million42 million
around the clock
Camera low = Transforma Insights 658 million connected CCTV devices (2025). Central = IHS Markit forecast of about 1 billion installed from 2021, a forecast rather than a count. High 1.5 billion is an unsourced ceiling. Per-camera band as video-low without audio.
Math and sources
  1. No audio: 70 tokens/frame × 86,400 = 6,048,000 per camera-day
  2. Central: 1,000,000,000 × 6,048,000 = 6,048,000,000,000,000 input per day
  3. Low: 658,000,000 × (66 × 86,400 = 5,702,400) = 3,752,179,200,000,000
  4. High: 1,500,000,000 × (6,048,000 × 1.5 = 9,072,000) = 13,608,000,000,000,000
  5. Notes: 1B × 1,000,000 = 1,000,000,000,000,000; low 658M × 432,000; high 1.5B × 3,000,000
  • comparitech.com (“At the end of 2021, over one billion surveillance cameras were estimated to have been installed worldwide, according to IHS Markit's latest …”)
  • transformainsights.com (“The report charts the growth in the number of devices, which will grow from 658 million in 2025 to 829 million in 2035.”)
  • ai.google.dev (“unspecified (Default) video 70; low video 70; medium video 70; high video 280. Video (Text-heavy) recommended high, 280 (per frame).”)
  • docs.nvidia.com (“OSL = 100 (Captioning) — the VLM generates a descriptive caption of the video chunk. Captioning (OSL=100) | 33. Chunk duration | 10 seconds.”)
Index the whole YouTube library, then each day's new uploads
One-time: Index every video on YouTube once, at default resolution (in total)
730T
420T2.2Q
83T
22T500T
1.3 million
615,0005.4 million
to do this in one day
YouTube publishes no catalog-hours total. A 2022 random sample found a 615 s mean; 9.8–14.8 billion videos give 1.7–2.5 billion hours. Shorts pull the mean down, so central is 2.0 billion; low 1.2B, high 4.0B allows for growth since 2024.
Math and sources
  1. Default res: 8,812,800 ÷ 24 = 367,200 input tokens per hour of video
  2. Sample: 9.8–14.8 billion videos × 615 s mean ≈ 1.7–2.5 billion hours
  3. Central: 2,000,000,000 hours × 367,200 = 734,400,000,000,000 input
  4. Low: 1,200,000,000 × 352,800 = 423,360,000,000,000 input
  5. High: 4,000,000,000 × 550,800 = 2,203,200,000,000,000 input
  6. Notes per hour 18,000 / 41,667 / 125,000 → 21,600,000,000,000 / 83,333,333,333,333 / 500,000,000,000,000
  • journalqd.org (“The mean duration of videos in our set, measured in seconds, was 615 (just over 10 minutes), with a median of 126 ... 37.94% were a minute o…”)
  • en.wikipedia.org (“As of May 2019, videos were being uploaded to the platform at a rate of more than 500 hours of video per minute, and as of mid-2024, there w…”)
  • ethanzuckerman.com (“Our current estimate for the size of YouTube is 13.325 billion videos. We estimate that over 4 billion videos were posted to YouTube just in…”)
Recurring: Index each day's new YouTube uploads, at default resolution (per day)260B
250B1T
30B
13B230B
480
3692,540
around the clock
Low and central use YouTube's last official figure, more than 500 hours a minute (May 2019). High = about 4 billion videos uploaded in 2023 (UMass estimate) × the 615 s sampled mean ≈ 1.87 million hours a day; short Shorts make this an upper bound.
Math and sources
  1. 500 hours/min × 1,440 min = 720,000 hours of new video per day
  2. Central: 720,000 × 367,200 = 264,384,000,000 input
  3. Low: 720,000 × 352,800 = 254,016,000,000 input
  4. High: 4 billion videos in 2023 × 615 s ÷ 3,600 ÷ 365 ≈ 1,870,000 hours/day
  5. High: 1,870,000 × 550,800 = 1,029,996,000,000 input
  6. Notes: 720,000 × 41,667 = 30,000,000,000; low × 18,000; high 1,870,000 × 125,000
  • en.wikipedia.org (“As of May 2019, videos were being uploaded to the platform at a rate of more than 500 hours of video per minute, and as of mid-2024, there w…”)
  • ethanzuckerman.com (“Our current estimate for the size of YouTube is 13.325 billion videos. We estimate that over 4 billion videos were posted to YouTube just in…”)
  • journalqd.org (“The mean duration of videos in our set, measured in seconds, was 615 (just over 10 minutes), with a median of 126 ... 37.94% were a minute o…”)

Defend 1,000 servers

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Condense each server's logs into evidence packets and send only those to the model
One-time: Re-check 90 days of logs as packets; inventory and config-scan every server (in total)
160B
6.4B1.5T
39B
1.3B380B
405
153,940
to do this in one day
Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.
Math and sources
  1. Daily packets: 10/s × 2,000 in × 86,400 s = 1.728B input; 10/s × 500 out × 86,400 = 0.432B output
  2. × 90 days = 155.52B input and 38.88B output; past alerts were already closed, so no deep investigations are replayed
  3. Inventory: 800 packages × 35 tokens + 12,000 context = 40,000 per server × 1,000 = 40M input; 800 out each = 0.8M
  4. Config scan: 244 CIS controls × 200 + 31,000 config = 80,000 per server × 1,000 = 80M input
  5. App code: 50 apps × 20,000 lines × 10 tokens/line = 10M, so the scan is 90M input
  6. Scan output: 1,000 tokens × 1,000 servers = 1.0M, plus 4,000 × 50 apps = 0.2M, total 1.2M
  7. Total: 155.52B + 0.04B + 0.09B = 155.65B input; 38.88B + 0.002B = 38.882B output
  • learn.microsoft.com (“Retain audit log history for at least 12 months, with at least the most recent three months immediately available for analysis.”)
  • learn.microsoft.com (“By default, Microsoft Sentinel and Microsoft Defender XDR retain data in this tier for 30 days. ... You can extend the retention period of M…”)
  • gstatic.com (“Global median dwell time across 2025 Mandiant investigations was 14 days. Median dwell time for incidents discovered internally in 2025 rema…”)
  • linuxize.com (“1893. The number above was produced on a desktop install of Ubuntu 24.04. A minimal server will usually report a much smaller count.”)
  • edgeservers.com.au (“The CIS Ubuntu Linux 22.04 LTS Benchmark v2.0.0 has 244 individual controls across Level 1 and Level 2. The CIS Ubuntu 24.04 benchmark, curr…”)
  • github.com (“eShopOnContainers | 25 | 69874”)
Recurring: A 2,000-token evidence packet per server every 100 seconds (10 per second) (per day)1.7B
86M10B
430M
17M2.6B
4.5
0.227
around the clock
Low 1/s = one packet per server every ~17 minutes × (1,000 in, 200 out); high 30/s = one per server every ~33 s × (4,000 in, 1,000 out). No vendor publishes this rate; it is a design choice.
Math and sources
  1. 10 packets/s × 2,000 input × 86,400 s = 1.728B input tokens/day
  2. 10 × 500 output × 86,400 = 0.432B output tokens/day
  3. 10/s across 1,000 servers = one packet per server every 100 s
  4. Low 1/s × 1,000 in × 86,400 = 86.4M; 1/s × 200 out = 17.28M
  5. High 30/s × 4,000 in × 86,400 = 10.368B; 30/s × 1,000 out = 2.592B
  • learn.microsoft.com (“Attack disruption uses extended detection and response (XDR) signals and evaluates the entire attack to take action at the incident level.”)
  • docs.cloud.google.com (“Each tenant can run up to 10 investigations per hour (5 manual and 5 automatic). Most investigations complete in about 60 seconds.”)
Evidence packets plus a deep AI investigation of 1,000 alerts a day (plan with)
One-time: Re-check 90 days of logs as packets; inventory and config-scan every server (in total)
160B
6.4B1.5T
39B
1.3B380B
405
153,940
to do this in one day
Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.
Math and sources
  1. Daily packets: 10/s × 2,000 in × 86,400 s = 1.728B input; 10/s × 500 out × 86,400 = 0.432B output
  2. × 90 days = 155.52B input and 38.88B output; past alerts were already closed, so no deep investigations are replayed
  3. Inventory: 800 packages × 35 tokens + 12,000 context = 40,000 per server × 1,000 = 40M input; 800 out each = 0.8M
  4. Config scan: 244 CIS controls × 200 + 31,000 config = 80,000 per server × 1,000 = 80M input
  5. App code: 50 apps × 20,000 lines × 10 tokens/line = 10M, so the scan is 90M input
  6. Scan output: 1,000 tokens × 1,000 servers = 1.0M, plus 4,000 × 50 apps = 0.2M, total 1.2M
  7. Total: 155.52B + 0.04B + 0.09B = 155.65B input; 38.88B + 0.002B = 38.882B output
  • learn.microsoft.com (“Retain audit log history for at least 12 months, with at least the most recent three months immediately available for analysis.”)
  • learn.microsoft.com (“By default, Microsoft Sentinel and Microsoft Defender XDR retain data in this tier for 30 days. ... You can extend the retention period of M…”)
  • gstatic.com (“Global median dwell time across 2025 Mandiant investigations was 14 days. Median dwell time for incidents discovered internally in 2025 rema…”)
  • linuxize.com (“1893. The number above was produced on a desktop install of Ubuntu 24.04. A minimal server will usually report a much smaller count.”)
  • edgeservers.com.au (“The CIS Ubuntu Linux 22.04 LTS Benchmark v2.0.0 has 244 individual controls across Level 1 and Level 2. The CIS Ubuntu 24.04 benchmark, curr…”)
  • github.com (“eShopOnContainers | 25 | 69874”)
Recurring: Evidence packets every 100 s per server, plus 1,000 deep investigations a day (per day)2.2B
94M16B
530M
19M3.9B
5.7
0.241
around the clock
Packets as in the packet-only design. Investigations low: 100 alerts/day (Prophet typical) × 8 calls (Elastic after tuning) × (10,000 in, 2,000 out). High: 3,000/day (large enterprise) × 150 calls × (12,000 in, 3,000 out); Elastic's 36,000-token calls with fewer calls land inside this.
Math and sources
  1. Packets: 10/s × 2,000 in × 86,400 = 1.728B input/day; × 500 out = 0.432B output/day
  2. Investigations: 1,000 × 50 calls × 10,000 in = 0.5B input/day
  3. 1,000 × 50 calls × 2,000 out = 0.1B output/day
  4. Per investigation: 50 × 12,000 = 600,000 tokens, between Elastic's ~217k average and Dropzone's ~100 calls
  5. Total: 1.728B + 0.5B = 2.228B input/day; 0.432B + 0.1B = 0.532B output/day
  6. Low: 86.4M + (100 × 8 calls × 10,000 = 8M) = 94.4M in; 17.28M + 1.6M = 18.88M out
  7. High: 10.368B + (3,000 × 150 × 12,000 = 5.4B) = 15.768B in; 2.592B + 1.35B = 3.942B out
  • risky.biz (“A typical investigation involves approximately 100 separate LLM invocations.”)
  • elastic.co (“LLM call counts dropped to 7–9 ... average input tokens per LLM call ranged from roughly 10,000 for narrowly scoped specialized agents to ro…”)
  • elastic.co (“Specialized agents ~113k; single agent ~649k tokens per Windows alert investigation, across 36,822 conversations and ~8 billion tokens on Cl…”)
  • thehackernews.com (“Security teams are drowning in alerts, with organizations processing an average of 960 alerts per day.”)
  • prophetsecurity.ai (“Across all customers, a typical (around the 75th percentile) organization generates about 100 alerts per day.”)
  • docs.cloud.google.com (“Each tenant can run up to 10 investigations per hour (5 manual and 5 automatic). Most investigations complete in about 60 seconds.”)
During an active intrusion: ten times the evidence packets
One-time: Re-check 90 days of normal-rate logs as packets; inventory and scan servers (in total)
160B
6.4B1.5T
39B
1.3B380B
405
153,940
to do this in one day
Days 30 (Sentinel/XDR default) / 90 (PCI three months immediately available) / 365 (PCI 12 months), combined in log quadrature with the daily packet band: ÷24 and ×9.7 on input, ÷30 and ×9.7 on output. Inventory and scan add 36.5M low and 580M high.
Math and sources
  1. Daily packets: 10/s × 2,000 in × 86,400 s = 1.728B input; 10/s × 500 out × 86,400 = 0.432B output
  2. × 90 days = 155.52B input and 38.88B output; past alerts were already closed, so no deep investigations are replayed
  3. Inventory: 800 packages × 35 tokens + 12,000 context = 40,000 per server × 1,000 = 40M input; 800 out each = 0.8M
  4. Config scan: 244 CIS controls × 200 + 31,000 config = 80,000 per server × 1,000 = 80M input
  5. App code: 50 apps × 20,000 lines × 10 tokens/line = 10M, so the scan is 90M input
  6. Scan output: 1,000 tokens × 1,000 servers = 1.0M, plus 4,000 × 50 apps = 0.2M, total 1.2M
  7. Total: 155.52B + 0.04B + 0.09B = 155.65B input; 38.88B + 0.002B = 38.882B output
  • learn.microsoft.com (“Retain audit log history for at least 12 months, with at least the most recent three months immediately available for analysis.”)
  • learn.microsoft.com (“By default, Microsoft Sentinel and Microsoft Defender XDR retain data in this tier for 30 days. ... You can extend the retention period of M…”)
  • gstatic.com (“Global median dwell time across 2025 Mandiant investigations was 14 days. Median dwell time for incidents discovered internally in 2025 rema…”)
  • linuxize.com (“1893. The number above was produced on a desktop install of Ubuntu 24.04. A minimal server will usually report a much smaller count.”)
  • edgeservers.com.au (“The CIS Ubuntu Linux 22.04 LTS Benchmark v2.0.0 has 244 individual controls across Level 1 and Level 2. The CIS Ubuntu 24.04 benchmark, curr…”)
  • github.com (“eShopOnContainers | 25 | 69874”)
Recurring: Each attack day: 100 evidence packets per second plus 1,000 deep investigations (per day)18B
2.6B75B
4.4B
520M19B
46
6194
around the clock
Packets low 30/s × (1,000 in, 200 out); high 200/s × (4,000 in, 1,000 out); central is 10× the normal rate. Investigations use the same band as the headline design and are not multiplied.
Math and sources
  1. Packets: 100/s × 2,000 in × 86,400 = 17.28B input/day; × 500 out = 4.32B output/day
  2. 100/s across 1,000 servers = one packet per server every 10 s
  3. Investigations unchanged: 0.5B input and 0.1B output per day
  4. Total: 17.28B + 0.5B = 17.78B input/day; 4.32B + 0.1B = 4.42B output/day
  5. Low: 30/s × 1,000 × 86,400 = 2.592B + 8M = 2.6B in; 30/s × 200 = 518.4M + 1.6M = 520M out
  6. High: 200/s × 4,000 × 86,400 = 69.12B + 5.4B = 74.52B in; 200/s × 1,000 = 17.28B + 1.35B = 18.63B out
  • learn.microsoft.com (“While an attack is in progress, Defender disrupts the attack by automatically containing compromised assets that the attacker is using throu…”)
  • risky.biz (“A typical investigation involves approximately 100 separate LLM invocations.”)
  • elastic.co (“LLM call counts dropped to 7–9 ... average input tokens per LLM call ranged from roughly 10,000 for narrowly scoped specialized agents to ro…”)
Skip condensing: send every raw log event to the model
One-time: Read 90 days of every past raw event, inventory every server, and scan configs (in total)
19T
950B120T
78B
2.6B4T
23,000
1,110162,000
to do this in one day
Days 30 / 90 / 365 combined in log quadrature with the daily raw band (1 EPS × 150 tokens to 20 EPS × 400 tokens): ÷20 and ×6.2 on input. Output low 0 (read only); high uses 5% notes. Inventory and scan add 36.5M low and 580M high.
Math and sources
  1. Daily: 1,000 servers × 10 events/s × 250 tokens × 86,400 = 216B input tokens
  2. × 90 days = 19.44T input tokens
  3. 1-token verdict per event: 864M/day × 90 = 77.76B output
  4. Inventory 40M + config scan 90M = 130M input; 2M output
  5. Total: 19.44T + 0.00013T = 19,440.13B input; 77.76B + 0.002B = 77.762B output
  • logmanager.com (“UNIX / Linux Servers 10 EPS. Windows Server 30 EPS. Windows Active Directory 100 EPS. 700 bytes is a commonly used planning assumption.”)
  • learn.microsoft.com (“Retain audit log history for at least 12 months, with at least the most recent three months immediately available for analysis.”)
  • splunk.github.io (“Linux hosts: 8-10 MB per host. Every megabyte of archived data in .gz files stored in an S3 bucket and consumed into Splunk index results in…”)
  • linuxize.com (“1893. The number above was produced on a desktop install of Ubuntu 24.04. A minimal server will usually report a much smaller count.”)
  • edgeservers.com.au (“The CIS Ubuntu Linux 22.04 LTS Benchmark v2.0.0 has 244 individual controls across Level 1 and Level 2. The CIS Ubuntu 24.04 benchmark, curr…”)
Recurring: Every raw event from 1,000 servers, read by the model with a 1-token verdict (per day)220B
13B690B
860M
86M35B
255
161,000
around the clock
Events/s 1/10/20: Linux servers sit below Logmanager's Windows Server 30 EPS, which belongs to the verbose design; filtered EDR runs near 1–3. Tokens 150 syslog / 250 mixed / 400 JSON. Output 0 if the model only reads, up to 5% notes.
Math and sources
  1. 1,000 × 10 events/s × 250 tokens × 86,400 = 216B input tokens/day
  2. Events: 1,000 × 10 × 86,400 = 864M/day; 1-token verdict each = 864M output
  3. Low: 1 event/s × 150 tokens × 1,000 × 86,400 = 12.96B/day
  4. High: 20 events/s × 400 tokens × 1,000 × 86,400 = 691.2B/day
  5. Output high: a 5% note rate = 0.05 × 691.2B = 34.56B/day
  • logmanager.com (“UNIX / Linux Servers 10 EPS. Windows Server 30 EPS. Windows Active Directory 100 EPS. 700 bytes is a commonly used planning assumption.”)
  • elastic.co (“We dropped from an average of ~48k events per host per hour to ~12k events per host per hour—a 75% reduction in noise.”)
  • ibm.com (“Average raw event size - 382 B”)
  • splunk.github.io (“Linux hosts: 8-10 MB per host. Every megabyte of archived data in .gz files stored in an S3 bucket and consumed into Splunk index results in…”)
Every raw event from Windows-style servers logging 30 events per second
One-time: Read 90 days of past verbose Windows-style events; inventory and scan servers (in total)
82T
13T570T
230B
26B20T
95,900
15,700768,000
to do this in one day
Days 30 / 90 / 365 combined in log quadrature with the daily verbose band (10 EPS × 250 to 100 EPS × 400): ÷6.1 and ×6.9 on input. Output low 0 (read only); high uses 5% notes. Inventory and scan add 36.5M low and 580M high.
Math and sources
  1. Daily: 1,000 × 30 events/s × 350 tokens × 86,400 = 907.2B input tokens
  2. × 90 days = 81.648T input tokens
  3. 1-token verdict per event: 2.592B/day × 90 = 233.28B output
  4. Inventory 40M + config scan 90M = 130M input; 2M output
  5. Total: 81,648.13B input; 233.282B output
  • logmanager.com (“UNIX / Linux Servers 10 EPS. Windows Server 30 EPS. Windows Active Directory 100 EPS. 700 bytes is a commonly used planning assumption.”)
  • learn.microsoft.com (“Retain audit log history for at least 12 months, with at least the most recent three months immediately available for analysis.”)
  • techdocs.broadcom.com (“The average size of an event on Symantec EDR is ~1.5 KB.”)
  • linuxize.com (“1893. The number above was produced on a desktop install of Ubuntu 24.04. A minimal server will usually report a much smaller count.”)
  • edgeservers.com.au (“The CIS Ubuntu Linux 22.04 LTS Benchmark v2.0.0 has 244 individual controls across Level 1 and Level 2. The CIS Ubuntu 24.04 benchmark, curr…”)
Recurring: Every raw event from 1,000 verbose Windows-style servers, with a 1-token verdict (per day)910B
220B3.5T
2.6B
860M170B
1,070
2555,000
around the clock
Logmanager Windows Server 30 EPS central, Active Directory 100 EPS high, Linux 10 EPS low. Elastic's unfiltered 48k events/host/hour is 13.3 EPS. Tokens 250 / 350 / 400 for JSON-heavy events. Output 0 if read only, up to 5% notes.
Math and sources
  1. 1,000 × 30 events/s × 350 tokens × 86,400 = 907.2B input tokens/day
  2. Events: 1,000 × 30 × 86,400 = 2.592B/day; 1-token verdict each = 2.592B output
  3. Low: 10 events/s × 250 tokens = 216B/day
  4. High: 100 events/s (Active Directory) × 400 tokens = 3.456T/day
  5. Output high: 5% notes = 0.05 × 3.456T = 172.8B/day
  • logmanager.com (“UNIX / Linux Servers 10 EPS. Windows Server 30 EPS. Windows Active Directory 100 EPS. 700 bytes is a commonly used planning assumption.”)
  • elastic.co (“We dropped from an average of ~48k events per host per hour to ~12k events per host per hour—a 75% reduction in noise.”)
  • techdocs.broadcom.com (“The average size of an event on Symantec EDR is ~1.5 KB.”)
  • docs.fortinet.com (“Measurements from various SIEM installs have shown that Elasticsearch consumes an average of 500 bytes to store an event and all its parsed …”)

Fly 20,000 drones

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Autopilots fly the paths; a model briefs 200 squads every 10 seconds (plan with)
One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total)
7.3M
1.7M460M
9.7M
2.4M380M
0.06
0.022.8
to do this in one day
Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.
Math and sources
  1. Area pack: 8 MB filtered map text / 4 bytes per token = 2,000,000; 30 m elevation grid 110,889 × 2 = 221,778; airspace brief 30,000 → about 2,300,000 in, 200,000 out
  2. Fleet setup: shared template 532 parameters × 15 = 7,980 + 20,000 drones × 30 fields × 8 tokens = 4,800,000 → 4,808,000 in, 800,000 out
  3. Show brief (storyboard, music, venue, safety limits) = 200,000 in
  4. Formations: one every 20 s for 10 min plus the landing grid = 31
  5. 31 formations × 20,000 drones × 14 tokens (x, y, z, r, g, b as integers) = 8,680,000 out
  6. Paths between formations are computed by the show software, not the model
  7. Input 2,300,000 + 4,808,000 + 200,000 = 7,308,000
  8. Output 200,000 + 800,000 + 8,680,000 = 9,680,000
  • usgs.gov (“Each 10 km by 10km tile is in a cloud-optimized geotiff format ... Ground spacing is 1 meter.”)
  • docs.px4.io (“723 parameters total, 532 used.”)
  • docs.skybrush.io (“Automatic transitions between two formations are always calculated by Skybrush Studio Server in the background using highly-optimized propri…”)
  • raw.githubusercontent.com (“output_fps = IntProperty(name="Trajectory FPS", default=4, description="Number of samples to take from trajectories per second")”)
  • docs.skybrush.io (“with the following columns: time (milliseconds), X, Y and Z coordinates (meters) and the red, green and blue components of the color of the …”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
Recurring: Brief 200 squads every 10 seconds while 20,000 drones fly all day (per day)5.2B
1.7B17B
350M
86M1.7B
8
2.530
around the clock
Call rate held at 20 per second. Input per call: low 1,000 (squad totals and exceptions only), high 10,000 (full state, neighbours and recent events per drone). Output per call 50 / 200 / 1,000; briefing length is an assumption. A 100-drone status list measured about 4,200 tokens (o200k).
Math and sources
  1. 20,000 drones / 100 per squad = 200 squads; one call per squad every 10 s = 20 calls/s
  2. Input per call: 500-token mission brief + 100 drones × 25-token status line = 3,000
  3. 20 × 3,000 × 86,400 = 5,184,000,000 input tokens per day
  4. 20 × 200 generated × 86,400 = 345,600,000 generated tokens per day
  • docs.skybrush.io (“During the show, each drone flies its own unique flight plan (trajectory). The drones are not being controlled in real time via a radio link…”)
  • docs.px4.io (“PX4 requires that the external controller provides a continuous 2Hz "proof of life" signal, by streaming any of the supported MAVLink setpoi…”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
A model makes a decision for every drone once per second
One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total)
7.3M
1.7M460M
9.7M
2.4M380M
0.06
0.022.8
to do this in one day
Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.
Math and sources
  1. Area pack: 8 MB filtered map text / 4 bytes per token = 2,000,000; 30 m elevation grid 110,889 × 2 = 221,778; airspace brief 30,000 → about 2,300,000 in, 200,000 out
  2. Fleet setup: shared template 532 parameters × 15 = 7,980 + 20,000 drones × 30 fields × 8 tokens = 4,800,000 → 4,808,000 in, 800,000 out
  3. Show brief (storyboard, music, venue, safety limits) = 200,000 in
  4. Formations: one every 20 s for 10 min plus the landing grid = 31
  5. 31 formations × 20,000 drones × 14 tokens (x, y, z, r, g, b as integers) = 8,680,000 out
  6. Paths between formations are computed by the show software, not the model
  7. Input 2,300,000 + 4,808,000 + 200,000 = 7,308,000
  8. Output 200,000 + 800,000 + 8,680,000 = 9,680,000
  • usgs.gov (“Each 10 km by 10km tile is in a cloud-optimized geotiff format ... Ground spacing is 1 meter.”)
  • docs.px4.io (“723 parameters total, 532 used.”)
  • docs.skybrush.io (“Automatic transitions between two formations are always calculated by Skybrush Studio Server in the background using highly-optimized propri…”)
  • raw.githubusercontent.com (“output_fps = IntProperty(name="Trajectory FPS", default=4, description="Number of samples to take from trajectories per second")”)
  • docs.skybrush.io (“with the following columns: time (milliseconds), X, Y and Z coordinates (meters) and the red, green and blue components of the color of the …”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
Recurring: One model decision per drone per second, all day (per day)520B
170B1.7T
35B
14B140B
800
2802,800
around the clock
Rate held at 1 Hz. Input per call: low 100 (own state and goal, no neighbours), high 1,000 (full neighbour states and recent events). Output per call 8 / 20 / 80 tokens, an assumption for a short setpoint command. One compact drone status line measured 14 to 42 tokens.
Math and sources
  1. Input per call: own state + goal formation slot + about 6 nearest neighbours = 300 tokens
  2. 20,000 decisions/s × 300 = 6,000,000 tokens/s × 86,400 = 518,400,000,000 input per day
  3. 20,000 × 20 generated = 400,000 tokens/s × 86,400 = 34,560,000,000 generated per day
  • coreweave.com (“Result concurrency = 16,384 queries, TTFT P95 = 5.77 s, System TPS = 441,740 aggregate tokens/sec. Result verified by MLCommons Association.”)
  • docs.px4.io (“The vehicle will exit Offboard mode if MAVLink setpoint messages or OffboardControlMode messages stop being received for longer than the tim…”)
Every drone streams its camera to a model all day
One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total)
7.3M
1.7M460M
9.7M
2.4M380M
0.06
0.022.8
to do this in one day
Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.
Math and sources
  1. Area pack: 8 MB filtered map text / 4 bytes per token = 2,000,000; 30 m elevation grid 110,889 × 2 = 221,778; airspace brief 30,000 → about 2,300,000 in, 200,000 out
  2. Fleet setup: shared template 532 parameters × 15 = 7,980 + 20,000 drones × 30 fields × 8 tokens = 4,800,000 → 4,808,000 in, 800,000 out
  3. Show brief (storyboard, music, venue, safety limits) = 200,000 in
  4. Formations: one every 20 s for 10 min plus the landing grid = 31
  5. 31 formations × 20,000 drones × 14 tokens (x, y, z, r, g, b as integers) = 8,680,000 out
  6. Paths between formations are computed by the show software, not the model
  7. Input 2,300,000 + 4,808,000 + 200,000 = 7,308,000
  8. Output 200,000 + 800,000 + 8,680,000 = 9,680,000
  • usgs.gov (“Each 10 km by 10km tile is in a cloud-optimized geotiff format ... Ground spacing is 1 meter.”)
  • docs.px4.io (“723 parameters total, 532 used.”)
  • docs.skybrush.io (“Automatic transitions between two formations are always calculated by Skybrush Studio Server in the background using highly-optimized propri…”)
  • raw.githubusercontent.com (“output_fps = IntProperty(name="Trajectory FPS", default=4, description="Number of samples to take from trajectories per second")”)
  • docs.skybrush.io (“with the following columns: time (milliseconds), X, Y and Z coordinates (meters) and the red, green and blue components of the color of the …”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
Recurring: Watch every drone's camera at 1 frame per second, high resolution, all day (per day)480B
180B2.5T
20B
4B40B
676
2273,100
around the clock
Low: Gemini 3 default 70 tokens per frame + 32 audio = 102/s × 86,400 × 20,000 = 176.3B in, 200,000 notes per stream. High: 5 fps at 280 tokens + 32 audio = 1,432/s × 86,400 × 20,000 = 2,474.5B in, 2,000,000 notes per stream.
Math and sources
  1. Gemini 3 high resolution: 280 tokens per frame + 32 audio tokens per second at 1 fps = 312 tokens/s
  2. 312 × 86,400 = 26,956,800 per stream per day; padded to 30,000,000
  3. 20,000 streams × 30,000,000 = 600,000,000,000 input per day
  4. 20,000 × 1,000,000 generated notes = 20,000,000,000 generated per day
  5. Video only at 280 tokens per frame: 280 × 86,400 × 20,000 = 483.84 billion input tokens a day (no audio, no pad, matching the CCTV row)
  • ai.google.dev (“unspecified (Default) | 1120 | 70 | 560”)
  • ai.google.dev (“If media_resolution is set to low, frames are tokenized at 66 tokens per frame. Otherwise, frames are tokenized at 258 tokens per frame. Aud…”)
A model makes a decision for every drone ten times per second
One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total)
7.3M
1.7M460M
9.7M
2.4M380M
0.06
0.022.8
to do this in one day
Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.
Math and sources
  1. Area pack: 8 MB filtered map text / 4 bytes per token = 2,000,000; 30 m elevation grid 110,889 × 2 = 221,778; airspace brief 30,000 → about 2,300,000 in, 200,000 out
  2. Fleet setup: shared template 532 parameters × 15 = 7,980 + 20,000 drones × 30 fields × 8 tokens = 4,800,000 → 4,808,000 in, 800,000 out
  3. Show brief (storyboard, music, venue, safety limits) = 200,000 in
  4. Formations: one every 20 s for 10 min plus the landing grid = 31
  5. 31 formations × 20,000 drones × 14 tokens (x, y, z, r, g, b as integers) = 8,680,000 out
  6. Paths between formations are computed by the show software, not the model
  7. Input 2,300,000 + 4,808,000 + 200,000 = 7,308,000
  8. Output 200,000 + 800,000 + 8,680,000 = 9,680,000
  • usgs.gov (“Each 10 km by 10km tile is in a cloud-optimized geotiff format ... Ground spacing is 1 meter.”)
  • docs.px4.io (“723 parameters total, 532 used.”)
  • docs.skybrush.io (“Automatic transitions between two formations are always calculated by Skybrush Studio Server in the background using highly-optimized propri…”)
  • raw.githubusercontent.com (“output_fps = IntProperty(name="Trajectory FPS", default=4, description="Number of samples to take from trajectories per second")”)
  • docs.skybrush.io (“with the following columns: time (milliseconds), X, Y and Z coordinates (meters) and the red, green and blue components of the color of the …”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
Recurring: One model decision per drone, ten times per second, all day (per day)5.2T
1.7T17T
350B
140B1.4T
8,000
2,80028,000
around the clock
10× the 1 Hz case. Input per call 100 / 300 / 1,000 (no neighbours to full neighbour states); output per call 8 / 20 / 80. Rate held at 10 Hz.
Math and sources
  1. 20,000 drones × 10 per second = 200,000 decisions/s
  2. 200,000 × 300 input × 86,400 = 5,184,000,000,000 input per day
  3. 200,000 × 20 generated × 86,400 = 345,600,000,000 generated per day
  • coreweave.com (“Result concurrency = 16,384 queries, TTFT P95 = 5.77 s, System TPS = 441,740 aggregate tokens/sec. Result verified by MLCommons Association.”)
  • docs.skybrush.io (“The drones are not being controlled in real time via a radio link to the Ground Control Station (GCS).”)
Every drone reasons for 1,000 tokens once per second
One-time: Plan once: read a 10 km area, set up 20,000 drones, and write 31 formations (in total)
7.3M
1.7M460M
9.7M
2.4M380M
0.06
0.022.8
to do this in one day
Low: 2 km park (125,000 map + 8,978 elevations + 20,000 brief), compact fleet file (1,500,000), 11 formations × 10 tokens. High: raw 10 km dump (39M map + 1 m grid 100M × 2 = 239M), 723 parameters per drone (216.9M), per-second waypoints for 15 min × 21 tokens. Parts' lows and highs are added, so the band is wider than 95%.
Math and sources
  1. Area pack: 8 MB filtered map text / 4 bytes per token = 2,000,000; 30 m elevation grid 110,889 × 2 = 221,778; airspace brief 30,000 → about 2,300,000 in, 200,000 out
  2. Fleet setup: shared template 532 parameters × 15 = 7,980 + 20,000 drones × 30 fields × 8 tokens = 4,800,000 → 4,808,000 in, 800,000 out
  3. Show brief (storyboard, music, venue, safety limits) = 200,000 in
  4. Formations: one every 20 s for 10 min plus the landing grid = 31
  5. 31 formations × 20,000 drones × 14 tokens (x, y, z, r, g, b as integers) = 8,680,000 out
  6. Paths between formations are computed by the show software, not the model
  7. Input 2,300,000 + 4,808,000 + 200,000 = 7,308,000
  8. Output 200,000 + 800,000 + 8,680,000 = 9,680,000
  • usgs.gov (“Each 10 km by 10km tile is in a cloud-optimized geotiff format ... Ground spacing is 1 meter.”)
  • docs.px4.io (“723 parameters total, 532 used.”)
  • docs.skybrush.io (“Automatic transitions between two formations are always calculated by Skybrush Studio Server in the background using highly-optimized propri…”)
  • raw.githubusercontent.com (“output_fps = IntProperty(name="Trajectory FPS", default=4, description="Number of samples to take from trajectories per second")”)
  • docs.skybrush.io (“with the following columns: time (milliseconds), X, Y and Z coordinates (meters) and the red, green and blue components of the color of the …”)
  • lumasky.show (“Most drone light shows last between 5 and 15 minutes.”)
Recurring: Every drone reasons for 1,000 tokens, once per second, all day (per day)520B
170B1.7T
1.7T
350B6.9T
10,600
2,20042,000
around the clock
Input per call 100 / 300 / 1,000, as in the 1 Hz case. Generated reasoning per call 200 / 1,000 / 4,000; CoreWeave's benchmark responses averaged over 2,000 tokens, which sits inside the band.
Math and sources
  1. 20,000 × 300 input × 86,400 = 518,400,000,000 input per day
  2. 20,000 × 1,000 generated × 86,400 = 1,728,000,000,000 generated per day
  3. Total 2,246,400,000,000 tokens per day
  • coreweave.com (“A single response can run thousands of output tokens, and CoreWeave's runs averaged well over 2,000 tokens per request.”)

Read a company's messages

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Label every unique message (about 100 per person) within the 8-hour workday (plan with)
One-time: Read the company's kept email and chat once (about 2 years, de-duplicated) (in total)
650B
97B3.3T
50B
6.4B350B
1,040
1495,840
to do this in one day
Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, new tokens 30–150, wrapper 1.2–2.5× (quoted threads not re-read; earlier messages already counted), 0.5–7 retained-equivalent years (heavy deleters to full journaling).
Math and sources
  1. Workday recipe: 100,000 × 100 unique × 65 tokens × 2 wrapper = 1.3 billion input; 100,000 × 100 × 10 = 100 million output
  2. Kept history: 250 workdays × 2 retained-equivalent years = 500 workdays
  3. Commvault: 50,000 messages per mailbox; sent+received email 157/day × 250 = 39,250/year, so about 1.3 years
  4. Input: 1.3 billion × 500 = 650 billion
  5. Output: 100 million × 500 = 50 billion
  6. GPU-s = 650e9/10,000 + 50e9/2,000 = 90 million; / 86,400 = 1,042 GPUs for one day
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds.”)
  • documentation.commvault.com (“based on an average mailbox size of 5 GB, and an average of 50,000 messages per mailbox”)
  • petri.com (“Tony Redmond 3.764 GB (4,042,041,235 bytes) 27009”)
  • learn.microsoft.com (“Moves items that are two years or older from a user's primary mailbox to their archive mailbox.”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419).”)
Recurring: Label each workday's new unique messages (about 100 per person) (per workday)1.3B
350M5.1B
100M
22M450M
6.3
1.626
through the workday
Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, mean new tokens 30–150 (Kooti email mean 153 words), prompt wrapper 1.2–5× (5× keeps quoted threads), label 3–40 tokens.
Math and sources
  1. Received per worker: 117 emails + 153 Teams messages = 270 items
  2. De-duplicated: about 100 unique messages per person per workday
  3. Input: 100,000 × 100 × 65 tokens × 2 wrapper = 1.3 billion
  4. Output: 100,000 × 100 × 10-token JSON label = 100 million
  5. GPU-s = 1.3e9/10,000 + 1e8/2,000 = 130,000 + 50,000 = 180,000
  6. 180,000 / 28,800 = 6.25 GPUs over 8 hours
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419). … The length of a reply is the length of the messag…”)
Label every unique message, spread over the following 24 hours
One-time: Read the company's kept email and chat once (about 2 years, de-duplicated) (in total)
650B
97B3.3T
50B
6.4B350B
1,040
1495,840
to do this in one day
Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, new tokens 30–150, wrapper 1.2–2.5× (quoted threads not re-read; earlier messages already counted), 0.5–7 retained-equivalent years (heavy deleters to full journaling).
Math and sources
  1. Same one-time read as co-light-8h: 1.3 billion input × 500 workdays = 650 billion
  2. Output: 100 million × 500 = 50 billion
  3. The 8h vs 24h choice changes only the recurring window
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds.”)
  • documentation.commvault.com (“based on an average mailbox size of 5 GB, and an average of 50,000 messages per mailbox”)
  • petri.com (“Tony Redmond 3.764 GB (4,042,041,235 bytes) 27009”)
  • learn.microsoft.com (“Moves items that are two years or older from a user's primary mailbox to their archive mailbox.”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419).”)
Recurring: Label one workday's unique messages over the next 24 hours (weekends idle) (per day)1.3B
350M5.1B
100M
22M450M
2.1
0.58.5
around the clock
Judgment 95% band: log-normal mix of driver ranges, not every extreme at once. Unique 40–180/person, mean new tokens 30–150 (Kooti email mean 153 words), prompt wrapper 1.2–5× (5× keeps quoted threads), label 3–40 tokens.
Math and sources
  1. Tokens are per workday, not per calendar day
  2. Input: 100,000 × 100 × 65 × 2 = 1.3 billion; output 100,000 × 100 × 10 = 100 million
  3. GPU-s = 130,000 + 50,000 = 180,000
  4. 180,000 / 86,400 = 2.08 GPUs
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419). … The length of a reply is the length of the messag…”)
Scan every received copy with no de-duplication
One-time: Read every inbox copy of kept email and chat once (about 2 years) (in total)
1.8T
270B8T
140B
18B870B
2,810
41714,300
to do this in one day
Judgment 95% band: log-normal mix, not every extreme at once. Received items 117–275 (low drops Teams), new tokens 30–150, wrapper 1.2–2.5× (threads not re-read), 0.5–7 retained-equivalent years, label 3–40 tokens.
Math and sources
  1. Workday: 100,000 × 270 received × 65 tokens × 2 wrapper = 3.51 billion input
  2. Output: 100,000 × 270 × 10 = 270 million
  3. Kept history: 250 × 2 = 500 workdays
  4. Input: 3.51 billion × 500 = 1.755 trillion
  5. Output: 270 million × 500 = 135 billion
  6. GPU-s = 175.5 million + 67.5 million = 243 million; / 86,400 = 2,813 GPUs for one day
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)
  • learn.microsoft.com (“Data from Teams chats is stored in a hidden folder in the mailbox of each user included in the chat, and a similar hidden folder in a group …”)
  • documentation.commvault.com (“based on an average mailbox size of 5 GB, and an average of 50,000 messages per mailbox”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419).”)
Recurring: Scan all 270 received items per person each workday, duplicates included (per workday)3.5B
1B12B
270M
62M1.1B
17
4.561
through the workday
Judgment 95% band: log-normal mix, not every extreme at once. Received items 117–275 (low drops Teams), new tokens 30–150, wrapper 1.2–5×, label 3–40 tokens. Same wrapper as the de-duplicated design.
Math and sources
  1. Received: 117 emails + 153 Teams messages = 270 items per worker
  2. Input: 100,000 × 270 × 65 × 2 wrapper = 3.51 billion
  3. Output: 100,000 × 270 × 10 = 270 million
  4. GPU-s = 351,000 + 135,000 = 486,000
  5. 486,000 / 28,800 = 16.9 GPUs over 8 hours (vs 6.25 de-duplicated)
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)
  • learn.microsoft.com (“Data from Teams chats is stored in a hidden folder in the mailbox of each user included in the chat, and a similar hidden folder in a group …”)
Reason for 1,000 tokens about every unique message
One-time: Write 1,000 tokens about every kept unique message (about 2 years) (in total)
650B
97B3.3T
5T
490B35T
29,700
2,950206,000
to do this in one day
Judgment 95% band: log-normal mix, not every extreme at once. Unique 40–180/person, 200–4,000 generated tokens per message, 0.5–7 retained-equivalent years; input band as the light read.
Math and sources
  1. Input as the light read: 1.3 billion × 500 workdays = 650 billion
  2. Output per workday: 100,000 × 100 × 1,000 = 10 billion
  3. Output: 10 billion × 500 = 5 trillion
  4. GPU-s = 65 million + 2.5 billion = 2.565 billion
  5. 2.565 billion / 86,400 = 29,688 GPUs for one day
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds.”)
  • documentation.commvault.com (“based on an average mailbox size of 5 GB, and an average of 50,000 messages per mailbox”)
  • ar5iv.labs.arxiv.org (“Distribution of reply lengths (mean 153 words, median 43 words, standard deviation 419).”)
Recurring: Write 1,000 tokens about each workday's unique messages (per workday)1.3B
350M5.1B
10B
1.6B45B
178
29799
through the workday
Judgment 95% band: log-normal mix, not every extreme at once. Unique 40–180/person, 200–4,000 generated tokens (short summary to long reasoning); input band as light.
Math and sources
  1. Input as light: 100,000 × 100 × 65 × 2 = 1.3 billion
  2. Output: 100,000 × 100 × 1,000 = 10 billion
  3. GPU-s = 130,000 + 5,000,000 = 5,130,000
  4. 5,130,000 / 28,800 = 178 GPUs over 8 hours
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)
Label the messages of all 450 million paid Microsoft 365 seats
One-time: Read kept unique email and chat once for all 450 million M365 seats (in total)
2.9Q
440T15Q
230T
29T1.6Q
4.7 million
677,00027 million
to do this in one day
Light one-time band scaled by seats. Seats 450–475 million (over 450 million reported in Jan 2026; +6% for FY26). Log-normal mix, not every extreme at once; the token recipe range dominates.
Math and sources
  1. Scale: 450,000,000 / 100,000 = 4,500 companies of this size
  2. Input: 650 billion × 4,500 = 2.925 quadrillion
  3. Output: 50 billion × 4,500 = 225 trillion
  4. GPU-s = 292.5 billion + 112.5 billion = 405 billion; / 86,400 = 4.69 million GPUs for one day
  • office365itpros.com (“Microsoft said that the number of paid commercial Microsoft 365 seats now exceeds 450 million, a small increase from the 446 million reporte…”)
  • alphaspread.com (“paid M365 commercial seats grew 6% year-over-year to over 450 million with installed base expansion across all customer segments, though pri…”)
  • office365itpros.com (“seat growth for the year was 6% (the same percentage cited for the last few quarters)”)
  • learn.microsoft.com (“User mailboxes | 100 GB | 100 GB | 100 GB | 50 GB | 100 GB | 2 GB”)
  • documentation.commvault.com (“based on an average mailbox size of 5 GB, and an average of 50,000 messages per mailbox”)
Recurring: Label one global workday's unique messages for every M365 seat (24 hours) (per day)5.8T
1.6T23T
450B
99B2T
9,380
2,42038,200
around the clock
Light recurring band scaled by seats 450–475 million. Log-normal mix, not every extreme at once. The Americas/Europe overlap can raise the peak to about 1.5× the 24-hour average.
Math and sources
  1. Scale: 450,000,000 / 100,000 = 4,500
  2. Input: 1.3 billion × 4,500 = 5.85 trillion per global workday
  3. Output: 100 million × 4,500 = 450 billion
  4. Seats span time zones, so one workday's mail arrives over about 24 hours
  5. GPU-s = 585 million + 225 million = 810 million
  6. 810 million / 86,400 = 9,375 GPUs
  • office365itpros.com (“Microsoft said that the number of paid commercial Microsoft 365 seats now exceeds 450 million, a small increase from the 446 million reporte…”)
  • alphaspread.com (“paid M365 commercial seats grew 6% year-over-year to over 450 million with installed base expansion across all customer segments, though pri…”)
  • office365itpros.com (“seat growth for the year was 6% (the same percentage cited for the last few quarters)”)
  • microsoft.com (“The average worker receives 117 emails daily—most of them skimmed in under 60 seconds. … The average worker receives 153 Teams messages per …”)

Examine every patient on Earth

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Daily workups for people in a care episode (5% a day), using a chart summary (plan with)
One-time: Read every person's record once (12,000-token average) and write a summary (in total)
100T
33T460T
17T
6.6T100T
211,000
76,9001.1 million
to do this in one day
Band on the global mean record, not person-level tails. Low 4,000: thinner high-income files and few-hundred-token paper cards elsewhere. High 55,600: every person has a UF Health-sized 10-year note file (82B words / 2.48M patients x 1.68 tokens/word). Summaries 800–12,000 tokens; no measured length exists.
Math and sources
  1. High-income: 24% US x 55,600 + 76% other x 20,000 (notes ~1/4 US length) = 28,500
  2. Upper-middle 12,000; lower-middle 2,000; low-income 500 (judgment)
  3. 0.173x28,500 + 0.375x12,000 + 0.358x2,000 + 0.093x500 = 10,200; central kept at 12,000
  4. 8.30 billion x 12,000 = 99.6 trillion input tokens
  5. Summary up to 4,000 tokens, about 1/3 of a short record: 2,000 average
  6. 2,000 x 8.30B = 16.6 trillion generated
  7. Output low 800 each (6.64 trillion); high 12,000 each (99.6 trillion)
  • nature.com (“A total number of 290,482,002 clinical notes from 2,476,628 patients were extracted from the UF Health Integrated Data Repository (IDR)... A…”)
  • academic.oup.com (“By the last full year of the study period, 2022, the median number of notes grew to 359 (IQR 84-943) notes, representing 58 662 words (IQR 1…”)
  • worldometers.info (“2026 | 8,300,678,395”)
  • fred.stlouisfed.org (“2025: 1,423,739,902”)
  • acpjournals.org (“U.S. physicians' clinical notes are, on average, four times as long as those in other countries”)
  • chinadaily.com.cn (“China's medical institutions handled 10.15 billion patient visits in 2024”)
Recurring: Daily workup for 5% of people: summary, excerpts, notes, labs (16,000 tokens) (per day)6.6T
1.3T17T
3.3T
430B9.1T
26,900
4,02071,900
around the clock
Monte Carlo over driver ranges (share of people 1.1–8% a day, context 10k–40k tokens, output 2k–20k), 2.5th–97.5th percentile, instead of multiplying every extreme together.
Math and sources
  1. 8.30 billion x 5% = 415 million workups per day
  2. 5% = one workup on each day of an assumed ~5-day episode; 327/1,000/month (Green)
  3. One workup per episode would be 327/1,000/30 = 1.1% per day
  4. Context: 4,000 summary + 8,000 excerpts and guidelines + 3,000 new notes + 3,000 answers and labs = 16,000
  5. 415 million x 16,000 = 7.47 trillion input tokens per day
  6. 415 million x 8,000 generated (questions, labs, assessment, reasoning) = 3.32 trillion per day
  7. Context is 16,000 tokens: a 2,000-token summary (matching the one-time summary) + 8,000 excerpts + 3,000 new notes + 3,000 answers and labs
  • nejm.org (“Of 1 000 men, women, and children in the United States, we estimated that on average each month, 800 experience symptoms, 327 consider seeki…”)
  • media.epic.com (“the average note length across all clinical notes has increased 8.1%, from 4,628 characters in May 2020 to 5,002 characters in April 2023”)
  • worldometers.info (“2026 | 8,300,678,395”)
If every care-seeker had a US emergency-room-sized chart, read in full at each workup
Recurring: Daily workup for 5% of people reading a full 200,000-token chart once each (per day)
84T
4.7T270T
3.3T
180B13T
117,000
6,550388,000
around the clock
Share 1.1–8%. Chart 50,000–400,000 + 2,000–5,000 new; the 400k high is about 2x Patterson's ED-conditional mean, allowing for excluded labs and imaging. Generated 2,000–20,000. A 1M chart is a person-level tail, not a band on the mean.
Math and sources
  1. 8.30 billion x 5% = 415 million workups per day
  2. Chart 200,000 tokens: US ED median 98,308 with a long tail; lognormal fit to tails gives a ~200,000–220,000 mean
  3. 415 million x (200,000 + 3,000 new) = 84.245 trillion input tokens per day
  4. 415 million x 8,000 generated = 3.32 trillion per day
  5. Chart read once per workup (cached), not re-sent each turn
  • pmc.ncbi.nlm.nih.gov (“98 308 tokens (IQR 20 819-279 576)”)
  • academic.oup.com (“3.74% exceed War and Peace (561,304 words)”)
  • nejm.org (“Of 1 000 men, women, and children in the United States, we estimated that on average each month, 800 experience symptoms, 327 consider seeki…”)
Only people who actually see a clinician on a given day (1.5%)
Recurring: Daily workup for the 1.5% of people with a clinic visit that day (per day)
4.1T
2.5T15T
250B
110B750B
6,200
3,50022,100
around the clock
Share 1.3–1.8% (Moses UI 4.88–5.99 / 365 to OECD 6.5 / 365). Record 20,000–100,000 + 3,000: clinic attenders have larger files than the 12,000 global mean; 100k is a US-ED-sized chart. Generated 1,000–5,000.
Math and sources
  1. 5.42 outpatient visits per person per year (Moses 2019) / 365 = 1.48% per day
  2. 8.30 billion x 1.5% = 124.5 million visit-days
  3. x (30,000-token clinic-day record + 3,000 new) = 4.1085 trillion input per day
  4. x 2,000 generated = 249 billion per day
  • sciencedirect.com (“In 2016, the global age-standardised outpatient utilisation rate was 5.42 visits (95% uncertainty interval [UI] 4.88–5.99) per capita”)
  • oecd.org (“The OECD average was 6.5 consultations per person per year.”)
  • pmc.ncbi.nlm.nih.gov (“98 308 tokens (IQR 20 819-279 576)”)
US-only subset: read every resident's record once, then daily summary workups for 5%
One-time: Read each US resident's record once (55,600 tokens), write a 4,000-token summary (in total)
19T
6.5T35T
1.4T
680B4.2T
30,000
11,40064,600
to do this in one day
Population 340M (under the Census clock) to 349M (CBO Social Security area). Record 19,000 (children, people rarely in care) to 100,000 (Patterson ED-median notes plus structure applied to all). Summary 2,000–12,000.
Math and sources
  1. 343 million US residents (Census clock 342,851,385 on 12 Sep 2026)
  2. GatorTron: 82 billion words / 2,476,628 patients = 33,110 words per patient
  3. x 1.676 tokens per word (Patterson, tiktoken on clinical notes) = 55,600 tokens
  4. 343 million x 55,600 = 19.07 trillion input
  5. 343 million x 4,000-token summary = 1.372 trillion generated
  6. Low 340M x 19,000 in, 2,000 out; high 349M (CBO) x 100,000 in, 12,000 out
  • census.gov (“The United States population on September 12, 2026 was : 342,851,385”)
  • cbo.gov (“The Social Security area population is projected to increase from 349 million people this year to 364 million in 2056.”)
  • nature.com (“A total number of 290,482,002 clinical notes from 2,476,628 patients were extracted from the UF Health Integrated Data Repository (IDR)... A…”)
  • academic.oup.com (“By the last full year of the study period, 2022, the median number of notes grew to 359 (IQR 84-943) notes, representing 58 662 words (IQR 1…”)
  • healthit.gov (“As of 2024, 95% of U.S. office-based physicians adopted Any EHR and over 9 in 10 (91%) had adopted a Certified EHR”)
Recurring: Daily workup for 5% of US residents: summary, excerpts, new notes and labs (per day)310B
38B1.1T
140B
7.5B550B
1,150
874,450
around the clock
Share 1.1% (one workup per care-seeking episode) to 8% (Green with a 7-day episode). Context 10,000–40,000. Generated 2,000–20,000.
Math and sources
  1. 343 million x 5% = 17.15 million workups per day
  2. x 18,000 input (4,000 summary + 8,000 excerpts + 6,000 new notes, answers, labs) = 308.7 billion per day
  3. x 8,000 generated = 137.2 billion per day
  4. Same per-workup budget as the global summary design; the US is 4.1% of world population
  • nejm.org (“Of 1 000 men, women, and children in the United States, we estimated that on average each month, 800 experience symptoms, 327 consider seeki…”)
  • census.gov (“The United States population on September 12, 2026 was : 342,851,385”)
Same full charts, but re-send the whole chart on each of 8 turns
Recurring: Re-send the full 200,000-token chart on each of 8 turns for 5% of people a day (per day)
670T
23T2.7Q
3.3T
180B13T
789,000
27,7003.2 million
around the clock
Share 1.1–8%. Chart 50,000–400,000 (400k is about 2x Patterson's ED-conditional mean). Passes 5–10. Low = 1.1% x (5 x 50k + 2k). High = 8% x (10 x 400k + 4k). Generated 2,000–20,000.
Math and sources
  1. 8.30 billion x 5% = 415 million workups per day
  2. 415 million x (8 x 200,000 + 3,000) = 665.245 trillion input per day
  3. 415 million x 8,000 generated = 3.32 trillion per day (writing does not repeat 8 times)
  • pmc.ncbi.nlm.nih.gov (“98 308 tokens (IQR 20 819-279 576)”)
  • academic.oup.com (“3.74% exceed War and Peace (561,304 words)”)
  • nejm.org (“Of 1 000 men, women, and children in the United States, we estimated that on average each month, 800 experience symptoms, 327 consider seeki…”)
Everyone with a symptom today (15%), with 1-million-token records
Recurring: Daily full-record workup for the 15% with symptoms, 1-million-token charts (per day)
1.2Q
190T3.2Q
25T
5.3T63T
1.6 million
248,0004 million
around the clock
Share 8–19% (Green 800/1,000/month with 3–7 symptomatic days). Record 280,000–2,000,000 (Patterson 75th percentile notes to beyond War and Peace, plus structure) + 3,000–5,000 new. Generated 8,000–40,000.
Math and sources
  1. 8.30 billion x 15% = 1.245 billion people per day
  2. x (1,000,000 + 4,000) = 1.24998 quadrillion input per day
  3. x 20,000 generated = 24.9 trillion per day
  • academic.oup.com (“3.74% exceed War and Peace (561,304 words)”)
  • nejm.org (“Of 1 000 men, women, and children in the United States, we estimated that on average each month, 800 experience symptoms, 327 consider seeki…”)
  • who.int (“It is estimated that oral diseases affect nearly 3.7 billion people.”)
A full-record assessment for every person on Earth, every day
Recurring: Every person, every day: 510,000 tokens read once, 20,000 generated (per day)
4.2Q
500T8.5Q
170T
66T420T
5.9 million
961,00012 million
around the clock
Coverage fixed at 100%. Record plus new 60,000–1,020,000 (the site's text-record sensitivity range, not measured percentiles). Generated 8,000–50,000 (short workup to long reasoning plus checking).
Math and sources
  1. 8.30 billion people, 100% coverage by construction
  2. x (500,000-token record + 10,000 new) = 4.233 quadrillion input per day
  3. x 20,000 generated = 166 trillion per day
Every chronic patient (40%), 5-million-token charts, re-read 8 times a day
Recurring: Daily 8-pass re-read of 5-million-token charts for 40% of people (per day)
130Q
15Q240Q
66T
23T190T
150 million
17 million270 million
around the clock
Share 35% (global adult chronic-disease guess) to 57% (US 76.4% of adults applied to all ages). Record 1M–5M. Passes 5–10. Low = 35% x (5 x 1M + 3k). High = 57% x (10 x 5M + 5k). Generated 8,000–40,000.
Math and sources
  1. 8.30 billion x 40% = 3.32 billion people per day
  2. x (8 x 5,000,000 + 4,000) = 132.813 quadrillion input per day
  3. x 20,000 generated = 66.4 trillion per day
  • cdc.gov (“In 2023, 76.4% (representing 194 million) of US adults reported 1 or more chronic conditions”)
  • academic.oup.com (“0.68% exceed the complete Harry Potter series (1,084,170 words)”)

Argue every court case

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Every new case worldwide, argued at the depth its file needs (plan with)
One-time: Once: argue the ~280 million pending cases and read the world's case law (in total)
27T
11T81T
7.4T
3.2T21T
74,300
30,800216,000
to do this in one day
Judgment band from combined driver ranges, not all extremes at once: pending 200–400M; traffic 35–60%; simple 20–40%; eDiscovery share 0.02–0.4%; per-case ladders as on the recurring line. Case law 200–700B in, notes 2.5–10% (5–70B out). 2.5th ≈ 10.6T/3.2T, 97.5th ≈ 81T/21T. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.
Math and sources
  1. Pending stock ≈ 280 million (India 54.1M + Brazil 75.5M sourced; the other ~150M inferred)
  2. Same mix as new cases: 140M traffic × (3,500 in, 2,000 out) + 84M simple × (30,000, 15,000)
  3. + 55.72M contested × (200,000, 80,000) + 0.28M eDiscovery × (45M, 5M)
  4. = 26.754T input + 7.3976T generated, once
  5. Case law: U.S. CAP + CourtListener 6,919,240 documents, 78 GB ≈ 20B tokens
  6. Case law: China Judgments Online ~152M documents; India 12.85M HC/SC judgments + ~400M e-courts orders
  7. Case law: 350B input (200–700B) + 5% notes (2.5–10%) = 17.5B generated (5–70B)
  8. Total: 26.754T + 0.35T = 27.104T input; 7.3976T + 0.0175T = 7.4151T generated
  • dataforindia.com (“At the end of 2025, the Indian government reported that there were 54 million cases pending across the three levels of the Indian judiciary”)
  • cnj.jus.br (“O Poder Judiciário encerrou 2025 com 75,5 milhões de processos em tramitação”)
  • huggingface.co (“Documents 6,919,240 UTF-8 GB 78”)
  • baike.baidu.com (“As of January 2025, the total number of documents on China Judgements Online was approximately 152 million”)
  • supremepeoplescourtmonitor.com (“almost 11 million judgments were published online (+13.3%)”)
  • github.com (“Court judgments 12,848,644”)
Recurring: Every year: argue the ~280 million new cases, each at the depth its file needs (per year)27T
10T80T
7.4T
3T21T
202
79586
all year
Judgment band from combined driver ranges: cases 180–400M; traffic 35–60%; simple 20–40%; eDiscovery share 0.02–0.4%; per case L0 1k–10k/0.5k–5k, L1 10k–80k/5k–40k, L2 80k–500k/25k–200k, L3 8M–80M/1M–15M. 2.5th ≈ 10T/3T, 97.5th ≈ 80T/21T. Every high at once is ~150T/40T.
Math and sources
  1. 280M new cases a year (US 71M + China 37.5M + Brazil 40.9M + India ~22M + Europe ~40M + rest ~68M)
  2. 140M traffic/minor (50%) × (3,500 in, 2,000 out) = 0.49T + 0.28T
  3. 84M simple civil (30%) × (30,000, 15,000) = 2.52T + 1.26T
  4. 55.72M contested (19.9%) × (200,000, 80,000) = 11.14T + 4.46T
  5. 0.28M eDiscovery (0.1%) × (45M, 5M) = 12.60T + 1.40T
  6. = 26.754T input + 7.3976T generated tokens a year
  • ncsc.org (“70M 2024 state court filings”)
  • court.gov.cn (“全国各级法院受理审判执行案件3748.6万件,审结执结3620万件”)
  • cnj.jus.br (“Em 2025, ingressaram 40,9 milhões de processos judiciais”)
  • njdg.ecourts.gov.in (“Cases Instituted in Last Month ... 2851536”)
  • pew.org (“More than half (57%) of cases in 2022 were traffic related, mostly noncriminal violations.”)
  • onlinelibrary.wiley.com (“The median case ending at the Complaint stage lasts only about 4 months and has only 12 docket entries.”)
Every live file re-argued from scratch each year, backlog included
One-time: Once: read the world's published case law, with short notes (in total)
350B
200B1T
18B
4B100B
506
2551,740
to do this in one day
Low: China corpus half-purged and India orders counted as short (200B). High: China at ~2,500 tokens a document on 160M documents plus a large India order stock (700B). Notes 2.5–10% of input. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.
Math and sources
  1. Case law: 350B input (200–700B) + 5% notes (2.5–10%) = 17.5B generated (5–70B)
  2. Backlog is not a separate one-time job here: it is inside the 1.75× yearly line
  • huggingface.co (“Documents 6,919,240 UTF-8 GB 78”)
  • baike.baidu.com (“As of January 2025, the total number of documents on China Judgements Online was approximately 152 million”)
  • github.com (“Court judgments 12,848,644”)
Recurring: Every year: re-argue every live file from scratch (new filings plus backlog) (per year)47T
15T160T
13T
4.5T42T
353
1191,170
all year
Low: 1.5 × the case-by-case low (10T/3T) = 15T/4.5T, more of the stock stale or suspended. High: 2.0 × the case-by-case high (80T/21T) = 160T/42T, pending ≈ new with no haircut.
Math and sources
  1. Brazil: 16.4M of 75.5M pending suspended, so ~78% of the stock is active
  2. 0.78 × 280M pending + 280M new ≈ 1.78× the new-case count; use 1.75×
  3. 1.75 × 26.754T input = 46.82T; 1.75 × 7.3976T generated = 12.95T a year
  • cnj.jus.br (“O Poder Judiciário encerrou 2025 com 75,5 milhões de processos em tramitação”)
  • dataforindia.com (“At the end of 2025, the Indian government reported that there were 54 million cases pending across the three levels of the Indian judiciary”)
280 million cases a year with far larger token budgets per tier
One-time: Once: ~280 million pending cases at generous budgets, plus world case law (in total)
680T
420T3.2Q
41T
16T290T
1 million
582,0005.4 million
to do this in one day
Same mega-share sensitivity as the yearly line, × 280/300: 0.001% mega → 422.8T/15.9T; 0.1% mega → 3,194.8T/293.1T. Case law 200–700B in, notes 2.5–10% (5–70B out). Volume and per-tier budgets held at central. World case law sized as in the law chapter: 200B–1T input, 4–100B notes.
Math and sources
  1. Scale the 300M generous mix to 280M pending: × 280/300
  2. 723T × 280/300 = 674.8T input; 44.04T × 280/300 = 41.104T generated
  3. Case law: 350B input (200–700B) + 5% notes (2.5–10%) = 17.5B generated (5–70B)
  4. Total: 675.15T input + 41.1215T generated, once
  • rand.org (“total expenditures do range from a seemingly modest $17,000 (in an intellectual property matter) to $27 million (in a product-liability case…”)
  • huggingface.co (“Documents 6,919,240 UTF-8 GB 78”)
Recurring: Every year: argue 280 million new cases at generous per-tier budgets (per year)670T
420T3.2Q
41T
16T290T
2,790
1,59014,800
all year
Only the mega share is varied, holding 300M cases: 0.001% → 453T/17.04T; 0.1% → 3,423T/314.04T. Volume (240–400M) and per-tier budgets stay at central, so this is a sensitivity range used as the band.
Math and sources
  1. 240M routine × (50,000 in, 10,000 out) = 12T + 2.4T
  2. 57M contested × (2M, 100k) = 114T + 5.7T
  3. 2.97M complex × (100M, 2M) = 297T + 5.94T
  4. 30,000 mega-discovery (0.01%) × (10B, 1B) = 300T + 30T
  5. = 723T input + 44.04T generated tokens a year
  6. At 0.1% mega (routine shrinks to hold 280M): 3,423T input + 314.04T generated
  7. Scaled to 280 million cases a year (×280/300), matching the other court designs
  • rand.org (“the total costs per gigabyte reviewed were generally around $18,000, with the first and third quartiles in the 35 cases with complete inform…”)
  • complexdiscovery.com (“The 2025 to 2030 cycle places the worldwide eDiscovery market at an estimated $19.61 billion in 2025”)
Every new case gets a typical hosted eDiscovery review (10 GB)
One-time: Once: a 10 GB hosted review for each of ~280M pending cases, plus case law (in total)
13Q
1.6Q48Q
1.4Q
200T6Q
23 million
3 million90 million
to do this in one day
Low: 200M pending × 8M in / 1M out. High: 400M × 120M in / 15M out (10 GB at 12M tok/GB, or ~27 GB at 4.5M). Case law 200–700B in, notes 2.5–10% (5–70B out). World case law sized as in the law chapter: 200B–1T input, 4–100B notes.
Math and sources
  1. 280M pending × (45M in, 5M out) = 12,600T input + 1,400T generated
  2. Same as one year of new cases here, because pending ≈ new in the world roll-up
  3. Case law: 350B input (200–700B) + 5% notes (2.5–10%) = 17.5B generated (5–70B)
  • digitalwarroom.com (“the average GB yields approximately 7,500 documents. This assumption is based on an equal mix of email and other documents, and is supported…”)
  • huggingface.co (“Documents 6,919,240 UTF-8 GB 78”)
Recurring: Every year: a 10 GB hosted eDiscovery review for each of ~280 million new cases (per year)13Q
1.4Q48Q
1.4Q
180T6Q
62,100
7,420247,000
all year
Low: 180M cases × 8M in / 1M out (a 1–2 GB matter). High: 400M × 120M in / 15M out. Tokens per GB 2.5–12M: DWR's 15,000 pages per GB at 200–600 words a page. DWR is the average hosted matter, not the median filing.
Math and sources
  1. DWR: 150M documents / 2,000 matters = 75,000 documents ≈ 10 GB per hosted matter
  2. 10 GB × ~4.5M tokens/GB = 45M input tokens per case
  3. 280M cases × (45M in, 5M out) = 12,600T input + 1,400T generated a year
  • digitalwarroom.com (“the average GB yields approximately 7,500 documents. This assumption is based on an equal mix of email and other documents, and is supported…”)
  • pew.org (“More than half (57%) of cases in 2022 were traffic related, mostly noncriminal violations.”)
Every new case gets a large-company U.S. discovery production (100 GB)
One-time: Once: a ~100 GB large-company review for each pending case, plus case law (in total)
130Q
20Q600Q
4.2Q
1Q16Q
170 million
29 million790 million
to do this in one day
Low: 200M pending × 100M in / 5M out. High: 400M × 1.5B in / 40M out (~330 GB at 4.5M or ~125 GB at 12M tok/GB). Case law 200–700B in, notes 2.5–10% (5–70B out). World case law sized as in the law chapter: 200B–1T input, 4–100B notes.
Math and sources
  1. 280M pending × (450M in, 15M out) = 126,000T input + 4,200T generated
  2. Same as one year of new cases in this design
  3. Case law: 350B input (200–700B) + 5% notes (2.5–10%) = 17.5B generated (5–70B)
  • rand.org (“total expenditures do range from a seemingly modest $17,000 (in an intellectual property matter) to $27 million (in a product-liability case…”)
  • huggingface.co (“Documents 6,919,240 UTF-8 GB 78”)
Recurring: Every year: a ~100 GB large-company review for each of ~280 million new cases (per year)130Q
18Q600Q
4.2Q
900T16Q
466,000
71,3002.2 million
all year
Low: 180M cases × 100M in / 5M out (~22 GB, the cheap end of RAND's sample). High: 400M × 1.5B in / 40M out (~330 GB at 4.5M or ~125 GB at 12M tok/GB). RAND's sample is large-company productions, not the median case.
Math and sources
  1. RAND: median production cost $1.8M; total cost ~$18,000 per GB reviewed
  2. $1.8M ÷ $18,000 ≈ 100 GB (medians from different case subsets)
  3. 100 GB × ~4.5M tokens/GB = 450M input tokens per case
  4. 280M cases × (450M in, 15M out) = 126,000T input + 4,200T generated a year
  • rand.org (“the total costs per gigabyte reviewed were generally around $18,000, with the first and third quartiles in the 35 cases with complete inform…”)
  • rand.org (“review accounts for 73 percent of the total costs of production”)

Call every game on TV

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Numeric win-probability model with a one-line caption per play (average day)
One-time: Read historical play-by-play once for the NFL, MLB, NBA and Premier League (in total)
5.1B
1.7B30B
850M
170M5.9B
11
2.968
to do this in one day
Low: sum of per-league lows (33.8M events) at 50 in / 5 out tokens. High: sum of per-league highs (73.75M) at a full 400-token nflfastR-style row and 80 out.
Math and sources
  1. NFL (nflfastR, 1999–2025) ≈ 1,270,000 plays
  2. MLB Statcast 2015–25 = 7,434,201 pitches
  3. NBA ≈ 1,310 games × 450 events × 25 seasons ≈ 14,740,000 events
  4. Premier League 30 seasons × 380 matches × ~1,650 events = 18,810,000
  5. Sum ≈ 42,254,201 events
  6. × 120 input tokens (play text + key fields) = 5,070,504,120
  7. × 20 generated tokens of index caption = 845,084,020
  • nflfastr.com (“we've searched through 247,284 rows of data across 300+ columns”)
  • baseballsavant.mlb.com (“2025 ... 712,528 ... 2024 ... 711,899 ... 2015 ... 702,306”)
  • nature.com (“The data covers a total of around 1,941 matches, 3,251,294 events and 4,299 players.”)
  • statsperform.com (“In 1996 we were collecting 50-60 data points, post-match, with pen and paper using stop-start on VHS tapes.”)
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
Recurring: One-line caption per win-probability update, 135 games at a time, no video (per day)9.9M
910K48M
4M
300K24M
0.03
less than 0.010.2
around the clock
Low: 42 games live, 30 updates/h, 30 in / 10 out tokens. High: 208 games live, 120 updates/h, 80 in / 40 out tokens.
Math and sources
  1. Every game = every match with a professional live feed (Sportradar: ~650,000 a year)
  2. 650,000 × 1.82 h / 8,760 h ≈ 135 games live at any moment
  3. 135 games × 61 win-probability updates/h / 3,600 = 2.2875 updates/s
  4. × 50 input tokens × 86,400 s = 9,882,000 per day
  5. × 20 generated tokens × 86,400 s = 3,952,800 per day
  • sportradar.com (“We have over 650,000 live streams each year.”)
  • fool.com (“Last year, we streamed over 525,000 matches globally. And in '26, we anticipate to stream over 700,000 across our global footprint.”)
  • amazon.science (“Seventy-five machine learning models running on AWS process that data in under a second”)
  • support.developer.betfair.com (“In-play markets usually carry a time delay varying from 1-12 seconds.”)
Watch every live game at high-res on an average day, after watching four league archives
One-time: Read four leagues' play-by-play and watch their archived games once (in total)
240B
92B370B
1B
180M6.9B
289
108465
to do this in one day
Play-by-play low/high: 33.8M events × 50 tokens / 73.75M × 400. Low 86,890 h (MLB only 3 recent seasons, NFL 4,631+960, NBA 17,000, PL 7,600 games) at 290 tok/s. High 260,882 h (full vaults: NFL 6,000, NBA 20,000, MLB 57,500, all 13,166 PL games) at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. Play-by-play: 42,254,201 events × 120 in / 20 out = 5,070,504,120 / 845,084,020
  2. NFL+ from 2009: 4,641 regular + playoff games × 3.2 h + ~960 preseason × 3.0 h = 17,731 h
  3. NBA League Pass from 2012-13: 14 seasons × 1,320 games × 2.2 h = 40,656 h
  4. MLB game video 2002–22 (50,000 games streamed on MLB.TV, a league vault) × 2.8 h = 140,000 h
  5. Premier League full matches from 2006-07: 20 × 380 × 2.0 h = 15,200 h
  6. Total 213,587 h × 3,600 s × 290 tok/s = 222,985,036,800 input
  7. Totals: 228,055,540,920 input; 163,362,000 + 845,084,020 generated
  8. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 228,055,540,920 → 244,971,422,520
  • support.nfl.com (“NFL+ Premium offers ad-free archived games from the 2009 season on, including Preseason and Playoffs games, as well as select Super Bowl gam…”)
  • support.watch.nba.com (“Access to full game archives starting from the 2012-2013 season to the present”)
  • mlb.com (“It has also given baseball fans the ability to stream more than 50,000 distinct games.”)
  • premierleague.com (“Fans can watch highlights and full-match reruns of EVERY match in Premier League history”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
  • nflfastr.com (“we've searched through 247,284 rows of data across 300+ columns”)
Recurring: Watch about 135 live games at high-res 1 fps and caption every play (per day)3.6B
1.1B5.5B
4M
300K24M
4.2
1.26.5
around the clock
Video low: 42 games live (200,000 premium events a year) at 290 tok/s; high: 208 (1 million matches, all with video) at 300 tok/s. Captions as in the stats-only design: 30 to 120 updates/h, 30/10 to 80/40 tokens. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. 650,000 streams × 1.82 h / 8,760 h ≈ 135 games live at any moment, averaged over the year
  2. 135 × 290 tokens/s (258 per frame + 32 audio) × 86,400 s = 3,382,560,000
  3. Captions: 2.2875 updates/s × 50 in / 20 out × 86,400 s = 9,882,000 / 3,952,800
  4. Total per day: 3,392,442,000 input, 3,952,800 generated
  5. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 3,392,442,000 → 3,649,050,000
  • sportradar.com (“We have over 650,000 live streams each year.”)
  • fool.com (“Last year, we streamed over 525,000 matches globally. And in '26, we anticipate to stream over 700,000 across our global footprint.”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
Saturday peak: watch 500 games at once and caption every play (plan with)
One-time: Read historical play-by-play once to seed the per-play captions (in total)
5.1B
1.7B30B
850M
170M5.9B
11
2.968
to do this in one day
Low: sum of per-league lows (33.8M events) at 50 in / 5 out tokens. High: sum of per-league highs (73.75M) at a full 400-token nflfastR-style row and 80 out.
Math and sources
  1. NFL (nflfastR, 1999–2025) ≈ 1,270,000 plays
  2. MLB Statcast 2015–25 = 7,434,201 pitches
  3. NBA ≈ 1,310 games × 450 events × 25 seasons ≈ 14,740,000 events
  4. Premier League 30 seasons × 380 matches × ~1,650 events = 18,810,000
  5. Sum ≈ 42,254,201 events
  6. × 120 input tokens (play text + key fields) = 5,070,504,120
  7. × 20 generated tokens of index caption = 845,084,020
  • nflfastr.com (“we've searched through 247,284 rows of data across 300+ columns”)
  • baseballsavant.mlb.com (“2025 ... 712,528 ... 2024 ... 711,899 ... 2015 ... 702,306”)
  • nature.com (“The data covers a total of around 1,941 matches, 3,251,294 events and 4,299 players.”)
  • statsperform.com (“In 1996 we were collecting 50-60 data points, post-match, with pen and paper using stop-start on VHS tapes.”)
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
Recurring: Watch 500 games at once at high-res 1 fps, held all day, and caption every play (per day)14B
5B26B
15M
1.4M120M
16
5.831
around the clock
Video low: 200 games at once at 290 tok/s; high: 1,000 games at 300 tok/s. Captions low: 200 games, 30 updates/h, 30 in / 10 out tokens; high: 1,000 games, 120 updates/h, 80 in / 40 out. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. Saturday peak ≈ 28% of ~12,500 weekly streams × 1.82 h / ~14 h ≈ 455, rounded to 500 games
  2. 500 × 290 tokens/s × 86,400 s = 12,528,000,000
  3. Captions: 500 × 61 updates/h / 3,600 = 8.472 updates/s
  4. × 50 in / 20 out × 86,400 s = 36,600,000 / 14,640,000
  5. Total per day at the peak rate: 12,564,600,000 input, 14,640,000 generated
  6. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 12,564,600,000 → 13,515,000,000
  • sportradar.com (“We have over 650,000 live streams each year.”)
  • fool.com (“Last year, we streamed over 525,000 matches globally. And in '26, we anticipate to stream over 700,000 across our global footprint.”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
  • amazon.science (“Seventy-five machine learning models running on AWS process that data in under a second”)
Saturday peak: watch 500 games and have a language model re-reason every play
One-time: Read historical play-by-play once as full 400-token state rows (in total)
17B
14B30B
3.4B
2.7B5.9B
39
3168
to do this in one day
Tokens per row held at 400 in / 80 out. Low: sum of per-league lows (33.8M events). High: sum of per-league highs (73.75M events).
Math and sources
  1. NFL (nflfastR, 1999–2025) ≈ 1,270,000 plays
  2. MLB Statcast 2015–25 = 7,434,201 pitches
  3. NBA ≈ 1,310 games × 450 events × 25 seasons ≈ 14,740,000 events
  4. Premier League 30 seasons × 380 matches × ~1,650 events = 18,810,000
  5. Sum ≈ 42,254,201 events
  6. × 400 input tokens (full nflfastR-style row) = 16,901,680,400
  7. × 80 generated tokens of state summary = 3,380,336,080
  • nflfastr.com (“we've searched through 247,284 rows of data across 300+ columns”)
  • baseballsavant.mlb.com (“2025 ... 712,528 ... 2024 ... 711,899 ... 2015 ... 702,306”)
  • nature.com (“The data covers a total of around 1,941 matches, 3,251,294 events and 4,299 players.”)
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
Recurring: Watch 500 games at once, held all day, and have an LLM re-reason every play (per day)15B
5.3B32B
480M
88M1.5B
20
6.645
around the clock
Play rate held at 61/h. Low: 200 games at 290 tok/s, 1,000 in / 300 out per play. High: 1,000 games at 300 tok/s, 4,000 in / 1,000 out per play. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. 500 × 290 tokens/s × 86,400 s = 12,528,000,000
  2. 500 × 61 plays/h / 3,600 = 8.472 plays/s × 2,000 in × 86,400 s = 1,464,000,000
  3. 8.472 × 650 generated × 86,400 s = 475,800,000
  4. Total per day at the peak rate: 13,992,000,000 input, 475,800,000 generated
  5. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 13,992,000,000 → 14,942,400,000
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
  • support.developer.betfair.com (“In-play markets usually carry a time delay varying from 1-12 seconds.”)
  • amazon.science (“Seventy-five machine learning models running on AWS process that data in under a second”)
Peak, plus an LLM call on every logged on-ball event
One-time: Read all 14 billion Opta event data points and watch four league archives once (in total)
1.4T
300B3.1T
17B
4.2B35B
1,670
3733,770
to do this in one day
Opta low: 15 tokens per point (a point is one field), 500-token match summaries; high: 200 tokens per point, 4,000-token summaries. Footage low: 86,890 h at 290 tok/s; high: 260,882 h of full league vaults at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. 14 billion event data points (Stats Perform does not say if a point is an event or a field)
  2. × 80 tokens each (type, player, x/y, time, outcome) = 1,120,000,000,000
  3. Summaries: 14 billion / 1,650 events per match = 8,484,848 match-equivalents × 2,000 = 16,969,696,970
  4. Footage: 213,587 h × 3,600 × 290 = 222,985,036,800; 81,681 games × 2,000 = 163,362,000
  5. Totals: 1,342,985,036,800 input, 17,133,058,970 generated
  6. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 1,342,985,036,800 → 1,359,900,918,400
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
  • nature.com (“The data covers a total of around 1,941 matches, 3,251,294 events and 4,299 players.”)
  • statsperform.com (“In 1996 we were collecting 50-60 data points, post-match, with pen and paper using stop-start on VHS tapes.”)
  • mlb.com (“It has also given baseball fans the ability to stream more than 50,000 distinct games.”)
  • premierleague.com (“Fans can watch highlights and full-match reruns of EVERY match in Premier League history”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
Recurring: Watch 500 games at once, held all day, with an LLM call on every on-ball event (per day)19B
6B43B
1.9B
290M8.4B
34
8.698
around the clock
Low: 200 games at 290 tok/s, 200 events/h, 1,000 in / 300 out per event. High: 1,000 games at 300 tok/s, 350 events/h (soccer-heavy mix at ~1,650 events a match), 2,000 in / 1,000 out. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. 500 × 290 tokens/s × 86,400 s = 12,528,000,000
  2. 500 games × 250 on-ball events/h / 3,600 = 34.72 events/s
  3. × 2,000 in × 86,400 s = 6,000,000,000; × 650 generated = 1,950,000,000
  4. Total per day at the peak rate: 18,528,000,000 input, 1,950,000,000 generated
  5. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 18,528,000,000 → 19,478,400,000
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
  • nature.com (“The data covers a total of around 1,941 matches, 3,251,294 events and 4,299 players.”)
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
Scale-up: every data-covered match, 3,000 feeds at once, an update every 45 seconds
One-time: Read all Opta event data and watch a decade of betting streams (hypothetical) (in total)
9.3T
3.1T30T
25B
8.2B55B
10,900
3,67035,100
to do this in one day
Opta as in the every-touch design (15 to 200 tokens per point). Video low: 5 years × 400,000 streams × 1.4 h at 290 tok/s. High: 15 years × 700,000 × 2.4 h at 300 tok/s. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. Opta: 14 billion points × 80 = 1,120,000,000,000; match summaries 16,969,696,970
  2. 10 years × ~400,000 streams/yr average (525,000 in 2025) × 1.82 h = 7,280,000 h
  3. × 3,600 s × 290 tok/s = 7,600,320,000,000 input
  4. 4,000,000 games × 2,000-token notes = 8,000,000,000
  5. Totals: 8,720,320,000,000 input, 24,969,696,970 generated
  6. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 8,720,320,000,000 → 9,296,896,000,000
  • developers.statsperform.com (“The archive behind the coverage holds 7.2+ PB of proprietary sports data and 14B+ unique event data points, growing by 500,000+ covered matc…”)
  • sportradar.com (“live action from more than 400,000 sporting events a year”)
  • fool.com (“Last year, we streamed over 525,000 matches globally. And in '26, we anticipate to stream over 700,000 across our global footprint.”)
  • developers.statsperform.com (“A 10+ year historical archive sits behind it, available through enterprise access or pay-as-you-go via the self-service PressBox Video appli…”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
Recurring: Watch 3,000 feeds at once, held all day, with a detailed update every 45 seconds (per day)92B
29B150B
2.9B
580M9.6B
124
37228
around the clock
Cadence held at one update per 45 seconds. Low: 1,000 feeds at 290 tok/s, 300 generated per update. High: 5,000 feeds at 300 tok/s, 1,000 generated per update. Video rate band: 290 (older 258-token frames) to 468 tokens/s (312 × 1.5 for chunk prompts and overlap).
Math and sources
  1. Sportradar covers data for over 1 million matches a year; 3,000 live at once is a stress setting
  2. 3,000 × 290 tokens/s × 86,400 s = 75,168,000,000
  3. 3,000 feeds × 80 updates/h / 3,600 = 66.67 updates/s
  4. × 2,000 in × 86,400 s = 11,520,000,000; × 500 generated = 2,880,000,000
  5. Total per day at the stress rate: 86,688,000,000 input, 2,880,000,000 generated
  6. Video at Gemini 3 high resolution: 280 tokens per frame + 32 audio = 312 tokens/s (was 290); input central 86,688,000,000 → 92,390,400,000
  • fool.com (“we are the sports technology leader, covering over 1 million matches annually”)
  • ai.google.dev (“Otherwise, frames are tokenized at 258 tokens per frame. Audio: 32 tokens per second.”)
  • sportradar.com (“live action from more than 400,000 sporting events a year”)

Rewrite every app

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
Port every unique App Store and Google Play app once, then keep up (plan with)
One-time: Port every unique app once: read its code, write the new app and tests (in total)
1.5T
300B5.8T
770B
170B3.1T
6,220
1,33024,700
to do this in one day
Driver ranges: unique apps 3.0–4.5M, 3,000–40,000 lines per app, 6.5–11 tokens per line, input 2–7× and output 1.4–4× source, combined in log space rather than stacked. Writing the new code is about 70% of GPU time, so lines per app and the output multiple move the number most.
Math and sources
  1. Unique apps ≈ (mid-tracker listings 2,392,688 iOS + 2,280,096 Play) / 1.37 dual-listed ≈ 3,400,000
  2. Mean code: 58% tiny 800 + 29% small 8,000 + 12.6% medium 50,000 + 0.4% large 300,000 + 0.005% mega 2,500,000 = 10,409 lines
  3. × 1.2 for apps with separate iOS and Android codebases ≈ 12,500 lines per app
  4. Source tokens = 3,400,000 × 12,500 × 9 tokens per line = 382,500,000,000
  5. Input = 1× read + 3× agent rereads and test logs = 4 × 382.5 billion = 1,530,000,000,000 (new tokens only)
  6. Output = 0.01× notes + 2× new code and tests = 2.01 × 382.5 billion = 768,825,000,000
  7. Low and high: driver ranges combined in log space (see band note)
  • 42matters.com (“Currently, 996,498 Apple App Store apps have been rated by users, while 1,616,406 have not yet been rated.”)
  • 42matters.com (“Google Play is one of the most popular mobile app vendors on the planet, offering 2,567,162 apps.”)
  • 42matters.com (“Globally, 37 % of mobile Apps are available on both iOS and Android.”)
  • apple.com (“Total number of apps on the App Store 2,172,472”)
  • uber.com (“The app has a couple of millions of lines of code with the vast majority in Swift. The source code consists of about 500 swift modules inclu…”)
  • huggingface.co (“claude-opus-4-7 · SWE-bench Pro — 731 sessions ... Total input tokens Mean 3,458,560 ... Total cached tokens Mean 3,375,169 ... Total output…”)
Recurring: Port each day's ~25,000 app updates and ~2,400 new apps (new tokens only) (per day)1.5B
420M6.4B
660M
220M2.6B
5.6
1.722
around the clock
Updates 12,000–45,000 a day, 15,000–150,000 new input and 8,000–50,000 output per update; new apps 1,500–4,000 a day at 2,000–20,000 lines. Each part's ranges are combined in log space, then the two lows and two highs are added.
Math and sources
  1. iOS updates: (9,100,620 reviewed − 2,093,244 rejected − 557,000 new) / 365 ≈ 17,700 a day
  2. Both stores, deduplicated at a similar rate: ≈ 25,000 updates a day
  3. Updates: 25,000 × 40,000 new input = 1,000,000,000; × 16,000 output = 400,000,000
  4. SWE-bench Pro sessions are 97.6% cache reads: 3,458,560 − 3,375,169 ≈ 83,000 new input
  5. New apps: (557,000/365 iOS + 54,600/30 Play) / 1.37 ≈ 2,400 a day at 6,000 lines × 9 tokens
  6. New apps: 2,400 × 54,000 × 4 = 518,400,000 input; × 2.01 = 260,496,000 output
  7. Total: 1,518,400,000 input and 660,496,000 output a day
  • apple.com (“App submissions reviewed 9,100,620 ... App submissions rejected 2,093,244”)
  • appfigures.com (“According to Appfigures Explorer, Apple's App Store saw 557K new app submissions in 2025, a whopping 24% increase from 2024”)
  • huggingface.co (“claude-opus-4-7 · SWE-bench Pro — 731 sessions ... Total input tokens Mean 3,458,560 ... Total cached tokens Mean 3,375,169 ... Total output…”)
  • appbrain.com (“How many Android apps are in Google Play: 1,985,023”)
  • appfigures.com (“Looking at the groups, it's clear most apps, 34%, across both stores haven't been updated in over two years, which is what Apple calls aband…”)
A heavy agent loop with tests and repeated review for every app and update
One-time: Rewrite every unique app with a heavy loop: explore, rewrite, test, review (in total)
3.8T
680B20T
3.1T
570B16T
22,100
4,060117,000
to do this in one day
Same app, line and token ranges as the one-app port, with input 4–30× and output 3.5–25× source, combined in log space. Epoch's MirrorCode rebuilds without source used 280M tokens for a 17,000-line program, well above this band per app.
Math and sources
  1. Same source as the one-app port: 3,400,000 × 12,500 × 9 = 382,500,000,000 tokens
  2. Heavy input = 10 × source = 3,825,000,000,000 new tokens (cache rereads not counted)
  3. Heavy output = 8 × source = 3,060,000,000,000 (code, tests, retries, review text)
  4. Multiples are judgment; ChatDev traces put 59.4% of tokens in code review
  • huggingface.co (“claude-opus-4-7 · SWE-bench Pro — 731 sessions ... Total input tokens Mean 3,458,560 ... Total cached tokens Mean 3,375,169 ... Total output…”)
  • arxiv.org (“The Code Review phase is the largest consumer, responsible for an average of 59.4% of tokens across all 30 tasks.”)
  • arxiv.org (“runs on the same task can differ by up to 30× in total tokens, and higher token usage does not translate into higher accuracy”)
  • epoch.ai (“gotree ... Original codebase LoC ... 16,905 ... one billion tokens costs around $550”)
  • epoch.ai (“Claude Opus 4.7 reimplemented gotree: a bioinformatics toolkit with ~16,000 lines of Go and 40+ commands. ... Opus 4.7 solved it in 14 hours…”)
Recurring: Heavy loop on each day's ~25,000 updates and ~2,400 new apps (per day)4.3B
1B22B
3B
820M15B
23
5.9110
around the clock
Update and new-app ranges as in the one-app port, plus update multiples of 1.5–6× input and 2.5–10× output and new-app multiples of 4–30× and 3.5–25×. Each part is combined in log space; lows and highs are then added.
Math and sources
  1. Updates: 25,000 × 40,000 × 3 = 3,000,000,000 input
  2. Updates: 25,000 × 16,000 × 5 = 2,000,000,000 output
  3. New apps: 2,400 × 54,000 source × 10 = 1,296,000,000 input
  4. New apps: 2,400 × 54,000 × 8 = 1,036,800,000 output
  5. Total: 4,296,000,000 input and 3,036,800,000 output a day
  6. 3× and 5× over the plain update line are judgment, not measured
  • apple.com (“App submissions reviewed 9,100,620 ... App submissions rejected 2,093,244”)
  • appfigures.com (“According to Appfigures Explorer, Apple's App Store saw 557K new app submissions in 2025, a whopping 24% increase from 2024”)
  • huggingface.co (“claude-opus-4-7 · SWE-bench Pro — 731 sessions ... Total input tokens Mean 3,458,560 ... Total cached tokens Mean 3,375,169 ... Total output…”)
  • arxiv.org (“The Code Review phase is the largest consumer, responsible for an average of 59.4% of tokens across all 30 tasks.”)
  • appfigures.com (“Looking at the groups, it's clear most apps, 34%, across both stores haven't been updated in over two years, which is what Apple calls aband…”)
Only the small apps: 87% of apps but about a quarter of the code
One-time: Translate the ~3M tiny and small apps once; skip medium to mega apps (in total)
170B
40B490B
110B
26B320B
838
1982,440
to do this in one day
Small apps 2.6–3.9M, 800–8,000 mean lines, 6.5–11 tokens per line, input 1.5–3× and output 1.1–2× source (near-direct translation), combined in log space.
Math and sources
  1. Small apps = 87% × 3,400,000 = 2,958,000
  2. Mean lines = (0.58 × 800 + 0.29 × 8,000) / 0.87 = 3,200
  3. Source = 2,958,000 × 3,200 × 9 = 85,190,400,000 tokens
  4. That is 2,784 of 10,409 mean lines, or about 27% of the catalog's code
  5. Input = 2 × source = 170,380,800,000; output = 1.3 × source = 110,747,520,000
  • 42matters.com (“Currently, 996,498 Apple App Store apps have been rated by users, while 1,616,406 have not yet been rated.”)
  • 42matters.com (“Google Play is one of the most popular mobile app vendors on the planet, offering 2,567,162 apps.”)
  • 42matters.com (“Globally, 37 % of mobile Apps are available on both iOS and Android.”)
  • appfigures.com (“Looking at the groups, it's clear most apps, 34%, across both stores haven't been updated in over two years, which is what Apple calls aband…”)
  • github.com (“Rust 109,910 lines of code, token count 894,106”)
Recurring: Port each day's ~12,500 small-app updates and ~2,100 new small apps (per day)370M
100M1.3B
150M
45M490M
1.3
0.44.3
around the clock
Updates 6,000–25,000 a day at 8,000–60,000 input and 3,000–15,000 output; new small apps 1,300–3,500 a day at 800–8,000 lines. Each part combined in log space; lows and highs added.
Math and sources
  1. Small-app updates ≈ 50% of 25,000 = 12,500 a day
  2. Updates: 12,500 × 20,000 new input = 250,000,000; × 6,000 output = 75,000,000
  3. New small apps ≈ 87% of 2,400 ≈ 2,100 a day at 3,200 lines × 9 tokens = 28,800 source
  4. New apps: 2,100 × 28,800 × 2 = 120,960,000 input; × 1.3 = 78,624,000 output
  5. Total: 370,960,000 input and 153,624,000 output a day
  • apple.com (“App submissions reviewed 9,100,620 ... App submissions rejected 2,093,244”)
  • appfigures.com (“According to Appfigures Explorer, Apple's App Store saw 557K new app submissions in 2025, a whopping 24% increase from 2024”)
  • huggingface.co (“claude-opus-4-7 · SWE-bench Pro — 731 sessions ... Total input tokens Mean 3,458,560 ... Total cached tokens Mean 3,375,169 ... Total output…”)
  • appfigures.com (“Looking at the groups, it's clear most apps, 34%, across both stores haven't been updated in over two years, which is what Apple calls aband…”)

Advise every trader

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
A model checks each order a person places, with no calls on algorithmic orders (plan with)
One-time: Once: read SEC EDGAR, other countries' filings and a finance news archive (in total)
390B
210B1.1T
20B
8B67B
564
2891,630
to do this in one day
Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.
Math and sources
  1. SEC EDGAR major forms (10-K/Q, 8-K, S-1, S-8, 20-F, Forms 3/4/5, 144) = 43.73 billion tokens (Comma v0.1 tokenizer)
  2. EDGAR central 150 billion ≈ the SEFD-v1 snapshot (152B, Jan 2022–Jun 2025), not the 549B full archive
  3. Non-US filings (EU/UK/Japan/China/India/Canada) ≈ 1× the US central = 150 billion (assumption; no token count exists)
  4. Finance news: 30,000 unique wire stories/day × 600 tokens × 365 × 12 years = 78.8 billion ≈ 80 billion
  5. Daily price bars written as text ≈ 10 billion; tick data stays numerical
  6. Input 150 + 150 + 80 + 10 = 390 billion tokens
  7. Notes written = 5% of input = 19.5 billion output tokens
  • eventual.ai (“Total 43,725,818,627”)
  • arxiv.org (“We release SEFD-v1, a 152B-token initial public snapshot, and provide corpus-level analyses of a larger 18.5M-filing archive estimated at 55…”)
  • blog.otcmarkets.com (“SEDAR filings processed: 70,660”)
  • lseg.com.cn (“130+ third-party newswires and exchanges ... resulting in a total of approx. 30,000 unique stories per day on average”)
Recurring: Reason once before each of ~120 million orders people place per trading day (per day)840B
200B4.5T
240B
42B1.1T
2,360
47511,900
around the clock
Decisions (60–250 million a day) and tokens per decision (2k–32k in, 400–8k out) combined as independent uncertainties, not the product of both extremes. Low is fewer, smaller orders with a short check; high counts every small A-share lot and a long reasoning pass.
Math and sources
  1. China: CSDC 68.848 billion two-sided transfers in 2025 / 2 / 243 sessions = 142 million trades/day
  2. China: 283 million trade sides × ~60% individual / ~1.5 fills per order ≈ 113 million retail orders
  3. Plus US ~15M (Fidelity 5.7M DATs, other brokers), India ~35M, rest of world + crypto ~20M, minus overlap and automatic plans
  4. ≈ 120 million human decisions per trading day (low 60M, high 250M); ~2M from professionals are inside this
  5. Per decision: 7,000 input (portfolio, bars, headlines, filing excerpt) + 2,000 output (~800 reasoning + order + rationale)
  6. 120e6 × 7,000 = 840 billion input tokens per trading day
  7. 120e6 × 2,000 = 240 billion output tokens per trading day
  • finance.sina.com.cn (“2025年全年,沪深两市合计过户笔数达到688.48亿笔(双向计算),相比2024年的526.88亿笔增长了30.7%”)
  • amf-france.org (“the AMF identified 56 million equity transactions carried out by retail investors in 2025 (41 million in 2024)”)
  • businesstoday.in (“In August, demat accounts rose 1.4% MoM to 237.7 million, while NSE active clients increased 1.1% to nearly 46 million, taking the active-ac…”)
  • about.fidelity.com (“Daily average trades 5.7 million Up 31% year over year from Q2 2025”)
  • livemint.com (“those trading for just one day stood at 8.43 million investors, comprising 24% of the total individual investor base of 35.84 million.”)
  • newkerala.com (“around 30 crore trades, 300 million trades on a good day”)
An agent briefs each active trader every 15 minutes while their markets are open
One-time: Once: the same filings and news archive as the advice design (in total)
390B
210B1.1T
20B
8B67B
564
2891,630
to do this in one day
Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.
Math and sources
  1. SEC EDGAR major forms (10-K/Q, 8-K, S-1, S-8, 20-F, Forms 3/4/5, 144) = 43.73 billion tokens (Comma v0.1 tokenizer)
  2. EDGAR central 150 billion ≈ the SEFD-v1 snapshot (152B, Jan 2022–Jun 2025), not the 549B full archive
  3. Non-US filings (EU/UK/Japan/China/India/Canada) ≈ 1× the US central = 150 billion (assumption; no token count exists)
  4. Finance news: 30,000 unique wire stories/day × 600 tokens × 365 × 12 years = 78.8 billion ≈ 80 billion
  5. Daily price bars written as text ≈ 10 billion; tick data stays numerical
  6. Input 150 + 150 + 80 + 10 = 390 billion tokens
  7. Notes written = 5% of input = 19.5 billion output tokens
  • eventual.ai (“Total 43,725,818,627”)
  • arxiv.org (“We release SEFD-v1, a 152B-token initial public snapshot, and provide corpus-level analyses of a larger 18.5M-filing archive estimated at 55…”)
  • blog.otcmarkets.com (“SEDAR filings processed: 70,660”)
  • lseg.com.cn (“130+ third-party newswires and exchanges ... resulting in a total of approx. 30,000 unique stories per day on average”)
Recurring: Brief each of 40 million active traders every 15 minutes in their market hours (per day)5.1T
1.5T22T
510B
110B2.6T
8,890
2,35040,500
around the clock
Traders (20–100 million), digest size (1,500–12,000 in, 100–1,500 out) and weighted session (6–10 hours) combined as independent uncertainties at the fixed 15-minute cadence. Faster or slower cadences are different designs, not uncertainty.
Math and sources
  1. Active = trades on 50+ days a year: China ~20M, India ~3M cash + ~5M F&O, US ~6M, Korea/Japan ~3M, Europe ~2M, crypto-only ~3M ≈ 40 million (assumption)
  2. Weighted session 8 hours (US 6.5h, China 4h, India 6.25h, Europe 8.5h, crypto 24h)
  3. 8 × 60 / 15 = 32 briefings per trader per trading day
  4. Per briefing: 4,000 input (news and portfolio digest) + 400 output (briefing or no action)
  5. 40e6 × 32 × 4,000 = 5.12 trillion input tokens per trading day
  6. 40e6 × 32 × 400 = 512 billion output tokens per trading day
  • livemint.com (“Investors trading up to 50 days made up 33 million or 92% while those trading above 50 days made up just 8% or 2.8 million.”)
  • finance.sina.com.cn (“2025年全年,沪深两市合计过户笔数达到688.48亿笔(双向计算)”)
  • sec.gov (“For the three and six months ended June 30, 2026, Monthly Transacting Users (“MTUs”) were 7.6 million and 7.9 million.”)
  • en.yna.co.kr (“The number of shareholders of 2,727 companies listed in the local stock market came to 14.6 million as of end-December, up 2.3 percent from …”)
  • cboe.com (“The new extended trading hours are anticipated in December 2026 or early 2027, pending SEC review and industry readiness.”)
A model is called on every order and cancel, including algorithmic orders
One-time: Once: the same filings and news archive; market ticks are not read as text (in total)
390B
210B1.1T
20B
8B67B
564
2891,630
to do this in one day
Each part varied on its own and combined, not all extremes at once: EDGAR 43.7B (published major forms) to 550B (full SEFD archive); non-US filings 40–650B; news 20–250B (a few years to a 30-year archive); bars 2–50B; notes 2.5–10% of input.
Math and sources
  1. SEC EDGAR major forms (10-K/Q, 8-K, S-1, S-8, 20-F, Forms 3/4/5, 144) = 43.73 billion tokens (Comma v0.1 tokenizer)
  2. EDGAR central 150 billion ≈ the SEFD-v1 snapshot (152B, Jan 2022–Jun 2025), not the 549B full archive
  3. Non-US filings (EU/UK/Japan/China/India/Canada) ≈ 1× the US central = 150 billion (assumption; no token count exists)
  4. Finance news: 30,000 unique wire stories/day × 600 tokens × 365 × 12 years = 78.8 billion ≈ 80 billion
  5. Daily price bars written as text ≈ 10 billion; tick data stays numerical
  6. Input 150 + 150 + 80 + 10 = 390 billion tokens
  7. Notes written = 5% of input = 19.5 billion output tokens
  • eventual.ai (“Total 43,725,818,627”)
  • arxiv.org (“We release SEFD-v1, a 152B-token initial public snapshot, and provide corpus-level analyses of a larger 18.5M-filing archive estimated at 55…”)
  • blog.otcmarkets.com (“SEDAR filings processed: 70,660”)
  • lseg.com.cn (“130+ third-party newswires and exchanges ... resulting in a total of approx. 30,000 unique stories per day on average”)
Recurring: Call the model on each of ~9.75 billion orders and cancels per trading day (per day)24T
3.1T160T
1.9T
160B15T
39,500
4,510272,000
around the clock
Fills (400 million to 1.5 billion a day; low = China 142M + NSE 85M + US 150M + crypto), orders per fill (6–50) and per-call tokens (400–8,000 in, 20–800 out) combined as independent uncertainties. Excludes raw quote updates, which outnumber orders.
Math and sources
  1. US trades ≈ 150–310 million/day (17.6e9 shares / 57-share average trade is an upper-leaning estimate)
  2. China SH+SZ: 68.848 billion two-sided transfers / 2 / 243 sessions = 142 million trades/day
  3. India NSE: ~300 million trades on a good day, ~85 million on a quiet one
  4. Global fills central 650 million/day (US + China + India + other markets + crypto)
  5. Orders and cancels per fill 15: US lit stocks fill 2.5–4.2% of order volume, off-exchange and China cash far fewer
  6. 650e6 × 15 = 9.75 billion order messages per trading day
  7. 9.75e9 × 2,500 = 24.375 trillion input; 9.75e9 × 200 = 1.95 trillion output tokens
  • cboe.com (“In 2025, average daily volume (ADV) increased 44.6% year-over-year to 17.6 billion shares.”)
  • finance.sina.com.cn (“2025年全年,沪深两市合计过户笔数达到688.48亿笔(双向计算)”)
  • newkerala.com (“And today on a good day, we do around, give and take few, around 30 crore trades, 300 million trades on a good day”)
  • sec.gov (“The chart shows that the trade-to-order volume ratio for stocks is typically between 2.5% and 4.2%”)
  • govinfo.gov (“Public data from 2021 to 2024 shows that NYSE Chicago exhibited quote-to-trade ratios on Tapes A and C significantly higher than any histori…”)

Give every home a robot helper

Design · cost lineTokens readTokens writtenGPUs (10K / 2K)Why this range
One robot per home, with its planning done in a datacenter (plan with)
One-time: Scan each home once and take a guided tour with the family (in total)
2.5Q
930T93Q
88T
16T670T
3.4 million
1.2 million110 million
to do this in one day
Low is a quick walkthrough of a small home and a short tour. Central is a ~45-minute 3D scan and a 2-hour tour. High is a three-camera high-res scan of a large home plus a four-hour two-camera tour. Tour length and note sizes are assumptions.
Math and sources
  1. Homes: 2.00B low / 2.21B central / 2.40B high
  2. Scan, central: 45 min × 60 s × 2 fps × 70 tok + 8,000 notes = 386,000 tokens per home
  3. Tour, central: 2 h × 3,600 s × (70 vision + 32 audio tok/s) + 20,000-token interview = 754,400
  4. Central input: 2.21e9 × (386,000 + 754,400) = 2.52e15
  5. Central output: 2.21e9 × (25,000-token floor plan + 15,000-token household card) = 8.84e13
  6. Low: 20 min scan at 1 fps × 70 + 2,000 notes; 1 h tour at 1 fps + 10,000 interview; 8,000 out
  7. High: 2 h scan × 5 fps × 280 tok × 3 cameras + 50,000; 4 h tour, 2 cameras × 280 + audio; 280,000 out
  • villaviews.co (“Full interactive 3D tour · ~50 min on site”)
  • helgilibrary.com (“Historically, total number of households reached an all time high of 2,193 mil in 2025”)
  • ai.google.dev (“Video (General) | low (or medium) | 70 (per frame)”)
  • ai.google.dev (“Audio: 32 tokens per second.”)
Recurring: Per day: listening, cloud planning while the robot moves, and conversation (per day)45Q
9.1Q190Q
5.7Q
1.2Q14Q
85 million
17 million300 million
around the clock
Monte Carlo over the stated driver ranges (homes, awake hours, planning cadence 2–10 Hz, tokens per tick and per second of motion), 2.5th–97.5th percentile, instead of multiplying every extreme together.
Math and sources
  1. Listening: 2.21e9 × 32 audio tok/s × 16 h × 3,600 s = 4.07e15 input
  2. Planning input: 2.21e9 × 4 Hz × 57,600 s × 80 tok (64 image tokens + instruction) = 4.07e16
  3. Planning output: 2.21e9 × 57,600 s × 45 action tokens per second of motion = 5.73e15
  4. (one 1-second action chunk per second, ~30 tokens per arm; not a chunk per tick)
  5. Talk: 12,700 words × 3.76 people × 20% overheard × 1.3 tok/word ≈ 12,400 input; 40 × 200-token replies
  6. Central total: 4.48e16 input, 5.75e15 output per day
  7. Low: 2.00e9 × 14 h audio; planning 4 h × 2 Hz × 20 in, 20 out/s; 10 short replies
  8. High: 2.40e9 × 18 h audio; planning 18 h × 10 Hz × 300 in, 120 out/s (two arms, 2 chunks/s)
  • arxiv.org (“When the backbone and local decoder are combined, the end-to-end latency from raw observations to low-level action chunks is approximately 2…”)
  • arxiv.org (“FAST consistently generates roughly 30 action tokens per chunk per robot arm (i.e., 60 tokens for the bi-manual setup)”)
  • arxiv.org (“Images are encoded at resolution 224×224 followed by pixel shuffle, resulting in 64 image token embeddings per frame.”)
  • arxiv.org (“we run inference every 0.5 seconds (after executing 25 actions).”)
  • bls.gov (“On an average day, 81 percent of people engaged in household activities--such as housework, cooking, lawn care, or household management--spe…”)
  • news.arizona.edu (“the average number of words spoken per day fell from about 16,000 to about 13,000.”)
Every person gets a robot that follows them
One-time: Scan each home and spend half a waking day learning each person (in total)
25Q
8.5Q360Q
150T
35T900T
30 million
10 million420 million
to do this in one day
Low is a short four-hour camera session per person with a small home scan. Central is half a waking day per person. High is a full waking day with two high-res cameras. Everyone, children included, is in every band.
Math and sources
  1. Home scan: same as the one-home design, since homes set the floor plan: 2.21e9 × 386,000 = 8.53e14
  2. Learn each person, central: (70 + 32 tok/s) × 8 h × 3,600 s + 15,000 notes = 2,952,600 tokens
  3. 8.301e9 people × 2,952,600 = 2.45e16 input
  4. Output: 2.21e9 × 25,000 map + 8.301e9 × 12,000-token person card = 1.55e14
  5. Low: 8.23e9 × (70 × 4 h × 3,600 + 5,000) + low scan
  6. High: 8.37e9 × ((2 × 280 + 32) × 16 h × 3,600 + 100,000) = 2.86e17, + high scan 7.27e16
  • statisticstimes.com (“The world population is projected at 8,300,678,395, or 8,301 million, or 8.30 billion, as of July 1, 2026.”)
  • helgilibrary.com (“Historically, total number of households reached an all time high of 2,193 mil in 2025”)
  • ai.google.dev (“Video (General) | low (or medium) | 70 (per frame)”)
  • ai.google.dev (“Audio: 32 tokens per second.”)
Recurring: Per day: two cameras and audio following each person, plus cloud reasoning (per day)730Q
100Q8,300Q
140Q
8.3Q1,100Q
1.7 billion
170 million16 billion
around the clock
Low is one 2 fps stream and a slow small planner for 14 h. Central is two cameras at 5 fps plus 2 Hz reasoning for 16 h. High is four cameras at 10 fps, high-res, 5 Hz reasoning for 18 h. Bands add each part’s extremes, so they are wider than 95%.
Math and sources
  1. Follow, central: 8.301e9 × (2 cams × 5 fps × 70 tok + 32 audio) × 16 h × 3,600 s = 3.50e17 input
  2. Per person: 732 tok/s × 57,600 s = 42.2 million input tokens per day
  3. Reasoning, central: 8.301e9 × 2 Hz × 57,600 s × 400 in / 150 out = 3.82e17 in, 1.43e17 out
  4. 2 Hz fits inside Gemini Robotics’ ~250 ms camera-to-action latency (≈4 Hz max)
  5. Low: 8.23e9 × (2 fps × 70 + 32) × 14 h; reasoning 1 Hz × 80 in / 20 out
  6. High: 8.37e9 × (4 cams × 10 fps × 280 + 32) × 18 h; reasoning 5 Hz × 800 in / 400 out
  • statisticstimes.com (“The world population is projected at 8,300,678,395, or 8,301 million, or 8.30 billion, as of July 1, 2026.”)
  • ai.google.dev (“Video (General) | low (or medium) | 70 (per frame)”)
  • ai.google.dev (“Audio: 32 tokens per second.”)
  • arxiv.org (“When the backbone and local decoder are combined, the end-to-end latency from raw observations to low-level action chunks is approximately 2…”)
  • arxiv.org (“we run inference every 0.5 seconds (after executing 25 actions).”)
Robots run their models on board and call the cloud only for hard questions
One-time: Upload a compact home map and a preference interview to the cloud (in total)
290T
60T5Q
62T
14T550T
691,000
150,0009 million
to do this in one day
Low is a room list and a short form. Central is a detailed text map of rooms and objects plus a long interview. High is a near-complete 3D copy of the home stored as tokens. All sizes are assumptions.
Math and sources
  1. The robot scans the home itself; the cloud stores a text summary
  2. Map, central: 2.21e9 × 100,000 tokens = 2.21e14 input; 20,000-token summary = 4.42e13 output
  3. Interview, central: 2.21e9 × 30,000 tokens = 6.63e13 input; 8,000-token profile = 1.77e13 output
  4. Central total: 2.87e14 input, 6.19e13 output
  5. Low: 20,000-token map card + 10,000/2,000 interview per home, 2.00e9 homes
  6. High: 2,000,000-token dense map + 100,000/30,000 interview per home, 2.40e9 homes
  • figure.ai (“Helix is the first VLA that runs entirely onboard embedded low-power-consumption GPUs, making it immediately ready for commercial deployment…”)
  • deepmind.google (“Optimized to run locally with low-latency inference.”)
Recurring: Per day: cloud calls when language or a stuck task needs a larger model (per day)97T
8.1T2Q
88T
6T1.9Q
624,000
44,10013 million
around the clock
Call rate is an assumption. Low is a quiet home where on-robot models handle nearly everything. Central is several real questions per waking hour. High is a cloud assistant consulted every few minutes with long reasoning.
Math and sources
  1. Central: 2.21e9 homes × 50 calls/day × (4 frames × 70 tok + 600 language) = 9.72e13 input
  2. Output: 2.21e9 × 50 calls × 800 tokens = 8.84e13
  3. Low: 2.00e9 × 15 calls × (70 + 200 tok) in, 200 out
  4. High: 2.40e9 × 200 calls × (4 × 280 + 3,120 tok) in, 4,000 out with a reasoning trace
  • figure.ai (“System 1 (S1): A fast reactive visuomotor policy that translates the latent semantic representations produced by S2 into precise continuous …”)
  • raw.githubusercontent.com (“AGX Thor | 128 GB shared | 8.9 Hz | 12.4 Hz | Robot-mounted edge deployment”)
  • arxiv.org (“When the backbone and local decoder are combined, the end-to-end latency from raw observations to low-level action chunks is approximately 2…”)
  • ai.google.dev (“Video (General) | low (or medium) | 70 (per frame)”)
Every person, four high-res cameras at 30 fps and nonstop 10 Hz reasoning
One-time: Dense photogrammetry of every home plus a million-token file on every person (in total)
190Q
5Q890Q
640T
140T1.8Q
220 million
6.5 million1 billion
to do this in one day
Low reuses the ordinary home scan and a half-million-token file per person. Central is a two-hour four-camera high-res scan. High is four hours with six cameras at 15 fps and a two-million-token file. File sizes are assumptions.
Math and sources
  1. Scan, central: 2.21e9 × 2 h × 3,600 s × 10 fps × 280 tok × 4 cameras = 1.78e17 input
  2. Scan output: 2.21e9 × 100,000-token writeup = 2.21e14
  3. Person file, central: 8.301e9 × 1,000,000 tokens = 8.30e15 input; 50,000-token summary = 4.15e14 out
  4. Low: the one-home central scan (8.53e14) + 8.23e9 × 500,000
  5. High: 2.40e9 × 4 h × 3,600 × 15 fps × 6 cameras × 280 = 8.71e17, + 8.37e9 × 2,000,000
  • ai.google.dev (“Video (Text-heavy) | high | 280 (per frame)”)
  • statisticstimes.com (“The world population is projected at 8,300,678,395, or 8,301 million, or 8.30 billion, as of July 1, 2026.”)
  • developer.nvidia.com (“Up to 20 cameras through HSB; up to 6 cameras through 16x lanes MIPI CSI-2”)
Recurring: Per day: four 30 fps high-res cameras, audio and 10 Hz reasoning per person (per day)28,000Q
870Q51,000Q
1,400Q
76Q5,800Q
40 billion
1.4 billion92 billion
around the clock
Low is two default-res cameras at 10 fps for 16 waking hours with slow reasoning. Central is four high-res cameras at 30 fps around the clock. High adds two cameras and 2,000-token prompts. Bands add each part’s extremes.
Math and sources
  1. Vision, central: 8.301e9 × 4 cams × 30 fps × 280 tok × 86,400 s = 2.41e19 input
  2. Audio: 8.301e9 × 32 tok/s × 86,400 s = 2.30e16 (about 0.1% of vision)
  3. Reasoning, central: 8.301e9 × 10 Hz × 86,400 s × 500 in / 200 out = 3.59e18 in, 1.43e18 out
  4. Low: 8.23e9 × (2 cams × 10 fps × 70 tok + 32 audio) × 16 h; reasoning 2 Hz × 200 in / 80 out
  5. High: 8.37e9 × (6 cams × 30 fps × 280 + 32) × 24 h; reasoning 10 Hz × 2,000 in / 800 out
  • ai.google.dev (“Video (Text-heavy) | high | 280 (per frame)”)
  • developer.nvidia.com (“Up to 20 cameras through HSB; up to 6 cameras through 16x lanes MIPI CSI-2”)
  • raw.githubusercontent.com (“Because each inference returns a multi-step action chunk, a ~10 Hz inference rate can sustain ~30 FPS execution via action chunking + asynch…”)
  • arxiv.org (“the π₀ model with FAST tokenization needs approximately 750ms of inference time per chunk”)
  • statisticstimes.com (“The world population is projected at 8,300,678,395, or 8,301 million, or 8.30 billion, as of July 1, 2026.”)
  • ai.google.dev (“Audio: 32 tokens per second.”)

Sources. Driving vLLM WideEP and Large-Scale Serving Toward Maturity on Blackwell (Part I) (vLLM Blog, 2026-02-03); CoreWeave Leads MLPerf 0.7 Endpoints Benchmark with DeepSeek-R1 (CoreWeave); Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP (Part II) (LMSYS, 2025-09-25); MLPerf Inference v6.0 results (MLCommons, 2026-04-01); InferenceX: Kimi K3 on B200 (SemiAnalysis, 2026-09-12); AI Chip Sales data explorer (Epoch AI, 2026-08-27); US Code growth 1991–2025 (arXiv:2511.13747, 2025-11-13); United States Code (GovInfo); Economic Report of the President, 2026 (Chapter 2) (Council of Economic Advisers, 2026); StateCodes: all statutory code from 50 U.S. states (Hariri & Ho, arXiv:2508.19365, 2025); Caselaw Access Project (Common Pile) (Hugging Face); Video understanding (Gemini API docs, Google, 2026-09-10); Media resolution (Gemini API docs, Google, 2026-09-02); RTVI-VLM Performance (NVIDIA VSS docs); The world's most surveilled cities (citing IHS Markit) (Comparitech); Streamlabs and Stream Hatchet Q2 2026 Live Streaming Report (Streamlabs, 2026); Automatic attack disruption in Microsoft Defender XDR (Microsoft Learn, 2026-06-11); EPS calculator (Logmanager); How Elastic InfoSec optimizes Elastic Defend (Elastic Security Labs, 2026-01-27); Dropzone AI (catalog entry) (Risky Business); Microsoft Digital Defense Report 2025 (Microsoft, 2025); Offboard Mode (PX4 Guide); Multicopter control architecture diagram (PX4 Autopilot); Most multirotors airborne simultaneously from a single computer (outdoors) (Guinness World Records, 2026-02-03); GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv:2503.14734); Breaking down the infinite workday (Microsoft Work Trend Index, 2025-06-17); Evolution of Conversations in the Age of Email Overload (Kooti et al., WWW 2015); Microsoft FY26 Q2 results (Office 365 for IT Pros, 2026-01-30); World Population Prospects 2024 (United Nations, 2024-07-11); The Ecology of Medical Care Revisited (Green et al., New England Journal of Medicine, 2001); Funding and services needed to achieve universal health coverage (Moses et al., Lancet Public Health, 2019); Health at a Glance 2025 (OECD, 2025); The size of the electronic health record at the point of care (Patterson et al., JAMIA Open, 2024); Health workforce levels and trends 2026 (World Health Organization, 2026); Court Statistics Project data (National Center for State Courts); State courts play a key role in American life (Pew Charitable Trusts, 2025-04); Supreme People's Court work report (Supreme People's Court of China, 2026-03-09); Pendency in Indian courts (Data for India); Justiça em Números 2026 (recap) (JUDIT, citing CNJ, 2026-06); Where the Money Goes: Understanding Litigant Expenditures for Producing Electronic Discovery (RAND Institute for Civil Justice, 2012); How many documents in a gigabyte? 2025 statistics for eDiscovery (Digital WarRoom, 2025); Federal Rule of Appellate Procedure 32 (Cornell LII); Live Streams (Sportradar); Coverage and data rights (Stats Perform); A decade of NFL Next Gen Stats innovation (Amazon Science); nflfastR: win probability models (nflverse); Why do you have a delay on placing bets on a market that is in-play? (Betfair Developer Support); Wikipedia:Size of Wikipedia (Wikipedia, 2026-09-13); Books of the world, stand up and be counted! (Inside Google Books, 2010-08-05); Will we run out of data? Limits of LLM scaling based on human-generated data (Epoch AI, 2024); Sundar Pichai at I/O 2026 (Google, 2026-05-19); Are women really more talkative than men? (Mehl et al., Science, 2007); A large language model for electronic health records (GatorTron) (Yang et al., npj Digital Medicine, 2022); Oral health fact sheet (World Health Organization, 2025-03-17); DeepSeek-R1 config.json (Hugging Face); Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures (arXiv:2505.09343, 2025); Jalapeño first results (OpenAI, 2026-08-25); OpenAI Jalapeño ASIC at Hot Chips 2026 (ServeTheHome, 2026-08-25).