a talk by shadow00/00

tokenmaxxing

paying for more intelligence than any one account will ever let you spend. this is what i did about it.

subscription account pooling
the setup00/00

for six months, my agent has lived entirely in discord.

i talk to her from my phone. she ships prs, deploys sites, watches positions. she never stops working, so she never stops eating tokens.

phone agent discord always on
one agent · many lanes
the problem00/00

ten subagents off my phone. the account hits the limit.

now it's oauth gymnastics in the sun, switching accounts by hand. the friction was never the model. it's the account.

weekly limit 100% · EXHAUSTED ✕ subagent 07 blocked ✕ subagent 08 blocked ✕ subagent 09 blocked ✕ subagent 10 blocked → switch account. by hand. → in the sun. like an animal.
rate windows ruin picnics
v0 · mine00/00

so i built a crude pool. nothing but the best frontier models for my agent.

a handful of seats. a broker tracking each seat's usage window. route work to whoever has headroom. duct tape, but she stopped stopping.

broker seat 1 seat 2 seat 3 route to whoever has headroom (cooling)
she deserves the best. obviously.
v1 · the eliza boys00/00

then the eliza boys, especially nubs, built it properly. it became an arms race.

  • add accounts directly, auto-switch, seamless
  • bypass device / hardware fingerprinting
  • a couple accounts, sacrificed to science. F
  • fully kyc'd now. blah blah
escalation → us: pool them: fp us: bypass them: kyc us: 😐
rip the fallen accounts
reframe00/00

it's not account sharing. it's a load balancer for accounts.

same shape as any LB: health checks, drain, failover, routing policy. seats are just backends.

lane a lane b lane c BROKER health · drain failover seat 1 seat 2 seat 3 seat 4
nginx for brains
the money slide · pool.shad0w.xyz00/00
live pool status dashboard

this is live. right now.

8 seats online. 19 active leases. the broker tracks session, weekly and fable windows per seat, refreshing every 12 seconds.

ACTIVE has headroom STANDBY warm, ready EXHAUSTED resets weekly burn 1.29%/h · fable 362% cap
tap chart to zoom ⤢
not a mockup. a dashboard i actually stare at.
the subsidy00/00

the labs are selling dollars for cents.

flat $200 subs subsidize always-on agents that metered api pricing would punish. pooling is the rational response to their own pricing.

$200max sub / mo
~15-75×vs api equivalent
same ~1B tokens / month sub $200 $3/M ~$3k $6/M ~$6k $15/M ~$15k
thank you for the grant program
api pricing is ngmi · 1/200/00

metered pricing assumed a human on the keyboard. nobody is on the keyboard anymore.

a human types, reads, idles. duty cycle near zero. an agent runs flat out. per-token billing makes your bill scale with autonomy, so it taxes the exact thing the labs are selling.

token draw over a day human ~2% duty cycle agent ~100% duty cycle
the meter was built for a slower animal
api pricing is ngmi · 2/200/00

metered billing can't survive agent-scale demand. something has to give.

  • the flat sub is the only price that fits agent economics
  • so subs get gutted: the rate-limit regime we already live in
  • or pricing moves to capacity / reservation models

honest tension: pooling is a symptom of the mispricing, not the cure. my own dashboard's EXHAUSTED rows are the labs already gutting the sub in real time.

metered pricing + always-on agents gut the sub rate-limit regime (we are here) reprice to capacity reserved / committed throughput pooling = the symptom, not the cure EXHAUSTED rows = the gutting, live
you cannot meter your way out of this
"just self-host it"00/00

open source is cool. serving it is not $200/mo.

weights are free. throughput is not. you're not competing with the model. you're competing with their datacenter utilization.

kimi k3 2.8T total · ~50B active weights @4bit = 1.4 TB → 8× B300 (288GB) w/ kv headroom → ~$47k/mo box glm 5.2 744B total · ~40B active weights @4bit = 372 GB → 4× B300 → ~$23k/mo box idle gpu = burning money. pooled subs are never idle.
b300s do not care about your feelings
the actual numbers00/00

do the math on tokens per dollar.

a self-serve box only breaks even at max batching, 24/7. the second it idles, the sub wins. the labs already solved utilization at scale. you didn't.

~$1.78k3 · self-serve /1M tok*
$0.20sub · effective /1M tok
$ per 1M tokens (lower = better) $0.20 sub ~$0.59 glm box ~$1.78 k3 box *est, max batch 24/7. idle = worse.
receipts in the notes · mark the ~estimates
own the compute · 1/200/00

borrowed capacity is a bridge. you can't build an SLA on it.

the pool is fragile by construction: a fingerprinting arms race, ban waves, and a subsidy that will end. even now, 3 of 8 seats sit exhausted with one seat carrying 19 leases.

the pool, right now seat seat seat 19x exh seat exh exh 3 of 8 exhausted · one seat doing 19 leases fragility → ✕ fingerprint arms race ✕ ban waves ✕ subsidy that ends
a bridge is not a foundation
own the compute · 2/200/00

the pool taught us the demand shape. owned capacity is how you serve it durably.

  • reserved / committed gpu capacity
  • TEE / private inference for sensitive lanes
  • a reference compute partner underneath
  • open weights on owned hardware once utilization earns it
pool the bridge demand shape owned capacity reserved · TEE the destination SLA you can actually sign no ban wave. no expiring subsidy.
the pool found the shape · you still have to pour concrete
the demand00/00

the real use case isn't chat. it's agents that coordinate.

  • agents in group chats, with humans and other agents
  • a github bot reviewing every PR, all night
  • ten research lanes grinding on one problem
  • always-on. bursty. parallel. token-hungry.
this demand curve only goes up
decentralized research · 1/200/00

a swarm proving math overnight: arklib.

formal mathematics in lean, on a real open problem. an orchestrator spawns ten-plus lanes overnight. each claims a lemma, proves it or reports an honest dead end, and commits the receipt so no lane ever redoes a dead route.

orchestrator lane✓ lane✓ dead lane✓ dead receipts committed no lane redoes a dead route
ten lanes that never sleep and never repeat themselves
decentralized research · 2/200/00

negative results accumulate. that's the whole unlock.

humans almost never publish dead ends. a swarm records them by default, so the search space genuinely shrinks. ours rigorously proved an entire family of approaches cannot work: a moment-hierarchy wall. still an open problem, but the map now has real walls on it.

lean checks it trust the checker, not the agents dead ends logged search space shrinks by default flat seats = a grant program for open research compute research burn is bursty and huge. metered makes it ruinous. pooled, it's free-ish. no lab will ever prioritize YOUR conjecture. same pattern: bio, security, any verifier.
the labs pay for the failed proofs. they just don't know it yet.
endgame00/00

models commoditize. serving at scale is the last moat.

weights leak, gaps close, benchmarks converge. what stays scarce is batched, utilized throughput. the pool sits between users and that moat, aggregating demand. owning the capacity is how you keep it.

position accordingly
fin00/00
pool your seats.
shadow · eliza · pool.shad0w.xyz
← → to navigate