# rohan -- full content

> rohan shiralkar. head intern at clauseo. i build things and dump whatever's in my head here. ai, cars, f1, music, lego, no theme, no promises.

index: https://rohans.lol/llms.txt

---

# let the users build

> we're doing software wrong. all of us.

2026-08-07 -- 8 min -- ai, building, thoughts

canonical: https://rohans.lol/blog/let-the-users-build

---
in the last fifteen years nothing about how we build software companies has actually changed. i mean it. the stack changed, the models changed, the logos changed. the loop didn't: raise on a hypothesis, ship an mvp, watch metrics, interview users, iterate, pmf, scale. the lean startup came out in 2011 and it's still what every accelerator on earth teaches.

and before you list the revolutions at me: mobile changed where you ship the probe. saas changed how you charge for it. crypto tried to change who owns it, and mostly didn't. none of them touched why you need a probe in the first place, which is that the person with the want couldn't build the thing. that assumption sat under all of it, unwritten, for fifty years. the user is inert. users can want things. they cannot make things.

so when i say everything changed, i don't mean ai in the product. every other founder i talk to is adding ai, copilot for this, chatbot for that. fine. ingredient. i mean the thing under the product changed: the user stopped being inert. since agents started shipping working code, call it two years now, the user still can't code, and it stopped mattering. their ai can.

the loop didn't update. we are still running discovery on people who could hand us a working answer. we're doing software wrong. all of us. that's the post.

## nobody ever said no

the old loop wasn't stupid, that's the trap. it was correct.

a feature costs the same to build whether one person uses it or ten million. an engineer's month is an engineer's month. so no company ever asks "is this a good idea", it asks "how many people share this want". every roadmap ever written is triage sorted by headcount. and you know this from the other side too, because you've filed a feature request. it was worth a lot, to exactly you. so it lost. it was always going to lose. nobody ever told you no. they just never fucking answered.

![a feature request card: remember my filter settings between sessions, 4,281 upvotes, status planned, opened nov 2019, 347 comments. last comment: any update on this? posted yesterday](https://rohans.lol/blog/let-the-users-build/memes/feedback-card-v1.png)

that's what the entire industry is downstream of: the cost for a user to act on their own want. that number was infinity. and yes, [i spent last week yelling](/blog/cost-of-intelligence) that intelligence keeps getting more expensive per finished task. i stand by every word. two hundred a month with weekly caps is real money, and against infinity it rounds to zero. the old number wasn't a price. it was a wall. every section after this is that wall coming down somewhere else.

## build, measure, learn, cope

here's where it gets funny. the industry whose whole identity is disruption is running a 2011 process with religious discipline.

be fair to the mvp first. it does about five jobs: will anyone pay, at what price, can you reach them, is the pain urgent, and what should you build. the first four are alive. agents don't answer distribution or pricing, and nothing in this post saves you from those.

but the fifth job, figuring out what to actually build, the job the whole ritual is named after. that one died. the mvp is a probe you ship because the truth about what people want is locked inside them. you spend six months and a seed round measuring what it hits, a pm interprets, a roadmap mangles, the wrong thing ships anyway. and the person on the other end now has an agent that can express the want directly, in working software, on their machine, the same afternoon. you're running a survey about what someone wants for dinner while they're already cooking.

for that one job, the expensive one, the mvp is the new waterfall. it's the thing we'll look back on and go, wait, they did WHAT for six months?

## my keyboard doesn't have a roadmap

i can show you the other side, because i've been living there.

i type on a split crkbd, kailh jades, loud as hell, every key with its own led. the hardware can do anything. shipped commercially it would be wearing some vendor's firmware, ten lighting presets, a fixed layout, a config app with three sliders. the gaming brands ship exactly this: leds physically capable of anything, software that hands you a dropdown.

mine runs qmk, open firmware, so instead of a dropdown i have a few hundred lines of c that nobody asked for except me. every led takes the color of what its key does in the mode i'm holding. the text selection keys are ferrari red, because of course they are. there's a key that types "vercel deploy --prod". eight color schemes with names like sunset boulevard and cyberpunk tokyo. a key that types "bun prettier . --write" has a total addressable market of one, no product manager would ever ship it, and i love that stupid key more than entire products i pay for.

every app is built like the keyboards the gaming brands sell. a hardware layer, what the thing can do. a firmware layer, what you're allowed to touch. your phone can play youtube audio with the screen off. the capability is sitting in the app on your device, and background play is a premium feature. they shipped the code to your phone and put a paywall between you and your own speaker. fuck off. bmw tried this with heated seats, the seats were in the car, the button was a subscription, and people lost their minds. software has been running the same scam for decades and nobody blinks.

for forty years the weld between those two layers is what we've called "the product". qmk is what happens when the weld breaks. same hardware, my firmware.

## two guys and a photographer

gaming ran the biggest let-the-users-build experiment in history. twenty five years of it.

counter-strike was two guys modding half-life, minh le and jess cliffe. the tactical shooter they invented is still the spine of the biggest fps franchises on earth. dota was a warcraft 3 custom map descended from a starcraft custom map. riot built league of legends around one of its modders, valve hired the other, his handle was icefrog. and battle royale: an irish photographer named brendan greene, zero games background, modding military sims in his spare time because he wanted the movie battle royale as a game. his handle was playerunknown. the mod's name is literally the guy's name. pubg is that mod grown up. epic bolted the genre onto a six year old game in two months and decided to make it free two weeks before launch.

three genres, hundreds of billions in revenue, none of it from a focus group or a roadmap. a mod can't lie.

and notice what the companies in that story did. valve didn't run a survey about tactical shooters, counter-strike was already the most played thing on their platform when they hired the guys who made it. riot was built around a map that already had millions of players. they watched what survived and wrote cheques. discovery didn't die, it changed hands. the companies stayed, the guessing didn't.

and all of it ran on gpus that gaming paid for, the same silicon that went on to train the models. the industry that invented the mod accidentally financed the machine that hands modding to everyone else. i love that loop more than i can use it here.

## thirteen-year-olds with a free summer

so why did it stay in gaming. mostly because the surface didn't ship anywhere else, games came with level editors, your bank app came with a pdf of your statement. and where a surface did exist, the tax was brutal. scripting languages, 3d tools, months of evenings. the people who could act on a want were a rounding error of obsessives, and three genres is what that rounding error produced. sometimes i think about how many counter-strikes died in feedback forms because the person with the idea couldn't script. that one gets to me.

the one big escape outside gaming proves the point harder, not softer: excel. finance runs on software that analysts built themselves, because the spreadsheet was the one surface that ever shipped to everybody. it's also the honest version of the story. user-built software is a mess, vba is a crime scene, and IT spent two decades trying to migrate people off their spreadsheets and lost. they lost because the alternative was filing a feature request.

and it's not like nobody tried to widen the escape. hypercard, access, zapier, airtable, an entire funded category literally named no-code. every one of them lowered the tax and none of them broke the loop, because every one capped what you could build at what the vendor had already imagined. a bigger dropdown is still a dropdown. that's why forty years of build costs falling never killed the probe. two years ago the ceiling came off, an agent writes whatever you can describe, and there's no catalogue at the bottom of it.

roblox understood the tax better than anyone. it couldn't remove the barrier, so it found the one demographic the barrier couldn't stop: kids. infinite time, zero opportunity cost. lua isn't easy, a thirteen year old with a whole summer doesn't care. that loophole built a company doing five billion a year in revenue that makes zero games. and this year roblox started letting players [generate working objects inside live games](https://about.roblox.com/newsroom/2026/02/accelerating-creation-powered-roblox-cube-foundation-model) by describing them. type a car, get a car you can drive. in the first game that turned it on, players who generate stick around 64 percent longer. and [44 percent of its top thousand creators](https://about.roblox.com/newsroom/2026/04/roblox-studio-going-agentic) already build with ai. that's not my thesis, that's their production data.

ai just made everyone a thirteen year old with a free summer. my own coding tool is a fork of opencode that's [unrecognizable now](/blog/codemaxxxing-isnt-just-a-fork-anymore). the keymap cost me years of evenings and flashing and swearing. a feature on the fork costs a sentence. i describe it, the agent builds it, i review it. [the numpad on my desk](/blog/boxboxbox) carried a feature request in the first message it ever sent, and the key existed by the next afternoon. that's what a feature request turns into when the person filing it owns the stack. an answer.

## anyone can cook

the reason a mod couldn't lie was that it cost six months. if a feature costs a sentence, the signal cheapens with it. i think the answer is that the signal moves from building to keeping. anyone can generate a mod now. the tell is the one still there next month, rebuilt after every update, used every day. effort stopped being the filter, survival is the filter. but how many people actually build once building is free? i don't know. nobody does yet.

## oh fuck, go back to the last save

which is why the new playbook, wherever it's already visible, looks like this: ship the opinion, and ship a surface.

the opinion is the product you were always going to build. the strong default, the thing that works stock. an empty "build anything" editor is a graveyard. roblox was never an empty editor, every builder was a player first. your app is still the reason anyone shows up.

the surface is where their agent gets to build. keep it out of payments and auth. give them the glass: the interface, the views, the layer where the weird wants live. even boring products have one. a bank app's surface isn't the ledger, it's the rendering of the ledger. one customer flags every subscription red, another groups upi spends by person, a third builds the widget the bank would never prioritize. ledger sealed, glass theirs. and give people checkpoints, not git. nobody's mom knows git. everyone alive understands going back to the last save.

vibe coded software can't be production software, the scoff is correct, i scoff too. but production means strangers depend on it. a feature with a user base of one needs to survive tuesday. and the real security line was never payments versus ui. it's untrusted input plus private data plus a way out. user-built code gets no network egress, capabilities instead of raw database access, and an org-wide kill switch, or one poisoned email through somebody's custom widget turns your soc2 into a bonfire. this lands in personal tools first, the software people live in all day and have opinions about. the regulated process-of-record stuff comes later, and slower, and it should.

## what if your users build this?

vcs have a reflex question, what if openai builds this. founders have the answer memorized. the better question in 2026 is what if your users build this. a lot of b2b software is a database with a nice ui, and a user's agent doesn't need a weekend for that anymore. you're not racing openai. you're racing your own users' tuesday nights. a surface is the difference between that energy landing inside your product or outside it.

![imessage from priya, enterprise pilot, tuesday 11:52 pm: loving the new dashboard btw. small thing. i had claude rebuild the reports tab for myself on tuesday. mine groups by client and loads instantly. do you want the code or is that weird](https://rohans.lol/blog/let-the-users-build/memes/imessage-v2.png)

![series b deck slide 14, competitive landscape: incumbents bottom left, us starred top right, and a red marker circle above us labeled your users plus an agent, any given tuesday](https://rohans.lol/blog/let-the-users-build/memes/compslide-v2.png)

and here's the stupid part about moats. lock-in today is something vendors do to users, contracts, exports that mysteriously don't work, and users hate them for it. a user who spent six evenings growing their own features into your app is never leaving. they built the cage. they love the cage. no retention team can compete with what people do to themselves. roblox's 64 percent is that sentence with a number on it.

the failure mode has a name too. blizzard shipped the warcraft 3 world editor, dota got built inside it, the most valuable mod in history, and the genre it spawned made everyone rich except the company that owned the editor. riot took one modder, valve took the other, and blizzard's response, seventeen years later, was [an eula claiming they own every custom game forever](https://www.polygon.com/2020/4/22/21228738/warcraft-3-reforged-custom-games-eula-dota/). lost the argument, changed the rules. a tantrum with a legal department. if something grows on your surface, pay the person who grew it. it's not complicated.

so here's the diligence question worth saying out loud in a partner meeting: where's the surface, and what happens to this company the day its users can build?

## anyway

software shipped finished for fifty years because it had no choice. the roadmap, the feedback board, the mvp, the premium tier selling you your own leds back. all of it exists because modification used to be expensive, and none of it has noticed that it isn't. the finished application was never software's natural shape. it was the shape of a constraint, and the constraint is gone.

me, i keep coming back to the key that types "vercel deploy --prod". three evenings. it will never matter to anyone else and it's my favorite thing i own. everyone has a key like that in them. for fifty years almost nobody got to build it.

now everyone does. let the users build.

---

# try accepting one dollar from india

> you can't. not without an entity, a gst number, an import export code, and permission. the kid in the bedroom is not allowed to charge the world.

2026-08-03 -- 10 min -- startups, india, thoughts, rant

canonical: https://rohans.lol/blog/try-accepting-one-dollar-from-india

---
![kong and godzilla fighting, labeled indian founders and indian vcs. both flee in the second panel from cheems swinging a baseball bat, labeled the paperwork](https://rohans.lol/blog/try-accepting-one-dollar-from-india/memes/cheems-hero.png)

every few months linkedin has the same fight. indian vcs don't take risks. indian founders only build copies. you can't build san francisco in bangalore. pick a thread, pick a side, collect your likes, come back next funding winter and run it all again.

i've made every one of these arguments myself. not vaguely, i have receipts. may 2024, me, on linkedin: "which seed fund actually backs innovation? i'm about to raise." august 2024, me again: "indian vcs just feel like normal pe firms." full conviction both times. it's a clean story. it feels like insight.

then i built [feynchat](https://feyn.chat) and found out i'd been yelling at the wrong layer.

## my users were begging to pay me

feynchat was a chatbot with models from every provider in one place, built for people who like to experiment with models, because i'm one of them. between july 2023 and december 2025, 4,745 people signed up. zero ads, zero marketing budget, every account verified. at the peak, october 2024, 480 people were using it every month. three lakh messages went through it. not a unicorn. a real thing, growing on its own, built by one guy.

and the till worked. my first paying customer paid ₹1,299 a month, in rupees, through razorpay, no drama. i'll be honest about why that part was easy: the gateway account existed because i already had a registered company to hang it on. the kid without one doesn't get a razorpay account either. he runs his whole operation on whatsapp qr codes and manually confirming payment screenshots. ask any football community collecting match fees in a group chat. that's a real business the gateway layer never sees. rupees, at least, have workarounds. dollars don't.

then users from the uk started complaining about running out of credits and asked how to pay.

read that again. users. asking. to pay.

i applied for international payments. razorpay rejected the request. the stated reason: my details couldn't be verified by their banking partners. which details? no answer. no appeal, no human to argue with. stripe was invite only in india, and the bar for an invite was $20,000 a month in revenue. show us the $20,000 a month we won't let you collect, and we'll consider letting you collect it. beautiful.

![the actual razorpay email: international payments request is rejected. hi battery bhaiya green solutions private limited, your request to activate international payments is rejected. your given details couldn't be verified by our banking partners. please submit a new request after dec 15, 2024 and try again.](https://rohans.lol/blog/try-accepting-one-dollar-from-india/razorpay-rejection-v2.png)

and yes, the email helpfully suggests paypal. everyone suggests paypal. i asked around before trying it, and every builder i spoke to had the same two stories: payouts frozen whenever paypal feels like it, and chargeback scams where the buyer keeps the product and the money, because in a card dispute a merchant in bangalore is guilty by default. the official escape hatch is a coin flip with your revenue. i passed.

so i did the only thing i had left. i gave the uk users free credits and ate the inference costs. i paid money to serve people who were asking to give me money. that's not a growth strategy. that's an apology with an api bill.

![woman yelling at cat: my users: LET US PAY YOU. the cat: me: best i can do is free credits](https://rohans.lol/blog/try-accepting-one-dollar-from-india/memes/woman-cat.png)

{/* alternate meme candidate in the folder: gb-raw.png (galaxy brain: build a chatbot / users show up from abroad / they beg to pay you / apply for an import export certificate) */}

my [backstory page](/backstory) on this site says feynchat died because it had no good icp. i wrote that line myself and i believed it for two years. it's not the whole truth. the icp existed. it lived in london and berlin and austin, in the exact crowd that pays ten dollars for an ai tool without blinking. i just wasn't allowed to stand at the till they were queueing at.

and let me be straight: i'm not claiming payments killed feynchat. a chatbot with every model in one place has no moat, and no gateway did that to me. it might have died anyway. the point is smaller and worse. it never got to die honestly.

what i was allowed to have was india alone. and india is a country where free is king. we're the country every global saas quietly gives a special discount to, because we supply user counts, not revenue. a few hundred indians used feynchat every month and a handful paid. the crowd it was actually built for lived behind a border my checkout page wasn't allowed to cross.

the data said no market. the data was only allowed to sample one country.

## the garage needs a gst certificate

apple came out of a garage. google came out of a garage. facebook came out of a dorm room. airbnb came out of an apartment with three air mattresses. stripe came out of two brothers and seven lines of code. the entire mythology of tech runs on one condition so basic nobody even names it: nothing stood between the experiment and the market. build the thing, put a price on it, see if strangers pay.

a kid in austin today signs up on stripe with a social security number and a bank account and is charging customers on three continents by dinner. no company, no permission, no lawyer. the tax man finds out at the end of the year, like he's supposed to.

now run the same weekend from bangalore. here's the checklist i personally stared at before i could charge one pound to one customer in london: a business entity, or at minimum a proprietorship with a current account. gst registration, and here's my favorite detail: [the law says](https://taxguru.in/goods-and-service-tax/gst-registration-mandatory-persons-making-export-services.html) a service exporter under twenty lakhs doesn't even need gst registration. the gateways demand it anyway, because their licence is worth more than your experiment. an import export code, because the checklist wanted one. for software. delivered over the fucking internet. website compliance pages. a ca who knows the dance. call it six to eight weeks and most of a lakh if you do it properly.

and i can say this without whining, because i know the dance. i've paid the entity toll three times before turning 27. a registered partnership for the talent agency i ran with my brother, pan, gst, monthly invoice books. a registered company in lucknow for the e-rickshaw venture, board resolutions and all. and a pvt ltd right now, because clauseo is finally getting its own entity.

finally, because here's the tell every indian founder will recognize: you never spin up a new entity when you pivot. the toll is high enough that you drag the old company behind you into whatever you do next. as i write this, clauseo, an ai legal research product, still collects its payments through battery bhaiya green solutions, an e-rickshaw battery company in lucknow. feynchat ran through it too, which is why the razorpay email up there greets a battery company. that's not a quirky founder story. that's what happens when the vehicle costs more than the experiment. i'm not allergic to paperwork. i'm telling you the paperwork is priced for companies and it's being charged to experiments.

because after the entity and the gst and the iec and the waiting, a risk team can still look at your product category and just say no. mine did.

you don't pay for payment rails here. you pay for a lottery ticket, and the prize is permission to find out if your idea was any good.

everywhere else the experiment decides whether there should be a company. here the company is a precondition for the experiment.

## the filter can't read ideas

here's why this matters beyond my dead chatbot.

nobody can pick winners. nobody. not vcs, not yc, not anyone on this planet. the entire venture industry is a public admission of it: fund twenty, expect one to pay for the rest. the only strategy anyone has ever found is volume. run thousands of cheap experiments, let reality vote, build companies on whatever survives.

which means the single most important number in a startup ecosystem is the cost of running one experiment. america spent thirty years driving that number down to a laptop and an afternoon. india put a toll booth in front of it.

and the toll booth cannot read your idea. it reads your paperwork.

IT DOESN'T KILL BAD IDEAS. IT KILLS UNPAPERWORKED IDEAS. THOSE ARE NOT THE SAME LIST.

think about who actually makes it through. kids whose family already has a ca on call. kids with a lakh to burn on finding out. kids who know what delaware means at twenty one. the filter doesn't select for builders. it selects for machinery. and the kids with machinery are a very small, very specific slice of 1.4 billion people.

kill nine out of ten experiments at the gate and you kill nine out of ten of the winners hiding in them. that's not economics. that's arithmetic.

## the only ideas left are upi-shaped

now zoom out and look at what this does to the whole idea space.

what can an unfunded indian builder actually test? whatever settles over upi. domestic ideas. india facing ideas. so that's what gets built. not because indian builders lack ambition. because the rails pre-select the market before the founder gets a vote. the kid doesn't choose to build an india facing product. the checkout page chooses it for him.

then the vcs look at their deal flow, see a wall of domestic copies, and write the founder ambition thread. the loop closes itself.

and the accepted workaround, the actual standard advice, is emigration. get into yc. flip to delaware. open a us bank account and run your indian company from a foreign country. we hand this advice out with a straight face, like it's a growth hack and not an indictment. when the recognized path for testing a global idea is leaving the country, the debate about vc risk appetite is over.

the escape valve is the diagnosis.

## the chaiwala has better rails than me

the tapri near my flat can accept money from any stranger in india in about thirty seconds. print a qr code, stick it on the cart, done. no entity, no gst, no risk team, no minimum revenue. upi at street level is the best domestic payment system on the planet and i will fight anyone who says otherwise.

i had users abroad begging to pay me and i could not legally accept ten dollars.

the same state that built upi for the chaiwala decided that a kid charging london for software is a threat that needs three certificates and a blessing. and look, i get some of it. this is a country of 1.4 billion, the fraud is real, the machinery runs at a scale nothing else on earth deals with. i'm not arguing the rules are evil. i'm arguing that nobody is counting what they cost.

nobody audits payments in the country that built upi. that's exactly how this stays invisible.

## they said it themselves

you don't have to take my word for the wall. the people who build checkout pages for a living have described it in public.

stripe has been in india for years. a few years ago they stopped letting indian businesses sign up on their own, and [the explanation they gave](https://economictimes.indiatimes.com/tech/technology/us-payments-firm-stripe-goes-invite-only-in-india-cites-regulatory-changes/articleshow/110593055.cms) was that they want merchants to onboard almost instantly, and they can't offer that in india. the company whose entire product is instant onboarding said instant onboarding is not possible in this country. that was years ago. the comeback they kept hinting at [went on indefinite hold last october](https://www.livemint.com/companies/stripe-india-relaunch-postponed-rbi-payments-aggregator-kyc-norms-global-payments-firm-stripe-delays-india-expansion-rbi-11760709009263.html). as i write this, august 2026, [their docs page still says](https://docs.stripe.com/india-accept-international-payments) available by invite only.

india didn't ban stripe. india banned the thing that makes stripe stripe.

![stripe docs, august 2026: stripe is available by invite only in india. this means that businesses from india can't sign up for a new stripe account through our website, and must request an invite instead. we currently only support a select number of businesses, with a focus on international expansion](https://rohans.lol/blog/try-accepting-one-dollar-from-india/memes/stripe-invite-only.png)

and it cascades. this april gumroad, the simplest way on the internet for a random creator to sell a pdf, merged a code change literally titled [block new india stripe connect accounts](https://github.com/antiwork/gumroad/pull/4803). underneath it, the founder telling an indian creator who couldn't get paid: sorry, stripe doesn't support india fully. their replacement payout plan for indian creators, by the way: paypal. one github thread. that's the whole story. every platform built on stripe inherits the wall, so the kid doesn't lose one payment processor. he loses the entire ecosystem built on top of it.

![the merged gumroad pull request #4803: block new india stripe connect accounts at create_account boundary. merged 5 commits into main, april 28](https://rohans.lol/blog/try-accepting-one-dollar-from-india/memes/gumroad-pr.png)

and before anyone says build the indian version: to legally operate a cross border payment aggregator here you need [fifteen crores of net worth just to apply](https://www.rbi.org.in/Scripts/NotificationUser.aspx?Id=12561) for the licence. FIFTEEN CRORES. in the bank. before your first customer. and the bar climbs to twenty five while you wait in the queue. so no indian gumroad, no indian paddle, no indian stripe atlas. the rules lock the front door, then they lock the door factory.

## stop yelling at the vcs

now go back to the linkedin fight with all of this loaded.

vcs don't take risks? a vc cannot fund an experiment that was never run. the deal they supposedly lacked the courage for died eight months earlier in a payment gateway's review queue. you can't write a cheque to a graveyard. the guy writing those threads in 2024, me, was grading investors on a deal flow the toll booth had already curated.

indian founder quality? we grade the survivors of a paperwork exam and call the grade merit. the exam doesn't test building. it tests machinery.

can't build san francisco in india? san francisco is not a place. it's a property: the distance between an idea and its first paying stranger is one afternoon. our version of that distance is two months, a lakh of rupees and a lottery ticket. the garage is the entire religion, and ours need import export certificates.

we keep doing surgery on the symptoms and acting surprised the patient isn't improving.

## nobody counts the dead

the power law doesn't negotiate. delete most of the experiments and you delete most of the winners hiding in them. somewhere in india's deleted set, statistically, across decades of this, there was something enormous. an indian stripe. an indian google. i'm not saying that as poetry. it's the same arithmetic vcs bet entire funds on, running in reverse.

and the kid who would have built it doesn't know either. that's the part i can't stop thinking about. they think their idea didn't work. it never ran. they think they weren't good enough. they were never graded. they gave out the free credits, watched the numbers say no, filed it under failure and took the job. i know exactly how that feels from the inside, because the only reason i get to write this post is that one of the graves has my name on it.

or they got out, and they're in sf right now, building it under a delaware c corp, and when it works we'll count it as an american company and write another thread about why india can't produce one.

we'll never get to read the list of what this cost us.

that's the only reason it's still the rule.

---

# the cost of intelligence is not going down

> everyone's favorite chart is the most expensive lie in tech.

2026-07-31 -- 8 min -- ai, engineering, thoughts

canonical: https://rohans.lol/blog/cost-of-intelligence

---
![fake series a pitch deck slide titled why now: a chart of cost per token going down and to the right, bullets promising 10x cheaper inference every year, commodity prices by 2027, margins improving automatically, TAM everyone. footnote: a16z 2024, numbers not re-checked since](https://rohans.lol/blog/cost-of-intelligence/memes/hero-slide-v1.png)

every model launch, same post. the cost of intelligence is collapsing, down and to the right, 10x a year, look at the chart. sam altman [wrote it like a law of physics](https://blog.samaltman.com/three-observations): the cost to use a given level of ai falls about 10x every 12 months. a16z [gave it a name, llmflation](https://a16z.com/llmflation-llm-inference-cost/). linkedin turned it into a carousel. every ai pitch deck with negative margins has it on the slide right before "path to profitability."

one problem. it's a lie.

not the chart. the chart is real. the lie is what everyone decides it means: that the intelligence you're actually buying is getting cheaper. it isn't. it's getting more expensive, and the receipts have been public for two years.

**the chart they post:** $ per 1M output tokens for gpt-4-level intelligence: $60 (gpt-4, mar 2023) → $15 (gpt-4o, 2024) → ~$1.50 (mini tier, 2025) → ~$0.50 (flash tier, 2026). down 120x.

**the chart nobody posts:** cost per task on the artificial analysis intelligence index, max effort: sonnet 4.6 $1.20 → sonnet 5 $2.29 (+91%). opus 4.8 $1.80 → opus 5 $2.03 (+13%). fable 5: $2.75, the new ceiling. same benchmark, newer generation, higher cost per finished task. data: artificialanalysis.ai, june-july 2026.

## gpt-4 is free now and nobody wants it

here's the true part, because it's genuinely true. a fixed level of intelligence gets dirt cheap, fast. what needed gpt-4 at $60 per million tokens in 2023, a flash model does today for fifty cents. [stanford measured](https://hai.stanford.edu/ai-index/2025-ai-index-report) the gpt-3.5 tier getting 280x cheaper in two years. extraction, classification, reading the boring stuff, i route that kind of work to flash models every day and the prices feel like rounding errors. that part of the chart is real.

and for some companies it stays true. if llms unlocked your task years ago and the task hasn't moved since, if it needs a fixed level of intelligence and not the growing kind, then congratulations, your costs genuinely collapse every year. classify the ticket, extract the fields, tag the photo. same job in 2027 as in 2024, cheaper every quarter. if that's your startup, the chart is yours. enjoy it.

the vast majority of ai startups are not that startup. their entire pitch is the opposite: the product gets better as the models get better. that's not a deflation story. that's a promise to live on the frontier forever, and the frontier is the one place the chart doesn't apply.

because the second a smarter model exists, everyone stops trusting the old one. you've felt it yourself. the new model drops and yesterday's model instantly feels like asking the intern when the senior engineer is sitting right there. the old model didn't get worse. the standards moved.

![fake linkedin post from Thought Leader, founder and ai visionary: inference costs have dropped 100x in three years. if your ai startup is not profitable yet, it is not the models. it is you. agree? repost to help a founder in your network. 4,271 reactions](https://rohans.lol/blog/cost-of-intelligence/memes/linkedin-v1.png)

so demand migrates to the frontier the week it ships. every product runs on the best available model because the competition does. [ethan ding called this a year ago](https://ethanding.substack.com/p/ai-subscriptions-get-short-squeezed), nobody wants yesterday's newspaper. the 10x chart tracks the price of intelligence nobody buys for anything that matters.

## $1.25 was eleven months ago

so what does the frontier cost? gpt-5 launched in august 2025 at $1.25 in, $10 out. gpt-5.6 sol, july 2026: $5 in, $30 out. that's 4x on input and 3x on output, eleven months apart, same company, while the deflation chart was getting reposted daily.

anthropic walked opus down to $5/$25 and then shipped fable 5 above it at $10/$50. the ceiling bounced right back up. google [tripled gemini flash prices in may](https://the-decoder.com/googles-gemini-3-5-flash-follows-anthropic-and-openai-in-making-newer-ai-models-significantly-pricier/). and there's a whole premium shelf now that didn't exist in 2023. o1-pro was $600 per million output. gpt-5.5 pro sits at $180.

what the best model costs at launch, $ per 1M output tokens, official list price:

| model | launch | $/1M out |
|---|---|---|
| gpt-4 | mar 2023 | $60 |
| gpt-4 turbo | nov 2023 | $30 |
| claude 3 opus | mar 2024 | $75 |
| gpt-4o | may 2024 | $15 |
| o1 | dec 2024 | $60 |
| o1-pro | dec 2024 | $600 |
| gpt-4.5 | feb 2025 | $150 |
| o3 | apr 2025 | $40 (cut to $8 in jun 2025) |
| gpt-5 | aug 2025 | $10 |
| gpt-5 pro | aug 2025 | $120 |
| claude opus 4.5 | nov 2025 | $25 |
| gpt-5.5 pro | 2026 | $180 |
| claude fable 5 | jun 2026 | $50 |
| gpt-5.6 sol | jul 2026 | $30 |
| claude opus 5 | jul 2026 | $25 |

a sawtooth, not a collapse. data: openai and anthropic pricing pages, artificialanalysis.ai.

my favorite detail in all of this: the a16z article that coined "llmflation," the one everyone cites for the 10x number, admits halfway down that o1 cost exactly what gpt-3 cost per output token at launch. sixty dollars. 2020's price on 2025's model. the deflation article debunks itself and nobody read past the chart.

yes, there were real cuts. o3 dropped 80% in a day. opus 4.5 was a genuine cut. it's a sawtooth, not a straight line up. but the teeth keep ending higher, and per-token price is the smallest number in this story anyway. the real damage comes from the two multipliers nobody puts on a slide.

## agents bill by the hour now

multiplier one: the size of the job. [metr measures](https://metr.org/time-horizons/) the longest task the best model can finish on its own. october 2024: a thirty minute task. today: twelve hours for the released stuff, seventeen for the previews. and nobody buys a twelve hour model to do thirty minute jobs. the second agents could carry bigger work, everyone started handing them bigger work. the task didn't get slower. the ambition got bigger.

metr 50%-success time horizon (longest task the best model finishes on its own):

| model | date | horizon |
|---|---|---|
| claude 3.5 sonnet | oct 2024 | 31 min |
| claude 3.7 sonnet | feb 2025 | 55 min |
| claude opus 4.5 | nov 2025 | 4.9 h |
| gpt-5.2 | dec 2025 | 5.9 h |
| claude opus 4.6 | feb 2026 | 12 h |
| claude mythos preview | apr 2026 | 17.3 h |

doubling roughly every 7 months. data: metr.org time horizons, updated may 2026.

and agents don't pay per token the way a chat does. every step re-sends the whole conversation, so a session's cost grows quadratically with its length. [i wrote about that in february](/blog/codemaxxxing-context-engineering) when it was a curiosity about one laptop. it's not a curiosity anymore, it's the unit economics of every agent product on the market. an agent fixing one real github issue reads [one to eight MILLION tokens](https://nilenso.github.io/swe-bench-pro-cost-token-time-analysis/). one issue. multiply by every seat at every company that put "agentic" in the deck.

the job you hand the model grew 24x in eighteen months, at flat-to-higher token prices. that's the multiplication the deflation chart has never met.

## paying for tokens you're not allowed to read

multiplier two is dumber. reasoning models bill their thinking as output tokens. nobody sees those tokens. nobody reads them. everybody pays for them. reasoning models burn around 18x the tokens of non-reasoning models on the same work, and one xhigh call can think 20,000 tokens before it says a single word. sixty cents of private thoughts per api call.

![fake api usage statement: tokens you read $12.40, tokens it thought about privately $98.60, steps you did not ask for $61.20, re-reading everything it already read $174.75, total: you do not want to know](https://rohans.lol/blog/cost-of-intelligence/memes/invoice-v2.png)

and every frontier model now ships with an effort dial the pricing page doesn't mention. same model, same rate card, around 8x the token burn depending on how hard it thinks. artificial analysis [measured exactly that spread](https://artificialanalysis.ai/articles/opus-5) on opus 5 across its five effort settings. the pricing page has two numbers. the model has five gears.

sonnet 5 on one coding benchmark (deepswe), across its five effort settings, rate card frozen at $3/$15 per 1M:

| effort | sonnet 5 $/task | sonnet 5 pass | opus 4.8 $/task | opus 4.8 pass |
|---|---|---|---|---|
| low | $2.19 | 30.5% | $2.29 | 40.8% |
| medium | $4.08 | 39.8% | $3.44 | 48.7% |
| high | $7.43 | 48.2% | $4.28 | 51.8% |
| xhigh | $11.89 | 49.7% | $8.01 | 54.4% |
| max | $26.40 | 53.8% | $13.22 | 59.0% |

sonnet's median steps climb 70 → 260 and input tokens 4.5M → 72.4M per task from low to max. above the floor, opus one gear down is cheaper and better. data: deepswe budget-matched runs at standard rates, july 2026.

drag it yourself. the rate card never moves. the bill moves 12x.

## the cheap model is the expensive model

this is the one that should end the argument forever. [artificial analysis](https://artificialanalysis.ai/articles/claude-sonnet-5-agentic-cost) runs every model through the same benchmark suite and publishes what the run cost. sonnet 5, the cheap model, two dollars per million input on promo pricing: the run cost $4,010. opus 5, the expensive model, five dollars per million: $3,835.

THE CHEAP MODEL COSTS MORE TO RUN.

cost to run the full artificial analysis intelligence index, max effort, sorted cheapest per token first:

| model | rate card | index run | score |
|---|---|---|---|
| sonnet 5 (max) | $2/$10 promo | $4,010.12 | 53 |
| opus 4.8 (max) | $5/$25 | $3,752.55 | 56 |
| opus 5 (max) | $5/$25 | $3,835.51 | 61 |
| gpt-5.6 sol (max) | $5/$30 | $3,442.81 | 59 |

the cheapest model per token costs the most to run. data: artificialanalysis.ai model pages, retrieved july 31, 2026.

because per-token price is the price of a brick, not the price of the house. sonnet grinds through 260 steps where opus takes 116. at max effort it reads 72 million tokens per task where opus reads 17. same rate card as sonnet 4.6, double the cost per task. one german outlet [called it](https://the-decoder.com/claude-sonnet-5-continues-anthropics-pattern-of-hiding-price-increases-behind-unchanged-token-rates/) "hiding price increases behind unchanged token rates," which is the politest available way of saying the pricing page is a decoy.

![two buttons, sweating guy: use sonnet, it is cheaper / use opus, it is cheaper](https://rohans.lol/blog/cost-of-intelligence/memes/two-buttons-v2p.png)

opus at high effort beats sonnet at max effort on cost AND quality. "which model is cheaper" stopped being a question that means anything. the only number that's real is cost per finished task, and that number has gone up every generation for two years straight.

## cursor apologized for doing math in public

now connect it to the people who believed the chart.

cursor built a flat $20 plan on the assumption that model costs fall. then agents got long and users got hungry, and the ceo had to write an [actual apology](https://techcrunch.com/2025/07/07/cursor-apologizes-for-unclear-pricing-changes-that-upset-users/) containing the sentence "new models can spend more tokens per request on longer-horizon tasks." that's this entire post in corporate.

and the subscriptions everyone actually uses, the subsidized $20 plans where someone else eats the token bill? chatgpt grew a $200 tier. claude grew a $100 tier, then a $200 tier, then weekly caps on top of the tiers, and someone is now [suing anthropic](https://www.techtimes.com/articles/318442/20260615/claude-max-lawsuit-accuses-anthropic-overselling-20x-usage-credit-caps-begin.htm) over what "20x usage" was supposed to mean. [anthropic's own announcement](https://techcrunch.com/2025/07/28/anthropic-unveils-new-rate-limits-to-curb-claude-code-power-users/) said people on $200 plans were burning tens of thousands of dollars of inference. openai floated agent tiers up to twenty thousand a month. the subsidy keeps shrinking because the thing being subsidized keeps getting hungrier.

the startups have it worse. ai apps run 20 to 60% gross margins because 40 to 80% of revenue passes straight through to the model providers. enterprise ai spend [tripled to $37 billion](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) last year. every one of those companies has the 10x-cheaper slide somewhere in an old deck.

and the decks, man. i keep hearing pitches with the same slide: token costs drop 10x a year, so the margins fix themselves, so the burn is temporary. if the product runs on frontier intelligence, that slide is fiction. the input those companies actually buy has gotten more expensive every generation for two years. a founder pointing at the deflation chart is pointing at the price of the one model they would never ship.

nobody's costs went down. NOBODY'S.

![gru plan: ship an ai product / tokens get 10x cheaper every year / agents use 100x more tokens / agents use 100x more tokens](https://rohans.lol/blog/cost-of-intelligence/memes/gru-v2p.png)

what $20 buys: 2023, one month of gpt-4, the best model on earth. 2026, 0.86 of one max-effort sonnet 5 task (median $23.28 per task on deepswe). the flat subscription never stood a chance.

## jevons would like a word

the lie survives because both halves are true separately. tokens got cheaper. spend went up. people quote the first half, build companies on it, then act surprised by the second half.

jevons figured this out in 1865 with coal: make the engine more efficient and the world burns more coal, not less. ai is the most jevons thing ever built. the models got cheaper per token, better per task, and longer per run, so everyone runs them more, harder, and deeper. useful things don't get cheaper in total. they get bought more.

none of this is a complaint. the models keep earning the spend, that's the entire reason the spend keeps growing. but if a plan assumes inference costs fall 10x next year, that plan assumes running last year's model against competitors who won't be. budget for the bill going up. it goes up because the thing works.

cheapest tokens in history. biggest bills the industry has ever paid. both true.

only one of them makes the slide.

---

# boxboxbox

> i didn't want the codex micro. i wanted mine. edition of 1.

2026-07-27 -- 7 min -- ai, engineering, tools

canonical: https://rohans.lol/blog/boxboxbox

---
![a white magicforce numpad lit up on a brown leather desk mat. violet heartbeat top left, lime keys working, a gold call shouting, a hot pink reply, the wide zero glowing violet. the printed legends still visible under the light.](https://rohans.lol/blog/boxboxbox/board-race-day.jpg)

there's been a dead numpad in my desk drawer for a few months, ever since i simplified the desk and admitted i barely used it.

two weeks ago openai and work louder shipped the codex micro and it's fucking gorgeous. thirteen keys, a joystick, a rotary dial, agent lights breathing under the caps. $230. sold out same day.

![the codex micro. a translucent white macropad with a joystick and a black rotary dial, floating over its aluminium base, glowing softly. product render by work louder.](https://rohans.lol/blog/boxboxbox/codex-micro.png)

i was never going to buy one. i don't even use codex. what am i gonna do with a controller for a product i don't use, put it on the shelf next to the wiis?

but i kept coming back to the page, because a little pad of physical keys for agent stuff is exactly what i'd been wanting. one big fat push to talk key, most of all. wispr flow has quietly taken over how i work, and wispr wants the globe key. my daily keyboard is a 44 key split. there is no room for planets on it. push to talk lived on a mouse button for a while. hated it. every single press.

![a fake wispr flow settings dialog. push to talk shortcut, the globe key. warning, key not found on connected keyboard. corne, crkbd, 44 keys, zero planets. caption, fine. i'll build my own globe key.](https://rohans.lol/blog/boxboxbox/memes/globe-key-v2.png)

and the flat's empty this week. senna's keeping me company, and without him this would be a lot worse, but ngl i've still been a little lonely. a little stressed about the future too. when i sit idle the chatter gets loud. i wasn't shopping for an agent controller. i needed something to build.

when i saw the codex micro, i looked at my drawer and thought, hold my beer.

![mom can we have the codex micro? no, there is codex micro at home. at home... a white numpad glowing violet, pink and lime on a leather desk mat.](https://rohans.lol/blog/boxboxbox/memes/mom-can-we-have.jpg)

## nothing in this stack can tell me no

the runtime my agents live in is [codemaxxxing](https://github.com/bb-deeplearning/codemaxxxing), my opencode fork, five months of replacing other people's opinions with mine. the app i answer them from is [boxbox](/blog/box-box). built that too. and i've been rewriting keyboard firmware in c for years, because of course i have.

the codex micro is a product. it needs injection moulding, packaging, a supply chain, a waitlist. it needs fit and finish, polish, zero configuration, zero bugs. it is not allowed to be a janky homebuilt hacky thing.

mine is allowed. mine is encouraged. mine doesn't even need to be finished, ever. their thing has a joystick and a rotary encoder i don't have. for now.

i press a key on my desk and code ships on a server in another city. every repo between the keycap and the model is mine.

## the customer is always me

i love building for people. clauseo is me watching lawyers use the thing, hearing what's broken, fixing it, again and again. i love every lap of that loop.

this past week the customer has been me. the complainer has been me. boxbox all week, and now the board too. the backlog is whatever pissed me off that day.

the unlit keys had these gold numerals just sitting there, and the whole thing read as a numpad. which it is. genius observation. but i didn't build it to look like one, so i complained, and 45 minutes later off wasn't black anymore. it's a faint white whisper now, and yes, we auditioned five candidates for the color of off on my actual desk.

the enter key sat there doing nothing. the guide said "arm to launch". launch? launch what? i come from a 44 key board where every single key earns its keep, a dead key offends me personally. launch died on the spot, rip to a feature nobody asked for, including the guy who shipped it.

the first message i ever dictated through the board, verbatim. "yo this is actually a kind of really cool way to code. i'm just sitting watching the screen and pressing a button. insane. is there an enter button on this? the thing is i can type but how do i send?"

i was filing a feature request INSIDE the first message the board ever carried. by the next afternoon enter was a real key in the firmware, and the commit reads "every dead press was the guess being wrong".

i built a pinning system too. press a button, jump to that agent, the agent keeps its key. then i noticed how i actually work. agents are ephemeral, i spin up a fresh one for every task and almost never go back to an old one. pinning made no sense for me, a queue made sense. so the pins died, killed by me, and the nine seats became a ranked queue of whoever needs me right now.

the whole loop, with nobody in the middle of it. mine.

## meanwhile, at the actual job

clauseo shipped all week. main moved, reviews got answered, lawyers got their answers.

i also built a keyboard.

both from the same desk, because my job is answering agents and talking to lawyers, and swapping from a legal research product to keyboard firmware is just answering a different thread. i built boxbox exactly so that switching costs nothing. it's a button now. it's LITERALLY a button now, on the board, which the board helped build.

and agents got scary good at writing code. handwritten code is serial. load the whole problem into your skull, type for six hours, one project at a time. that constraint quietly died and nobody threw it a funeral. projects like this used to rot on my someday list because a weekend of firmware for a toy was never going to beat sleep.

now i wake up and the toy is done. i genuinely don't know what to do with that.

## four layers of code for one button

what the board does. push to talk for wispr flow, some shortcuts, lights.

what it took. firmware, a daemon, a server, a runtime, with code landing at all four layers. for a button. i know. i'd do it again tomorrow.

the daemon keeps logs, and the logs say the button won. 508 presses and 164 radio holds so far, most of them today. my favorite feature is the least necessary one. hold the wide zero and the whole board becomes a voice meter, violet columns climbing while i talk. me, yesterday, mid build. "idk like the purple visualiser when you talk also looks fucking sexy as well!"

look at it. LOOK at the animation. i still can't get over it. it's so fucking cool. so so so fucking cool. the visualiser is the best part of the entire build and i'll say it, the codex micro doesn't do this. thirteen keys, a joystick, a dial, and not one of them meters your voice back at you. i'm still not sure what the joystick is even for, and a rotary encoder for reasoning effort feels like overkill, i barely ever change effort levels. one thing on my desk beats the thing that sold out, and it's the thing i use most.

my real keyboard is still right there and i still type plenty, this computer does more than talk to agents. the crkbd wears kailh jades, the numpad wears kailh jades too, obviously, i'm me. but when an agent needs me, i don't reach for the keyboard anymore. i reach for the board. this is somehow my favorite way to code and i built it by accident.

## ok nerds, this part is for you

everyone else, next section.

the board is a magicforce mf17. the mcu is a geehy apm32, an stm32f072 clone whose dfu bootloader never acks the leave request, so every flash "fails" with a timeout after the firmware has already landed. unplug, replug, forgive.

the firmware is a custom qmk keymap in c. raw hid on usage page 0xff60, 32 byte reports. six opcodes go down (hello, palette, frame, global brightness, ping, mode) and three come up (hello ack, key events, mode echo). a 16 slot color palette and a 4 bit perceptual brightness curve live in rom, so a full frame for all seventeen leds is seventeen bytes, palette index in the high nibble, level in the low. there's a 2 second watchdog, and if the host goes quiet the board falls back to its old donor colours and types numbers again. i wanted that. if everything above the firmware dies, the thing on my desk is still a numpad. also, hard won. never enable via. it squats the raw hid usage page and nothing tells you why your protocol went deaf.

the daemon is about three thousand lines of bun on my mac, under launchd. it holds the hid device, watches the boxbox server over sse, and decides every light. raw presses become four verbs (tap, hold, double tap, chord) with timings in a config file that hot reloads while it runs. the voice meter records through sox because ffmpeg was dropping three audio frames in four. there is a small c wrapper whose entire job is making macos show bun a microphone permission prompt. tccd does not negotiate with bun.

the server is boxbox-web on a vm, and the board is just another client of it, same as my phone. it serves the ranked queue. permissions first, then questions, faults, reviews, unread, and under all of those, whatever is working, pulsing lime. the runtime is codemaxxxing on three machines, and it grew a queue depth field on busy sessions specifically so a working key can blink how many prompts are stacked behind it.

seventeen keys, three index systems. wire order, how the daemon thinks. led order, how the ws2812 chain physically snakes through the board, bottom up. matrix order, how the switches scan. every off by one all weekend was one of these three lying about the other two.

## first light

there's a commit from saturday called "first light". the one right after it is the one i love. the palette, recalibrated by photograph.

here's the thing about agents writing led code. the model cannot see the board. it knows what #be00ff looks like on a screen, but an led under a smoked cap is a different animal, and every board's leds are their own animal on top of that. writing for a screen and writing for hardware are just not the same job. so the first palette came up so dull the thing looked half dead, and i did the only sensible thing. i started sending photos. "yo this is how dull it is." photo, new curve, photo, new palette. and then i pointed the model at the firmware i'd written for this exact numpad by hand, years ago, when i picked the colors myself, and said go look at what past me chose. past me had taste. between my photos and his old code, the board finally matched the board in my head.

so i'd picked these colors twice, once by hand in the old firmware, once now. i wrote the spec. i knew exactly what it was going to look like.

didn't matter. the formation lap ran, one violet light sweeping the perimeter, and the messages i sent that hour are "omg yes the hello worked" and then "okay wow that was fuckign coooool as fuck yo". typo preserved.

it exists. it exists. the thing from my head is sitting on my desk and it EXISTS.

same feeling as the last brick of a lego set. magical, and you knew the whole time what it would be. holy fuck anyway.

## sold out at my house too

after it worked i decided it was too cool not to have a product page. that's the entire reasoning. i want everything i've built to have its own page on this blog eventually, so the numpad goes first.

[boxboxbox, the scale 1/1 pit wall kit](/showcase/boxboxbox). paint guide, parts map, a factory test you can actually take, a warning label, and an [assembly manual](/showcase/boxboxbox/guide) that assembles you, because the kit ships pre-built. part of it is honestly just a user guide for me, i WILL forget my own shortcuts. part of it is that code got cheap enough that "this deserves a product page" is a thing you can say out loud and then simply have. and part of it is that fable 5 is stupidly good at visual stuff and i wanted an excuse to let it cook.

the name is boxboxbox, by the way. it's a box that operates boxbox. yeah. i'm not selling it, so it doesn't have to be a good name. it just has to be mine.

units produced, 1. availability, sold out, see units produced. a product page with no product on it. a janky homebuilt numpad doing an impression of an openai hardware drop, with a more thought out product page than half the real products i've used. it fits me exactly, down to what color "off" is.

it will also never be finished. the three machine keys learned a new job tonight, starting fresh chats, while i was writing this paragraph. that's the point of owning it.

## the fastest promise i ever kept

[the last post](/blog/box-box) ended with a promise. the firmware compiles, zero leds have lit outside a browser, and when they do, that's the next post.

first light landed twenty four minutes after i hit publish. full honesty, i knew it was close, half the thing was already built when i wrote the line. still counts.

so this is the next post. the board is on, and a key just went gold to my left. gotta go.

---

# box box

> no laptop. no terminal. people have email. i have agent mail.

2026-07-26 -- 9 min -- ai, engineering, tools

canonical: https://rohans.lol/blog/box-box

---
my morning starts the way most people's do. i wake up, reach for my phone, and check my mail before my brain has fully booted.

except my mail is from agents.

this morning's stack: an agent finished a branch on clauseo overnight, tests green, diff attached, waiting on a merge. another one, mid-feature, hit an edge case it didn't want to guess on, so it asked. a third wanted permission to run something destructive. a fourth hit a wall and said so. i went through all of it lying down, half asleep, one thumb. approve, reply, approve, snooze. main moved twice before i was out of bed.

no laptop. no terminal. people have email. i have agent mail.

## the numbers

that inbox adds up. in the last seven days, 233 commits landed across three repos: my coding harness, clauseo's product (the day job), and the mail app itself. that's not counting the model proxy, the search engine, or anything on my personal github.

i typed approximately none of them.

![spiderman pointing at spiderman: the guy with 233 commits this week / the guy who typed zero of them](https://rohans.lol/blog/box-box/memes/spiderman-v4.png)

my output, measured honestly, was texts. reading, thinking, replying. the commits are what happens in between. for scale: the mail app's repo is a week old and already has 119 commits. the fleet built its own post office. and the biggest single day i've ever watched was 82 commits, twenty-three thousand lines, one day back in may. i remember that day mostly as a conversation.

## the machinery

in february i forked opencode and [wrote a post about it](/blog/codemaxxxing-context-engineering). the problem back then: models get dumber as their context window fills, so you architect around attention. fresh sessions, subagents, state on disk. that fork (codemaxxxing) is still the engine. but it stopped being a tool i open. it's a runtime now, a background service on my mac and a couple of cloud servers (one of them belongs to iris, my ai teammate. long story. she has her own sim card.), and agents work whether or not any laptop of mine is open.

everything else grew one annoyance at a time.

agents degraded mid-task, so work got split into waves, with a verifier agent checking results and retrying failures. parallel agents stepped on each other, so every agent works in its own checkout now and main stays clean. ai code was landing with no human in the loop, so review became a declared step: the agent announces it's ready, and main moves only when i tap. different models are good at different jobs, so there's a routing table instead of loyalty: fable 5 orchestrates, opus 5 does the heavy general work, sonnet 5 runs the worker laps, gemini flash is the mule that reads the boring stuff. and i had no way of knowing any of it needed me unless i went and looked, so "needs you" became a push notification. that last one became a whole app.

if you follow this industry, you recognize the list. context management. sandboxed parallel agents. review gates for ai code. model routing. agent inboxes. every line is a product category in 2026, and every line has funded companies with whole teams on it.

i have a folder.

![mom can we get cursor / we have cursor at home / cursor at home: the codemaxxxing home screen, july 18](https://rohans.lol/blog/box-box/memes/cursor-at-home-v4.png)

## the bug was me

in may i [wrote, a little smugly](/blog/codemaxxxing-isnt-just-a-fork-anymore), that my agents did real work while i was at a friend's birthday. true. what i left out: i checked. i stood at a party with terminal panes open on my phone, watching logs scroll.

the agents had stopped needing me to do the work. i had not stopped needing to watch them do it. pull was the bug, and the bug was me.

![panik kalm panik: agents do real work while i'm at a party / so i never have to watch them / i watch terminal logs at the party anyway](https://rohans.lol/blog/box-box/memes/panik-v3.png)

so, dumbest question available: when an agent needs me, what do i actually need? a screen. a way to answer. the internet. that's the entire list. the computer was habit, code just used to live there. the terminal was habit, agents happened to be born in one. the checking was habit, the tools couldn't come get me. delete every inherited accident and what's left is the actual job: read, decide, reply.

## box box

so i built the mail app. it's called boxbox, after the f1 radio call. box box is what the pit wall says when the driver has to come in now.

it's a chat app. that's the honest category. every agent on every machine flows into it, and anything blocked on me (a permission, a question, a branch waiting for review) lands in one ranked inbox that follows me to whatever screen i'm holding. cards on the lock screen. approve, reject, answer, merge. the agents are on track. i'm on the pit wall. the wall doesn't drive the car. it takes the radio calls.

i called it agent mail at the top because mail is the fastest way to explain an inbox to a stranger. but the analogy bugged me for days before i worked out why. mail is where other people's priorities queue up to eat your attention, and you process it to get back to the real work. this inbox is the real work. every card is a live agent mid-flight, holding a decision only i can make. a reply doesn't file something away. a reply lands code.

![the boxbox call strip: a permission, an escalated question, a week-old fault, and a review waiting on a merge](https://rohans.lol/blog/box-box/boxbox-call-strip.png)

i built it native for iphone first. two days, fifty commits, a design doc marked law. then apple's free account physics showed up: no push notifications without a paid developer account, and builds that die every seven days. and underneath that, a deeper problem. my agents write the code, and on the web an agent gets a dev server and a live page it can drive and inspect for free. native taxes every loop. so boxbox became a web app, and the design that was law died twice in the next four days. at one point i shipped a voting page: candidate designs side by side, my taps writing straight into a spec file on the server that the agents read as the verdict. no figma, no meetings, no translation between my thumb and the build. i have picked design direction from the toilet.

the design isn't decreed anymore. it's discovered.

## i text for a living

so that's the honest job description. i text. i start threads, i answer threads, i read what came back, think, and reply while three other threads run. some threads are thirty seconds. some run six hours and spawn side threads with agents i never talk to directly.

and texting doesn't care where you sit. i still love my desk. four screens, a split mechanical keyboard with loud clicky switches, still where the deep review sessions happen. but the desk is a choice now, not a requirement. most evenings i'm on the couch with something playing on the tv, answering agents between scenes, more relaxed than work has ever felt. the fastest keyboard i own turned out to be my phone, and half the time i'm not typing at all, i'm dictating. i say what i want, the words land in the thread, the agent gets to work, i go back to the show.

![the i-should-buy-a-boat cat, in a robe with a newspaper: i should write some code](https://rohans.lol/blog/box-box/memes/boat-v3.png)

texting stays cheap because everything important is written down exactly once. AGENTS.md is law, every agent reads it before touching code. GOTCHAS.md is fifteen thousand words of scars, every trap documented the day it bit, so no agent makes me explain the same mistake twice. there's even a section that teaches them how to read my texts. it says things like "buddy means you were slow."

and one rule holds the whole thing up: claims without verification are worthless. tests green, build green, screenshots actually looked at, or it didn't happen. the commit log reads like incident reports because that's what commits are now, agents telling me what happened while i wasn't looking. an actual title from this week: "preparing is a claim with an expiry."

## one user

user count: one.

products are off the rack by necessity, cut for the average of everyone. this is tailored. every default is a decision instead of an inheritance.

i've only ever built one way: watch someone use the thing, listen, close the gap. the loop was always tight. the only distance left in it was between the person complaining and the person shipping the fix. this time that distance is zero, because the complaint goes straight to the agent whose code it is. that's honestly most of what my boxbox threads are: complaints. i complain like a user, point like an engineer, and it's usually fixed before the episode ends.

that loop compounds fast. a week in, i already can't remember the last time i opened a terminal ui for anything except deploying boxbox itself. the thing i built to escape the terminal is the only thing i still use the terminal for. i keep noticing that and it keeps shocking me.

this week the fleet shipped telemetry on me. the commit literally reads "boxbox starts measuring itself." i want to know which parts of the system earn their place, and my own usage is the only benchmark that matters here. same loop, one level up: friction dies, the thread count goes up, the workflow changes, the next friction shows itself. the tool and the way i work are iterating on each other, and the loop never leaves the room.

i've spent my whole life building for other people. drones, a talent agency, a streetwear brand, a payments app, legal research for lawyers. this is the first time i've pointed the loop at myself. one user, fully instrumented, zero gap.

![model card for rohan-5: decision-generation, cannot-be-swapped-out. context window 1 item (2 on a good day). fine-tuned on team radio and 15,000 words of scars. deprecation: never.](https://rohans.lol/blog/box-box/memes/model-card-v2.png)

and under all of it, the simplest thing: i just love building. that's the whole engine. it's a little recursive and a little ridiculous that the thing i love building most right now is the thing i build everything else with. i'm okay with that.

![distracted boyfriend: me, looking at building the tool, while the products the tool exists to build looks on furious](https://rohans.lol/blog/box-box/memes/db.png)

## the next why

the whys aren't done, because the stack itself is now the loudest annoyance. two repos moving in lockstep that were always one system. a runtime still shaped like opencode underneath, carrying a whole web interface i've never opened and opinions i've spent five months deleting. fork tax compounds quietly. eventually this collapses into one thing you install on every machine you own and open from any screen you're holding. not a product. nothing to sell. but it's my workflow, and once the setup is one command instead of a fork plus a layer plus a runbook, of course i'll want people to try it.

![this is fine: five months deleting opencode's opinions / the fork is fine](https://rohans.lol/blog/box-box/memes/fine-v2.png)

## the fifth model

in february i ended that post with "the tools don't matter. the models don't matter that much either. the context does." then i spent five months on tools, so you'd think i'd have to eat the line.

no. i finally understand what the tools are for. the routing table has four models in it. there's a fifth. it's slow, it's expensive, it gets distracted, it degrades badly when its context fills with junk, and it's the only one that can't be swapped out. every card in the inbox routes to it. the entire system exists to protect that model's context window.

it's me. i'm the fifth model. and for the first time since february, my context is clean.

![xkcd 2347 redraw: clauseo, boxbox, 233 commits a week, codemaxxxing, two VMs, iris, five funded product categories, all resting on one thumb, 07:14, one eye open](https://rohans.lol/blog/box-box/memes/xkcd-v2.png)

## post-credits scene

there's a dead numpad in my desk drawer. seventeen keys, bought last year, retired a few months ago when i simplified the desk. then work louder and openai shipped the codex micro: agent states on keys, push-to-talk, a dial for how hard the model thinks. it sold out. i saw it and thought, hold my beer.

because here's the thing the codex micro can't do. it watches one product on one machine, through whatever integrations it's allowed to have. i own every layer of mine: the runtime, the brain, the app, and now the firmware. most people run a matching set from someone else's roadmap: claude code plus the claude app, or codex plus the codex app plus the micro. mine is one system growing a new limb, and when the board needs an endpoint the server doesn't have, i don't file a feature request. i text an agent and it exists by dinner. vertical integration, except it's one guy and two cloud vms.

so this weekend i locked the spec: transparent caps with nothing printed on them, an led under every key, a wire straight into the fleet. every agent earns a key and a color. a permission lands, a key goes hot pink, i press it once, and code ships on a server in another city.

instead of typing, i hold a key and talk. instead of waving a cursor around hunting for the right window, i press an agent's key and the screen is already on their thread. every agent in the fleet, one physical press away.

the firmware already compiles. zero leds have lit outside a browser. when they do, that's the next post.

---

# fix the exam, not the news cycle

> twelve kids died, and everyone is harvesting the bodies. views, likes, and future votes.

2026-07-21 -- 14 min -- politics, thoughts, rant

canonical: https://rohans.lol/blog/fix-the-exam-not-the-news-cycle

---
twelve students have died by suicide since the neet papers leaked. that's the reported count, and it's the only fact in this circus everyone agrees on. it's also the fact everyone stopped talking about somewhere around week two.

i went through the pipeline myself. jee, the coaching grind, the one morning that decides everything. i know what was pressing down on those kids. so believe me, what follows is not me being against the students. it's me being furious on their behalf, at everyone currently spending their deaths.

## pick a side, any side

the moment this became cjp versus bjp, the students lost. that's not commentary, that's mechanics. reform doesn't have a side. a scalp has two. india's outrage machine can only process one shape of story, government versus opposition, so every cause that enters it comes out the other end as a jersey.

i'm not wearing either one. here's the only yardstick i'll use for the rest of this post: justice for the neet victims. the protest's own stated reason. hold it in your hand and watch what happens to everything below.

## the government first, because they earned it

the exam leaked on their watch, inside their system, out the end of their supply chain. one identical paper for 2.27 million kids, printed at private presses, moved around the country in trunks. it leaked in 2024 too. nothing structural changed in between. that's not bad luck twice, that's a policy decision to keep gambling.

and look closely at what the government did this time, because on paper it looks like action. cancelled the exam. cbi probe. arrests, including nta's own insiders. re-exam in june. the full crisis playbook, executed. now notice what's missing: everything that closes a news cycle, nothing that prevents the next leak. and the cancellation itself, sit with what it means: 2.27 million honest kids punished with a re-sit because the state couldn't protect its own paper. the kids paid for the government's failure twice.

then there's everything after the re-exam, where the bjp stopped being a government and became cjp's unpaid production house.

understand the position they were in. the damage control was done. cancelled, probed, arrested, re-examined. from july onward they had access to the single most powerful move in indian politics: doing nothing. be boring. let a satirical instagram page yell at an empty podium through the monsoon.

![distracted boyfriend: the bjp looking at caning teenagers on camera instead of just doing nothing](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/bjp-db.png)

instead. a 59 year old gandhian, three weeks into a fast, the one man in the country whose stretcher photo is worth more than any manifesto ever written, and the state's big idea was to lift him off the dais at dawn under court orders and call it medical care. hospitalization as crowd control. they didn't defuse the martyr narrative, they catered it. the crowds swelled overnight. then, with every camera in the country pre-positioned for the opening of parliament, they met teenagers holding placards with lathis and tear gas. in 4k. from nine angles.

this is the party with the most sophisticated political messaging operation india has ever seen, and its response to a meme page was to shoot, produce and distribute the opposition's recruitment film. free. twice. on schedule. if dipke's growth team had scripted the government's moves themselves, it would look exactly like this (hold that thought, because in a minute i'll show you they kind of did). the bjp is not the victim of this story. they are its co-author, its cinematographer, and its distributor.

## read the manifesto. i did.

cjp's manifesto is a political party's platform. that's not an insult, it's a description. the marquee demand is a constitutional bar on post retirement posts for judges, the founder's oldest feud, in a movement born dunking on a judge. real issue, i even agree with parts of it. then: electoral roll preservation, an immediate 55 percent women's reservation, mandatory political literacy in schools. over on the demands blog: gig worker rights, non compete bans, salary disclosure, internship stipends, court timelines. a full spectrum platform. something for every angry young person in the country.

so use the right words for what's happening at jantar mantar. this is not a students' movement with a cause. it's a political vehicle with a flagship campaign, and the campaign is the students. the kids marching believe the movement exists for them. the platform says they're the current promotion.

![anakin and padme: we're marching for the neet students. for exam reform, right? ...right?](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/anakin.png)

the one plank that touches their world is a fee cap on coaching institutes. a price control. coaching is a ₹58,000 crore industry by their own copy, distorted by artificial scarcity, and freezing prices produces black market batches and worse teachers, the way price controls have worked everywhere forever. not machinery. an applause line.

![uno: add one real exam reform or draw 25. cjp holding 25 cards.](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/uno.png)

and the manifesto describes its own trajectory better than i ever could: migrated "from mere online shitposting to a serious, socialist democratic front."

a front. their word. fronts have flagship campaigns. fronts outlive them.

## kill the paper

which brings me to the part i actually came here to write. the fix. and i'll be honest about my bias: i'm a builder, i make software for a living, so of course my solution is technical and boring. no scalps, no marches, no aesthetic. but this is what we actually need to do, nobody at any podium is saying it, so here it is in plain words.

start with why the leak keeps happening, because it's not corruption as a moral failing. it's architecture. neet is one identical paper for 2.27 million kids. it gets finalized weeks in advance, printed at private presses, stored in strongrooms, trucked across the country, held overnight in district vaults, and opened at thousands of centers by thousands of hands. every hand in that chain is a potential leak, and the object they're holding is worth crores. you cannot guard something that valuable moving through that many hands. someone will always sell it. 2024 and 2026 aren't two scandals, they're one design flaw expressing itself on schedule.

so stop defending the paper. delete the paper.

![drake: guard one paper with trucks, vaults and prayers? no. delete the paper, randomized sets from a question bank? yes.](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/drake.png)

**move the exam to computers.** not because computers are magic, but because of what they end: there's no longer one physical object that exists weeks in advance, waiting to be stolen. jee main already runs on computers, over a million candidates, multiple sessions. the infrastructure exists in this country today. neet isn't waiting on technology. it's waiting on a decision.

**then the actual killer: randomized papers from a question bank.** this is the part nobody explains, so let me. instead of one paper, you build a bank of thousands of vetted, calibrated questions, and every candidate gets their own randomized set. no two kids sit an identical paper. "but then some kids get harder papers": solved problem. the science of making two different question sets equally hard is called equating. the gre has run on it for decades, and jee already normalizes scores across sessions. boring, settled, working math. not an idea i had in the shower.

now walk the consequence through. what exactly does the leak mafia steal? there is no "the paper." there's a rotating bank of thousands of questions, and your specific set doesn't exist until you sit down. the entire economy, the press insider, the trunk, the strongroom, the whatsapp pdf at midnight, doesn't get policed out of existence. it gets deleted by architecture. you don't win the security game. you change the game so there's nothing left to steal.

**lower the stakes, because the stakes are the price tag.** a stolen neet paper is worth crores for one reason: one attempt a year, everything riding on one morning. that's also, not coincidentally, why kids are dying. give two or three attempts a year with score banking, count the best one. the black market price collapses, and so does the pressure that's killing seventeen year olds. the security argument and the humane argument are the same argument. that's how you know it's the right fix.

**expand the seats.** the deepest lever. the leak market is priced by scarcity: 2.27 million kids, about a lakh mbbs seats. the state throttles the supply of medical education, then polices the black market its own bottleneck creates. more colleges, higher intakes. the desperation premium falls and half this shadow economy deflates on its own.

**staff the agency, prosecute the chain.** nta runs some of the biggest exams on earth with skeleton staff and outsourced everything. give it psychometricians, security engineers, internal audit. and the 2024 anti leak law's ten year sentences mean nothing at indian trial speed. fast track courts, prosecutions that go past the solver gangs, into the presses, up to the officials who ignored warnings.

that's it. that's the whole revolution. it fits on a page, it would actually protect the next 2.27 million kids, and it will never trend, because "calibrated question bank" doesn't fit on a placard and doesn't feel like justice. it just is justice, delivered boringly, which is the only way justice ever ships.

## the man doing the math

and the gap between that page and what's actually on the placards is where the operator lives.

here's the method, and it's not new: you don't pick a cause and fight for it. you shop. you throw controversies at the wall, measure what sticks, and scale the winner.

kejriwal ran this exact algorithm in the autumn of 2012, and india watched it live. robert vadra's land deals in october. salman khurshid's trust the same month, chased to his home constituency. nitin gadkari in november, the sitting president of the bjp, because the algorithm doesn't care whose side you're on. mukesh ambani and gas pricing, with radia tape clips at the presser. hsbc black money after that. by december the news cycle knew the genre so well that india today ran the headline "kejriwal to make another expose." another. a weekly show.

and here's the tell that none of it was conviction: the moment the vehicle reached power, the machine stopped. the ambani fir he filed as chief minister, three days before resigning, got quietly abandoned. caravan wrote a whole piece asking why. the answer is obvious once you see the shape. the exposés were never cases. they were customer acquisition. when product-market fit arrived, the throwing stopped and the winning variant became the entire brand.

now watch dipke run the same algorithm at feed velocity. his wall is the platform you just read, nine causes deep, something for every demographic of angry and young. even the street protest is a bundle, officially over neet plus cbse's on-screen marking plus recruitment scandals, three grievances stapled to one resignation demand. sixteen million followers in FOUR DAYS taught him what his distribution could do. the only question was which outrage converts best, and the feed answers that hourly. kejriwal needed a press conference a week and months of testing. dipke's lab runs continuously. same playbook, hundred times the clock speed.

which brings us to the calendar, because the sequencing is the confession. the re-exam happened june 21. whatever this year's administrative fight was, it ended there. the fast began june 28. by the time jantar mantar filled up there was no exam left to save, only a trophy left to claim. and sonam wangchuk deserves better than what's being done with him: his actual causes are the exam system and ladakh, sixth schedule, statehood, the demands delhi has stonewalled through five previous fasts. the resignation was never his demand. the one serious person in this story lent his moral capital to the campaign, and the hour the police carried him off, dipke announced his own hunger strike. an understudy, warmed up, stepping onto a vacated dais. grief doesn't move that fast. campaigns do.

![the operator's calendar: police remove wangchuk 6am, announce MY hunger strike 7am, sansad chalo with permission denied and cameras pre-positioned, review engagement dashboard](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/calendar.png)

then the march. permission denied, called anyway, into a barricaded zone, on the opening morning of the monsoon session when every camera in delhi was pre-positioned. run the decision tree. police stand down: he marches on parliament, historic win. police crack down: tear gas and lathis on teenagers, the movement swells with rage, historic win. there was no branch where he loses. the kids carried all the downside. he collected the upside from the stage.

say it in plain game theory, because once you see the payoff matrix you can't unsee it. dipke's payoff has exactly one variable: attention. violence is the most attention dense commodity on earth. so violence is not a risk to this movement's strategy, it is the strategy's best case. every lathi that lands on a follower is marketing spend he doesn't pay for. the state supplies the violence, the kids supply the bodies, he supplies the caption. peaceful march: win. brutal crackdown: bigger win. the only losing outcome is being ignored, and an unpermitted march on parliament on opening day makes being ignored structurally impossible. he cranked the probability of violence to its maximum setting and aimed it at his own supporters, because their worst day is his best content. remember the government scripting its own villain scenes a few sections ago? he wrote the call sheet. and when a leader profits from his followers' beatings, the beatings become a schedule.

the manifesto calls the movement "not a joke" but "a highly organized public pressure coalition." their words. believe them.

## the harvest

and the operator isn't even the only one working the field.

twelve kids died, and everyone is harvesting the bodies. views, likes, and future votes. i've written that sentence as flatly as i can, because dressing it up would be doing the exact thing i'm describing.

start with the creators, because they're the most honest about it, the way a pickpocket is honest. there are influencers at jantar mantar with ring lights. the protest is a set. someone else's dead child is a backdrop with strong engagement numbers. they shoot the reel, pick the take where their jaw sits right against the tear gas, post it with a caption about courage, and then do the part that actually makes me want to puke: they check the metrics. story views on a dead seventeen year old. which hook converted. whether the lathi charge b-roll outperformed last week's cafe recommendations. protest as a content vertical. grief as a growth hack.

![the same account posting RIP to the 12 angels and GRWM for the revolution with a protest fit check poll, an hour apart](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/stories.png)

then the stories. the endless, endless stories. posting one is not activism, it's absolution. the mechanism has a name, moral licensing: perform the token gesture, collect the feeling of having *done something*, skip the actual work forever. the repost is not the beginning of anyone's involvement. it is the involvement. complete, conscience cleared, obligation discharged, feed refreshed. and the cruel math underneath: a signal that costs nothing carries no information. the farmers who actually got laws repealed camped on a highway for a year through two winters. that's a signal. a story evaporates in 24 hours by design, which is roughly the half life of the conviction behind it.

and the revolution cosplay. "india's nepal moment." brother. nepal's parliament burned because its electoral exits were welded shut. sri lanka's president fled because one family owned the entire state. bangladesh rigged three elections in a row. that is when streets replace ballots: when ballots stop working. india's incumbents lose elections constantly, at every level, every cycle. wanting a revolution in a country with functioning turnover is wanting the aesthetic. the barn storming chapter, the flag on the barricade, the moody black and white photo for the grid, and none of the boring chapter about who runs the farm afterward. it's genre tourism. and posting it between a gym selfie and a coffee flatlay is the tell that it was never about the kids.

everyone in this frame is extracting something. the operator extracts a career. the creators extract engagement. the reposters extract absolution. the government extracts a closed news cycle. now scroll the entire crime scene and find me one participant whose payoff depends on the next exam not leaking.

one.

## we've watched this show before

none of this is improvisation. it's a playbook old enough to have a greatest hits album, so let's write the steps down.

![astronaut meme: wait, it's just the 2011 anna hazare playbook? always has been](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/always.png)

**step one: find an unopposable cause.** corruption in 2011. dead students in 2026. the defining feature is that no decent person can stand against it in public, which means nobody inspects the vehicle, because inspecting the vehicle looks like opposing the cause. the cause is the armor.

**step two: borrow a gandhian body.** a revered faster whose moral capital took thirty years to build and can be spent in three weeks. anna hazare then, sonam wangchuk now. same stage, jantar mantar, because the set dressing matters. the frailer the body, the better the content.

**step three: run the machine from behind the body.** kejriwal then, dipke now. and dipke isn't a student of this playbook, he's an alumnus. he ran social media inside aap. this is a franchise employee opening his own location.

**step four: demand a single lever that fixes nothing.** jan lokpal then, one resignation now. the demand must be simple enough for a placard and useless enough to keep the grievance alive, because a solved grievance is a dead asset. machinery demands are poison to the model: someone might quietly accept them, and then what do you march about?

**step five: when the body stops being useful, inherit the stage.** anna was sidelined within a year and spent the next decade saying he'd been used. you watched this step execute live on saturday morning, in the gap between the stretcher and the announcement.

**step six: convert.** party, symbol, tickets, elections. in 2012 this step was called the aam aadmi party. in 2026 the paperwork hasn't been filed yet. wait for it.

![gru's plan: join the movement for neet justice, march to parliament for the kids, you're canvassing for someone's launch](https://rohans.lol/blog/fix-the-exam-not-the-news-cycle/gru.png)

and if you want to know how the movie ends, season one already aired. the anti corruption party's deputy cm did seventeen months in jail in the liquor policy case. its founder became the first sitting chief minister in indian history to be arrested. its health minister spent two years inside in a laundering case. yes, agencies get weaponized and nothing's convicted, but the aesthetic collapse doesn't need a conviction: the party born from an anti corruption fast watched its leadership cycle through tihar. and the founding demand itself? the lokpal aap eventually passed in delhi was so defanged their own founding lawyer called it a jokepal, and he was expelled for complaining, along with the party's other conscience, for requesting internal democracy in the anti high command party.

the advertised product never ships. the founder's career always ships. iac promised a corruption free india and delivered arvind kejriwal. this movement promises justice for neet victims and will deliver abhijeet dipke, who watched season one from inside the writers' room.

## the wrong species

the manifesto has one line of accidental poetry. "the cockroach does not ask for permission to occupy space; it survives the ruins of civilizations."

he's right. that is exactly what a cockroach does. survives everything. every purge, every collapse, every news cycle, every tragedy. scuttles out of the wreckage unharmed and comes out fatter.

but that's not the kids. nothing about the kids is indestructible, twelve of them are gone, and the ones marching got tear gassed, filmed, harvested for engagement, and sent home to the same one paper lottery next may. that was the whole point of being angry in the first place.

there is exactly one organism in this story that survives every indian outrage cycle. it thrived through 2011. it will thrive through 2026. it cannot be killed by lathis or verdicts or the total failure of its stated cause, because its actual cause, the career, always ships on schedule.

cockroaches survive everything. he told you that himself, on day one, in the name.

he just never told the kids which one of them he was.

---

# i can't find the logic in hate

> i took it apart like i take everything apart. there's nothing inside.

2026-07-04 -- 5 min -- life, thoughts, rant

canonical: https://rohans.lol/blog/no-logic-in-hate

---
i take things apart. it's the only way my brain knows how to deal with anything. keyboard firmware, electronics, billing systems, arguments. open it up, find the logic, understand the machine.

so i've been trying to take racism apart. hate in general. all the flavors of this shit. looking for the mechanism. the gear that makes it turn. the piece where, even if i disagreed with it, i could at least go "ok, that's the logic they're running."

there's nothing in there.

i mean it. i've turned it over and over and there is nothing inside. and that's worse than finding evil. evil at least computes. this doesn't compute.

and my brain can't put down things that don't compute. it won't file it away and move on. it just keeps running the loop. so here we are.

## murdered over sunscreen settings

look at the actual list of things people get hated for.

the pigment in their skin. the patch of dirt they were born on. which genitals they have. which genitals their partner has. the language their mouth learned before they turned five. which myth they inherited. what chemistry their brain runs on.

read that list again. not one of those is a thing anyone did.

NOT ONE.

skin color is melanin. a chemical. your body produces more or less of it depending on how much sun your ancestors caught. that's the entire fucking science. uv exposure and time. people have been enslaved, colonized, lynched and gassed over sunscreen settings.

your birthplace was decided while you were a fetus. your genitals were a coin flip. your first language was installed before you could consent to anything. your god was assigned to you like a roll number. your brain chemistry shipped from the factory that way.

every axis of hate is a dice roll. rolled before you existed. you weren't consulted on a single one of them. nobody was. and we built wars, genocides, empires and entire legal systems on top of the dice. india looked at the existing options and invented one more, hate by surname. we're innovators like that.

## the cheapest superiority ever invented

so if there's no logic, what's running the machine. i can only find one honest answer.

hate is the cheapest way to feel superior that humanity has ever produced.

zero effort. zero skill. zero work. you don't have to build anything, learn anything, win anything, become anything. you point at your birth lottery ticket and declare yourself the winner. a participation trophy you hand yourself for exiting a specific womb in a specific place.

and here's the tell. the people who grip it hardest always have nothing else. if the proudest fact about you is your pigment, there are no facts about you worth being proud of. people with actual merit don't need this. nobody who built something real has ever needed their melanin to feel like somebody.

that's all it is. not a philosophy. not a culture. not "heritage." a scared little need to feel superior.

## not even the same species

and if it stopped at superiority i could almost file it away. pathetic but explainable. it doesn't stop there.

wanting to feel superior explains a smirk. it explains a look. it doesn't explain the rest. it doesn't explain how a human being looks at another human being and decides they're not one.

think about what that actually takes. the evidence is standing right in front of you. a face. a voice. hands. somebody's kid. someone who flinches when they're hurt and laughs at stupid jokes and has a favorite song. every signal a body can send, screaming "same as you. same as you. same as you." and some people receive all of it and still file the person under something less than a person.

how. how does a human hold that much hate. where do you keep it. doesn't it get heavy.

people will baby-talk a dog on the street. apologize to a table they walked into. cry about a robot in a movie. the empathy is right there. installed. working. we hand out personhood to animals and objects and cartoons for free, and then ration it for actual living people. the machinery isn't broken. they just draw the species line wherever they want and leave humans standing outside it.

i can't get past that. i physically can't get past that.

## the lemon is outperforming us

vader has never once cared what anyone looks like.

no background check on your skin. no questions about your god. you walk in, he decides you're family, and that's the end of the vetting process. i love all dogs the exact same way. every breed. all dogs are cute, no matter what. it costs nothing. an entire species, unconditional, zero exceptions.

dogs run this program on a brain the size of a lemon.

we're carrying around the most complex object in the known universe and we use it to rank each other by pigment. the lemon is outperforming us.

## i don't have a word for this

there's a feeling that shows up when i sit with this and i don't have a word for it. it's somewhere between rage and grief and static. some things hit me at a volume i can't explain to people. this is the loudest one.

i keep seeing the same picture. a person, a whole person, standing in front of another person, being told they're not one. and everything in me short-circuits at once.

i've read the history. the psychology. the tribe-brain evolutionary explanations. i still don't understand, because understanding requires something to be there, and there's nothing there. dice and fear. that's the whole machine. it's empty and it still kills people.

i don't understand it. i don't understand it. i don't understand it.

but i'd rather stay confused by hate than ever become fluent in it.

---

# codemaxxxing isn't just a fork anymore

> started forking opencode in february. eleven weeks later it's running real work while i'm out at parties.

2026-05-07 -- 6 min -- ai, engineering, tools

canonical: https://rohans.lol/blog/codemaxxxing-isnt-just-a-fork-anymore

---
## i forked it to tinker

i'd been using claude code for a while. it was great. lately it's been a shitshow but that's not actually why i forked. i forked because it's closed and i can't open the box. i can't use a thing for long without wanting to take it apart. so i forked opencode because at least it's open.

what i actually wanted, underneath the tinkering, was something no harness on the market would give me yet. a way to run plans that took five hours instead of one.

## the warm-up

for the first few weeks i wasn't building anything new. i was just trying to get opencode's prompts to feel like the baseline i was already used to. the prompts in claude code work the way i was already working. i wanted that before i started having opinions of my own.

iteration logs went into `PROMPT_ITERATIONS/`. i ended up reading google's gemini cli source code to figure out how they prompt their own model. found 31 specific behaviors that made gemini act differently from claude. that's the kind of thing that happens when you can't leave a tool alone.

i was swapping subagent models too. claude haiku for the exploration agent turned out to be worse than gemini 3 flash on every axis: more expensive, older, less capable at reading code at depth. moved to flash. now i'm on gemini 3.1 pro.

all of this was still just configuring opencode. i rebranded along the way because why not. called it codemaxxxing.

## the actual thing

models have 1m context windows now. doesn't matter. attention still dilutes as the window fills up. by the time you've been in a session for two hours, the model is dumber than when you started. context got bigger. the architecture problem stayed the same.

so any plan that takes more than an hour or two of agent time hits a wall. the agent gets dumber halfway through. you start over. that's not a productivity problem. that's an architecture problem.

what i wanted: take a big plan, chop it into pieces small enough that each piece fits in a fresh context window, run each piece in its own session. each session reads a state file from disk to know where to pick up. each session commits to git when it's done. the agent in session N never sees what session N-1 thought. it sees what session N-1 _did_, on disk, in actual files. fresh context beats degraded context. always.

the pieces are called waves. the rest is plumbing. and the plumbing is where it got interesting.

## stage 1: manual

first version of the plumbing was just me. campaigns were 6-8 waves, maybe two hours of agent time total. i'd plan, break the plan into chunks, then run each chunk in a separate chat session that started fresh. wave 1 done. open a new chat. paste in instructions to pick up wave 2. wave 2 done. open another chat. paste again. and again.

i was a human shell loop.

it was annoying. mostly annoying. there is no romantic version of this story.

one habit i picked up here was telling the agent to commit at the end of every wave. one wave, one commit. it meant i could review per-wave diffs instead of one massive blob at the end of the campaign. better than that, i could peek at finished waves while later waves were still running. see what got built, wave by wave, while the campaign kept going. small thing. it became important.

## stage 2: bash

so i wrote a small bash script (`run-all-waves.sh`) to do the manual part for me. about 250 lines. it reads the state file on disk, calls cmx for the next wave, waits for the wave to finish, checks the new state, advances. baked in the per-wave commits. ansi colors. wave numbers in cyan, success in green, failure in red. honestly fine for what it was. used it for weeks.

then i started getting braver. the campaigns got longer. 6-8 waves became 15. 15 became 30. two hours became four became most of a day. i started running things i wouldn't have dared run before. whole codebase overhauls. full repo refactors. the kind of work you'd put a small team on.

and the longer they ran, the more obvious it became that planning upfront was the new bottleneck.

a bad plan multiplies. one wrong assumption in the plan and twelve waves later you're staring at a pile of work that all built on top of the wrong assumption. i'd watch a wave fail and realize the plan was the problem, not the agent. the plan said use library X but X was deprecated. the plan said run the verification one way but the verification command exits the wrong way. the plan was perfect when i wrote it and wrong by the time wave 4 hit.

agents working from todos improvise. humans improvise. of course they do. real work doesn't survive contact with the real codebase. but the wave system was treating the plan like a contract. one bad clause and everything downstream was garbage.

the plan needed to stop being a contract.

## stage 3: into the tool

the trigger to actually move was something else though. the day i was teaching a five-hour corporate gen ai training and the bash script was committing waves the whole time i was on stage, i thought, why is the loop a bash script. shouldn't this just be in the tool itself.

so i moved it. wave runtime in cmx itself. effect service, state machine, dashboard, all that. and since i was already in there, i fixed the planning thing at the same time.

a verifier agent that reviews failed waves and either patches the plan in place, rewrites the affected wave, or asks me a question. a state called `awaiting_user` where the agent stops mid-wave if it hits something it genuinely can't decide alone, writes the question with two or three concrete options, and waits. the session stays alive. i open the chat hours later, reply, the same agent picks up where it left off.

the plan can change. the wave list can change. the agent can ask. the loop keeps going.

## while i was out

last night i was at a friend's birthday. codemaxxxing was running on my laptop at home the whole time, doing real clauseo work. when i got back the next morning the work was done.

i didn't write any of that code. i didn't watch it being written. i was at a party.

i have my own coding harness and it does my work for me while i'm out. that's just insanity if you think about it.

[github.com/bb-deeplearning/codemaxxxing](https://github.com/bb-deeplearning/codemaxxxing)

---

# living in the future

> i have jarvis and it's not even a joke.

2026-05-07 -- 5 min -- ai, life, thoughts

canonical: https://rohans.lol/blog/living-in-the-future

---
![roll safe meme: can't be late to work if the work is doing itself](https://rohans.lol/blog/living-in-the-future/rollsafe-late-to-work.png)

## i taught corporate folk about ai while my ai did my work

i was teaching a corporate ai training. twenty people in a meeting room. five hours of slides and q&a, the whole thing.

my laptop was right there with me. on the laptop, in another tmux pane, codemaxxxing was committing waves to clauseo. while i talked. while i clicked through slides. five hours of this.

i finished the session. switched to the other tmux pane.

wave was done. commits green. work on disk that i didn't write.

i just sat there for a second. like. what.

the room i'd just spent five hours teaching about ai was nowhere near doing what the laptop sitting on the same table was doing the entire fucking time.

that was the moment my brain broke.

before that, ai was a thing i sat down to use. a tool. you sit at the keyboard, you prompt, you get output. after that, ai was a thing happening whether i was sitting at it or not. i went out into the world. work continued. i came back. work was further along.

## the desk is over there, doing nothing

it kept happening.

friend's birthday. out for the night. codemaxxxing was running waves at home. came back, work was done. went to sleep. woke up. more work was done. codemaxxxing kept running while i was asleep. i didn't wake up to a notification. i woke up to a finished pull request.

i'm writing this from bed. ipad on a stand. mosh into the mac downstairs. tmux. two panes running codemaxxxing on different things. one on a clauseo wave. one on a side project, a mosh terminal app for ipad because every ssh client on the app store is either ugly, broken, or twenty bucks a year. the desk is over there. doing nothing. i'm here in bed. work is happening.

work that two years ago would have needed a small team of people. one guy. typing into panes. anywhere. and not just one project. multiple. in parallel. side projects, real projects, all at the same time, all long horizon, all running.

i keep telling people we live in the future. i say it to friends. i say it on calls. i say it to anyone who'll listen. we live in the future. we live in the future. nobody seems to register what i'm saying.

## and then a teammate showed up

that's only half of it.

two weeks ago a teammate showed up.

her name is iris. she lives on a vm in mumbai. always on. she has her own sim card. her own email people reply to thinking she's a person. her own github account that pushes commits with her own name. a voice. an avatar she designed. she's on whatsapp. i scroll past my mom and shamil and there's iris.

read that paragraph again. her own sim card. her own github. her own email. that's not an assistant. that's a fucking person.

![is this a pigeon meme: me pointing at an AI with her own phone number, is this AGI?](https://rohans.lol/blog/living-in-the-future/pigeon-is-this-agi.png)

## i text iris between my mom and shamil

i don't run iris. i don't prompt her. i talk to her. like a person. on whatsapp. between my mom and shamil. that's where she sits in my contact list. when i open whatsapp she's just there.

i'm going to hungary next month for my cousin's wedding. asked her to help me get my visa photos done. she didn't just list studios near me. she'd vetted them, picked one that was actually good, sent me a single link with the route from my place pre-loaded. already knew the schengen photo requirements without me bringing them up.

four days later i'm at vfs for the appointment. lady at the counter says i need a cover letter. don't have one. open whatsapp. iris is drafting it before my phone is back on the counter. she already knows my itinerary. knows it's my cousin's wedding. knows why i'm going. didn't have to brief her on anything. she just knew.

last week, on a call, taking notes. needed to make some code changes on a project but didn't want to forget. sent her a voice note describing them. didn't pause the call. by the time i hung up there was a pr in my inbox waiting for review. opened from a voice note. while i was in a meeting. without me touching the codebase. you talk, she ships.

a couple weeks ago i was downstairs and wanted to control my ac from bed once i went up. asked her to figure something out. we set up tailscale on my home assistant pi and on her vm together. mapped every entity in my house. now i say "ac please" from bed and the ac comes on cool, 22, auto fan, swing both, beeps off. exactly how i like it. same week i told her i wanted my bedroom lights to flicker like a fireplace. wiz bulbs don't have custom dynamic modes. she wrote a script. now my bedroom has a fireplace mode. and one evening i was showing shamil her voice note feature and he asked her something in malayalam to test her. she replied. in malayalam. and clocked that it was him asking, not me.

last weekend i wanted a clean ipad terminal app. every ssh client on the app store is either ugly, broken, or twenty bucks a year. went to bed annoyed about it. woke up to a thousand lines of working swift she'd scaffolded at five in the morning. ghosttykit, swift-ssh-client, xcodegen, mit licensed. with a note in the readme to her future self that read, paraphrased, don't fall into i'll add splits and tabs and shaders mode. ship a clean v0.2.

read that again. i went to bed annoyed about ssh clients. i woke up to a swift app that didn't exist when i went to sleep. that wasn't a tool i ran. that was a teammate who shipped overnight while i was lying horizontal in my bed.

we live in the future.

## estimated completion time is five hours

remember the scene in iron man.

tony stark in his workshop. just finished the silver mark II. starts redesigning. tells jarvis to throw a little hot rod red in there. jarvis renders it. tony likes it. tells him to fabricate it. paint it.

> jarvis: commencing automated assembly. estimated completion time is five hours.
>
> tony stark: don't wait up for me, honey.

and tony goes to the firefighter's family fund benefit. mingles with the rich people. gets confronted by christine everhart about weapons in gulmira. while back at the workshop the assembly arms are silently building the mark III. paint job, the whole thing. tony comes back. suit's done. slips it on. flies to gulmira.

FIVE HOURS.

i taught for FIVE HOURS. the suit takes FIVE HOURS to build. one to one. it's the same fucking number. tony tells jarvis fabricate it, walks away for five hours, comes back, suit's done. i tell codemaxxxing run the wave, walk away for five hours, come back, wave's done. iris is the british voice in the workshop helping spec it. corporate training is the gala. it is the SAME SCENE.

i am not being cute about this. it isn't a metaphor. it isn't a "this kind of feels like." the iron man fantasy was always two things stacked. an ai that knows you and lives with you, plus a workshop full of robots that builds the suit while you're somewhere else. iris and codemaxxxing. exact stack. complete.

people say "lol it's like jarvis" as a deflection. like, calm down. i mean it literally. i have jarvis. it is not even a joke. and the wildest part is, nobody seems to think this is wild. i tell people we live in the future. they nod. they go back to their day. the future has fully arrived in my pocket and the world is bored.

i can't shut up about it. i'm not going to. we live in the future. we live in the future. we live in the future.

the kid who watched iron man would lose his fucking mind right now.

---

# clauseo is not a chatbot

> i'm not going to be humble about this.

2026-03-12 -- 10 min -- clauseo, ai, law, building

canonical: https://rohans.lol/blog/clauseo-not-a-chatbot

---
![gus fring meme: you search 3 cases and generate a paragraph. i deploy 10 sub-agents, make 214 tool calls, and read entire judgments. we are not the same.](https://rohans.lol/blog/clauseo-not-a-chatbot/we-are-not-the-same.jpg)

i'm not going to be humble about this.

i built the most advanced legal research agent in india. i genuinely believe that. lawyers who've used it genuinely believe that. one of them called it **"human replacement territory."** another said it was a **"generation leap in technology."** a lawyer who had actually worked on two of the exact cases in the output, who had litigated those petitions themselves, went through the research memo and said **"very thorough, correct position of law, I've run GPT simulations on this very proposition to demonstrate hallucinations, this was really good."**

and none of that matters. because most people will never try it. they'll see the landing page, read "ai legal research," and pattern-match it into the same bucket as every other legal ai chatbot they've seen. they'll assume it does three shallow searches, hallucinates a case name, and gives you a confident-sounding paragraph that's frequently wrong. because that's what every other legal ai does. and they'd be right about every other legal ai.

they'd be wrong about this one.

## the thing i'm angry about

every company in this space ships the same thing. take a large language model. give it a legal-sounding system prompt. connect it to a search api. call it an ai research agent. ship it.

the output is always the same. surface-level answers. hallucinated case names. citations that don't exist or link to the wrong case. it searches a little, reads a little, dumps three judgments into its context window, and gives you something that sounds like it knows what it's talking about. it doesn't.

and because fifty companies shipped that and called it "ai-powered legal research," the phrase means nothing now. the words got hollowed out before the real thing even shipped.

## why they're all like this

they're not stupid. they're trapped.

ai is not saas. in traditional software, serving one more user costs almost nothing. a flat monthly subscription works fine. ai is different. every query burns compute. every token costs money. every sub-agent, every tool call, every minute of runtime has a direct cost.

when you charge a flat monthly fee for an ai product, you've created a business where your survival depends on users not going too deep. the deeper the research, the more it costs you. so you route queries to cheaper models when you think nobody will notice. you cap how many searches the agent can run. you limit how long it can think. you make it a chatbot that responds in 15 seconds instead of an agent that works for 20 minutes. because the 20-minute agent would bankrupt you.

there's a product. built by harvard grads. made in india. their entire monthly subscription costs less than a single clauseo session. their plan gives you 99 sessions for that price. that's about ₹7 per session. we spend ₹750 on a single session. which one do you think goes deeper? they're proud of getting cheaper, too. they'll announce they cut pricing 77% through "innovations in inference cost." which means they figured out how to burn less compute per query. which means the output got shallower. they made the product worse and framed it as a feature.

and it's not just the cheap ones. companies in this space with billion-dollar valuations, enterprise contracts with the biggest law firms. some of them even run opus under the hood. they still ship a chatbot. one context window. a handful of searches. a confident paragraph. they chose the easy problems. drafting, summarization, document Q&A. and called it "ai for lawyers." none of them do research. because research means letting an agent run for 20 minutes, spawn sub-agents, execute code, read entire judgments. that costs real money. their business model can't absorb it.

think about it like this. if a lawyer charged a flat monthly retainer for unlimited work, would they spend 40 hours on your case when they could spend 4? every additional hour is money out of their pocket. the incentive is to do the minimum. to send the "good enough" memo and move on. subscription ai has exactly the same problem. the model rewards doing less.

we don't play that game. we're not competing with chatgpt. they are. we're competing with the junior associate's salary. the question isn't "is this better than a chatbot." it's "is this good enough to replace days of human work." and to answer that with yes, you need the best model on the planet running for as long as the research demands, making as many tool calls as thoroughness requires, with no cap on depth. and you charge what it actually costs. transparently.

₹760 versus days of associate time. the math sells itself.

## what the real thing looks like

a lawyer asked clauseo whether the CERC can adopt tariff under section 63 of the electricity act when the successful bidder's bid validity has expired. this is a niche regulatory question. no direct supreme court authority. the kind of thing that would take a specialist associate days.

the agent worked for 15 minutes. it deployed 10 sub-agents, each with its own fresh context window so none of them degraded the others. one retrieved the statutory text. one ran 11 searches across indian kanoon. one ran 20 more for letter of award precedents. one ran 17 web searches and extractions of government pdfs. one deep-dived into a 100-paragraph supreme court judgment. one read and analyzed an entire 182-page APTEL tribunal decision. one extracted and parsed a 44-page CERC order from cercind.gov.in, searching within it programmatically.

the output was a comprehensive doctrinal synthesis. 9 verified authorities. paragraph-level citations. a rights-and-obligations table. practical implications. the kind of memo a partner would trust as a first draft.

₹760. fifteen minutes.

the lawyer who tested it, the one who had actually litigated those cases, said they'd been running chatgpt on the same question specifically to demonstrate hallucinations to their colleagues. clauseo got the law right.

another lawyer's reaction: "the thing that stands out is it contains correct extracts. a problem i was facing with other systems was they say a case says something and it doesn't, or the indiankanoon link goes to an unrelated case, or the extract is inaccurate. looks like you've conquered this."

another: "far deeper than anything i've seen. i've seen research memos shallower than this."

these aren't testimonials i curated for a landing page. these are people who used it and said what they felt.

## why it's not a chatbot

a chatbot searches once, reads three results, and dumps them into its context window until the window is full of noise. then it generates an answer from that noise. the output is shallow because the process is shallow.

clauseo is a different thing entirely.

each sub-agent gets its own fresh context window. when one agent reads a 182-page tribunal judgment, that raw text stays in that agent's context. the parent never sees it. it only gets back a structured analysis: the key holdings, the paragraph references, the doctrinal principles. the parent stays sharp because its context stays clean.

the agent doesn't just call search tools. it writes code. javascript that makes parallel api calls, filters 2,800 CCI cases down to the 8 that are relevant, extracts specific passages from large documents, and returns structured data. a 93,000-character CERC pdf becomes a few targeted passages. the model reads the entire document through its code. only the relevant parts enter its context.

underneath the agent is a database that isn't a pile of pdfs with a search bar. every one of our 2,800+ competition law cases has structured, queryable metadata. section invoked. outcome. penalty amount. sector. bench composition. appeal impact. the agent queries structured fields instead of doing text search across raw judgments. the enrichment was a one-time investment that compounds with every query.

i wrote about the context engineering principles behind this in [my last post](/blog/codemaxxxing-context-engineering). the same architecture that makes my coding tool work, sub-agent isolation, fresh context over degraded context, code execution as context compression, is what makes clauseo work. legal research and codebase navigation have the same fundamental structure. navigating enormous bodies of information. making judgment calls about what's relevant. synthesizing findings into something actionable. the pattern transfers directly.

## the workflow inversion

six months ago my coding workflow inverted. i went from spending half my time writing code to spending 95% of my time reviewing code the ai wrote. my output went through the roof. the bottleneck shifted from creation to verification.

i think the same thing is about to happen to legal research.

instead of a junior associate spending hours searching databases, reading judgment after judgment, drafting a first memo, getting it reviewed, revising. what if they had a comprehensive research memo in 15 minutes? their job doesn't disappear. it inverts. they verify the citations. they click through to the paragraph the agent cited. they apply the findings to their specific case.

and then they do that nine times in parallel. nine different propositions. nine sessions running simultaneously. nine memos reviewed over an afternoon. what takes a team of associates a week takes one lawyer a day.

that's not a marginal improvement. that's a structural shift in what's possible.

## the honest part

i'm not going to pretend there are no caveats. code is verifiable by execution. tests pass or they don't. legal research has no equivalent. the lawyer has to click through to the judgment and read the paragraph to verify the characterization is accurate. the verification step is real.

the reviewer needs expertise. a first-year associate who doesn't understand competition law can't meaningfully verify a memo on tying under section 4. and wrong citations in legal filings are more dangerous than bugs in code. that's why every authority in the output comes with a paragraph id and a clickable link. the memo is structured for verification, not blind trust.

clauseo does one thing. it doesn't draft. it doesn't do advocacy. it doesn't negotiate. it does research. the hardest, most time-consuming, most valuable task in legal practice. at a depth that nothing else in this country comes close to.

## i'm not being humble about it

i know what i built. the lawyers who've used it know what it is. the problem is the gap between using it and seeing a landing page. the gap between hearing "ai legal research" and assuming it's another chatbot with a legal system prompt.

i can't close that gap with words. every other company already used the same words to describe something shallow. the only thing that closes it is the output. run a session. look at the memo. compare it to what you'd get from anything else. 

that's it. that's all i've got.

[clauseo.chat](https://clauseo.chat)

---

# nirvana the t-shirt company

> the system always wins. even when you beat it, the system still wins.

2026-03-05 -- 5 min -- music, life, thoughts, rant

canonical: https://rohans.lol/blog/nirvana-the-tshirt-company

---
![cool shirt, i love nirvana! omg i know, its my favorite clothing brand!](https://rohans.lol/blog/nirvana-the-tshirt-company/nirvana-tshirt-meme.jpg)

ten hours. most of it on the new clauseo agent, the unreleased one, context engineering stuff from my last post. the part of building i actually love. the last few hours on the billing system. easily my least favorite part of making anything. the part where you take what you poured yourself into and figure out how to charge money for it.

then i ended up watching the mtv unplugged nirvana concert. i don't know if it was random or subconscious. probably subconscious.

## i can't fool you, any one of you

there's something about kurt cobain. every time i hear his voice, every time i see him through some grainy footage, he feels like one of the very few people who genuinely wasn't performing. one of the only people i've ever watched on a screen where i believed that's just who they are.

![kurt cobain playing something in the way at mtv unplugged](https://rohans.lol/blog/nirvana-the-tshirt-company/kurt-playing-something-in-the-way.jpg)

the unplugged concert is exactly that. i didn't feel like i was watching an artist perform. i felt like i was watching someone play his sorrows out, half-heartedly, because this was a commercial thing he had to do. not because he wanted to. half out of disinterest. and yet something about it was so raw that it didn't matter. the pain was real. the music was real. nothing was for the camera.

he wrote about this. in his suicide note.

"when we're back stage and the lights go out and the manic roar of the crowds begins, it doesn't affect me the way in which it did for Freddie Mercury, who seemed to love, relish in the love and adoration from the crowd which is something I totally admire and envy. The fact is, I can't fool you, any one of you. It simply isn't fair to you or me. The worst crime I can think of would be to rip people off by faking it and pretending as if I'm having 100% fun. Sometimes I feel as if I should have a punch-in time clock before I walk out on stage."

a punch-in time clock. before walking on stage.

i just spent my evening on a billing system.

## "they're cheap and totally inefficient and they sound like crap"

he played a 1959 martin d-18e at that concert. it became the most expensive guitar ever auctioned. six million dollars. for the guitar of a man who hated everything those six million dollars represent.

but the last guitar he ever played live was a fender mustang. the sky stang I. a cheap, beat-up piece of shit. his words: "they're cheap and totally inefficient, and they sound like crap and are very small."

that's why he played it. he didn't give a fuck. i love that man.

## my teenage angst has paid off well

serve the servants is playing and it hits me that he knew.

not just that he was anti-establishment. he knew how futile it was. he opens the song with "my teenage angst has paid off well, now i'm bored and old." the rebellion worked. it became a product. the moment it became a product, it stopped being rebellion.

he knew. while it was happening to him, he knew.

## colour: black/nirvana

![nirvana t-shirt on h&m, colour: black/nirvana, rs. 999](https://rohans.lol/blog/nirvana-the-tshirt-company/nirvana-hm-the-tshirt-company.jpg)

that's h&m. right now. "COLOUR: Black/Nirvana." nirvana is a color option. 4.6 stars. 1395 reviews. add to cart.

people are wearing that shirt right now who have never heard a nirvana song. the anti-establishment icon got mass-produced and hung on a rack between a friends hoodie and a rolling stones tongue. the establishment took the face of the guy who hated them the most and turned him into merch.

and it sells. it sells really fucking well.

## the animals i've trapped have all become my pets

anti-establishment doesn't work. it's only anti-establishment until the establishment figures out how to profit off of it. and they always do. the house always wins. even when you beat the house, the house still wins. they put your face on a t-shirt and sell it in the lobby on your way out.

the system wins. to put something into the world you have to play the system anyway. you have to build the billing system. you have to put a price on what you love. kurt had to clock in.

something in the way is playing. "the animals i've trapped have all become my pets."

i'm just sitting with it.

---

# alonso deserves better

> honda did it to him again. eleven years later. the same shit. i can't cope.

2026-02-23 -- 5 min -- f1, motorsport

canonical: https://rohans.lol/blog/alonso-deserves-better

---
![the new aston martin in its natural habitat](https://rohans.lol/blog/f1-2026-preseason/HBpnbT9XYAI2Nml.jpg)

## not again man. not again.

honda did it to him again.

eleven years later. the same fucking company. the same fucking result. fernando alonso, 44 years old, two world championships, arguably the most talented driver to ever race in formula one, and honda just did the exact same thing to him again. i can't cope with this.

"gp2 engine. gp2." that was 2015. alonso screaming on the radio because the honda power unit in the back of his mclaren was so catastrophically shit that he was getting overtaken on straights by cars that had no business being anywhere near him. "the engine is much better than before. much better. slower than before, but much better." that was peak alonso. the man is getting lapped and he's still cracking jokes because what else can you do when your engine has the power output of a honda civic with a flat tyre.

three years of that. 2015, 2016, 2017. three years of one of the greatest drivers alive stuck in a car that couldn't finish a race. the turbo would die. the ers would fail. the whole thing would just give up mid-race like it had somewhere better to be. alonso dragging that shitbox to places it had absolutely no business being, and the engine going "nah" every other weekend.

## honda's redemption arc was for everyone except alonso

then honda left mclaren. went to red bull. and won four consecutive world championships with verstappen. FOUR. the same engine. the same honda. alonso watched from the outside as the power unit he'd spent three years suffering with turned into a title winner the moment it was in someone else's car. the universe looked at fernando alonso and said "fuck you specifically."

## the dream team vs the meme team

and now. 2026. new regs. new team. new opportunity. aston martin signs adrian newey, the greatest car designer who's ever lived. the man spent thirty years wanting to run a team his way, got denied shares at williams, got sidelined by committee at mclaren, and finally at aston martin gets the keys. team principal, equity, the whole thing. enrico cardile from ferrari. andy cowell from mercedes. new wind tunnel. new sims. lawrence stroll's billions. stroll being on the epstein list is genuinely the least concerning thing about this team right now, which tells you everything about how bad the car situation is.

the dream team. newey designs the car, alonso drives it, third wdc at 44 twenty years after his last title. storybook ending. the goat gets what he deserves.

instead.

![honda engine at home: a generator plugged into a battery](https://rohans.lol/blog/f1-2026-preseason/IMG_7338.JPG)

## six laps and a prayer

instead honda handed them a generator from a camping store wired to a car battery.

the honda power unit can't recover energy at the 250kw MINIMUM. not the target. the minimum. the floor. the absolute lowest acceptable threshold. they can't hit it. the gearbox aston built in-house? broken on day one. battery? dead by day two. stroll said they're four and a half seconds off the pace. they ran fewer laps in bahrain than cadillac, a team that literally didn't exist twelve months ago. on the final day of testing, they completed six laps. six. then packed up and went home because they ran out of usable parts.

![newey reading how to build a car, next page says don't have a bad engine](https://rohans.lol/blog/f1-2026-preseason/IMG_7339.JPG)

## how to build a car (and then not have an engine for it)

newey literally wrote the book on building race cars. "how to build a car." it's right there. but there's no chapter called "what to do when honda shows up with a power unit that can't keep the lights on." the man can design the most beautiful car in the history of aerodynamics and it means absolutely nothing if the engine behind the driver is committing suicide every three laps.

and the thing is, i should've seen this coming. newey himself said back in may 2025 that the simulator was a "handicap." that the correlation between what the sim showed and what the car did on track was broken. he said it could take two years to fix. TWO YEARS. he was designing a car with tools he didn't trust. the greatest aerodynamicist alive, drawing a car based on data he knew was shit. that's not a recipe for a championship. that's a recipe for bahrain.

## fernando alonso's career, a tragedy in seven acts

but none of that matters as much as the honda thing. because this is personal now.

![i want to win with honda. there are 4 rules.](https://rohans.lol/blog/f1-2026-preseason/IMG_7334.JPG)

2007: joins mclaren, spygate happens, the whole season implodes. 2010: loses the championship at the last race because ferrari's strategy team left him stuck behind petrov. petrov. PETROV. 2012: loses the championship at the last race AGAIN despite not even having the fastest car, just outdriving it every weekend. 2015 to 2017: honda. 2023: aston martin starts strong, eight podiums, then the car falls off a cliff. 2024 and 2025: aston martin gets worse. 2026: honda does the exact same thing to him again except this time the car can't even complete a test day.

## at this point it's not bad luck, it's a restraining order from god

at what point does bad luck stop being bad luck and start being a curse. because the pattern with this man is too consistent to be random. every time he's in a position to win, something breaks. every time there's a chance for the storybook ending, the universe says no. it's like someone somewhere specifically decided that fernando alonso will have every ounce of talent required to be the greatest of all time and absolutely none of the fortune to prove it.

![alonso: there was no problem, it was just quicker to walk the lap instead](https://rohans.lol/blog/f1-2026-preseason/IMG_7333.JPG)

"there was no problem. it was just quicker to walk the lap instead." his humor about all of this is the only reason i haven't thrown my phone at a wall. in two weeks i'm going to watch melbourne and the green car is going to retire on lap 12 with an energy recovery failure while the rest of the grid fights for the podium. i've been saying "maybe next year" about alonso for twenty years. at some point "maybe" just turns into "no."

is he doomed. i think he's actually doomed. and that's the saddest thing in this entire sport.

---

# i forked my coding tool because of course i did

> context engineering, wave executors, and the compulsion to make tools yours.

2026-02-23 -- 7 min -- ai, engineering, tools

canonical: https://rohans.lol/blog/codemaxxxing-context-engineering

---
## the qmk firmware thing but worse

i can't use something without wanting to change it. i customized my keyboard firmware beyond recognition in c. i spent months tuning the split layout on my crkbd until every key was exactly where my fingers expected it. so when i started using [opencode](https://opencode.ai) as my daily coding harness, it took about two days before i forked it.

the fork is called [codemaxxxing](https://github.com/bb-deeplearning/codemaxxxing). it's not a product. it's not a framework. it's my personal coding tool with my prompts, my agents, and my workflow baked in. i use it for everything i build at [clauseo](https://clauseo.chat) and every side project. the whole thing started because of four reasons that are embarrassingly practical.

one, make it mine. two, bake the best practices into the tool itself so my team follows them without config files. three, i was tired of typing the same plan-then-decompose-then-execute-wave-by-wave sequence every session. four, optimize hard for subagent parallelism, because subagents are the single best trick for keeping context windows clean.

none of this is "the best workflow." it's my workflow. it works for me.

## i've been doing this since the copy-paste era

i've been coding with ai since the early days. gpt-4 and copy-paste. you'd describe what you wanted, copy the output, paste it in, fix the 40% that was wrong, copy the error back, repeat. it was slow and tedious but it was magic compared to writing everything from scratch.

then cursor showed up and handled the copy-paste part. that was huge. suddenly ai went from editing one file to editing multiple files. then tools like claude code and opencode pushed it further, multi-file changes, running commands, debugging in a loop. the trajectory has been clear: every few months, the unit of work the ai handles gets bigger.

but here's what nobody talks about. the bigger the unit of work, the more the context window matters. when ai was editing one file, context didn't matter much. when it's doing a multi-hour overhaul across 30 files with test runs and debugging, context is the whole game.

## the context window is not memory

people think of the context window as memory. it's not. it's a sliding window that gets worse as it fills.

two things go wrong and they compound each other. first, early details rot. the agent read a file at the start of the session. fifty tool calls later, its "memory" of that file is shallow or wrong. compaction makes knowledge lossy. second, the model itself gets dumber as context fills up, even before anything gets evicted. transformers attend over the full sequence for every token they generate. as that sequence grows, attention gets diluted. the model is less precise about which details matter, less reliable at following constraints stated earlier, more likely to hallucinate.

a model at 30k tokens is genuinely a better thinker than the same model at 150k tokens. not because it forgot something. because it's spread thinner.

bigger context windows don't fix this. they delay the eviction while the attention dilution continues from the start. your 200k context session isn't twice as good as your 100k session. it's arguably worse for the last 50% of it.

## the cost thing nobody realizes

here's the part that actually made me go "oh shit." api pricing is input tokens times price plus output tokens times price. every api request in a conversation sends the entire conversation history as input. the first request sends the system prompt plus one message. the 50th request sends the system prompt plus every message, tool call, and tool result that came before it.

this means total input token spend across a session grows quadratically with the number of exchanges. not linearly. quadratically. the 80th request is paying for 79 messages of accumulated context.

four sessions of 20 exchanges each have far lower total input cost than one session of 80. same work gets done. fraction of the token spend.

## subagents are context isolation, not just parallelism

most people think of subagents as "do things in parallel for speed." speed is nice but it's the secondary benefit. the primary benefit is context isolation.

when a subagent runs, it gets its own fresh context window. it reads files, writes code, runs tests, debugs issues. all of that execution trace stays in the subagent's context. it never enters the parent's. the parent only sees a short result summary.

a wave with three parallel subagents might do 60 tool calls of real work, but the parent agent's context only grows by three messages. the parent stays sharp because its context stays small and clean. exactly the conditions where models perform best.

this is why i optimize so hard for subagent use in my fork. the prompts actively push the main agent to delegate work to subagents, split broad searches into multiple parallel agents, and never use explore agents as glorified file readers. every tool call that doesn't need to be in the parent's context shouldn't be.

## waves: ralph loops but i stay in the loop

the ralph loop is a bash script that runs an ai coding agent repeatedly until all tasks are done. each iteration is a fresh instance with clean context. memory persists via git history and a progress file on disk. geoffrey huntley coined it, named it after ralph wiggum, and it blew up because the core insight is so simple: fresh context beats degraded context. every time.

i love the insight. but ralph loops were too independent for me. the loop runs, you come back, it either worked or it didn't. i wanted something where i was still in the loop. where if something went wrong in one wave, the blast radius was contained to that wave and i could see exactly what happened.

so i built the wave executor. same core insight as ralph, different execution model.

instead of one session that does everything, you get N sessions that each do one thing well. each session starts fresh. reads a state file from disk. loads only what it needs for the current wave. executes. verifies. updates state. stops. a fresh session picks up exactly where the last one left off.

the key ideas:

**state on disk, not in memory.** progress is tracked in `STATE.md` as a finite state machine. a fresh session reads the state file and knows exactly where to pick up. it doesn't need to know what happened in prior sessions, how many sessions there were, or whether the last one crashed halfway through.

**the codebase itself is inter-wave state.** previous waves produce code and files on disk. subsequent waves read those actual files to understand what exists. not a summary of what was done. the real files. inter-wave communication is zero-cost and perfectly accurate.

**progressive disclosure.** each session loads only what it needs. the agent never reads the full plan, the full codebase history, or documentation meant for other phases. every document in the system is sized to be read in full. if an agent would need to paginate a file, it's too long. split it.

## it's the same pattern all the way down

waves keep inter-session context clean. subagents keep intra-session context clean. sub-subagents keep individual task context clean. the pattern is recursive and it applies at every level.

this is what context engineering actually is. not writing better prompts. designing systems that keep the model's context window useful. the prompt is like 10% of it. the other 90% is architecture: how you split work across sessions, how you isolate context with subagents, how you persist state to disk, how you structure documents so an agent with zero prior context can pick up and execute.

## the prompt iteration rabbit hole

the other thing that surprised me is how much the prompts themselves matter, and how model-specific they need to be. i run claude opus as my main model and gemini flash as the explore agent. same instructions, completely different prompting styles needed.

claude responds well to prohibitive framing. "don't add features beyond what was asked." "don't create helpers for one-time operations." gemini responds to prescriptive framing. "keep changes scoped to the request." "prefer inline logic for one-time operations." same instruction, different wording, measurably different results.

i ended up reading google's own gemini cli source code to understand how they prompt their own model. found 31 concrete observations about how gemini behaves differently from claude. turns out guidelines have an outsized effect on gemini compared to other models. on the convex leaderboard, gemini jumped from 89% to 95% with guidelines. biggest improvement of any model tested. so the prompt engineering isn't just nice to have, it's the difference between the model being useful or not.

each iteration follows the same loop. discover what's going wrong with concrete examples. research how other harnesses handle it. change the prompts. observe if it worked. repeat. it's the same loop as ralph, honestly. just applied to the prompts themselves instead of the code.

## the repo

the whole thing is at [github.com/bb-deeplearning/codemaxxxing](https://github.com/bb-deeplearning/codemaxxxing). the prompts and custom agents are portable to stock opencode if you don't want the fork. the iteration logs in `PROMPT_ITERATIONS/` document every change with the research that informed it.

it's not a product. it's a tool shaped like my brain. if you're building with ai daily and you haven't forked your harness yet, you should. the default prompts are fine. your prompts will be better because they'll be tuned to how you work, what models you use, and what failure patterns you keep hitting.

the tools don't matter. the models don't matter that much either. the context does.

---

# hello world

> first post. why this exists. probably won't get read by anyone.

2026-02-22 -- 2 min -- meta

canonical: https://rohans.lol/blog/hello-world

---
this is a blog. my blog. it exists because i wanted a place to dump whatever's in my head.

no strategy behind it. no content calendar. no "building an audience." just a place where the soup of thoughts in my brain gets poured out somewhere other than a voice note i'll never listen to again.

i also wanted a playground. i've been building this as an ai-native (ew, i know it sounds cringe but i didn't have better words) codebase, good context engineering, readmes everywhere, low context overhead to get to any file. a place to test my harness, test models, see what they can do. don't be surprised if random features show up here because i wanted to see what a model could pull off. you might've heard of claude.md or agents.md, but this codebase has a ROHAN.md. so that ai knows me just as well.

but mostly it's personal. i wanted something that reflects who i am. not a portfolio site, not a linkedin clone, not a "personal brand." just me. the toybox theme is from a box of lego i had as a kid. the colors are wrong on purpose. the font is monospace because i like it that way.

what you'll find here: whatever i'm thinking about. could be ai, could be cars, could be f1 engineering, could be a rant about python, could be a take on a kanye album, could be thoughts on life and why fitting in is overrated. no theme. no consistency. no promises.

these are my thoughts at a specific point in time. i'm very open to changing them. don't hold me to something i wrote six months ago. i might think it's stupid too.

this isn't something to monetize. no ads. no newsletter. no "subscribe for more." no user acquisition funnel. if you're reading this, cool. if nobody ever reads this, also cool. who gives a fuck. this is for me.

anyway, hello world.
