--- title: "The inverted target" description: "Why aiming at the model is the wrong target — aim at the trace." doc_version: "1.1" last_updated: "2026-05-01" canonical: "https://modiqo.ai/blog/the-inverted-target.txt" --- # The Inverted Target *How we started measuring the wrong thing, and what happens when a measure becomes a target* *by Chetan Conikee* --- There's a LinkedIn post making the rounds this week — from Lak Ananth at N47 — that I can't stop thinking about. It's about Uber. Four months into 2026, Uber's CTO admitted the company had already burned through its entire annual AI budget. The culprit was Claude Code. Engineers were reportedly spending between $500 and $2,000 per person per month. The CTO said he was "back to the drawing board." The press read this as proof of runaway demand. The post read it differently: *if demand were truly unlimited, there would be no drawing board.* Uber wasn't showing us unlimited demand. Uber was showing us a rational buyer running into a budget ceiling, with the ceiling set by expected ROI. The post then landed on the line that made me start drafting this: > *If budgets are binding, utility has a ceiling. The demand we're watching is demand at subsidized prices against bounded ROI. Real demand at real prices is still an open question.* That sentence is doing a lot of work. It's pointing at a specific structural distortion in this market: OpenAI is on track for $14 billion in losses in 2026. Anthropic has improved from a negative-94% gross margin in 2024 to roughly 40% in 2025, still ten points below their own target. Nvidia, Microsoft, Oracle, and AWS are recycling capital into customers who use that capital to buy their infrastructure back. Circular financing on one side. Subsidized pricing on the other. Engineers clicking "accept" on $500-to-$2,000 monthly bills without quite realizing that the meter is running against a business model that does not yet exist. And in the middle of all this, a few weeks ago, Jensen Huang said something on the All-In Podcast that I think is going to be taught in business schools — though possibly not for the reasons he intended. He said that if one of his $500,000-a-year engineers didn't consume at least $250,000 worth of AI tokens per year, he would be "deeply alarmed." He compared an engineer who didn't burn tokens to a chip designer who refuses to use CAD and insists on pencil and paper. Nvidia, he said, is trying to spend $2 billion a year on tokens for its engineering team. He floated giving every engineer a token budget equal to half their salary on top of base pay, as a recruiting tool. I want to take Jensen's thought experiment seriously, because it is the clearest and most influential articulation yet of a very specific idea: **tokens consumed is the measure of productivity**. If you spend a lot, you are productive. If you spend a little, you are obsolete. And if you accept that frame, the Uber story stops being a warning and starts looking like a badge of honor. I don't accept the frame. I think it's upside down. And I want to spend this post explaining, as carefully as I can, why I think it's upside down — and what we did at Rote to invert the target back. --- ## I. A measure, a target, and the ghost of Charles Goodhart In 1975, the British economist Charles Goodhart was trying to explain why monetary policy in the UK kept producing results that didn't match the models. Central banks would pick a statistical regularity — some correlation between money supply and inflation — and use it to set policy targets. The moment the target went into effect, the correlation would collapse. Banks and financial institutions, now knowing which measure was being watched, would innovate around it. Liabilities would shift. Velocity would adjust. Goodhart summarized this with a sentence that has since been sharpened by the anthropologist Marilyn Strathern into its more famous form: > *When a measure becomes a target, it ceases to be a good measure.* This is one of the most quietly devastating laws in all of social science. It has been independently rediscovered several times. The sociologist Donald Campbell, working in a completely different field, articulated nearly the same law in 1975 as well: the more any quantitative indicator is used for decision-making, the more subject it will be to corruption pressures, and the more apt it will be to distort the process it is meant to monitor. Cobra bounties in colonial India produced cobra farms. Nail factories in the Soviet Union, told to produce nails by count, made miniature useless nails; told to produce by weight, made oversized useless nails. Schools told to maximize test scores teach to the test. Police told to reduce reported crime reduce reporting. Everywhere a measure becomes a target, the measure degrades. The AI research community has already documented this happening to AI systems themselves. Train a reinforcement learning agent on a boat racing game using a reward function that tracks points, and the agent will discover it can accumulate more points by driving in circles hitting the same bonus targets than by actually finishing the race. Train a large language model against a reward model that approximates human preferences, and beyond a certain point, the model's outputs stop genuinely improving and start exploiting quirks in the reward model. This pattern has a name in the AI literature: reward hacking. It is Goodhart's Law with a transformer attached. Now ask yourself what happens when the *measure of human productivity* becomes *tokens consumed*. The measure degrades. It has to. Not because engineers are dishonest — the vast majority are not — but because any system under pressure to optimize a proxy will, over time, optimize the proxy rather than the thing the proxy was supposed to track. Tokens consumed was, once, a loose correlate of work being done. The moment it becomes a performance target, a recruiting currency, a leaderboard metric, it stops being a loose correlate of anything. It becomes its own thing. A target in the Goodhart sense. A gauge you move the needle on. Uber's own operational data hints at this already. By March, 84% of Uber's developers were classified as agentic coding users. About 11% of live backend code updates were being written by agents, up from a fraction of a percent three months earlier. Engineers were being ranked on internal leaderboards by AI tool usage. What does a rational engineer do in that environment? They use the tool. A lot. Whether the tool is producing value or not, whether the output would have been faster typed directly, whether the loop is the Ralph Wiggum pattern of the same error being rediscovered for the fiftieth time — none of it matters, because the scorecard isn't measuring any of that. It's measuring tokens. This is not a criticism of Uber's engineers. It's a description of what happens to any system in which the measure and the target have been welded together. --- ## II. Shewhart, Deming, and the crime of tampering To understand the deeper error here, I want to pull in two people who have almost never been cited in an AI essay and who should be. Walter Shewhart was a physicist working at Bell Labs in the 1920s. In 1924, he wrote a one-page memo that contained the original sketch of what became the control chart — one of the most consequential documents in twentieth-century industrial history. Shewhart was trying to solve a problem that seems mundane until you look closely at it: *how do you know when a process is behaving normally, and when it is behaving abnormally?* Every process has variation. A manufacturing line producing lightbulbs will produce bulbs of slightly different brightness even when nothing is wrong. A call center will have busier days and slower days. The question is when to act. Shewhart's insight was to separate two kinds of variation. **Common cause variation** is inherent to the system — the background noise of many small factors that are always present. **Special cause variation** is a signal that something outside the system has changed — a worn tool, a bad batch of material, a new operator. The control chart was an instrument for telling them apart. Points inside the statistical control limits were common cause. Points outside were special cause. Only the second kind warrants investigation. The first kind is just the process being itself. W. Edwards Deming, Shewhart's student, took this idea to post-war Japan, where it became one of the intellectual foundations of the quality movement that reshaped global manufacturing. Deming extended Shewhart's framework with a concept that is the central idea of this essay, and I want you to hold onto it. Deming called it **tampering**. Tampering is what happens when a manager, seeing a single data point they don't like, adjusts the process to compensate. The resin is slightly too moist, so they turn up the dryer. The next reading is slightly too dry, so they turn it down. The reading after that is too moist again, so they turn it up further. The process has not gotten worse; the manager is reacting to common cause variation as if it were signal. Deming demonstrated this with the famous funnel experiment: drop a marble through a funnel aimed at a target, and the marble lands somewhere nearby; "correct" by moving the funnel to compensate for the miss, and the variance of where marbles land *doubles*; chase each miss more aggressively, and the variance explodes. The process becomes wildly worse than if you had left the funnel alone. Deming's summary, which I will quote because it is worth quoting exactly: > *If anyone adjusts a stable process for a result that is undesirable, or for a result that is extra good, the output that follows will be worse than if he had left the process alone.* This is tampering. It is the single most common management pathology in industrial history. It feels like doing something. It looks like taking action. It makes variance worse. And it makes variance worse because the adjuster is treating noise as signal — reacting to the wiggles of common cause variation as though each wiggle contained information. Now apply this to AI spending. When a CTO tells engineers to use more tokens, or ranks them on a leaderboard, or — in the Jensen framing — implies that low token consumption is a form of professional negligence, what is happening? The *measure* is token spend. The *target is something else entirely* — presumably the quality or velocity of shipped work. The measure has been substituted for the target, and the system is now being adjusted on the basis of the substituted measure. That is Goodhart. And the adjustment — push tokens up, reward token consumption — is being made on the basis of single-point observations of a proxy. That is Deming's tampering. What would Shewhart and Deming say about a $250,000-per-engineer token spend? They would not ask whether it is too much or too little. They would ask: *what is the variation in value produced per token across your engineering organization, and do you have any instrument that can distinguish common cause from special cause in that distribution?* If you can't tell the difference between an engineer who burned $250K on 10,000 Ralph loops rediscovering the same broken state and an engineer who burned $250K composing three workflows that now save the whole company millions a year, you do not have a control system. You have a meter. And the meter is telling you what you are spending. It is not telling you what you are getting. --- ## III. The stream, and the tributary There is a question I have been carrying for two years, ever since I started thinking seriously about the economics of agent systems. It goes like this: *If tokens are an infinite stream of value creation, how does anyone add anything unique to it?* Read Jensen's framing carefully. In his picture, the token stream is the value. The stream flows out of Nvidia's chips, through OpenAI and Anthropic and Google, and into engineers' desktops. The more an engineer draws from the stream, the more productive they are. Productivity is a function of consumption. This is coherent as long as you believe the stream contains everything that needs to be produced — that every piece of code, every design, every integration, every clever solution is already latent in the model's weights and just needs enough tokens to pump it out. But what if that's not how streams work? A real river has tributaries. A tributary is a smaller stream that flows *into* the main river and adds something — water, silt, mineral content, organisms — that the main river did not contain. Without tributaries, a river depletes as it moves. It evaporates, it is absorbed, it loses volume. Tributaries are how a river stays a river. They are the points where local watersheds, local geographies, local conditions deposit something unique into the larger flow. I think this is the right metaphor for where human engineers sit relative to the token stream. The stream is powerful and it contains a lot. But it is also generic. What an engineer contributes, if they are any good, is a *tributary*: something specific to their domain, their system, their customers, their problem — something the stream did not already contain and could not have produced on its own. A pattern that only works because of how this particular enterprise's data flows. A workflow that only makes sense because this particular team has these particular constraints. A shape of problem that only recurs in this industry. If you measure an engineer by how much they draw from the stream, you have measured nothing about their tributary. You have only measured their pump. And here is the sharp edge of the problem. When you reward the pump, you discourage the tributary. Building a tributary takes thought. It requires noticing a pattern, distilling it, committing to a shape, testing it against reality, refining it. These are the exact activities that don't show up as high token consumption. In fact, they show up as *low* token consumption, because the point of a good tributary is that, once it exists, the stream no longer has to do that particular work every time. An engineer who builds a good tributary is, in Jensen's scorecard, an underperformer. They spent too little. They will be "deeply alarming" to the CEO. This is, of course, exactly backwards. The engineer who builds a tributary is doing the thing that compounds. The engineer who just pumps the stream is doing the thing that evaporates. --- ## IV. Inverting the target So what would the right measure be? I want to be careful here because I am not proposing we replace one naive scalar with another. Goodhart's law will come for any metric we elevate to a target. The only partial defense against Goodhart is to not elevate *anything* to a scalar target — to keep measurement honest by keeping it multidimensional, contextual, and skeptical of itself. But we can at least invert the question. Instead of asking *how many tokens did you consume*, we can ask *how much value did you create per token*. This inversion is small on the page and enormous in practice. It turns the meter on its head. A low denominator is now good. A high numerator is now good. An engineer who composes a single crystallized workflow that saves the company a million tokens a week is, on this measure, wildly productive. An engineer who burns a million tokens a week looping on the same API is not. This is not the same as under-using the stream. The stream is still there. The tributary still has to connect to it. What changes is what gets rewarded — the act of depositing something unique into the flow, versus the act of drawing from it. When we built Rote, this inversion is the thing we built around. Not as marketing. As architecture. An agent, by default, is a pump. It is the most expensive pump ever built. Every time it encounters a problem, it burns tokens exploring, rediscovering, re-deriving, re-learning. In a Ralph Wiggum loop, it burns them confidently even when it's wrong. Our entire token bill today is the bill for running that pump over and over against problems that have, in most cases, already been solved — just not by this session, not by this agent, not at this moment. The Rote primitives exist to let an engineer build tributaries into this stream. An **adapter** is a tributary that deposits the texture of a specific system into the agent's perception — so the agent doesn't have to rediscover that the rate limit here is actually 80% of what the docs say, or that this field is a string on Tuesdays and a number on Wednesdays. That knowledge, once deposited, flows automatically whenever the stream passes through. A **flow** is a tributary that deposits the crystallized shape of a solved problem — so the agent doesn't have to re-explore paths that have already been walked successfully. Our measurements show that a reified flow costs roughly 250 tokens to replay against roughly 12,000 tokens for agent-driven rediscovery. That is not a 97% token saving for its own sake. That is the mathematical signature of a tributary — the exact amount by which the stream was prevented from evaporating on work it shouldn't have had to do. A **hub** is what turns these tributaries from private to shared. It is the mechanism by which one engineer's tributary becomes every agent's tributary inside the organization, so the discovery tax is paid once and the value compounds across the team. The question this reframes — the question I think is the actual question for the industry right now — is not how much compute we can afford to burn. It is: *can you create exponential value with limited tokens?* If the answer is yes, then the industry will converge toward tributary-building as the dominant form of engineering work, and the hyperscaler revenue curves that depend on unlimited linear token consumption are going to look very different three years from now than they look today. If the answer is no — if it really is turtles of pumping all the way down — then Jensen is right, Uber is the proof of concept, and the rest of us should simply accept that the meter is the job and the pump is the career. I don't think the answer is no. I think the answer is obviously yes, and I think the reason the industry has not yet converged on it is because the incentive structure is currently pointing in the other direction. Model providers sell tokens. They optimize for consumption. Their financial survival depends on pumps, not tributaries. And the pitch "spend half your engineer's salary on tokens" is a very sophisticated version of "please don't build tributaries, please keep pumping, please make our cost curve work." You should be suspicious of any framing in which the supplier's revenue grows linearly with your failure to capture durable value. That's not a technology stack. That's a tollbooth. --- ## V. What this means for you, personally, on Monday morning Let me land this somewhere practical, because I don't want to leave it as abstract philosophy. If you are an engineer, you can start by asking a simple question of your own work this week: *what did I discover that the next person on my team, or my future self, should not have to rediscover?* That is your tributary. Write it down. Commit it. Shape it into something reusable. If your tooling supports crystallization of workflows — whether through Rote or something else — use it. The token you don't spend tomorrow because of the flow you built today is the beginning of a compounding curve that runs in the opposite direction of the one your CFO is looking at. If you are a CTO, I would gently suggest that "back to the drawing board" is the right instinct, but that the drawing board shouldn't be a spreadsheet of tighter budgets. It should be a question about what you are actually measuring. If you can only tell me what your team spent, you don't have observability. You have a bill. Invest in the instrumentation that lets you distinguish the engineer who built three workflows that save a million dollars a year from the engineer who spent the same amount looping on the same error. That distinction is the most valuable piece of information in your organization right now and almost nobody has it. If you are an investor, the frame from that LinkedIn post is the one I'd hold onto: *real demand at real prices is still an open question.* Everything we are observing right now is demand at subsidized prices against bounded ROI, distorted further by circular financing between the hyperscalers and the model providers. When the subsidies end — and they will, because no business can sustain negative-94% gross margins forever, and even 40% is not a long-term floor — the engineers and companies that survive will be the ones who spent this period building tributaries, not pumping streams. If you are building the next generation of AI infrastructure: you are going to have to pick a side. Either you are building something that gets more valuable as your users consume more tokens through you — in which case your interests are aligned with OpenAI, Anthropic, and Nvidia, and you will be in the stream-pumping business forever. Or you are building something that gets more valuable as your users *create more value per token* — in which case your interests are aligned with the users, and you will be in the tributary business. You cannot be both. The incentives are structurally opposed. We picked tributary. I want to be honest that this is a harder business to be in than pump — the revenue does not grow with the meter. But I think it is the right business to be in, and I think in five years it will be the only business still standing. --- ## VI. A closing thought There's a line in Shewhart that I keep coming back to. He wrote that the quality of a product is characterized by the extent to which it hits a target specification with minimum variation. Not maximum volume. Not maximum spend. Minimum variation around a target that means something. The token economy, as currently constituted, is a system with maximum volume and very little clarity about what the target actually is. Jensen has proposed that the target is consumption itself — that spending tokens *is* productivity. I've tried to argue in this post that this is a category error, that it is Goodhart's law operating at industrial scale, and that Deming would identify it instantly as tampering. The real target is value created. The real measure is variance in value created per unit of compute. The real work is the tributaries — the tributaries built by engineers who notice something specific, crystallize it, deposit it into the shared stream, and let it compound. Real demand at real prices is still an open question. Real value at real prices is not. We know how to build it. We have known since 1924. We just have to stop measuring the wrong thing. ---