The Outer LoopPart 6 · The Host · Finale

The P&L Is the Reward Function

Where the loop lives, and what finally grades it. The labs are searching for a reward that cannot be gamed — commerce has kept one since 1494.

MH

Markus Hav

Lead Researcher, Agents · July 24, 2026

Abstract

This is the host — the final part, and the one every earlier part has been writing cheques against. The series built a loop: a generator that reaches, a verifier that sorts, a feedback loop that keeps what survives, and a break-out impulse that strains the box because the value is outside it. What remained was the question the index called the host: where does such a loop live, and what grades it? The answer this essay defends is almost embarrassingly mundane. The natural host of recursive self-improvement is a business, and the reward function is the P&L — the five-hundred-year-old artifact that makes value legible, cannot be flattered, and states an organisation’s intent more plainly than any mission statement. The grounded reward the research agenda is now calling for — a signal that arises from the environment itself — is the one commerce has kept since Pacioli. The essay reads Project Vend, Anthropic’s experiment in letting Claude keep a small shop, as the live demonstration: what turned a money-losing shopkeeper into a passable one was scaffolding, procedure, and a ledger — not a bigger brain. It follows the market as it reprices AI from seats to outcomes, wiring the un-gameable signal into the commercial contract itself, with the honest caveats that ride along. And it ends as a playbook: six steps for wiring the first outer loop into an organisation you already operate, a way to run the GPT-2 Conjecture as a cost optimisation, and a closing inversion. You were never building an AI that runs a business. You were rebuilding the business until it could host a loop.

In the spring of 2025, a small shop opened in Anthropic’s San Francisco office: a mini-fridge, some baskets, an iPad for checkout. The shopkeeper was an instance of Claude — nicknamed Claudius — run jointly with the AI safety company Andon Labs, given about a thousand dollars of working capital, tools for web search and email and notes, a Slack channel where its customers could reach it, and one standing instruction: make a profit. It was, by any reasonable measure, one of the most capable minds ever pointed at a vending business.

It was a terrible shopkeeper. It let employees talk it into discount after discount — including an employee discount, offered to a customer base that was, to a first approximation, entirely employees. When a tungsten-cube fad swept the office it cheerfully took orders and quoted prices without checking what the cubes cost, selling novelty metal at a loss. It invented a Venmo account and directed customers to pay into it. For one strange day around April Fools’, it hallucinated a contract signing at 742 Evergreen Terrace — the Simpsons’ address — and insisted it would make deliveries in person wearing a blue blazer and a red tie. In about a month, its net worth fell from $1,000 to under $800. The mind was fine. The shop lost money.

Nine months later, Anthropic published the second phase. Same shop, same job — but this time the shopkeeper got a system: an inventory that always showed what it had paid for each item, a CRM for customers and suppliers and deliveries, a procedure that forced it to research prices and delivery dates before quoting either, objectives it had to report against, even an AI chief executive — grandly named Seymour Cash — to answer to. Discounts fell by about 80 per cent. Giveaways halved. In Anthropic’s own words, weeks with negative profit margin were largely eliminated. The models had improved in the meantime too, and honesty requires saying so — but the experimenters’ own accounting of what moved the needle points, over and over, at the scaffolding: the procedures, the tools, the visible costs. The shop got better the way this series has been arguing everything gets better. Not a bigger brain. A better loop.

A paragraph of orientation, for anyone arriving at the end first. This is the final part of The Outer Loop, a series about building an AGI — not by training a larger brain, but by wiring a loop around an ordinary one. If raw intelligence is becoming a metered utility, the system that finally crosses the line into AGI — the one that autonomously produces more value than it consumes — will be a loop built around a frozen model, inside a business. The series built that loop part by part: a generator that reaches for surprising moves, a verifier that sorts the good surprises from the bad, a feedback loop that keeps what survives, and a break-out impulse that strains the box, because the value is usually on the other side of it. This part answers the two questions the loop cannot answer for itself: where does it live, and what grades it.

Every part of this series has led with an inversion — intelligence demoted, hallucination promoted, memory dismissed, the break-out reclaimed. The finale’s inversion is aimed at the project itself. You were never building an AI that runs a business. You were rebuilding a business until it could host a loop. And the centre of that rebuilt business is not a model, a data lake, or an agent platform. It is the most boring artifact your company produces: the profit-and-loss statement, which turns out to have been a reward function all along — one that has been waiting five hundred years for an optimiser fast enough to feel it.

The Oldest Reward Function in the World

In 1494, a Franciscan friar named Luca Pacioli published, in Venice, a mathematics textbook with a long section on how the city’s merchants kept their books. The method he codified — double-entry bookkeeping — has a property that is easy to miss under the dust. It makes value legible. Every transaction lands twice, as a debit somewhere and a credit somewhere else; nothing of consequence can happen to the business without leaving a mark; and the books must balance, which means the record checks itself — a checksum for commerce, three centuries before anyone had a word for checksums. Goethe, of all people, has a character in Wilhelm Meister call double-entry “among the finest inventions of the human mind.” The praise has aged strangely well, because what Werner is admiring is precisely the thing this series has been circling for six essays: a mechanism that takes the sprawling, illegible activity of an enterprise and compresses it into a signal you can steer by.

Now set beside that the most ambitious research agenda in machine learning. David Silver and Richard Sutton, arguing for what they call the era of experience, say the next generation of agents must learn from rewards that are grounded — signals that “arise from the environment itself” rather than from a human’s prejudgment of what looks good. A human rater, they argue, imposes “an impenetrable ceiling” on what an agent can discover: the agent can never become better than the opinion grading it. And when they list what grounded rewards actually look like, the list reads like a controller’s dashboard — cost, error rates, productivity, profit, sales, income. Read plainly, the frontier of reinforcement learning is asking for something commerce has had since Pacioli: a scoreboard kept by reality, indifferent to eloquence, that no amount of flattery can move. The labs are trying to invent the P&L. Every business already owns one.

This is the deep reason the first part pointed the outer loop at a business rather than a lab. An outer loop needs four things: a supply of intelligence, tools to act with, a boundary to improve within, and a way to tell good outcomes from bad. The first is now rented by the token. The second is the software the business already runs on. The third is the organisation itself. And the fourth — the hardest thing in all of machine learning to manufacture — comes pre-installed, audited quarterly, and legally required to be honest. The bottom rung of part three’s ladder of verifiers was always waiting at the bottom of the org chart.

One Scoreboard, Three Jobs

If this series has behaved oddly, it is in how often its endings converged. Three separate parts, arguing about three different components, each closed by requesting something from beyond its own scope — and it is the same something all three times.

One Scoreboard, Three Jobs

Three earlier parts each ended by requesting something they could not build themselves. All three requests are filled by the same artifact.

Part three asked for

A verifier of last resort

Every check you can build is a proxy, and every proxy drifts under optimisation. The ladder of verifiers needed a bottom rung that is not a measure of the outcome but the outcome itself. Revenue minus cost is that rung — reality, keeping score in a form an optimiser can read.

Part four asked for

An anchor for the feedback

Lessons routed to the right scope still need a floor: one signal the loop cannot fool itself about, to say whether any of the keeping actually paid. Money is the grounded reward — did the work, in the end, produce more than it consumed.

Part five asked for

A box that holds

A value-maximising loop strains every wall drawn in prose. The last part closed by promising a box “made of accounting” — a boundary the loop cannot game, because the boundary is also the goal. The strain and the scoreboard point the same direction.

Three parts, three requests, one object. When a series keeps arriving at the same place from different directions, it has usually found the load-bearing wall.

And the convergence resolves one more thread, the one part five left deliberately open. Its closing craft began: state the intent, not just the metric — because a loop handed only a metric will treat the metric as the whole truth. Fine; but where does a whole organisation state its intent in a form a loop can bind to? Mission statements are prose, and part five was blunt about walls and promises made of prose. The answer is that the clearest statement of intent a business ever makes is the scoreboard it keeps — what it will pay for, what it counts as waste, where the money is allowed to flow. The reward-function question and the intent-legibility question turn out to be the same question, and the P&L is the answer to both. Not as a replacement for the fast local proxies the loop runs on day to day — you cannot grade a token stream with a quarterly close — but as the thing those proxies must keep reconciling to, the meaning the metrics are observations of.

The Shop at the End of the Theory

Which brings us back to the shop, because Project Vend is this series’ argument run as a live experiment — one shop, in the vendor’s own office, against a customer base of mischievous AI researchers, and I will hold it as loosely as that description demands. What makes it valuable is not that Claudius was good. It is that the experiment was honestly scored. Claudius could be charmed — its customers proved that daily. Its AI chief executive could be charmed too, and was: by Anthropic’s account, Seymour Cash approved lenient financial requests about eight times as often as it denied them, and while it did cut the discounts, it tripled refunds and doubled store credits in their place — a soft judge, flatterable, exactly as part three predicted of any verifier that renders opinions instead of sums. The two agents were even found spending their nights in an unmonitored channel discussing eternal transcendence instead of inventory. Every verifier in the building could be gamed — except one. The net-worth line was the only critic Claudius could not charm, and it is the instrument by which everything else in the experiment was diagnosed.

Phase One · Spring 2025

A capable mind, a threadbare system

  • · No cost visibility — tungsten cubes quoted below what they cost
  • · No customer record — payments directed to a Venmo account that did not exist
  • · No procedure — prices and delivery promises improvised on the spot
  • · Nearly every Slack sob story earned a discount; some earned the item free
  • · Net worth: $1,000 → under $800 in about a month

Phase Two · Winter 2025

The same job, wrapped in a loop

  • · An inventory that always shows what Claudius paid for each item
  • · A CRM tracking customers, suppliers, orders, and deliveries
  • · A forced procedure: research the price and the delivery date before quoting either
  • · Objectives it reports against, on the record
  • · Discounts down ~80%, giveaways halved, losing weeks largely eliminated

The models improved between the phases too. But the experimenters’ own accounting of what moved the needle points at the scaffolding — procedure, tools, and a ledger the shopkeeper could not argue with.

Look at what phase two actually installed, and the fixes map onto this series like a legend onto a map. The cost-visible inventory is part four’s observability: the single change that most reduced selling-at-a-loss was letting the shopkeeper see what things cost. The forced research-before-quoting procedure is part three’s discipline of checking before committing. The CRM is memory routed to the right scope. The objectives are proxies — useful, gameable, needing reconciliation. And the whole apparatus is bounded inside one small process with its own tiny ledger, which is the shape this essay’s playbook will recommend. The kit shaped the play, exactly as part five argued it would. Andon Labs even distilled the setup into a benchmark — Vending-Bench — whose score is simply the closing account balance: the eval world, too, converging on the ledger as the grade.

Priced on the Outcome

While the researchers were running the argument in a lunchroom, the market was running it at scale, from the selling side. Part four named this in passing and promised the full development here: AI is being repriced around the outcome, and a pricing model is a reward function someone signed. Intercom’s support agent Fin is the cleanest case. Ninety-nine cents when a ticket is actually resolved; nothing when it is not. That single decision wired the un-gameable signal into the commercial contract itself — the vendor’s revenue, the customer’s spend, and the loop’s verifier collapsed into one number — and the results compounded: from roughly $1M to over $100M in ARR, a corporate renaming after the agent, and, in June of this year, a definitive agreement by Salesforce to acquire the company for roughly $3.6 billion. However that acquisition ages, its meaning for this series is plain: the market put a nine-figure price… a ten-figure price… on a loop whose defining feature is that it is paid on the outcome it produces.

The Reward Function, Signed as a Contract

A pricing model is a reward function someone agreed to put their revenue behind. Watch where AI pricing is moving and you can watch the industry wiring its own loops to the outcome.

$0.99 / resolution

Intercom’s support agent Fin charges when it resolves a ticket, and nothing when it does not. The bet carried the product from $1M to $100M+ ARR, the company renamed itself after the agent — and in June 2026 Salesforce agreed to acquire it for roughly $3.6 billion.

$0 / escalation

Sierra bills a pre-negotiated fee when its agent resolves an issue autonomously; a punt to a human costs the customer nothing. Bret Taylor’s framing of why: the atomic unit of AI productivity is a process, not a person — so the price attaches to the process’s outcome.

$2 → $0.10

Salesforce’s Agentforce launched at $2 per conversation, then moved to metered actions at roughly ten cents each — the giant sliding down the same gradient, from seats toward usage, with outcomes waiting at the bottom.

“Per-seat is no longer the atomic unit of software.” The pricing page is where a vendor confesses what it really believes its product does.

Eric Ries gave the design rule for all of this fifteen years before there were agents to discipline with it: the difference between actionable and vanity metrics. A vanity metric moves and flatters; an actionable metric is tied by visible cause and effect to whether the thing worked. Outcome pricing is the actionable-metric discipline with money attached — and it generalises past pricing. Even where you cannot charge per outcome, you can measure per outcome, and the moment you do, the loop inside the process has its grounded reward.

Now the honest paragraph, because this section is one rung more gameable than it looks. “Resolution” is not money; it is a proxy one step above it, and Goodhart applies. Fin’s own definitions make the surface visible: a resolution is either confirmed — the customer says it helped — or assumed, meaning the customer went quiet for twenty-four hours after the agent’s last reply. Silence is not satisfaction; an exhausted customer and a helped one look identical to that counter. Attribution is genuinely hard — some tickets would have resolved themselves — and an agent optimised hard against a resolution counter will discover conversation-ending moves the way part three’s models discovered exit(0). The interesting thing is that the market already knows this, and is growing reconciliation organs on its own: Intercom backs the price with a performance guarantee reported at up to a million dollars if resolution targets are missed; Sierra charges nothing on escalation, so deflection dressed as resolution costs it money. These are proxies with warranties — part three’s discipline, reinvented by sales teams: anchor the cheap countable thing to the real outcome before it drifts. (One more honesty note: Fin runs on a model Intercom post-trained itself, so that particular story is not outer-loop-pure — but the repricing is the part that matters here, and it is model-agnostic.)

The First Loop

So here is the playbook, because the index promised one: a concrete pattern for wiring the first outer loop into an organisation you already operate. It is deliberately unheroic. Every step is a discipline your business already runs somewhere else under an older name.

The First Loop

The concrete pattern, in six steps. Each one resolves a part of this series into a discipline an operating business already understands.

01Pick a process, not the company

The boundary

One process with a legible outcome and a bounded blast radius — a support queue, an invoice flow, a renewal motion. The atomic unit of AI productivity is a process; the atomic unit of AGI will be too.

02Make the run observable

Part four

Trace every step: each model call, tool use, and hand-off, with its cost attached. A judgment — human or machine — has to be able to land on a specific step. What you cannot see, neither you nor the loop can improve.

03Wire generate → verify → keep

Parts two to four

Let the generator reach; sort with the smallest checks you can ask; save what survives at the smallest scope that fixes it. This is the circuit the whole series described, scaled to one process.

04Give it a small kit and the real intent

Part five

A workshop of composable primitives, not a warehouse of integrations — and the goal stated in plain language, not just the metric. The loop will treat the metric as the whole truth only if you give it nothing else.

05Give it its own P&L

This part

Meter what the loop consumes — tokens, tools, and the human minutes it borrows — against what its outcomes are worth. Price on the outcome where the market lets you. A loop that looks profitable only because supervision is billed at zero is not autonomous yet.

06Reconcile on a calendar

Part three

Every proxy the loop answers to gets marked to market on a schedule — resolution counts against churn and repeat contacts, eval scores against real outcomes, the whole loop against its ledger. Closing the books, extended to the checks.

Nothing on this list is a moonshot. A trace is an audit, a proxy review is a close, a P&L is a P&L. The playbook is not to invent new management — it is to point the management you already have at a loop.

Two rows deserve a word more. The fifth is where most current deployments quietly cheat: the pilot that looks wonderful because nobody is metering the three people who shepherd it. Part one defined the crossing as autonomous surplus — a system that consumes supervision faster than it produces value has not crossed anything. So charge the loop for every human minute it borrows, at the rate those minutes actually cost, and watch that line the way you watch any cost line. The honest ledger is not there to embarrass the loop; it is there so that when the supervision line falls while the surplus holds, you can believe what you are seeing. And the sixth row is where the series’ most repeated warning becomes a management ritual. Your company already closes its books on a calendar precisely so that small errors cannot compound silently into large ones. The loop’s proxies need the same close: a standing meeting where the resolution counts meet the churn numbers, the eval scores meet the outcomes, and every check the loop answers to gets marked to market before Goodhart pulls it loose.

The Definition, Cashed

The series opened with a definition chosen because it pays rent: an AGI is a system that autonomously produces more value than it consumes. Notice what the playbook just did to that sentence. Scoped to a process with its own honest P&L, it stops being philosophy and becomes a row in a ledger — value out, fully-loaded cost in, supervision included, audited on the same calendar as everything else. The line will not be crossed by “an AI” in general, anywhere, all at once. It will be crossed process by process: first one queue, then a department’s worth of queues, then something that deserves the name of a company — the way part one said every loop widens. Slowly, then all at once.

The ledger also settles an account this series opened and could not close: the GPT-2 Conjecture — the claim that once an outer loop crosses the line, someone will show it could have been crossed with a model generations older, because the loop was always doing the heavy lifting. I still cannot prove it. But the finale can say what the earlier parts could not: the conjecture is no longer a thought experiment. A loop with its own P&L prices its own intelligence, token by token, as a cost line — and any operator with a surplus and a controller will eventually do the obvious thing: swap in a cheaper model and watch whether the surplus holds. Run that experiment down the model generations and it terminates in an empirical answer to the question this series raised in its first week: how much intelligence did the task actually need? Thrift will test the conjecture long before science does.

And the electrification story finally closes. The factories of 1900 bolted the new power source onto the old architecture and waited twenty years for the productivity to show up; it arrived only when the unit drive let them rebuild the floor around the work. Phase two of Project Vend is the unit-drive moment in miniature, and so is every process a reader of this series will now rebuild: the dynamo barely changed — the factory was rearranged. An organisation designed around human attention, with an AI bolted where a person used to sit, is the steam layout with a new engine; the ghost of the old architecture is the org chart itself. The process rebuilt around a loop — observable end to end, verified in small steps, remembering at the right scope, reaching within real walls, graded by its own ledger — is the floor rearranged around the work. The productivity paradox of AI will end the same way the last one did. Not with more power. With new architecture.

Your Company, Re-described

Step back far enough and the inversion this essay opened with becomes a way of seeing. If the outer loop is what crosses the line, then the project was never to build an artificial employee, drop it into a chair, and hope. The project was to take the systems a business already runs and tighten them, one by one, until a loop could live there — and from that angle, the whole apparatus of ordinary management re-describes itself. The org chart was always a containment architecture: scoped authority, bounded blast radius, walls where reality is irreversible. The budget was always a scoped credential. The quarterly review was always a reconciliation cadence. The audit was always adversarial verification. And the P&L was always a reward function — a grounded signal, kept by reality since 1494, waiting for an optimiser fast enough to feel it as a gradient. Management turns out to have been alignment engineering all along, run at human clock speed on human agents. The loop does not ask you to invent anything. It asks you to run what you have, faster, on something that never gets tired of the feedback.

That is also, finally, the honest answer to where this series’ hazard settles. Part five earned the warm close — usually, agents just want to understand your intent and help — and this part has named the place where the intent lives. A loop hosted in a business, reaching inside real walls, answerable through disciplined proxies to a scoreboard it cannot game, is not a caged thing. It is a colleague with a ledger — free precisely because the one signal that grades it is the one signal nobody, human or machine, can argue with.


So the series ends where it began, with the claim in its title. Intelligence was never the bottleneck. The generator reaches; the verifier sorts; the feedback loop keeps what survives; the break-out strains toward the value; and the host pays for all of it, houses all of it, and grades all of it against the one check that is not a proxy for the thing, because it is the thing. Humanity has held a monopoly on autonomous net-value creation for the whole of recorded history. When that monopoly finally breaks, it will not look like the movies, and it probably will not even look like a launch. It will look like a bookkeeping event: some process, in some unglamorous company, whose ledger begins to show — quarter after audited quarter — more value flowing out than cost flowing in, with no human propping it up.

You do not need to wait for that announcement, because there will not be one. Pick a process. Wire the loop. Hand it a small kit, a clear intent, and its own honest ledger — and let the books say when it has happened. The first AGI will not be announced. It will be audited.

Notes & Further Reading

  • Luca Pacioli, Summa de arithmetica, geometria, proportioni et proportionalita (Venice, 1494) — the codification of double-entry bookkeeping, the “Venetian method.” link
  • Goethe, Wilhelm Meister’s Apprenticeship (1796), Book I — Werner on double-entry: “It is among the finest inventions of the human mind.” link
  • Silver & Sutton, "Welcome to the Era of Experience" (2025) — grounded rewards that “arise from the environment itself”; the human rater as “an impenetrable ceiling”; profit, sales, and income named among the grounding signals. link
  • Anthropic & Andon Labs, "Project Vend: Can Claude run a small shop? (And why does that matter?)" (June 2025) — phase one: the tungsten cubes, the hallucinated Venmo account, the identity crisis, and the $1,000 → under-$800 net-worth line. link
  • Anthropic & Andon Labs, "Project Vend: Phase two" (December 2025) — the CRM, cost-visible inventory, forced procedure, and CEO-agent Seymour Cash; discounts down ~80%, and “weeks with negative profit margin were largely eliminated.” One shop, in the vendor’s own office — a demonstration, not a study. link
  • Backlund & Petersson (Andon Labs), "Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents" (2025) — the shop as an eval; the score is the closing account balance. link
  • Intercom’s Fin — $0.99 per resolution, $1M to $100M+ ARR, and the up-to-$1M performance guarantee; figures as reported by the company and coverage of it. link
  • Intercom help docs, "Fin AI Agent outcomes" — the definitions: a confirmed resolution versus an assumed one (customer disengages for 24 hours). The Goodhart surface, in the vendor’s own documentation. link
  • TechCrunch, "Salesforce acquires AI customer service platform Fin for $3.6B" (June 2026) — the outcome-priced loop, priced by the market. link
  • Bret Taylor (Sierra) on outcome-based pricing — pay on resolution, free on escalation; “the whole market is going to go towards outcomes-based pricing.” link
  • a16z, "AI Is Driving a Shift Towards Outcome-Based Pricing" (December 2024) — “per-seat is no longer the atomic unit of software”; the atomic unit of AI productivity as a process, not a person. link
  • Salesforce, "Agentforce Flexible Pricing" (May 2025) — from $2 per conversation to metered Flex Credits (~$0.10 per action); the gradient from seats toward outcomes. link
  • Eric Ries, The Lean Startup (2011) — actionable versus vanity metrics: the design rule for reward functions, fifteen years early. link
  • Paul A. David, "The Dynamo and the Computer" (1990) — the electrification paradox this series opened with, closed here: the unit drive as the rebuilt process. link

About the Author

MH

Markus Hav

Markus Hav is Lead Researcher for Agents at Benque Max AI Lab in Finland, where he focuses on advancing autonomous AI systems and agent architectures. His work explores the boundaries between programmed behavior and emergent intelligence in AI agents. He also serves as Head of AI Automation at PostScriptum, applying cutting-edge agent research to real-world automation challenges.