The AI Team Member: Enthusiastic, Error-Prone, and Off Every Invoice
Every professional team now has a fifth, unofficial member. In an agency campaign, it works the project end to end, runs up nearly $3,000 in costs every time, and appears on no invoice anywhere. Clients are about to find out.
Nadim A. Massih10 June 2026 · 11 min read
In late 2025, Microsoft gave an AI coding assistant called Claude Code to thousands of its engineers. Five months later it cancelled most of the licences. The tools had become 98 per cent cheaper per use over three years, yet the heaviest users were still costing $500 to $2,000 a month each, because they were using them vastly more than anyone had planned. Advertising agencies now run the same kind of tools on every campaign. Microsoft could read its bill and act on it. Most agencies cannot even find theirs.
The brief lands on a Tuesday morning: a fashion retailer wants the full relaunch. New strategy, a new look, new advertising across every channel, twelve weeks to deliver. The director of the advertising agency builds the team on paper. A strategy lead. A senior designer. A copywriter. A project manager. Four people, their rates, their hours, and a $3,000 production budget with every photo licence accounted for. The proposal goes out at six that evening. The client signs.
By the end of week one a fifth member has joined the team, and it never clocks off. By the first coffee break each morning it has already researched the competition. Through weeks three and four it redrafts the campaign’s big idea fourteen times before the strategy lead settles on one. Twice it invents a competitor that does not exist. Twice the strategy lead catches it. By week nine it is recording test voiceovers for the brand film. On the morning of delivery it is still re-cutting social media videos. And all the while a meter is turning, because AI is not bought once like a camera. It charges as it is used, the way electricity does. Twelve weeks of work. $2,861 in computing costs on that meter. Almost exactly the size of the production budget the director wrote down so carefully.
It is not in the proposal. That missing line is the mistake quietly eating agency profits this decade.
The Meter That Doesn’t Stop
To see why the meter runs so hot, start with the most public AI advert ever made. When Kalshi (a prediction market platform where people bet on real-world events) aired an AI-generated commercial during the NBA Finals in June 2025, the team behind it generated roughly 400 video clips to find 15 it could broadcast. A 3.75 per cent yield. Everything else was money spent discovering what did not work.
That gap between what gets generated and what gets delivered is where the cost lives. Professional teams do far better than 3.75 per cent: with experienced people choosing and discarding at every stage, a team typically generates 1,600 seconds of footage to deliver 400 finished seconds. One second in four survives. The other three are the cost of finding the one. Any budget priced only on delivered seconds ignores everything it took to get there.
The problem is built into how the tools are sold. A photographer’s day rate is fixed. A software licence is fixed. AI compute, the computing power consumed every time the tool generates something, is variable: the price of one usable output depends on how many attempts it took to find the one worth keeping. Nor is this a quirk of video. Lightrun (a software engineering tools company) surveyed engineers in 2026 and found that 43 per cent of AI-generated code changes need human debugging once they reach production, even after passing every automated test.
So the failure rate is not a quality disclaimer. It is a cost driver. Every failed generation is compute already spent. The retry is compute spent again. The revision round is a third pass. A budget that counts only the final output is paying for all three without saying so. That arithmetic is what made Microsoft blink, as the technology news site The Next Web reported: the licences were cancelled not because the tools failed, but because nobody had decided who owns a meter that never stops.
Here is what that meter looks like on one project.
A ten-person boutique agency. A three-month brand campaign. The proposal totals $75,000, and most of it is people: 60 hours of strategy direction at $450 an hour, 90 hours of senior design at $280, 60 hours of copy at $220, and 30 hours of account and project management at $220. That is $72,000. The last $3,000 is the production budget: stock imagery, music licensing, font licences. Every dollar is itemised, and the client can audit every line.
There is a second production budget on this campaign. It appears nowhere.
It starts with strategy research: 25 agentic sessions through Claude Fable 5 (Anthropic’s 2026 flagship model) to produce three client-ready briefs. Agentic means the AI works through multi-step research on its own, without being prompted at each step. Each session runs two to three hours and costs about $18, and that billing lands as API overage: the Claude Teams plan covers everyday usage, and sessions like these exhaust its allocation early. The 25 sessions come to $450. Copy iteration running past the plan under revision pressure adds $150.
Then the visuals: Midjourney Pro seats for two designers at $60 per user per month across three months, $360, plus $60 in fast-hour top-ups, the extra processing time Midjourney sells when the monthly allocation runs out during crunch weeks. An agency that already has Midjourney on its rate card, the standard price list it bills from, carries only the top-ups. An agency procuring it just for this campaign carries the full $420.
The video stack is project-specific no matter what the agency already subscribes to, because no standing plan absorbs meaningful video generation volume. The production team generates 1,600 seconds of Sora 2 Pro footage to deliver 400 finished seconds of film and social content: $1,120. Runway Gen-4 cuts the social format adaptations for $225. Seedance 2.0 (ByteDance’s video model) for motion content: $44. Higgsfield (an AI image generation platform) renders 500 campaign images at published per-image rates: $117. ElevenLabs voice prototypes, including the redos that always happen in the final week: $30. Deck iteration through Claude Sonnet 4.6 adds $305 across three months.
Total project-specific AI compute: $2,861. The second production budget, line by line. Every dollar of it ran on behalf of this one client. One of the two budgets is itemised and recoverable. The other does not exist on paper.
The standing subscriptions run in parallel: Claude Teams for ten staff at $300 a month, Microsoft Copilot Business for five seats at $90, ChatGPT Plus for three staff at $60. That is $450 a month, $1,350 across the campaign, and it is true overhead. It belongs in the rate card, and the rate card is how it gets recovered. The $2,861 has no such home. It is 3.8 per cent of the whole engagement, consumed for a named client, and invisible to them.
If those overage lines look pessimistic, the industry’s own data says otherwise. Zylo (a firm that tracks company software spending) found in its 2026 SaaS Management Index that 78 per cent of IT leaders reported unexpected charges from consumption-based AI pricing in the past year. Not outliers. The majority. ServiceNow’s chief information officer, Kellie Romack, disclosed that the company burned through its full-year Anthropic coding budget in the first months of 2026, joining Uber in the same admission. Deadline compression, revision rounds and parallel generation runs are not exceptional events. They are the last two weeks of every project.
Now scale it. Five concurrent projects, the standard operating state for a boutique, put unrecovered AI spend at just under $5,000 a month. On Patient Comet’s modelling, a 25-person agency running sixty projects a year, where a typical brief carries a second video deliverable and roughly twice the research load, lands near $4,800 per project. Across sixty projects, that is $288,000 a year. For an agency billing $7.2 million on a typical 20 per cent target margin, that is one-fifth of the year’s planned profit, absorbed without a single client conversation.
The arithmetic repeats across every sector in that table. The tool names change by industry. The billing gap does not.
The Two Costs You Are Treating as One
A gap this large, repeated this widely, raises an obvious question: why is nobody billing for it? The most common answer is a history lesson, and it comes from the legal profession.
In the 1990s, law firms billed clients for legal database research line by line; Westlaw and LexisNexis charges passed straight through to the client file. Flat-rate contracts spread through the 2000s, and after the 2008 recession clients began refusing the line item altogether: 97 per cent of firms still passed research costs on in 2008, 73 per cent by 2011, and within a few years firms were absorbing roughly half of those costs themselves. The lesson the industry drew: research tools end up in the overhead column, the general running costs no client pays for directly. AI, the argument goes, will follow the same path and disappear from client invoices.
This argument is half right. The half it misses is where the money goes.
Two types of AI cost are running at the same time in every agency, and the industry is treating them as one.
Type one: flat-rate subscriptions. A Claude Teams licence at $30 per user per month. A Microsoft Copilot (Microsoft’s AI assistant) seat at $18 to $30 depending on tier. A ChatGPT Plus plan at $20. These costs do not change based on how complex any individual project is. They are infrastructure, exactly like the legal database contract. They belong in the rate card.
Type two: project-specific consumption. When a strategy team runs a three-hour agentic session for a named client’s brief, that compute was consumed on behalf of that client. When a production team renders eight minutes of video for one campaign, the per-second cost is exact and traceable. One boundary rule keeps the categories clean: a subscription procured for a single campaign is Type two for that campaign; if the agency keeps it afterwards, it moves to the rate card.
Type two is not the legal database after the flat-rate contract. It is the database before it: real, variable, per-use costs that belong in the client budget. Every firm treating Type two as flat overhead has mistaken infrastructure for project cost.
The legal story, read with that split in mind, makes the same point. What clients killed was a fee that looked identical on every invoice whatever the matter: Type one’s profile exactly. What they have never killed is the production expense. The stock library, the location fee, the photography day rate all still sit on invoices, because each one maps to a deliverable the client can point at. Type two has that shape. The 1,600 seconds of generated footage map to this campaign’s film and to nothing else.
And the legal profession’s own story has a sting in its tail. In 2024, Stanford’s RegLab (its regulation and technology research lab) tested the AI research assistants that Westlaw and LexisNexis now sell on top of those classic databases. Westlaw’s AI was accurate on 42 per cent of queries. LexisNexis’s did better at 65 per cent. Both numbers mean the same thing: most AI-generated legal analysis still needs human verification before it reaches a client. The tool whose history is used to argue that AI becomes invisible overhead now ships an AI layer that, in Westlaw’s case, fails more often than it succeeds. That verification runs on someone’s account. The question is whether it appears in the proposal.
Why More AI Means More Judgment, Not Less
There is a reason agencies flinch from that question, and it is not only the fear of looking expensive next to a rival that bills nothing. Underneath sits a sharper fear: if the machine now does this much of the work, perhaps the people are worth less, and putting the machine on the invoice only advertises it.
The numbers say the opposite. The yield rates from the meter section are not evidence of a failing tool. They are a description of the job. The professional running AI is not doing less work. They are reviewing more outputs, making more judgment calls about which ones are worth developing, and closing the gap between what AI produces and what a client can actually use. That is not automation. It is skilled editorial work performed at a volume that was not possible before.
Research confirms the size of that gap. BetterUp Labs (the research arm of the BetterUp employee development platform) published a study with Stanford University in the Harvard Business Review in September 2025: 41 per cent of professionals had received AI output requiring significant rework in the prior month, and each instance took almost two hours to fix. That rework is not optional, and it is not running itself.
Which is why one distinction settles the fear. The fifth presence on that Tuesday brief behaves like a colleague. Commercially, it is equipment. A film production company renting cinema-grade cameras does not absorb the rental into the director of photography’s day rate. The camera is the cost of delivering work to that standard. The director is the value. Neither is optional, and the rental goes in the client budget because it was consumed on behalf of that client’s production.
AI model costs follow the same logic. The strategist is not the AI, and the judgment that filters the AI’s output has never been worth more. So the open question is not whether that judgment deserves its rate. It is who pays for the tools it runs on. And on that question, the industry has already split in two.
The Agencies That Already Bill for AI
In March 2026, Digiday (a publication covering the marketing industry) asked media houses, creative agencies, production shops and consultants how they handle AI compute costs. The answers split cleanly.
On one side, the agencies already billing. Full-service agency Merge meters AI costs and bills them per project on a case-by-case basis. Production company Big Spaceship treats compute like catering or equipment hire; its chief executive, Taryn Crouthers, put it directly: “We’re treating it similar to a production cost.” Lerma/ (a Dallas agency that styles its name with the slash) charges AI token costs as a transparent line item, no different from other project expenses. Tokens are the base unit AI uses to measure work, roughly one word of text per token. Pencil (the AI marketing platform owned by marketing services group Brandtech) charges generation credits: one per image, one per second of video.
On the other side, the agencies absorbing everything. Chris Neff, global chief AI officer at creative agency Anomaly, named the fear behind that choice: “It feels like a money grab.”
Sit with that fear for a moment and three objections emerge, each harder to answer than the last. The first is administrative: nobody wants to send a client an invoice with a $44 line for motion content. Nobody has to. Production expenses have always rolled up; no agency itemises gaffer tape. One AI production line in the proposal, backed by a metered report if the client asks for one, is exactly how the agencies above are handling it.
The second is the tempting middle path: skip the line item and quietly fold 4 per cent into next year’s blended rates. It feels safer. It also misprices every project on the books, because AI intensity varies enormously between briefs, and even the same task can swing thirty-fold in token use from one run to the next. A rate set in January is chasing a curve that multiplied eighteen times in nine months, on the latest engineering data. Burying a cost this variable in a flat rate is not recovering it. It is guessing.
The third objection cuts deepest, because transparency runs both ways. Each of the boutique’s client-ready briefs took roughly eight agentic sessions, call it $150 of compute, against the $450 one afternoon of billed junior research time used to swallow on a brief like this. Show a client those numbers side by side and a hard-nosed procurement team will ask why the research line still exists at all. The honest answer: it survives wherever a human must be accountable for what the brief claims, and it shrinks where one need not be. That repricing is coming either way, and a price move an agency leads reads as efficiency, while one forced by an auditor reads as concealment. The discarded generations invite the same scrutiny: why should a client fund the 1,200 seconds that never made the film? They always have. Photography shoots bill the full day, not the eleven frames that made the campaign. Selects have outtakes. The ratio is the craft.
That scrutiny is coming whether agencies invite it or not, because AI is entering the contracts. Lerma/ already line-items tokens. Pencil bills by generation credit. Once contracts start naming AI, procurement templates learn to ask about it everywhere, and Ruben Schreurs, chief executive of marketing analytics consultancy Ebiquity, has already told Digiday how that ends: clients will audit agency AI usage the way they audit media spending today. “If it’s part of the contract, then yes, we’d audit that.” The agencies that meter first get to answer the procurement question with their own numbers, on their own framing. The agencies that chose silence will have it answered for them, by an auditor, with no goodwill in the room.
“The camera was never free. The agency just stopped sending the bill.”
Cheaper Models, Bigger Bills
Every argument for silence rests, in the end, on one hope: that the cost is melting. Models improve, prices keep falling, so why build billing machinery around a problem that will shrink on its own? The research says the problem is growing instead.
In April 2026, researchers at Stanford and the University of Michigan, a team including the economist Erik Brynjolfsson, published a study of what AI agents actually cost when deployed at scale (arXiv:2604.22750). Two findings carry the argument. Agentic workflows consume up to 1,000 times more tokens than simple back-and-forth chat. And the same task can vary by up to 30 times in token consumption from one run to the next. The meter is not just running. It is unpredictable while it runs.
The usage data tells the same story at company level. Jellyfish (an engineering analytics firm), in a figure reported by TechCrunch in June 2026, measured per-engineer AI token consumption rising 18.6 times in nine months as agentic tools became standard workflow. EY’s analysis of agentic task costs found that a single AI task that cost $0.04 in 2023 costs $1.20 in 2026, a 30-fold rise over the same years in which raw processing prices collapsed. That is why large companies’ AI bills tripled even as efficiency gains were expected to shrink them: volume swallowed every saving. Bryan Catanzaro, Vice President of Applied Deep Learning at Nvidia, put it plainly in April 2026: “For my team, the cost of compute is far beyond the costs of the employees.”
Cheaper models do not mean smaller bills, because every price drop so far has been swallowed by appetite. When image generation got cheap, clients started expecting video. When chat got cheap, agencies started running three-hour agentic research sessions. Goldman Sachs projects 24-fold growth in total token consumption by 2030. The Federal Reserve Bank of Atlanta’s survey, published in May 2026, already puts AI spending in professional and business services at $3,470 per employee per year, around $289 a month. Set that average against the $500 to $2,000 a month of a heavy agentic engineer and the migration is already visible: agencies are moving from the first group to the second, one workflow at a time. The 25-person agency absorbing $288,000 a year runs at roughly three times the Fed’s sector average per head, and the arithmetic explains why: that is $960 a month per professional, sitting between the creative-agency and media-production bands of the table above, the blend a shop lands on when every brief carries a video deliverable, and inside the $500 to $2,000 a month range already measured for heavy agentic users. If agentic adoption spreads at anything like the rates the usage data shows, the average plausibly runs two to four times higher by 2028.
Which brings the story back to Microsoft. Its answer to its own runaway bill was to cancel the licences and move its engineers to a cheaper tool it happens to own. That is the luxury of owning the alternative. An agency holding a signed scope of work has no such exit. The film is due whether the meter ran hot or not. The only open question is who pays for it.
The market has already produced three answers. Cost passthrough: AI compute as a transparent line item in proposals, at cost, the way production expenses appear today. Capability packaging: live at the top of the market since June 2025, when digital services company Globant launched AI Pods at approximately $20,000 per month, bundling 100 million tokens with dedicated human experts who direct the work. And outcome pricing: premium rates that build AI capability into the deliverable itself, which is where the market heads once packaging is normal. These are not distant phases. They are a ladder, and the firms on the upper rungs started climbing from the first one.
The agencies that skip the first rung do not arrive at the third with stronger margins. They arrive having absorbed years of compute silently, having set no commercial precedent with clients, and having trained the market to expect AI as a free inclusion. The conversation starts in the next proposal. Or it starts in the next audit, on someone else’s terms.
Where to Start
If it is going to start in the next proposal, it starts with four moves.
Separate your AI costs by type first. Flat-rate subscriptions are overhead. They go in the rate, not the proposal. Consumption-based production costs are project costs. These include Claude research sessions, image runs above plan allocations, Sora or Runway video seconds, and ElevenLabs voice prototypes. They belong in the brief.
Track what you spend on a project before you decide what to bill. You cannot have a pricing conversation without a number. Run one project end-to-end with full AI cost visibility: what gets consumed, what it costs, what fraction of engagement value it represents. The number will either validate your instinct or surprise you.
Frame it as production transparency, not a surcharge. AI production costs belong in the same section of the proposal as software licences, photography, and motion graphics rendering. It is the cost of the tools your team used to produce the work. That is a different conversation from adding a new fee.
Start with one client you trust. Not a new pitch. Not every project. One existing relationship where you can have an honest conversation about what your team deployed and what it cost. The agencies already doing this are not losing that client. They are having a better commercial conversation with them.
The AI team member was on Tuesday’s brief. It will be on Monday’s. The only thing that changes is the proposal.
Sources & references
~400 Veo 3 generations produced 15 broadcast-usable clips: a 3.75% yield. AdMonsters and NPR, June 2025.
43% of AI-generated code changes require manual debugging in production, even after passing QA. VentureBeat and GlobeNewswire, April 2026.
Westlaw AI: accurate on 42% of queries. Lexis+ AI: 65%. Stanford RegLab, 2024; peer-reviewed in the Journal of Empirical Legal Studies, 2025.
41% of professionals received AI output needing significant rework in the prior month; average rework time 1 hr 56 min per instance. Harvard Business Review, September 2025.
Agentic workflows consume up to 1,000× more tokens than chat; the same task varies up to 30× in token consumption across runs. Stanford / University of Michigan, including Brynjolfsson. April 2026.
Per-engineer AI token consumption rose 18.6× in nine months. Enterprise AI bills tripled despite a 98% drop in raw token prices. TechCrunch and The Next Web, June 2026.
A single agentic task that cost $0.04 in 2023 costs $1.20 in 2026: a 30-fold rise while raw processing prices collapsed.
Per-engineer costs of $500–$2,000/month drove Microsoft to cancel most Claude Code licences five months after rollout. The Next Web, Windows Central, June 2026.
Uber consumed its full annual AI coding budget by April 2026 and introduced spending caps. TechCrunch, June 2026.
“For my team, the cost of compute is far beyond the costs of the employees.” Catanzaro, VP Applied Deep Learning, Nvidia. Fortune and Axios, April 2026.
24-fold growth in total token consumption projected by 2030. Goldman Sachs Research, 2026.
Planned 2026 AI spending in professional and business services: $3,470/employee/year (~$289/month). Fed Atlanta macroblog, 6 May 2026.
Bradley, S. “Agencies grapple with economics of a new marketing currency: the AI token.” Digiday, 3 March 2026. Merge, Big Spaceship, Lerma/, Pencil/Brandtech, Anomaly, Ebiquity named.
~100 million tokens bundled with dedicated human experts at ~$20,000/month. Globant investor release, 5 June 2025.
78% of IT leaders reported unexpected charges from consumption-based AI pricing in the past year. Zylo, January 2026.
CIO Kellie Romack: ServiceNow burned through its full annual Anthropic coding budget in the first months of 2026. Reported June 2026.
Claude Fable 5: $10/M input, $50/M output (launched 9 June 2026). Claude Teams: $30/user/month. Copilot Business: $18–$30/seat. ChatGPT Plus: $20/month. Midjourney Pro: $60/user/month. Sora 2 Pro: $0.70/second. Runway Gen-4: $0.12/second. Seedance 2.0: ~$0.24/second. Higgsfield: ~$0.23/image. ElevenLabs: $0.12–$0.30/1,000 chars.
More articles

Seven Trends Reshaping Design: AI and the Next Three Years of Creative Work
The AI on your team will execute everything and decide nothing. Here are the seven shifts that follow.

AI Anatomy: Most Companies Built the Brain. Almost None Built the Body.
Most companies built a brain and called it a strategy.

The Taste Problem: AI Can Match Anyone’s Output. It Cannot Match Their Judgement.
When everyone rents the same intelligence, judgment becomes the moat.

The Second Customer: Your Product Has Two Users Now. One Cannot Read Your Homepage.
AI-sourced traffic now converts 42% better than human traffic.

Anyone Can Make It Now: When the Mona Lisa Took Eleven Seconds
Google made its film studio free. WPP cut a third of its creative headcount. The tools gap closed.

The Last Human Reader: How AI Became Your First Audience
The pages you publish are no longer primarily read by people.

LLMflation: The AI Gets Cheaper. The Bill Keeps Growing. Neither Is Your Fault.
Microsoft cancelled its Claude Code licences after engineers burned through its entire annual AI budget in weeks.

The Cheap Code Problem: What Snap’s Memo Got Right About AI and Engineering
Snap fired a thousand people because AI writes 65% of its code.

Own Your AI: Why Renting Intelligence Is About to Look Foolish
AI is shifting from a service you subscribe to, to a feature you ship.

Everyone Can Build It Now. Building It Is the Easy Part.
An AI-built social network was fully breached three days after launch.