How much should a Claude Code content pipeline cost per article?
A Claude content pipeline usually costs cents to a few dollars per article in model fees, and the spread depends on three things: how much context each step rereads, which model each step uses, and whether you cache. Output tokens are the visible cost. Input tokens, reread on every turn, are the real budget.
That last point surprises most people the first time they look at a bill. You think you are paying for the words the model writes. You are mostly paying for the words it reads, again and again, as an agent loops through research, drafting, and checking.
This article is about setting a token budget before you scale a pipeline, not after the first invoice. I will use Anthropic's published API prices from its pricing docs as of October 2026, and I will label every worked number as an illustration built on my own assumptions.
What is a token budget, and why does a pipeline need one?
A token budget is a cap on how many input and output tokens one unit of work may use, such as one article. It turns a vague worry about AI costs into a number you can monitor. Without one, an agent that loops too long or rereads too much can multiply your costs silently.
Anthropic's docs give a handy rule of thumb: one token is roughly four characters, or about 0.75 words, of English. So a 1,600-word article is a bit over 2,000 output tokens. That part is cheap. The research pages, instructions, and earlier turns the model rereads are where the volume hides.
There is one more wrinkle. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. If you built a cost model on an older model, rerun it. The same article can count as more tokens today.
If you want a plain-language refresher on what a token is before going further, I covered the basics in what a token is and what it costs for website content.
Where do the tokens actually go in a content pipeline?
Tokens go to four places: instructions, research material, conversation history, and the draft itself. In an agent loop, the first three are resent on many turns. A long instruction file and a few fetched web pages can easily outweigh the article the pipeline is producing.
Research is the obvious one. Anthropic's web fetch docs estimate an average 10 kB web page at about 2,500 tokens. Fetch four sources and you have added roughly 10,000 tokens of input, before the model writes a single sentence.
Instructions are the sneaky one. A detailed style guide, a list of rules, and a few examples can run to thousands of tokens. That block rides along with every request. If the pipeline makes twenty calls for one article, you pay for those instructions twenty times unless you cache them.
History compounds both. Each turn in a conversation includes what came before. By the end of a long run, the model may be rereading every tool result it saw along the way. That is why a pipeline that feels light per step can be heavy per article.
What do the current Claude prices mean for a real budget?
On Anthropic's pricing page, Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens. Claude Haiku 4.5 costs $1 and $5. Claude Opus 5.5 costs $4 and $20. Those rates, multiplied by your real token counts per article, are your base budget before any discounts.
Here is a worked illustration, using my own assumptions rather than measured data. Say one article's run reads 300,000 input tokens across all its turns and writes 15,000 output tokens, including drafts and checks. On Sonnet 5.5, that is $0.60 of input plus $0.15 of output, or about $0.75 per article.
Now change one assumption. If 80% of that input is served from the prompt cache, the math shifts a lot. Cache reads cost 0.1x the base input price on Sonnet 5.5, so 240,000 cached tokens cost about $0.05, and the remaining 60,000 uncached tokens cost about $0.12. Add the same $0.15 of output and you land near $0.32.
The lesson is not the exact figure. Your counts will differ. The lesson is that input volume and cache rate move the budget far more than the length of the final article does.
How does prompt caching change the math?
Prompt caching stores a repeated prefix, like your instructions, so later requests read it at a fraction of the price. Anthropic prices a five-minute cache write at 1.25x the base input rate and a cache hit at 0.1x on most models. Its docs say caching pays off after one cache read at the five-minute duration.
For a content pipeline, that is close to free money. Put the stable parts first: the system prompt, the style rules, the examples. Put the changing parts last: the topic, the research, the draft. Then the stable prefix gets cached and reused across every call in the run.
There is also a one-hour cache option, priced at 2x the base input rate for the write. That makes sense when runs are spread out across an hour, such as a pipeline that drafts one article, waits for a review step, and then continues.
The common mistake is putting something that changes, like today's date or a run ID, at the top of the prompt. That one line breaks the cache for everything after it. Keep volatile values at the end.
When should you use the Batch API instead of live calls?
Use the Batch API for steps that do not need an answer right away. Anthropic says batch processing gives a 50% discount on both input and output tokens. Outlines, metadata drafts, internal link suggestions, and overnight rewrites are good fits. Anything a person is waiting on in real time is not.
In practice, I recommend splitting a pipeline into interactive steps and batch steps. Research and drafting often run live, because each step depends on the last. Bulk tasks, like generating excerpt options for fifty existing posts, can wait a few hours and cost half as much.
Anthropic's docs also say the batch discount and prompt caching can be combined. For large, repetitive jobs, stacking both is the single biggest cost lever you have.
Which model should each step of the pipeline use?
Match the model to the difficulty of the step, not to the importance of the article. Anthropic's own cost guidance suggests Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning. Classification and formatting rarely need the most capable model. Judgment calls often do.
In a content pipeline, the cheap steps are things like extracting slugs, checking for banned characters, or tagging a category. The expensive steps are deciding whether a claim is supported, or whether a draft actually answers the question it promises to answer.
A sensible default is to draft and check on a mid-tier model, route simple housekeeping to a smaller model, and reserve the largest model for a final review or for steps where errors are costly. Then measure. If a cheaper model passes your quality checks on a step, keep it there.
What should you monitor once the budget is set?
Monitor tokens per article, cache hit rate, and cost per published piece, not just the monthly total. A rising token count per article usually means context is bloating or an agent is looping. A falling cache hit rate usually means someone changed the prompt order. Both are fixable once you see them.
Log the usage numbers the API returns with every response, keyed to the article being produced. A simple table with input tokens, cached tokens, output tokens, and model per step is enough. Review it weekly while the pipeline is new.
Set an alert threshold too. If any single article run crosses, say, three times your normal token count, stop it and look. Runaway loops are one of the easiest ways for a cheap pipeline to become an expensive one. I wrote about the wider cost picture in what a small automation stack costs per month.
What should you do next?
Pull the usage numbers from one recent article run and split them into instructions, research, history, and output. Move stable instructions to the top of the prompt and turn on caching. Shift non-urgent steps to the Batch API. Then set a per-article token cap with an alert, and review it weekly.
Do this before you scale, not after. A budget set at ten articles a month will catch problems cheaply. The same problems found at a thousand articles a month cost real money to discover.
If you are building a content pipeline with Claude Code and want help structuring it so costs stay predictable, reach out. This is the kind of system I build, and I am happy to look at how yours is set up. For a broader view of the build itself, see how I automate content ops with Claude Code and Airtable.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.