AI

How Do I Decide Whether an AI Draft Is Good Enough to Publish?

Written by
Pravin Kumar
Published on
Sep 22, 2026

How do I decide whether an AI draft is good enough to publish?

I score it on five things: truth, specificity, stance, structure, and whether anyone else could have written it. Truth is a gate, so a failure there ends the conversation. The other four are judgments, and a draft that passes truth while failing two of the rest gets rewritten rather than edited.

Most advice about AI content stops at fact-checking, which is necessary and nowhere near sufficient. A draft can be entirely accurate and still be unpublishable, because accuracy is not the same as usefulness and neither is the same as having something to say.

This is the rubric I actually apply, in the order I apply it, and the rough thresholds I use to decide between editing and starting again.

Why is fact-checking only the first gate?

Because it is the only test with a binary answer, so it should run first and cheaply. Every checkable claim either traces to a primary source or it does not. Anything that does not gets cut, and if cutting it removes the point of the piece, the piece dies there rather than after two hours of editing.

Running truth first also protects you from a specific trap. A well-written draft is persuasive, and the better it reads the less inclined anyone is to check it. Doing the checking before the reading keeps your judgment honest, because you are evaluating claims rather than prose.

What makes this manageable is treating claims as a list rather than as a feeling. I go through the draft and pull out every sentence that asserts something checkable, then verify each one. The mechanics of that are their own discipline, which I set out in fact-checking AI-written content before it reaches a client site.

What does Google actually say about AI-generated content?

Google published guidance on this on February 8, 2023, and the position has been quoted loosely ever since. What it actually says is that Google's ranking systems aim to reward original, high-quality content that demonstrates E-E-A-T, meaning expertise, experience, authoritativeness, and trustworthiness, and that the focus is on the quality of content rather than how it is produced.

Google is equally clear on the other side. Using automation, including AI, to generate content with the primary purpose of manipulating ranking in search results violates its spam policies, and Google names its SpamBrain system as part of the spam-fighting effort. The guidance also says plainly that not all use of automation is spam, giving sports scores, weather forecasts and transcripts as long-standing examples of helpful automated content.

The useful takeaway for a rubric is that the production method is not the question. Google frames its helpful content system as ensuring that searchers get content created primarily for people rather than for search ranking purposes. That is a test about intent and outcome, and it is exactly what the remaining four dimensions are trying to measure.

How do you test a draft for specificity?

Count the sentences that could be deleted without losing information. In a weak AI draft that number is enormous, because the model fills space with restatements of the heading and with general observations that nobody would dispute and nobody needed to read.

The faster version of the test is to look for the nouns. A specific draft names tools, standards, numbers, versions, and situations. A vague draft talks about solutions, strategies, approaches, and best practices. If you can swap the topic of the article for a different topic and most sentences still work, there is no specificity in them.

I hold this test strictly because it is where AI drafts fail most predictably and where the failure is hardest to spot on a quick read. Generic writing is smooth. Smoothness reads as competence, and competence reads as quality, and none of that survives a reader who was looking for an actual answer.

How do you test whether a draft has a position?

Ask what the piece would look like if you argued the opposite. If the opposite is obviously absurd, the draft has no position, it has a consensus. If the opposite is a thing a reasonable person believes, the draft is taking a side, which is the minimum for being worth reading.

Models default to balance, because balance is safe and because it reflects the average of what has been written. The average of everything written about a contested question is a shrug. Publishing the shrug adds nothing to a topic that already has ten thousand shrugs on it.

The fix is not to manufacture a contrarian take. It is to decide what you actually think before drafting, and to treat any sentence that hedges the position as a defect. A draft that will not commit is usually a sign that the person commissioning it has not committed either.

What does the "anyone could have written this" test catch?

It catches the absence of a source of authority. If nothing in the piece depends on who wrote it, then the piece has no reason to exist on your site rather than on anyone else's, and no reason to be attributed to a person at all.

The test is simple to run: remove the byline and the branding, and ask whether a reader could guess where this came from. A good piece carries fingerprints, which are usually a specific method, a stated preference, a limit the writer has hit, or a thing they got wrong and changed their mind about.

This is the dimension where a human has to contribute rather than review. A model can organise what you know. It cannot know what you have done, and no amount of editing turns a draft with no experience in it into a draft with experience in it. That material has to be added.

Why do I score structure separately from writing?

Because structure determines whether the piece can be found and reused, and writing determines whether it is pleasant once found. They fail independently. A beautifully written piece with one heading and eleven paragraphs is close to invisible to anything that retrieves passages rather than pages.

The structural questions are mechanical. Does every heading name a real question. Does the paragraph under it answer that question completely without depending on what came before. Could a section be lifted out and still make sense. Those are all checkable in a few minutes.

Structure matters more than it used to, because of how passages get selected and scored before anything gets quoted. I went into the mechanics of that in what a reranker is and why it decides whether your page gets quoted.

What separates an edit from a rewrite?

Failing one dimension is an edit. Failing two is a rewrite. Failing truth is a delete. The reason for a hard rule is that editing a draft you have already read three times feels cheaper than starting again, and it almost never is.

Specificity and stance are the two most expensive to fix by editing, because both require adding material rather than changing it. If a draft is vague and non-committal, the edit is really a rewrite performed one sentence at a time, and it usually produces something with the seams still visible.

Structure and writing are the cheap fixes. Reordering sections, rewriting headings as questions, and tightening sentences are all mechanical work that genuinely does go faster as an edit. That is why I score them last, after the expensive dimensions have already decided the draft's fate.

Does the rubric change when the stakes are higher?

The thresholds move, not the dimensions. For a client site I will not publish a draft that fails anything. For my own writing I will publish something that is a little loose on structure if the thinking is genuinely mine, because I can fix structure later and I cannot manufacture the thinking later.

The one dimension that never moves is truth. A draft with an unverifiable claim does not get published anywhere at any threshold, because the cost of being wrong in public is borne by whoever is named on the page, and there is no version of that cost that is worth the time saved.

The other reason to keep truth absolute is that models fail at it in a very specific way. They produce plausible detail that has no source, which is a different failure from being mistaken, and it is worth understanding on its own terms. I wrote about the mechanism in what causes AI hallucinations in website content.

What should you do next?

Take the last AI-assisted draft you published and score it on the five dimensions honestly. Most people find it passes truth and structure and fails specificity and stance, which tells you exactly where your process needs a human rather than a better prompt.

Then write your own version of the rubric down, with your own thresholds, before you next commission a draft. A rubric written in advance is a standard. A rubric invented while looking at a draft you want to publish is a justification.

If you are running AI-assisted content at any volume and you are not sure where your quality is actually leaking, send me one published piece and I will tell you which dimension I think is failing. Let's chat.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.