Industry News

OpenAI Shipped Three GPT-6 Models in One Month. What Should You Actually Do?

Written by
Pravin Kumar
Published on
Sep 29, 2026

Why did OpenAI ship three GPT-6 models in one month?

Because different jobs need different economics. In September 2026 OpenAI released GPT-6 Astra on the 3rd, then GPT-6 Sol and GPT-6 Luna on the 22nd. Astra is positioned for the hardest work. Sol and Luna are reasoning models priced twenty times apart from each other.

I read release notes the way I read a pricing page, because that is what they are. A model release tells you what a vendor thinks the work is worth. When three models land inside four weeks, the interesting question is not which one is best. It is which one belongs in which part of your stack.

Most of the advice you will read this week will tell you to upgrade. That is the wrong frame. The right frame is routing: deciding, job by job, which model earns the token spend.

What exactly shipped in September 2026?

Three things, on OpenAI's published changelog. On 3 September, GPT-6 Astra, described by OpenAI as its most capable model, built for the hardest end-to-end work. On 22 September, GPT-6 Sol and GPT-6 Luna, both reasoning models that accept text and image inputs and generate text.

Sol and Luna are reachable through the Responses API and the Chat Completions API. Astra carries documented limitations that matter if you are wiring it into something. OpenAI's notes state that Astra does not support a reasoning effort setting of none, does not accept custom temperature or top_p values, and does not return logprobs. Tool calling with Astra requires the Responses API.

Those limitations are not complaints. They are design decisions, and they tell you Astra is meant to be pointed at hard problems and left to think, not tuned into a cheap high volume worker.

Why does the price gap between Sol and Luna matter so much?

Because it is exactly twenty times, in both directions. OpenAI lists GPT-6 Sol at two dollars per million input tokens and ten dollars per million output tokens. GPT-6 Luna is listed at ten cents per million input tokens and fifty cents per million output tokens. Cached input is cheaper on both.

Divide those numbers and you get the same ratio twice. Two divided by one tenth is twenty. Ten divided by one half is twenty. That is not a rounding artefact. It is a deliberate pricing ladder, and it is the single most useful fact in the whole release.

Here is why it changes decisions. If a job runs a thousand times a month, the difference between Sol and Luna is the difference between a line item you notice and one you do not. If a job runs twice a month and the output goes in front of a buyer, the difference is rounding error. Price only becomes an argument at volume, and most marketing teams have no idea which of their jobs are high volume until they look.

Does a cheaper model mean worse marketing output?

Not automatically, and that assumption costs people real money. Price tracks the cost of serving a model, not the quality of any one answer on your specific task. The only honest way to know whether the cheaper model is good enough for a job is to run that job through both and read the results yourself.

I have written more than three hundred and fifty articles on AI answer engines, schema and E-E-A-T, and the pattern I keep seeing is that people pick a model once, wire it everywhere, and never revisit it. Then a release like this lands and they upgrade everything at once, which is the same mistake in the other direction.

A classification job that sorts inbound form submissions into three buckets does not need the most capable model available. A piece of positioning copy that a founder will argue with their board about probably does. The work is deciding which of your jobs is which, and that decision has more to do with consequences than with benchmarks. I went through this reasoning in more detail when I wrote about how to choose between a large model and a small one for a marketing job.

How should you decide which model does which job?

Sort your jobs by two questions: how often does this run, and what happens if it is wrong? High frequency and low consequence goes to the cheap model. Low frequency and high consequence goes to the capable one. The jobs that are both frequent and consequential are the ones that need a human in the loop, whatever you spend.

That grid is simple enough to draw on paper, and it survives every model release, which is the point. The names change. The question of what a mistake costs you does not.

In practice I find three or four jobs in most marketing stacks that have been quietly running on an expensive model for no reason. Tagging. Summarising. Extracting a company name from a form. These are the jobs where a twenty times price gap is real money, and they are almost always the jobs nobody has looked at since the day they were built.

The reverse case is rarer but worse. Someone routes their outbound copy or their pricing page draft through the cheapest thing available because a spreadsheet told them to, and the saving is thirty dollars a month against a page that is supposed to close business.

What does a model release month tell you about your own cost baseline?

That you probably do not have one. If you cannot say what your AI spend was last month, broken down by job rather than by vendor, then a twenty times price difference is not actionable information. It is trivia.

Before you change a single model name in a config file, get the baseline. Which jobs run, how often, and what does each one cost. I keep this in a simple table and check it monthly, because automation cost has a way of growing quietly while everyone is looking at output quality. I wrote up the version of this I actually use in what a small automation stack costs to run each month.

Once the baseline exists, a release like September 2026 becomes a ten minute exercise instead of a project. You look at your two or three highest volume jobs, you test the cheaper model on them, and you either move them or you do not.

Should you rewrite your prompts every time a model ships?

No, and if you feel you have to, that is a signal about your prompts rather than about the model. A prompt that only works on one specific model version is a liability you will pay for on every release. Write instructions that describe the job clearly enough that a competent stranger could do it.

This is the discipline that saves the most time over a year. Model names churn. Three arrived in September alone. If each one costs you a week of prompt archaeology, you will spend a meaningful fraction of the year on maintenance that produces nothing a customer can see.

I have written about this at length in writing prompts that survive a model upgrade, and the short version is this: put the stable stuff in the prompt and the unstable stuff in the config. Then a release is a one line change and a test run, not a rewrite.

What does the image encoding fix tell you about release day?

That release day is not the same as production ready. On 25 September, three days after shipping Sol and Luna, OpenAI published a fix for a bug in image encoding that had degraded image understanding in both models. The vendor found it, disclosed it in the changelog, and corrected it.

I read that as a good sign about the vendor and a useful warning about timing. Reasoning models that accept image input are genuinely useful for marketing work, and an encoding bug in the first days is exactly the kind of thing that would have shown up as inexplicably poor results if you had migrated everything on day one.

My own rule is to let a new model sit for a couple of weeks on anything that touches a client deliverable, while testing it freely on work only I will see. That is not caution for its own sake. It is the recognition that the first week of a release is when the unknown problems are still unknown, and that the changelog is where they surface.

What should you do next?

Pick your three highest volume AI jobs and find out what each one costs you this month. Run the two cheapest suitable models against the same real inputs and read the outputs yourself. Move the jobs where quality holds. Leave the ones where it does not, and write down why.

That is the whole exercise, and it takes an afternoon. It will also outlast GPT-6, because the method is about your work rather than about a model name. The next release will come, and the month after that, and the teams that handle it well will be the ones who already know which of their jobs are expensive and which are consequential.

If you want a second pair of eyes on where your AI spend is going, or on whether your content is actually being read and cited by these engines, reach out. I work on this every day and I am happy to tell you what I would change first. Let's chat.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.