OpenAI just shipped always-on agents. Should your sales team care?
Yes, but not because the agent will run your pipeline next week. OpenAI's dots are agents that keep working between your conversations, with their own cloud computer and browser. For a B2B go-to-market team, that changes which tasks are worth handing to software, and it raises the bar on what you let software touch.
I build go-to-market systems for B2B companies: the outbound, enrichment, scoring, routing, and reporting that sit behind a sales team. Every few months a new tool promises to replace half of that stack. Most of them are a chat window with a nicer coat of paint. Dots are different in one specific way, and that one way is what this article is about.
Below I cover what OpenAI has actually said about dots, where I think they fit in a revenue team, where I would keep them out, and the small test I would run before anyone gives one access to a CRM.
What exactly is a dot, according to OpenAI?
OpenAI's own documentation describes a dot as an always-on agent that keeps work moving across your tools and projects. It runs on GPT-6 Astra, lives in the cloud, and has its own computer and browser. That means it can keep going on a task when your laptop is closed and you are doing something else.
The documentation lists four headline jobs: research, data analysis, preparing documents, and building software. A dot can hand pieces of work to background agents so several things run in parallel. It can reach connected apps, plugins, and your local computer. It draws on the current conversation, relevant ChatGPT memory, and its own saved notes, so it does not start from zero each time.
You can talk to a dot in ChatGPT, Slack, Microsoft Teams, or on a voice call. OpenAI says the messages stay in each channel while the dot keeps one memory across all of them. You can also take control of its cloud browser when you need to step in.
Who can use dots right now?
Access is limited for now. OpenAI's documentation says dots are available on Pro plans for users over 18 outside the EEA, the UK, and Switzerland, and are rolling out to Business Premium and Enterprise. You create a dot on desktop, then you can use it on mobile once it is set up.
That matters for planning. If your team sits on a standard Business seat, you are not testing this tomorrow. If you sell into Europe and your operators are based there, check the regional limits before you build a plan around it. OpenAI's own help pages are the place to confirm current availability, because rollouts like this tend to change week to week.
My advice for most small teams is simple. Let one person with access run a two-week test on low-risk work before anyone else asks for a seat. You will learn more from fourteen days of real use than from any launch thread.
Why does an always-on agent matter more than a smarter chatbot?
Because most go-to-market work is not one question and one answer. It is a chain of small steps spread over days: research an account, wait for a reply, update a record, follow up, report. A chatbot answers when asked. An agent that keeps state and keeps working can carry a chain like that without you restarting it.
Think about how a founder runs outbound today. They pull a list, research twenty accounts, write notes, draft emails, wait, check replies, and update HubSpot or Salesforce. Each step is easy. The problem is the gaps between steps, where things get forgotten. Tools like Clay, Apollo, and HubSpot workflows exist mostly to close those gaps with rules.
An always-on agent tries to close the same gaps with judgment instead of rules. That is powerful and risky in equal measure. Rules are boring but predictable. Judgment is flexible but harder to audit. The teams that get value from dots will be the ones that know which gaps need which.
Where would I actually use a dot in a GTM team?
I would start with work that is research heavy, read only, and easy to check. Account research before a call, a weekly summary of changes at your top target accounts, and a first draft of a proposal from discovery notes all fit. In each case a person reviews the output before it touches a buyer.
Account research is the obvious first job. A dot can browse a prospect's site, read their careers page and recent posts, and write a short brief. Today a rep or an enrichment table does that. The difference is that the dot can come back to the same accounts next week and tell you what is new, using its own notes.
Proposal drafting is the second. Most proposals I see are built from the same three inputs: discovery notes, a pricing sheet, and a past proposal. A dot that can read all three and produce a first draft saves real time. The human still owns scope and price. That line should not move.
Reporting is the third. If your pipeline data lives in a CRM the dot can reach through a connected app, a plain English weekly summary for the founder is a fair job. Just treat the numbers as a draft until you have checked them against the CRM yourself a few times.
Where would I keep a dot away from for now?
Anything that sends to a buyer without a human look, anything that writes to core CRM fields, and anything that touches billing. Those three areas are where one confident mistake costs real money or a real relationship. I would also keep agents out of lead routing until you can see exactly why each lead went where it did.
I have written before about what an automation should never be allowed to do, and the same list applies here with more force. A Zapier step that misfires does the same wrong thing every time, so you spot it fast. An agent that misjudges might do a different wrong thing each time, which is harder to catch.
Sending email is the big one. An agent that drafts outbound is useful. An agent that sends outbound from your domain is a deliverability and brand risk, and the buyer has no idea a model wrote it. Keep a person in the send path until you have weeks of reviewed drafts that you would have sent unchanged.
How do the approval controls change the risk?
They reduce it but do not remove it. OpenAI says that before a dot takes an action that affects your accounts or shares information, a review step decides whether it can proceed, needs your approval, or must hand the step back to you. That is a sensible default. It is still a default you should test, not trust blindly.
Approval prompts only work if the person approving reads them. In a busy week, approvals turn into a reflex click. I have seen the same thing happen with human sign-off in content pipelines, which is why I wrote about where human sign-off belongs in an automated publishing pipeline. The checkpoint has to sit where a mistake is costly, not everywhere.
So decide in advance which actions always need your approval in your setup. Writing to the CRM, sending anything external, and sharing files outside the company should all sit on that list. Then look at what the dot actually asks you about in the first week and compare it with your list.
How would I test a dot before trusting it with pipeline work?
Run it in parallel with your current process for two weeks, on read-only tasks, and score every output. Pick three jobs, such as account briefs, a weekly pipeline summary, and proposal first drafts. Compare each one with what a person or your existing automation produced. Keep a simple log of errors, not impressions.
The log matters more than the tool. For each output, write down whether it was correct, whether it was useful, and how long it took you to check it. If checking a brief takes as long as writing it, the dot is not saving you time yet. If it catches things your team missed, that is a real signal.
This is the same discipline I use for every automation I put into production. The Airtable and WhaleSync system I built for Ajust has delivered over 25,000 cases and saved over 50,000 hours, and the HubSpot flows I built for Kismet Health through Zapier are in production. None of that came from trusting a tool on day one. It came from watching outputs until the failure modes were boring and known.
Will dots replace Clay, Apollo, or your GTM stack?
Not soon, and probably not in the way the hype suggests. Data tools like Clay and Apollo give you structured records at volume. An agent gives you judgment on a small number of tasks. Most teams will end up using an agent on top of their stack, not instead of it.
Here is how I see the split. Rules-based systems are best for volume and repeatability: enrich ten thousand rows, route every demo request, score every lead the same way. Agents are best for the messy middle: interpret a reply, research the odd account, write the brief nobody has time for. If you have read my take on which buying signals are noise in outbound, the same logic holds. Volume without judgment creates noise, and judgment without structure does not scale.
The useful question is not which one wins. It is where the handoff sits between your rules and your agent, and who checks the work at that handoff.
What should you do next?
If you have access, pick one read-only job and run a two-week parallel test with a written error log. If you do not have access yet, use the time to write down which actions in your stack need human approval. That list will be useful whichever agent you end up trusting with your pipeline.
The teams that do well with tools like this are rarely the ones that adopt first. They are the ones that already know where their process breaks, so they can point the agent at the gaps and keep it away from the cliffs. If you want help mapping that out for your own go-to-market stack, reach out. I am happy to talk it through.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.