The 75% figure is evidence, not a forecast
OpenAI says Presence now resolves 75% of inbound issues on its English-language phone support line without human assistance. It also says a Codex-powered improvement loop reduced human handoffs by 15 percentage points in ten days. Those are substantial claims. They are also OpenAI’s figures, measured against its own support benchmarks, on its own products and systems.
A bank, insurer or retailer should not copy 75% into a savings spreadsheet. Computerworld cites analysts who expect many large companies to start lower because they have fragmented legacy systems, uneven knowledge and heavier compliance demands. The harder the organization is for an employee to navigate, the harder it will be for an AI agent too.
There is another wrinkle. A lower handoff rate does not automatically mean better service. It can mean the agent solved more cases. It can also mean the agent held onto cases it should have released. A serious pilot should measure repeat contacts, corrected actions, customer effort and escalation quality beside the automation rate.
Ask what arrives on the human side after the easy requests disappear. Support staff may get fewer password resets and a much denser queue of angry, unusual or legally sensitive cases. That can be a better job if staffing, training and breaks change with it. If they do not, automation has concentrated the hard part and called it efficiency.
The cheap token is not the expensive part
AI buying conversations still drift toward model prices because tokens are easy to compare. Presence points at the larger bill: integration, policy cleanup, access design, testing, monitoring, legal review and ongoing change management.
Pricing is not public. OpenAI says the scope and price are specific to each deployment. That makes a simple return-on-investment comparison difficult, but not impossible. Start with the full cost of one narrow job: vendor fees, integration work, internal subject-matter experts, security review, support-team training, exception handling and maintenance after a policy changes.
Then compare that total with an outcome customers can feel. Did billing disputes end faster? Did people repeat themselves less often after transfer? Did the human queue get calmer, or merely harder? Did the old tool, report or contractor expense actually go away? If the deployment adds a permanent AI-review team while preserving every old process, the automation may be technically impressive and financially decorative.
The managed model also creates a dependency question. OpenAI says exact models, channels, capacity, data handling and service commitments are defined per deployment. Buyers should know which parts of their process map, tests, policy definitions and performance history remain portable if they change vendors later. Engineers attached to the product are useful. Engineers attached forever are a business model.
Ivy wants the implementation budget named. Mina wants the hard cases protected.
Ivy Chen sees the managed deployment as a useful honesty test for buyers. If the project needs support leads, security staff, legal review and system owners, put their hours in the budget before approving it. Calling that work “enablement” does not make it free. Her first pilot would stay with one case type whose owner, baseline and stop condition already exist.
Mina Torres starts on the other side of the phone. If the agent handles routine requests, the person who takes over will hear more of the calls where the policy failed, the account is tangled or the customer is already exhausted. The handoff has to carry the story forward, and the human team needs enough time to use judgment rather than race a new productivity target.
Those views are not anti-automation. They are what serious adoption looks like after the demo. Count the labor honestly before launch. Protect the people who inherit the exceptions after launch.
What to ask before buying an enterprise AI agent
Begin with one repeated job, not “customer service.” Name where it starts, what a good ending looks like, which systems it touches and which cases should go straight to a person. If the job cannot be explained cleanly to a new employee, an AI vendor will not discover the missing policy by magic.
Bring real failures into the test set. Use duplicate accounts, outdated documentation, unavailable systems, policy conflicts, upset callers, accessibility needs and requests that should be refused. A polished happy path only proves that the launch team can stage a polished happy path.
Set a human-quality floor before setting an automation target. Track whether the issue stayed solved, whether the customer had to contact support again, whether account changes were correct and whether the receiving employee got usable context. A handoff avoided is not automatically a customer helped.
Finally, write the exit plan while the vendor is still courting you. Know which records, tests, policies and integration definitions you can export. Decide who can pause the system and how service continues in a simpler mode. The boring plan is not a lack of ambition. It is the part that lets a company use an AI agent without making the agent the only person who understands the new process.
Presence may turn out to be a strong enterprise product. Its best contribution this week is simpler: it makes it harder to pretend that production AI agents are just smarter chatbots. They are organizational changes with software in the middle. Buy them that way.