OpenAI's Presence Bets Enterprise Agents Win Narrow, Not Fast
OpenAI introduced Presence, an enterprise product for putting AI agents into business-critical, customer-facing work — billing resolution, insurance claims support, IT service requests. Every deployment is scoped to one job, with the agent given only the knowledge and system access that job requires, and companies set their own thresholds for when it needs approval or must hand off to a human. Presence ships with simulation and evaluation tools for pre-launch testing, a Codex-driven improvement loop for iterating after launch, and production monitoring with quality-signal tracking. OpenAI’s own support line is the proof point: it now resolves 75% of inbound issues without a human, and Codex-driven changes cut human handoffs by 15 percentage points within 10 days. Early users include BBVA (AI voice support for banking in Mexico), SoftBank (Japanese-language customer conversations), and IAG (support for high-demand weather events). Presence isn’t self-serve — it ships through OpenAI’s Forward Deployed Engineers and select systems integrators.
That scoped, escalation-gated design maps onto a category Forrester analyst Craig Le Clair named the same week. Le Clair splits the enterprise agent market into Hares (fast, horizontal SaaS agents from Microsoft, Salesforce, ServiceNow, SAP, and Google, racing toward commoditization on similar foundation models), Sloths (small edge agents near data sources), Tortoises (custom agents behind the firewall solving one substantive process), and Super-Tortoises (industry-specific systems coordinating whole business processes). His argument — “the fastest runner doesn’t always win” — is that Tortoises close the “action gap” horizontal agents can’t, because they’re built for one company’s actual compliance and ROI constraints. Presence’s job-scoped deployments, sold through FDEs rather than self-serve signup, read like a Tortoise product.
The BBVA voice deployment is worth watching against a second data point: Hugging Face’s Real World VoiceEQ benchmark, built from over a million human ratings across 40-plus voice models, found speech-to-speech systems show the widest quality variance of any voice category, and many still miss paralinguistic cues — tone, hesitation — that change what a response means in something like a fraud check. Presence is aiming voice agents at exactly that kind of high-stakes exchange. For consulting clients evaluating agent vendors, the real question isn’t whether the agent resolves the ticket — it’s whether it hears what wasn’t said.