Test the seam between agent infrastructure and domain configuration
Use one representative workflow to expose the work left to buyers of managed runtimes and domain-specialist agents.
OpenAI and Zendesk have drawn product boundaries around different parts of agent work.
OpenAI introduced the Agents API in public beta as managed infrastructure for long-running agents. The announcement lists session orchestration, context compaction, tool use, subagent coordination, and recovery. Compute can run in an OpenAI-hosted sandbox, buyer infrastructure, or a sandbox partner (source).
Zendesk introduced Industry Agents and Custom Agents grounded in an organization’s workflows, knowledge, policies, and connected systems. Zendesk describes them as specialists intended to complete complex service work from start to finish (source).
Those descriptions do not establish equivalent deployments, implementation effort, operating cost, or performance. They do show why category labels are insufficient for procurement: one offer emphasizes the execution environment, while the other emphasizes configured service work.
Put the same workflow through both product boundaries
Choose one representative workflow with a policy condition, a connected-system action, an exception, and an escalation. Write down its expected result before showing it to either vendor.
Then ask each proposed configuration to perform the workflow. The useful output is not a winner. It is a marked-up account of where the product stops and buyer work begins.
For the managed runtime, observe how the workflow exposes orchestration, tool calls, recovery behavior, compute placement, and retained evidence. Record which workflow definitions, tool contracts, policy translations, integrations, evaluation cases, and operational controls the buyer would still have to design or operate. That list is an architectural inference to validate against the proposed implementation, not a claim about every Agents API deployment.
For the domain-specialist product, observe how policies, knowledge, connected systems, exceptions, and escalation fit the packaged service model. Record what can be configured, what is fixed product behavior, and what requires additional integration. Zendesk’s announcement does not establish how much work a particular deployment will require.
Inspect what remains after the demonstration
The residue of the workflow is more informative than a generic feature comparison.
If the unresolved work clusters around orchestration, tool topology, recovery, and runtime controls, the organization is effectively taking responsibility for an execution harness. That may be appropriate when those mechanisms are part of what the buyer intends to shape.
If the unresolved work clusters around policy encoding, knowledge maintenance, workflow exceptions, and system mapping, the organization is taking responsibility for domain configuration inside a packaged service model. That may be appropriate when service rules and organizational context are where the buyer expects continued change.
Both configurations can mix vendor-managed and buyer-managed layers. The exercise should therefore capture the actual work left behind, rather than forcing the products into a simple build-versus-buy classification.
Price the unfinished work
Before comparing commercial terms, attach an owner and a verification method to every unresolved item from the workflow run. Include the integrations that must be maintained, the configuration that must be revised when policy changes, and the tests needed before expansion.
This also makes portability concrete. A buyer shaping a custom harness needs to know whether its behavior and tool contracts can move independently of the runtime. A buyer encoding service knowledge and workflows needs to know whether that configuration can be exported, reconstructed, or tested elsewhere. Neither announcement establishes those capabilities, so they require evidence from the proposed configuration.
Claims about lower cost, faster deployment, completion quality, or recovery should receive the same treatment. Product positioning cannot answer them. The buyer needs results from the representative workflow and the operating arrangement under consideration.
A useful procurement record is the annotated workflow, its unfinished work, the named owners, and the evidence required for expansion. That record lets a team evaluate a managed runtime or a domain specialist without pretending that the categories determine the outcome. Teams that need help turning the workflow into a limited implementation can begin with a scoped build.
