Plan a harness change as a system migration
A harness switch can change agent behavior, retained state, controls, and consumption, so evaluate the full system configuration.
UiPath’s advanced-agent preview changes more than the execution wrapper around a model. The documented preview adds file-backed working memory, context compaction, sub-agent delegation, and programmatic tool calls. It is limited to Automation Cloud, does not support guardrails, and does not retain completed autonomous-run workspaces by default. UiPath also warns that converting an existing agent permanently removes configured tool values and per-model settings, while delegation and revision can increase LLM calls (UiPath documentation).
Those changes alter state, control coverage, configuration, execution paths, and potential consumption. For an operator, the useful unit of comparison is therefore the complete agent configuration—not merely the model selected inside it.
That judgment also fits the available independent research. Workspace-Bench evaluates four harnesses and seven foundation models on realistic tasks with large file dependencies. Its authors report substantial performance variation between harnesses using the same underlying model, alongside differences in interaction turns and token consumption; performance also declines as task complexity rises (Workspace-Bench paper). The paper does not establish how a particular UiPath agent will behave. It does show why holding the backbone model constant is insufficient evidence that a harness change is behaviorally neutral.
The migration begins before conversion
Because UiPath documents destructive configuration effects, the first migration artifact is a recoverable record of the current agent. Before changing the harness, export or otherwise capture every configured tool value and per-model setting that the team may need to reconstruct. Record the current harness, model identity, tool inventory, prompts, environment, and evaluation version in the same snapshot.
The snapshot is not proof that restoration will work. A separate rollback rehearsal can establish whether the old configuration can actually be rebuilt within an acceptable interval. If the platform conversion cannot be reversed directly, define rollback as reconstruction of a known configuration rather than as a toggle.
Preview status belongs in this record too. UiPath labels the advanced harness a preview and documents it as Automation Cloud-only (UiPath documentation). That status does not establish instability, but it does bound the approval: the team is evaluating a preview configuration with its currently documented limitations, not a generally available equivalent of the standard harness.
Rebuild the control case
A control present in the old system cannot be assumed to survive under a new execution design. UiPath explicitly says guardrails are unsupported in the advanced-agent preview (UiPath documentation). The migration decision should identify each control that depended on that feature and choose a disposition: replace it elsewhere, narrow the agent’s authority, add a human decision point, or reject the migration for that workflow.
This is a control-equivalence check, not a feature-name comparison. For each consequential action, trace:
- what initiates the action;
- which agent or sub-agent can invoke the tool;
- where authorization is enforced;
- what evidence records the attempt and result;
- how execution is stopped or recovered after a failure.
Delegation makes that trace especially important. UiPath documents sub-agent delegation and programmatic tool calls in the preview (UiPath documentation). A derived migration test should therefore cover both direct tool use and delegated tool use, because they are distinct paths through the proposed system even when they pursue the same task.
Test the workspace as an operating boundary
File-backed working memory can change what information is available during a run. Context compaction can change what remains available to later steps. UiPath further documents that completed autonomous runs do not retain workspaces by default (UiPath documentation). These facts create separate questions for the migration:
- Which files enter working memory, and at what point?
- What information survives compaction?
- What remains after successful, failed, cancelled, and timed-out runs?
- What evidence is available after a workspace is no longer retained?
- Which recovery procedures depend on artifacts that may disappear?
The answers should come from tests in the target tenant and configuration. Documentation identifies the expected product boundary; it is not execution evidence for a team’s particular workflow.
Representative tasks should include realistic file counts, dependencies, and multi-step edits. Workspace-Bench reports lower performance as task complexity rises, so a small single-file exercise cannot stand in for a complex workspace workload (Workspace-Bench paper). This result does not prescribe a universal task suite. It supports sampling across the actual complexity range rather than testing only the easiest path.
Compare behavior and consumption together
A useful migration comparison runs the old and proposed configurations against the same task set and scoring rules. Capture task completion, incorrect or missing changes, tool outcomes, recovery behavior, interaction turns, model calls, and token consumption. Keep the model fixed where possible, but retain the full configuration identity beside every result.
The reason is empirical but bounded: Workspace-Bench reports that harnesses using the same model can differ substantially in results, turns, and token consumption (Workspace-Bench paper). UiPath separately notes that delegation and revision can increase LLM calls in its advanced harness (UiPath documentation). Neither source predicts a specific cost increase for a given deployment. Direct measurement is needed to determine whether additional calls occur, how large the change is, and whether any task-quality difference justifies it.
Use distributions and failure categories rather than a single average score. A migration can preserve aggregate completion while moving failures into more consequential actions, or reduce tokens while adding interaction turns. Those are different operating changes and may lead to different decisions.
Define a bounded release decision
The release record can stay compact if it answers four questions:
- Can the original configuration be reconstructed from the captured snapshot?
- Does every relied-upon control have a tested equivalent or an explicit compensating decision?
- Do representative workspace tasks meet the team’s predeclared behavioral and recovery thresholds?
- Are measured call and token changes acceptable for the expected workload?
Add the preview status, unresolved limitations, evidence date, approver, and rollback trigger. Approval then applies to one named configuration and workload envelope. It does not imply that the advanced harness is generally better, that the same result will transfer to other tasks, or that preview behavior will remain unchanged.
Teams building agent systems can use this migration record to keep model identity, harness behavior, control coverage, retained state, and consumption in one decision. The central point is narrow: when the harness changes how an agent remembers, delegates, acts, and accounts for work, the resulting system deserves a fresh evaluation.
