Durable Agent Harness Runtime
ai4j-harness is an optional SDK module. It ports the core management ideas of Harness Anything into a Java runtime boundary — without bringing in the ha CLI, and without hardcoding any business workflow into the SDK.
Its positioning can be summarized in one sentence:
The
Agentowns thinking and execution inside one slice; theHarnessowns work state, recovery, and governance across requests, processes, and time.
Harness Outer-Loop Mechanism Diagram
Interactive architecture diagram: AgentHarness acts as the durable outer loop around an Agent — each run executes one bounded slice (claim lease → adapter.open → session.run → status mapping → persistOutcome), while the Agent's ReAct/tools/permissions/sandbox run unchanged; CommandGateway is the only write path, and Wait/Wakeup/Checkpoint all land in HarnessStore.
1. It Does Not Replace Your Existing Agent
Without Harness configured, the existing Agent, AgentSession, ToolExecutor, MCP, Function Call, Skill, A2A, Subagent, Agent Team, Memory, context compaction, Sandbox, Permission, Hook, and Plugin behaviors stay unchanged.
With Harness enabled, the existing Agent still owns:
| Existing Agent capability | Responsibility boundary |
|---|---|
| Model calls and protocol adaptation | ai4j / ai4j-agent |
| Per-run strategies such as ReAct, CodeAct, Workflow | ai4j-agent |
| Tool declarations, MCP, Function Call, Skill | ai4j-agent and the business side |
| Memory, context projection, context compaction | ai4j-agent / ai4j-coding |
| Sandbox, Permission, Hook, Plugin, Subagent, Agent Team | Existing Agent/Coding Runtime |
| When one Agent invocation stops | The Agent's step, token, and wall-clock settings, plus the Harness budget for this run |
Harness only adds, on the outside:
| Harness capability | What it does |
|---|---|
| Task | Dynamically records a long-running goal at runtime; no fixed task list must be declared up front |
| Execution | A durable execution instance for one input or one resumption |
| Checkpoint | Persists the Agent/Adapter state and summary needed to resume across slices |
| Wait/Wakeup | Persists waits for user input, approvals, async services, external events, or timers |
| Lease/Fencing | Prevents multiple workers from mutating the same Execution or Session concurrently |
| Task relation | Expresses parent/child tasks and dependency DAGs; rejects cyclic dependencies |
| Fact/Decision/Evidence/Relation | Turns long-running facts, decisions, evidence, and relations into queryable records |
| Submission/Review/Gate | Separates "the Agent says it is done" from "the system accepts it as done" |
| File/JDBC store | Persists all of the above to disk or a database instead of JVM memory only |
2. Core Object Relationships
Harness deliberately does not force-bind Task to Session:
business input message / external event / worker dispatch
|
v
HarnessRunRequest
|
v
Execution (one slice)
| |
| +--> Session ID: restores Agent context
|
+--> Task ID: optional; the Agent may create and bind one at runtime
|
v
Agent or HarnessExecutionAdapter
|
v
one bounded Agent slice
|
v
checkpoint + durable outcome + wait/wakeup
The distinction between these identities matters:
| Identity | Typical meaning | Binding rule |
|---|---|---|
| Task | Long-running work like "complete the refund investigation", "maintain the payment module", "produce one episode" | One Task can have many Executions |
| Execution | The concrete slice for this incoming message, resumed wait, or worker continuation | An Execution has at most one current Task, but may have none |
| Agent Session | The Agent's memory, event log, and run identity | One Session can participate in many independent Executions |
| scopeKey | Partitions tenants, projects, or workspaces inside one Harness ledger | Optional; it is neither a Conversation nor a Session |
So a new customer-service message normally creates a new Execution; it can reuse the customer's Agent Session, but reusing a Session does not automatically inherit the previous message's Task. A Coding Agent's Task usually belongs to the project and can be continued by Executions from different CLI, TUI, ACP, or background workers.
Entity Lineage: What Anchors to What
Every governance entity is anchored to a Task or an Execution through foreign keys, forming an auditable lineage graph:
| Entity | Anchored by | Additional references | Lifecycle |
|---|---|---|---|
| Fact | taskId + scopeKey | source, confidence, provenance | Invalidatable: invalidateFact sets valid=false with a trail; never deleted |
| Decision | taskId + scopeKey | proposer / arbiter | PROPOSED → settled by an entitled Actor via resolveDecision |
| Evidence | taskId + executionId | location, contentRef, kind | Append-only, cannot be invalidated — an artifact is a historical fact |
| Submission | taskId + executionId | submitter, evidenceIds[] | On creation it points task.submissionId at itself; Task → IN_REVIEW |
| Review | submissionId + taskId | reviewer, verdict | CHANGES_REQUESTED bounces the Task back to ACTIVE |
| Gate/Acceptance | taskId + executionId + submissionId | Evaluation result and reason | Evaluated at completeTask; all PASS required for DONE |
| HarnessRule | decisionId + scopeKey | kind, payload, preset, promotedBy | Promoted via promoteLesson (human Actor only by default); revokeLesson disables with lineage kept |
Entities can also be linked by RelationRecord — a generic directed edge whose endpoints are (EntityKind, id) pairs (TASK/FACT/DECISION/EVIDENCE/EXECUTION/CHECKPOINT/WAIT/REVIEW). Four built-in types: PARENT_OF (parent/child tasks), DEPENDS_ON (task dependency; the Gateway rejects cycles), SUPPORTS (evidence backing a conclusion), DERIVED_FROM (a conclusion derived from evidence).
Execution State Machine Diagram
Interactive state machine: Execution migrates among READY / RUNNING / WAITING / SUCCEEDED / FAILED / UNKNOWN — when a lease is lost or the persisted outcome is uncertain it is marked UNKNOWN rather than FAILED; a WAITING execution can only return to READY after deliver atomically writes answer+wakeup; terminal states cannot transition further.
3. How One run Works
AgentHarness.run(...) executes one bounded slice each time:
- The host puts its own business input object into
HarnessRunRequest.input, along with a stablesessionId, an optionaltaskId, ascopeKey, and a message-level idempotency key. - Harness creates or loads a durable Execution. Even when no Task exists yet, it can first record the input and the execution outcome.
- Harness acquires the Execution lease, and acquires a session lease for the shared Agent Session.
- Harness restores the existing checkpoint or Session snapshot, then overlays the management tools and enforced execution boundary onto the existing Agent.
- The Agent runs inside its own model, tool, MCP, Memory, compaction, and permission semantics; it may decide — based on the current input — whether to create, split, update, or relate Tasks.
- When the slice ends, Harness atomically persists the Agent/Adapter state, checkpoint, tool invocations, wait states, and the Execution outcome.
- The host then
resumes, callsdeliver, waits for external events, or hands runnable Tasks to the next worker.
// message, messageId, and sessionId are your own business-side concepts and fields.
HarnessRunResult result = harness.run(HarnessRunRequest.builder()
.scopeKey("shop-A")
.sessionId("customer-A-agent-session")
.idempotencyKey("message:msg-1001")
.input(message)
.build());
if (result.getStatus() == HarnessRunStatus.WAITING) {
// Persist result.getWaitId() on the business side and hand the "awaiting user answer / processing"
// state to your own channel layer.
publishWaitingReply(result.getOutputText(), result.getWaitId());
} else if (result.getStatus() == HarnessRunStatus.CONTINUATION_REQUIRED) {
// Work remains but this slice hit its boundary; a worker can continue it — do not treat it as done.
enqueueExecution(result.getExecution().getExecutionId());
}
message does not have to implement any SDK-fixed interface. It can be an e-commerce message DTO, an HTTP request, an event object, a CLI prompt, a ticket object, or any business input. Harness only persists the input summary and the run state the Agent needs to resume; whether and how the full business object is persisted or masked is the business system's decision.
Single-Run Internal Flow Diagram
Interactive flow diagram: run/resume entry → claimExecution lease+fencing → adapter.open(ctx,budget,prevState) → bounded Agent/ReAct slice → snapshot → status mapping → persistOutcome / ensureWait; the deliver loop atomically writes answer+wakeup, then continues the next slice from prevState+answer.
Open the interactive diagram in a new window4. Tasks Are Created Dynamically at Runtime
Developers do not need to write fixed TaskDefinitions for "refund", "change address", or "query order", nor create a single Task at application startup.
When the current input does not yet correspond to a Task, the Agent can call the auto-injected:
harness_task_manage {
"operation": "create",
"title": "Verify the customer's refund request",
"goal": "Confirm the order, refund eligibility, and execution result"
}
If the current Execution has no Task, the first create binds the new Task to that Execution; afterwards, within the same long-running work, the Agent can:
splitout subtasks such as order verification, policy verification, refund submission, and result notification;- use
add_dependencyto express "refund submission depends on order verification"; updatethe Task's goal and plan as new facts arrive;- record Facts, Decisions, and Evidence;
- save a checkpoint at slice boundaries;
- submit for external review via
harness_submission_requestinstead of declaring completion itself.
The host only passes taskId when it already knows the long-running work identity — for example, a background retry of a known refund ticket, or continuing a project-level coding Task. A Task can still have multiple Sessions and multiple Executions.
Submission-Review-Gate Flow Diagram
Interactive flow diagram: the Agent cannot mark a Task IN_REVIEW directly — it must submitTask → reviewSubmission (APPROVED / CHANGES_REQUESTED) → AcceptanceCoordinator evaluates evidence and completion gates → only then can completeTask move it to DONE; repair(parentExecutionId) creates a child execution under the parent's lineage.
Open the interactive diagram in a new window5. What the Agent, the Host, and the Harness Each Own
The Agent owns
- understanding the input and choosing business tools;
- deciding when to create or split Tasks based on work complexity;
- writing important facts, decisions, and evidence into the Harness;
- requesting a Wait when it needs the user, an approval, or an external system;
- producing staged answers or submission material;
- continuing work within its Agent step budget.
The business developer owns
- defining input DTOs, output DTOs, message idempotency keys, and external event formats;
- deciding how customers, orders, tickets, projects, repositories, and other entities map to
scopeKey, Task metadata, or the business database; - implementing business Tools and async service calls;
- deciding which tools require an existing Task and which require approval;
- deciding the time, round, token, and cost budget of one slice;
- defining human-takeover, conversation-close, retry, and timeout rules;
- configuring Gates for completion submissions and providing human or system review entry points;
- choosing File or JDBC, and owning the database, backups, workers, callbacks, and operations.
The Harness Runtime owns
- atomically saving and restoring long-running state;
- leasing, idempotency, and concurrency isolation for Executions, Sessions, and tool invocations;
- maintaining waits, wakeups, checkpoints, dependencies, evidence, and completion gates;
- marking outcomes as
UNKNOWNwhen a lease expires and the result is uncertain — never assuming an external side effect succeeded or failed; - quarantining late Agent/async results that arrive after cancellation or human takeover;
- leaving an existing Agent untouched when Harness is not configured.
Do not stuff business rules into the fixed fields of HarnessTaskSpec or HarnessContract. For example, "close a Conversation after 24 hours" is a customer-service rule, not a generic Harness Task state; Harness only provides durable Waits, Executions, and event records on which business rules can be built.
6. Configuring a Harness
For a standard Agent, use AgentHarness:
AgentHarness harness = AgentHarness.builder()
.agent(existingAgent) // existing ai4j Agent; its capability assembly is preserved
.persistence(HarnessPersistence.file(
projectRoot.resolve(".ai4j/harness")))
.contract(HarnessContract.builder()
.taskRequiredTool("submitRefund")
.approvalRequiredTool("submitRefund")
.build())
.build();
try {
HarnessRunResult result = harness.run(HarnessRunRequest.builder()
.scopeKey("shop-A")
.sessionId("customer-A-agent-session")
.input(message)
.build());
} finally {
// In a long-running service the Harness is usually an application-lifecycle bean,
// closed once on shutdown.
harness.close();
}
HarnessContract is a governance rule set, not a business task template. The rules above say: submitRefund cannot be called without a Task, and calling it requires approval; they say nothing about the Task's title, order fields, or conversation fields.
HarnessContract.builder() defaults to a strict posture — relaxing it requires an explicit call:
| Guard | Default | Meaning |
|---|---|---|
requiresApprovedReview | true | Task completion requires an APPROVED Review first |
requiresCompletionEvidence | true | A Submission must reference Evidence to pass the completion gate |
allowSystemApproval / allowSystemCompletion / allowSystemReconciliation | true | system Actors may approve/complete/reconcile; set false to restrict these to human Actors |
allowSystemPromotion | false | whether system Actors may promote decisions into rules; by default only human Actors can |
mayApprove / mayComplete / mayReconcile | !actor.isAgent() | Hard-coded: no configuration ever lets an Agent identity approve, complete, or reconcile its own work |
mayPromoteLesson | actor.isHuman() | promoting a rule is stricter than resolving a decision — rules constrain every future execution |
In other words, "the Agent cannot approve itself" is enforced by an identity check inside the Gateway, not by prompt discipline.
Task-level presets: different acceptance chains per task type
A single HarnessContract is instance-wide — when one Harness serves different task types (refund vs FAQ vs logistics lookup), PresetHarnessContract routes contract queries by task.metadata["preset"]:
HarnessContract contract = PresetHarnessContract.builder()
.preset("refund", p -> p
.requiresApprovedReview(true)
.requiredAcceptanceCheck("payment-verified")
.completionGate(refundReconciledGate))
.preset("faq", p -> p
.requiresApprovedReview(false)
.requiresCompletionEvidence(false))
.base(HarnessContract.builder().build()) // tasks without a preset land here
.onUnknownPreset(UnknownPresetPolicy.FAIL_CLOSED) // default: a typo'd preset fails
.build();
A task declares its preset at creation: harness_task_manage passes metadata through on create/update, and the host can set preset in HarnessTaskSpec.metadata directly. Deliberate boundary: tool-level rules do not enter presets — taskRequiredTool/approvalRequiredTool stay global, because whether an action needs approval should not depend on the task type.
LEARN reflow: turning a lesson into a runtime rule
A failure should not leave only an error message — it should pay rent. The full pipeline: failure → Fact ("this input class times out") → Decision proposal → an entitled Actor resolves it ACCEPTED → promoteLesson turns it into a HarnessRule → subsequent executions are constrained by the new rule:
// The Agent may propose; only a human may resolve and promote.
DecisionRecord decision = gateway.proposeDecision(HarnessDecisionSpec.builder()
.scopeKey("shop-A")
.question("Must refunds check dispute status first?")
.chosenOption("yes — that is where the last chargeback came from")
.build());
gateway.resolveDecision(decision.getDecisionId(), DecisionStatus.ACCEPTED,
"chargeback incident INC-88", HarnessActor.human("ops-lead"));
gateway.promoteLesson(decision.getDecisionId(), HarnessRuleSpec.builder()
.kind(RuleKind.REQUIRE_ACCEPTANCE_CHECK) // adds a mandatory acceptance check
.payload("dispute-status-checked")
.preset("refund") // optional: only refund-preset tasks
.build(), HarnessActor.human("ops-lead"));
Three RuleKinds: REQUIRE_ACCEPTANCE_CHECK (payload = check id), REQUIRE_APPROVAL_FOR_TOOL (payload = tool name), REQUIRE_TASK_FOR_TOOL. Rules merge with the static contract as a union — they can only tighten, never weaken what was declared at build time; relaxing a static rule means changing code and redeploying, on purpose. revokeLesson(ruleId, actor) disables a rule while keeping the record and its lineage; the Agent sees the rules currently in force through harness_context_get's learnedRules field, so it knows what it must satisfy.
For ai4j-coding, use CodingAgentHarness, which preserves workspace tools, CodeAct, compact, processes, MCP, subagents, and the existing approval semantics — see Coding Agent Harness Integration.
7. No CLI Is Brought In
Harness Anything's ha command suits humans and external Coding Agents operating a project governance directory; the SDK does not duplicate that CLI.
What the SDK provides:
- Java
AgentHarness/CodingAgentHarnessentry points; - a Java
HarnessCommandGatewayfor hosts, workers, webhooks, human back-offices, and tests; - optional Harness Function Call tools so the Agent can manage its own Tasks, facts, and waits mid-run;
- File/JDBC persistence implementations.
Developers can call these APIs from their own HTTP services, message consumers, CLIs, TUIs, background workers, or schedulers. The same Harness ledger can be shared by multiple Agent workers and modified by a human back-office that updates Tasks or delivers Waits — without requiring users to install ha.
8. Other Runtimes
Standard Agent and CodingAgent already have direct adapters. If your business has an irreplaceable runtime of its own, implement HarnessExecutionAdapter, which only needs to:
- open or resume its own run state from a checkpoint;
- execute one bounded slice;
- export serializable Adapter state;
- apply Wait results delivered by the host.
Task, Execution, lease, wait, checkpoint, dependency, review, and completion gating remain managed by the Harness. This extension point exists to onboard existing runtimes — ordinary business developers are not asked to implement another Agent.