Skip to main content

Durable Agent Harness Runtime

ai4j-harness is an optional SDK module. It ports the core management ideas of Harness Anything into a Java runtime boundary — without bringing in the ha CLI, and without hardcoding any business workflow into the SDK.

Its positioning can be summarized in one sentence:

The Agent owns thinking and execution inside one slice; the Harness owns work state, recovery, and governance across requests, processes, and time.

Harness Outer-Loop Mechanism Diagram​

Interactive architecture diagram: AgentHarness acts as the durable outer loop around an Agent — each run executes one bounded slice (claim lease → adapter.open → session.run → status mapping → persistOutcome), while the Agent's ReAct/tools/permissions/sandbox run unchanged; CommandGateway is the only write path, and Wait/Wakeup/Checkpoint all land in HarnessStore.

Open the interactive diagram in a new window

1. It Does Not Replace Your Existing Agent​

Without Harness configured, the existing Agent, AgentSession, ToolExecutor, MCP, Function Call, Skill, A2A, Subagent, Agent Team, Memory, context compaction, Sandbox, Permission, Hook, and Plugin behaviors stay unchanged.

With Harness enabled, the existing Agent still owns:

Existing Agent capabilityResponsibility boundary
Model calls and protocol adaptationai4j / ai4j-agent
Per-run strategies such as ReAct, CodeAct, Workflowai4j-agent
Tool declarations, MCP, Function Call, Skillai4j-agent and the business side
Memory, context projection, context compactionai4j-agent / ai4j-coding
Sandbox, Permission, Hook, Plugin, Subagent, Agent TeamExisting Agent/Coding Runtime
When one Agent invocation stopsThe Agent's step, token, and wall-clock settings, plus the Harness budget for this run

Harness only adds, on the outside:

Harness capabilityWhat it does
TaskDynamically records a long-running goal at runtime; no fixed task list must be declared up front
ExecutionA durable execution instance for one input or one resumption
CheckpointPersists the Agent/Adapter state and summary needed to resume across slices
Wait/WakeupPersists waits for user input, approvals, async services, external events, or timers
Lease/FencingPrevents multiple workers from mutating the same Execution or Session concurrently
Task relationExpresses parent/child tasks and dependency DAGs; rejects cyclic dependencies
Fact/Decision/Evidence/RelationTurns long-running facts, decisions, evidence, and relations into queryable records
Submission/Review/GateSeparates "the Agent says it is done" from "the system accepts it as done"
File/JDBC storePersists all of the above to disk or a database instead of JVM memory only

2. Core Object Relationships​

Harness deliberately does not force-bind Task to Session:

business input message / external event / worker dispatch
|
v
HarnessRunRequest
|
v
Execution (one slice)
| |
| +--> Session ID: restores Agent context
|
+--> Task ID: optional; the Agent may create and bind one at runtime
|
v
Agent or HarnessExecutionAdapter
|
v
one bounded Agent slice
|
v
checkpoint + durable outcome + wait/wakeup

The distinction between these identities matters:

IdentityTypical meaningBinding rule
TaskLong-running work like "complete the refund investigation", "maintain the payment module", "produce one episode"One Task can have many Executions
ExecutionThe concrete slice for this incoming message, resumed wait, or worker continuationAn Execution has at most one current Task, but may have none
Agent SessionThe Agent's memory, event log, and run identityOne Session can participate in many independent Executions
scopeKeyPartitions tenants, projects, or workspaces inside one Harness ledgerOptional; it is neither a Conversation nor a Session

So a new customer-service message normally creates a new Execution; it can reuse the customer's Agent Session, but reusing a Session does not automatically inherit the previous message's Task. A Coding Agent's Task usually belongs to the project and can be continued by Executions from different CLI, TUI, ACP, or background workers.

Entity Lineage: What Anchors to What​

Every governance entity is anchored to a Task or an Execution through foreign keys, forming an auditable lineage graph:

EntityAnchored byAdditional referencesLifecycle
FacttaskId + scopeKeysource, confidence, provenanceInvalidatable: invalidateFact sets valid=false with a trail; never deleted
DecisiontaskId + scopeKeyproposer / arbiterPROPOSED → settled by an entitled Actor via resolveDecision
EvidencetaskId + executionIdlocation, contentRef, kindAppend-only, cannot be invalidated — an artifact is a historical fact
SubmissiontaskId + executionIdsubmitter, evidenceIds[]On creation it points task.submissionId at itself; Task → IN_REVIEW
ReviewsubmissionId + taskIdreviewer, verdictCHANGES_REQUESTED bounces the Task back to ACTIVE
Gate/AcceptancetaskId + executionId + submissionIdEvaluation result and reasonEvaluated at completeTask; all PASS required for DONE
HarnessRuledecisionId + scopeKeykind, payload, preset, promotedByPromoted via promoteLesson (human Actor only by default); revokeLesson disables with lineage kept

Entities can also be linked by RelationRecord — a generic directed edge whose endpoints are (EntityKind, id) pairs (TASK/FACT/DECISION/EVIDENCE/EXECUTION/CHECKPOINT/WAIT/REVIEW). Four built-in types: PARENT_OF (parent/child tasks), DEPENDS_ON (task dependency; the Gateway rejects cycles), SUPPORTS (evidence backing a conclusion), DERIVED_FROM (a conclusion derived from evidence).

Execution State Machine Diagram​

Interactive state machine: Execution migrates among READY / RUNNING / WAITING / SUCCEEDED / FAILED / UNKNOWN — when a lease is lost or the persisted outcome is uncertain it is marked UNKNOWN rather than FAILED; a WAITING execution can only return to READY after deliver atomically writes answer+wakeup; terminal states cannot transition further.

Open the interactive diagram in a new window

3. How One run Works​

AgentHarness.run(...) executes one bounded slice each time:

  1. The host puts its own business input object into HarnessRunRequest.input, along with a stable sessionId, an optional taskId, a scopeKey, and a message-level idempotency key.
  2. Harness creates or loads a durable Execution. Even when no Task exists yet, it can first record the input and the execution outcome.
  3. Harness acquires the Execution lease, and acquires a session lease for the shared Agent Session.
  4. Harness restores the existing checkpoint or Session snapshot, then overlays the management tools and enforced execution boundary onto the existing Agent.
  5. The Agent runs inside its own model, tool, MCP, Memory, compaction, and permission semantics; it may decide — based on the current input — whether to create, split, update, or relate Tasks.
  6. When the slice ends, Harness atomically persists the Agent/Adapter state, checkpoint, tool invocations, wait states, and the Execution outcome.
  7. The host then resumes, calls deliver, waits for external events, or hands runnable Tasks to the next worker.
// message, messageId, and sessionId are your own business-side concepts and fields.
HarnessRunResult result = harness.run(HarnessRunRequest.builder()
.scopeKey("shop-A")
.sessionId("customer-A-agent-session")
.idempotencyKey("message:msg-1001")
.input(message)
.build());

if (result.getStatus() == HarnessRunStatus.WAITING) {
// Persist result.getWaitId() on the business side and hand the "awaiting user answer / processing"
// state to your own channel layer.
publishWaitingReply(result.getOutputText(), result.getWaitId());
} else if (result.getStatus() == HarnessRunStatus.CONTINUATION_REQUIRED) {
// Work remains but this slice hit its boundary; a worker can continue it — do not treat it as done.
enqueueExecution(result.getExecution().getExecutionId());
}

message does not have to implement any SDK-fixed interface. It can be an e-commerce message DTO, an HTTP request, an event object, a CLI prompt, a ticket object, or any business input. Harness only persists the input summary and the run state the Agent needs to resume; whether and how the full business object is persisted or masked is the business system's decision.

Single-Run Internal Flow Diagram​

Interactive flow diagram: run/resume entry → claimExecution lease+fencing → adapter.open(ctx,budget,prevState) → bounded Agent/ReAct slice → snapshot → status mapping → persistOutcome / ensureWait; the deliver loop atomically writes answer+wakeup, then continues the next slice from prevState+answer.

Open the interactive diagram in a new window

4. Tasks Are Created Dynamically at Runtime​

Developers do not need to write fixed TaskDefinitions for "refund", "change address", or "query order", nor create a single Task at application startup.

When the current input does not yet correspond to a Task, the Agent can call the auto-injected:

harness_task_manage {
"operation": "create",
"title": "Verify the customer's refund request",
"goal": "Confirm the order, refund eligibility, and execution result"
}

If the current Execution has no Task, the first create binds the new Task to that Execution; afterwards, within the same long-running work, the Agent can:

  • split out subtasks such as order verification, policy verification, refund submission, and result notification;
  • use add_dependency to express "refund submission depends on order verification";
  • update the Task's goal and plan as new facts arrive;
  • record Facts, Decisions, and Evidence;
  • save a checkpoint at slice boundaries;
  • submit for external review via harness_submission_request instead of declaring completion itself.

The host only passes taskId when it already knows the long-running work identity — for example, a background retry of a known refund ticket, or continuing a project-level coding Task. A Task can still have multiple Sessions and multiple Executions.

Submission-Review-Gate Flow Diagram​

Interactive flow diagram: the Agent cannot mark a Task IN_REVIEW directly — it must submitTask → reviewSubmission (APPROVED / CHANGES_REQUESTED) → AcceptanceCoordinator evaluates evidence and completion gates → only then can completeTask move it to DONE; repair(parentExecutionId) creates a child execution under the parent's lineage.

Open the interactive diagram in a new window

5. What the Agent, the Host, and the Harness Each Own​

The Agent owns​

  • understanding the input and choosing business tools;
  • deciding when to create or split Tasks based on work complexity;
  • writing important facts, decisions, and evidence into the Harness;
  • requesting a Wait when it needs the user, an approval, or an external system;
  • producing staged answers or submission material;
  • continuing work within its Agent step budget.

The business developer owns​

  • defining input DTOs, output DTOs, message idempotency keys, and external event formats;
  • deciding how customers, orders, tickets, projects, repositories, and other entities map to scopeKey, Task metadata, or the business database;
  • implementing business Tools and async service calls;
  • deciding which tools require an existing Task and which require approval;
  • deciding the time, round, token, and cost budget of one slice;
  • defining human-takeover, conversation-close, retry, and timeout rules;
  • configuring Gates for completion submissions and providing human or system review entry points;
  • choosing File or JDBC, and owning the database, backups, workers, callbacks, and operations.

The Harness Runtime owns​

  • atomically saving and restoring long-running state;
  • leasing, idempotency, and concurrency isolation for Executions, Sessions, and tool invocations;
  • maintaining waits, wakeups, checkpoints, dependencies, evidence, and completion gates;
  • marking outcomes as UNKNOWN when a lease expires and the result is uncertain — never assuming an external side effect succeeded or failed;
  • quarantining late Agent/async results that arrive after cancellation or human takeover;
  • leaving an existing Agent untouched when Harness is not configured.

Do not stuff business rules into the fixed fields of HarnessTaskSpec or HarnessContract. For example, "close a Conversation after 24 hours" is a customer-service rule, not a generic Harness Task state; Harness only provides durable Waits, Executions, and event records on which business rules can be built.

6. Configuring a Harness​

For a standard Agent, use AgentHarness:

AgentHarness harness = AgentHarness.builder()
.agent(existingAgent) // existing ai4j Agent; its capability assembly is preserved
.persistence(HarnessPersistence.file(
projectRoot.resolve(".ai4j/harness")))
.contract(HarnessContract.builder()
.taskRequiredTool("submitRefund")
.approvalRequiredTool("submitRefund")
.build())
.build();

try {
HarnessRunResult result = harness.run(HarnessRunRequest.builder()
.scopeKey("shop-A")
.sessionId("customer-A-agent-session")
.input(message)
.build());
} finally {
// In a long-running service the Harness is usually an application-lifecycle bean,
// closed once on shutdown.
harness.close();
}

HarnessContract is a governance rule set, not a business task template. The rules above say: submitRefund cannot be called without a Task, and calling it requires approval; they say nothing about the Task's title, order fields, or conversation fields.

HarnessContract.builder() defaults to a strict posture — relaxing it requires an explicit call:

GuardDefaultMeaning
requiresApprovedReviewtrueTask completion requires an APPROVED Review first
requiresCompletionEvidencetrueA Submission must reference Evidence to pass the completion gate
allowSystemApproval / allowSystemCompletion / allowSystemReconciliationtruesystem Actors may approve/complete/reconcile; set false to restrict these to human Actors
allowSystemPromotionfalsewhether system Actors may promote decisions into rules; by default only human Actors can
mayApprove / mayComplete / mayReconcile!actor.isAgent()Hard-coded: no configuration ever lets an Agent identity approve, complete, or reconcile its own work
mayPromoteLessonactor.isHuman()promoting a rule is stricter than resolving a decision — rules constrain every future execution

In other words, "the Agent cannot approve itself" is enforced by an identity check inside the Gateway, not by prompt discipline.

Task-level presets: different acceptance chains per task type​

A single HarnessContract is instance-wide — when one Harness serves different task types (refund vs FAQ vs logistics lookup), PresetHarnessContract routes contract queries by task.metadata["preset"]:

HarnessContract contract = PresetHarnessContract.builder()
.preset("refund", p -> p
.requiresApprovedReview(true)
.requiredAcceptanceCheck("payment-verified")
.completionGate(refundReconciledGate))
.preset("faq", p -> p
.requiresApprovedReview(false)
.requiresCompletionEvidence(false))
.base(HarnessContract.builder().build()) // tasks without a preset land here
.onUnknownPreset(UnknownPresetPolicy.FAIL_CLOSED) // default: a typo'd preset fails
.build();

A task declares its preset at creation: harness_task_manage passes metadata through on create/update, and the host can set preset in HarnessTaskSpec.metadata directly. Deliberate boundary: tool-level rules do not enter presets — taskRequiredTool/approvalRequiredTool stay global, because whether an action needs approval should not depend on the task type.

LEARN reflow: turning a lesson into a runtime rule​

A failure should not leave only an error message — it should pay rent. The full pipeline: failure → Fact ("this input class times out") → Decision proposal → an entitled Actor resolves it ACCEPTED → promoteLesson turns it into a HarnessRule → subsequent executions are constrained by the new rule:

// The Agent may propose; only a human may resolve and promote.
DecisionRecord decision = gateway.proposeDecision(HarnessDecisionSpec.builder()
.scopeKey("shop-A")
.question("Must refunds check dispute status first?")
.chosenOption("yes — that is where the last chargeback came from")
.build());
gateway.resolveDecision(decision.getDecisionId(), DecisionStatus.ACCEPTED,
"chargeback incident INC-88", HarnessActor.human("ops-lead"));

gateway.promoteLesson(decision.getDecisionId(), HarnessRuleSpec.builder()
.kind(RuleKind.REQUIRE_ACCEPTANCE_CHECK) // adds a mandatory acceptance check
.payload("dispute-status-checked")
.preset("refund") // optional: only refund-preset tasks
.build(), HarnessActor.human("ops-lead"));

Three RuleKinds: REQUIRE_ACCEPTANCE_CHECK (payload = check id), REQUIRE_APPROVAL_FOR_TOOL (payload = tool name), REQUIRE_TASK_FOR_TOOL. Rules merge with the static contract as a union — they can only tighten, never weaken what was declared at build time; relaxing a static rule means changing code and redeploying, on purpose. revokeLesson(ruleId, actor) disables a rule while keeping the record and its lineage; the Agent sees the rules currently in force through harness_context_get's learnedRules field, so it knows what it must satisfy.

For ai4j-coding, use CodingAgentHarness, which preserves workspace tools, CodeAct, compact, processes, MCP, subagents, and the existing approval semantics — see Coding Agent Harness Integration.

7. No CLI Is Brought In​

Harness Anything's ha command suits humans and external Coding Agents operating a project governance directory; the SDK does not duplicate that CLI.

What the SDK provides:

  • Java AgentHarness / CodingAgentHarness entry points;
  • a Java HarnessCommandGateway for hosts, workers, webhooks, human back-offices, and tests;
  • optional Harness Function Call tools so the Agent can manage its own Tasks, facts, and waits mid-run;
  • File/JDBC persistence implementations.

Developers can call these APIs from their own HTTP services, message consumers, CLIs, TUIs, background workers, or schedulers. The same Harness ledger can be shared by multiple Agent workers and modified by a human back-office that updates Tasks or delivers Waits — without requiring users to install ha.

8. Other Runtimes​

Standard Agent and CodingAgent already have direct adapters. If your business has an irreplaceable runtime of its own, implement HarnessExecutionAdapter, which only needs to:

  • open or resume its own run state from a checkpoint;
  • execute one bounded slice;
  • export serializable Adapter state;
  • apply Wait results delivered by the host.

Task, Execution, lease, wait, checkpoint, dependency, review, and completion gating remain managed by the Harness. This extension point exists to onboard existing runtimes — ordinary business developers are not asked to implement another Agent.