For the complete documentation index, see llms.txt. This page is also available as Markdown.

Human-in-the-Loop (HITL)

The HITL framework lets you insert human approval gates into AI agent conversations. It works on both surfaces:

Surface
Trigger
Pause state
Use case

Regular 1:1

PAUSE_CONVERSATION action in behavior rules

AWAITING_HUMAN

Single-agent pipelines needing human checkpoints

Group conversations

requiresApproval: true on a phase

AWAITING_APPROVAL

Multi-agent orchestration needing oversight


Why Two Surfaces?

The two surfaces exist because regular and group conversations have fundamentally different execution models — and therefore different natural "pause points."

Regular 1:1: Action-Driven

A regular conversation runs a linear pipeline of lifecycle tasks (parse → behavior rules → output → templating → …). The natural pause point is between tasks in the pipeline, triggered by a behavior rule emitting the reserved PAUSE_CONVERSATION action.

This is opt-in per turn: whether to pause depends on runtime conditions (user intent, slot values, context). Behavior rules are EDDI's existing mechanism for conditional actions — reusing them keeps the logic in configuration, not Java.

For a rule pause, hitlConfig's approvalTimeout/timeoutPolicy/pauseReason only control timeout behavior — the actual triggering is in the behavior rules. Without a rule emitting PAUSE_CONVERSATION, none of them do anything. (hitlConfig.toolApprovals is a separate trigger that needs no behavior rule — see Tool-Level Approval Gating.)

Group: Config-Driven

A group conversation runs a structured pipeline of phases (Opinion → Challenge → Synthesis). "Should this phase be reviewed?" is a design-time decision — a boolean on the phase definition (requiresApproval: true).

Granularity (group only)

Granularity
Semantics

PHASE (default)

Pause after the phase's repeats fully complete. Resume advances to the next phase.

TASK

For EXECUTE phases only: each task result is individually submitted for approval; resume re-enters the same phase so approved/retried tasks and their dependents can continue.

[!IMPORTANT] granularity: "TASK" only changes behavior for EXECUTE phases (they have a task list). A non-EXECUTE phase with requiresApproval: true falls back to a PHASE-style pause — the gate is never silently skipped.


Surface 1: Regular 1:1 Conversations

How It Works

While paused, user input is rejected without being consumed: POST /agents/{conversationId} returns 409 Conflict immediately with a body directing to the resume endpoint — the message is not queued and not processed. The same applies if a pause commits in the narrow window after a message was accepted (the turn is skipped and answered with 409 instead of executing against the paused conversation). Chat clients should render "awaiting approval" and disable input until the decision lands.

Configuration

Field
Type
Default
Description

approvalTimeout

ISO-8601 duration ("PT30S", "PT15M")

null

Time before the timeout policy fires. Required for finite policies.

timeoutPolicy

AUTO_APPROVE | AUTO_REJECT | ABORT | WAIT_INDEFINITELY

WAIT_INDEFINITELY

What happens when the timeout expires.

pauseReason

string, ≤ 500 chars

null

Approver-facing text shown in pending-approval listings and approval-status — answers "what am I approving?". Falls back to a generic reason when absent.

Validation happens at save time (and on ZIP import): a finite policy without a valid positive approvalTimeout (or a malformed duration like "30m", or an overlong pauseReason) is rejected with an actionable 400 — it never degrades silently to wait-forever at runtime.

Behavior Rule (the actual trigger)

PAUSE_CONVERSATION is a reserved action (like CONVERSATION_START, CONVERSATION_END, STOP_CONVERSATION). Do not reuse it as a quick-reply expression.

REST API

Method
Path
Description

POST

/agents/{conversationId}/resume

Submit a decision ({"verdict": "APPROVED"|"REJECTED", "note": "..."} — verdict is case-insensitive)

GET

/agents/{conversationId}/approval-status

Pause summary: state, pausedAt, pauseReason, timeoutPolicy, approvalTimeout (render the auto-decision deadline as pausedAt + approvalTimeout), pauseDetails — structured, computed at read time (see below). ?detail=full for the whole snapshot — approver-only callers get full only while the conversation is paused, 403 otherwise

GET

/agents/pending-approvals?limit=200

List paused conversations as summaries (bounded, max 1000). Non-admin/non-approver callers get their own conversations, filtered inside the query — the limit applies after the owner restriction, so a personal inbox is never starved by other users' backlog

POST

/agents/{conversationId}/cancel

Cancel a paused (or same-pod running) conversation

Status codes are discriminating: 400 invalid body (missing verdict, note > 4 KB — the pause is not consumed), 404 unknown conversation, 409 not in a resumable state (body names the current state), 200 success. HITL operations on a conversation whose descriptor is missing (legacy/corruption) fail closed — only admins/approvers may act on them.

pauseDetails shape

pauseDetails is computed at read time from the snapshot (nothing new is persisted) and is null once the conversation is no longer paused. Its shape depends on hitlPauseType:

  • TOOL_CALL (a gated tool call paused for approval):

    arguments is always the redacted, capped value (argumentsRedacted) — the raw arguments never appear in this response. outcomeUnknown lists callIds that have an EXECUTING journal entry — i.e. a prior approval crashed mid-execution and the outcome is genuinely unknown; it is empty in the common case.

  • RULE (a behavior-rule PAUSE_CONVERSATION, including legacy snapshots with no hitlPauseType recorded):

    actions is the ACTIONS data of the paused step.

GET /agents/pending-approvals entries also carry pauseType (RULE/TOOL_CALL; a pause stored without a type — rule pauses before 6.4 — is reported as RULE) and toolNames (names only, no arguments) so inbox UIs can badge tool-call pauses without a second round trip.

What the Agent Sees After a Decision

Where
Key
Notes

Templates (same turn, APPROVED)

{memory.current.hitlDecision}, {memory.current.hitlDecisionNote}

Conversation output written at resume time

Properties (this + later turns)

{properties.hitlVerdict}

Conversation-scoped property — next-turn behavior rules can react via a property matcher

Raw step data (pipeline tasks)

hitl:decision_verdict, hitl:decision_note, hitl:decision_by

Not template-accessible (colon keys)

On REJECTED, the remaining pipeline tasks are skipped (the actions that would have triggered API calls are still in the step — they must not run) and a public output message with the reviewer's note is emitted so UIs render feedback.


Surface 2: Group Conversations

Configuration

Field
Type
Default
Description

granularity

PHASE | TASK

PHASE

approvalTimeout / timeoutPolicy

as above

WAIT_INDEFINITELY

Validated at save time

onTaskRejection

FAIL | RETRY

FAIL

FAIL: rejected task is terminal. RETRY: task is re-queued (ASSIGNED) with the reviewer's note as feedback — the agent re-executes it addressing the rejection reason.

phases[n].requiresApproval

boolean

false

Per-phase approval gate

REST API

Method
Path
Description

POST

/groups/{groupId}/conversations/{gcId}/approve

Approve/reject (optionally per-task)

POST

/groups/{groupId}/conversations/{gcId}/approve/stream

Same, with SSE progress events

GET

/groups/{groupId}/conversations/{gcId}/approval-status

Pause summary (state, pausedAt, phase, pauseType, reason, timeoutPolicy, awaiting task ids); ?detail=full returns the whole conversation incl. transcript — approver-only callers get full only while paused, 403 otherwise

GET

/groups/{groupId}/conversations/pending-approvals?limit=100

This group's paused conversations as summaries (bounded, max 1000)

GET

/groups/pending-approvals?limit=100

Global approval inbox across all groups (bounded, max 1000): admins/approvers see everything, other callers their own conversations

POST

/groups/{groupId}/conversations/{gcId}/cancel

Cancel the discussion — 409 if it is already in a terminal state

The {groupId} path segment is validated: a conversation that does not belong to the given group returns 404 on read/delete/cancel/approve/approval-status. A discussion deleted concurrently with an approve/cancel also returns 404 (never a misleading 409 state conflict).

Approval bodies:

Unknown task IDs, tasks not awaiting approval, and unknown decision values (anything other than case-insensitive APPROVED/REJECTED) are rejected with 400 before anything is applied — the pause and its timeout schedule survive a bad request. An explicit empty taskApprovals map ({}) behaves like the approve-all shortcut. Resume-vs-resume and cancel-vs-approve races are arbitrated by a DB-level compare-and-set; the loser gets 409.

SSE events: awaiting_approval (stream closes; re-attach via /approve/stream), hitl_resume (emitted once the approve commits — the stream stays open for the resumed discussion's events), cancelled and group_complete/group_error (stream closes), plus the standard discussion events. Every terminal outcome — including REJECTED verdicts, config-drift aborts, and resume failures — emits a closing event, so /approve/stream clients never hang.

Config drift protection

Pauses can last days. On resume, the bookmarked phase (index + name) is verified against the current group config; if phases were added, removed, or reordered in between, the resume is aborted and the pause is restored (state back to AWAITING_APPROVAL, timeout schedule re-armed, group_error emitted with the mismatch details) — fix the config and approve again, or cancel. Transient failures before the resumed execution starts restore the pause the same way instead of failing the discussion terminally.

No-progress protection (TASK granularity)

A TASK-gate pause caused by turn-budget exhaustion could previously loop forever under AUTO_APPROVE (approve → zero work → identical re-pause). Resumes now fingerprint the pause (phase + non-terminal task states): an explicit human APPROVED grants a fresh turn budget; if a resumed leg re-pauses with an identical fingerprint (no progress), the discussion fails with an actionable error instead of pausing again — guaranteed termination. Automated (system:*) approvals never grant fresh budget, so a timeout policy cannot sustain the loop.


Tool-Level Approval Gating

The two surfaces above pause turns (a behavior-rule PAUSE_CONVERSATION) or phases (requiresApproval). Tool-level HITL is a third, finer gate: it pauses when the LLM invokes a matching tool, before that tool executes. It works for all seven tool sources — built-in @Tool, http (httpcall), mcp, a2a, dynamic (dynamic agents), memory, and recall.

The gate lives in the tool-execution loop (AgentOrchestrator), so it is fail-safe: a gated call is intercepted before execution. It reuses the existing pause machinery — same AWAITING_HUMAN state, same POST /agents/{id}/resume endpoint, same timeout/audit/Slack/crash-recovery paths — with hitlPauseType = "TOOL_CALL" distinguishing it from a RULE pause. A single resume verdict resolves either pause type.

Configuration

toolApprovals has two homes, both optional:

  • Agent-level defaultAgentConfiguration.hitlConfig.toolApprovals (applies to every LLM task in the agent).

  • Per-task overrideLlmConfiguration.task[n].toolApprovals (a langchain task). How it combines with the agent-level block is decided by eddi.hitl.tool.task-approvals.mode: under the default strict the task block can only strengthen the agent gate (patterns are united, task-level exempt is ignored, task-level AUTO_APPROVE is demoted); under replace it is a full replace and the agent-level block is ignored for that task — the pre-6.3.0 behavior. See Precedence for the per-field rules.

Field
Type
Default
Description

requireApproval

list of glob patterns

absent/empty = gate off

Tools whose calls must be human-approved. An absent or empty list disables tool gating entirely (backward compatible).

exempt

list of glob patterns

null

Exemptions — always beat requireApproval. A pattern in exempt with no requireApproval is rejected at save time (no effect).

maxPausesPerTurn

integer, 1..10

3

Max tool pauses in one turn. Fail-closed at the cap: once reached, the remaining gated calls in that turn are auto-error-resulted (hitl_pause_cap) rather than executed — never silently run.

maxAutoApprovalsPerTurn

integer, 0..10

2

Max consecutive system:* (timeout) auto-approvals per turn before the no-progress guard applies onNoProgress.

onNoProgress

WAIT_FOR_HUMAN | AUTO_REJECT | ABORT

WAIT_FOR_HUMAN

What happens when a tool pause re-pauses with an identical fingerprint after a system decision (loop protection). WAIT_FOR_HUMAN demotes the re-pause to WAIT_INDEFINITELY; AUTO_REJECT reject-alls; ABORT cancels.

approvalTimeout

ISO-8601 duration

inherits hitlConfig.approvalTimeout

Tool-pause timeout override. A finite policy requires a positive duration (validated at save time).

timeoutPolicy

AUTO_APPROVE | AUTO_REJECT | ABORT | WAIT_INDEFINITELY

Tool-pause timeout policy override.

pauseReason

string, ≤ 500 chars

generic

Approver-facing reason; the literal {toolNames} placeholder is substituted with the gated tool names.

pendingMessage

string, ≤ 500 chars

null

End-user-facing message stored as public output at pause commit (so a chat UI shows the user why the turn stalled). {toolNames} is substituted.

inGroupTurns

REJECT

REJECT

Behavior when a member agent's tool call is gated inside a group turn. INBOX is reserved (rejected with a 400 in v1). See Group members.

rules

list of objects

null

Per-tool friction overrides. See Per-tool approval rules.

All values are validated at save time and on ZIP import (HitlConfigValidation.validateToolApprovals): bad patterns (with the offending index), a pattern in both requireApproval and exempt, duplicates, exempt without requireApproval, out-of-range integers, a reserved INBOX, a finite policy without a valid approvalTimeout, and overlong reason/message all yield an actionable 400 — the gate never degrades silently at runtime.

Per-tool approval rules

The fields above are single scalars for every gated tool, so "deploy an agent" and "create an agent" cannot differ in how long a reviewer has or what the approval card says. rules fixes that:

Field
Type
Description

match

pattern

Required. Same pattern language as requireApproval — bare name, source:name, or source.method:path.

timeoutPolicy

AUTO_APPROVE | AUTO_REJECT | ABORT | WAIT_INDEFINITELY

Overrides toolApprovals.timeoutPolicy for this rule.

approvalTimeout

ISO-8601 duration

Overrides toolApprovals.approvalTimeout.

pauseReason

string, ≤ 500 chars

Overrides toolApprovals.pauseReason. {toolNames} is substituted.

pendingMessage

string, ≤ 500 chars

Overrides toolApprovals.pendingMessage. {toolNames} is substituted.

A rule tunes friction; it never gates or ungates. Whether a call needs approval is decided only by requireApproval/exempt. This is deliberate: the gate allows an unmatched call, so it only survives by gating broadly and exempting narrowly — a rule able to ungate would let a config grant capability by adding an entry.

Resolution (ToolApprovalRules):

  1. Per call, most specific wins. Rules are tried fewest-wildcards-first, then longest-pattern-first, so http.post:/agentstore/agents beats http.post:* whatever order they are listed in.

  2. Per batch, strictest wins. A model can emit several gated calls in one message and they pause together, under one policy — so the matched rules are reduced to one: the rule assuming the least human authority on timeout, ordered WAIT_INDEFINITELY > ABORT > AUTO_REJECT > AUTO_APPROVE > (no policy stated). Bundling a lenient call into a batch can therefore never soften a stricter rule. Ties break on specificity.

  3. Fields fall back individually to the toolApprovals scalars, so a rule that sets only pauseReason keeps the configured timeout policy.

The governing rule is resolved at gate time and persisted on the pending batch (PendingToolCallBatch.effectiveRule), because the persisted batch keeps tool names and sources but no endpoints — an endpoint-addressed rule could not be re-matched after the pause. Each gated call also records the rule that tuned it (matchedRule), so an approver can tell which call brought the batch's policy.

Validation mirrors requireApproval: match goes through the same ToolApprovalPatterns.validate, duplicates are refused, rules without requireApproval is refused, a finite timeoutPolicy with no duration at either level is refused, and a match string-identical to an exempt pattern is refused — an exempt call is never gated, so such a rule could never apply. Only exact equality is refused: a broader rule may legitimately overlap an exemption (http.*:* alongside an exempt http.get:*) while still covering gated calls. The counter eddi.hitl.rule.matched{match="<pattern>"} records which rules actually fire (deduplicated per pause; the tag is the configured pattern, never a URL or argument).

Pattern language

Patterns are matched by ToolApprovalPatterns / ToolApprovalGate:

  • * is the only wildcard — it matches any run of characters (including empty). Every other character is a quoted literal, so compilation is ReDoS-safe.

  • Source-qualified or bare. A pattern may carry a known source prefix (mcp:read_*, http:*) or match the bare tool name (delete_account). A call is tested against the endpoint-qualified form first (where one exists, see below), then source:name, then the bare dispatch name — a tool with an unknown source still matches bare-name patterns (fail-safe).

  • Endpoint-qualified (http tools only). A pattern may also address what a tool calls rather than what it is named: http.post:* matches every POST, http.post:/agentstore/agents matches exactly one endpoint. This is the robust form for tools generated from an OpenAPI spec, whose names come from operationId or a slug and change when the spec does — http.post:* keeps gating every mutation even when a name changes or a new endpoint appears. The path is matched as the httpcall config declares it, normalised to a leading slash (an absolute URL contributes only its path).

  • Known sources (the only accepted prefixes): builtin, http, mcp, a2a, dynamic, memory, recall. An unknown prefix is rejected at save time with a typo suggestion (Levenshtein ≤ 2). Only http may carry a method qualifier: it is the only source whose tools record an endpoint, so mcp.post: is rejected rather than saved as a pattern nothing could match.

  • Allowed characters: A-Za-z0-9_-.:* plus / { } for endpoint path templates. Braces are matched literally, not as regex quantifiers.

  • Case-sensitive. Patterns match tool names exactly; Delete_* does not match delete_account. An endpoint path is matched as the httpcall config declares it, so its case must match too; the method half is normalised to lower case.

  • Patterns may not start or end with :, and are capped at 256 characters.

Precedence

Rule
Behavior

Task combines with agent per eddi.hitl.tool.task-approvals.mode

strict (default): the task block can only STRENGTHEN the agent gate — requireApproval is the union of both lists, task-level exempt entries are ignored (exempt beats require, so a task-added exemption is precisely the ungating vector), task-level AUTO_APPROVE (scalar or rule) is demoted to WAIT_INDEFINITELY unless the agent-level policy itself grants it, and maxAutoApprovalsPerTurn takes the minimum. replace (legacy): a present task block is a full replace of the agent-level block — the pre-6.3.0 behavior, for designs that deliberately loosen one task. See TaskToolApprovalsResolver.

Exempt beats require

If a call matches any exempt pattern, it runs — even if it also matches requireApproval.

Absent/empty requireApproval = gate off

No requireApproval patterns → the gate is fully inactive; every call runs.

Any match suffices

A call is gated if it matches any requireApproval pattern (and no exempt pattern).

Already-approved never re-gated

A callId the human already approved is pre-cleared, so a model that reissues the same call in a continuation does not re-pause on it.

Effective timeout policy

Tool pauses resolve their effective timeout policy with a deliberate rule (ConversationHitlService.applyEffectiveToolTimeoutPolicy):

  1. The governing rule's timeoutPolicy wins — the most specific statement in the config, so it is honoured verbatim (AUTO_APPROVE included).

  2. Otherwise an explicit toolApprovals.timeoutPolicy wins verbatim — including an explicit AUTO_APPROVE (the designer opted in at the tool level). Its approvalTimeout is used, falling back to the outer hitlConfig.approvalTimeout.

  3. Otherwise the outer hitlConfig.timeoutPolicy/approvalTimeout is inherited, except an inherited AUTO_APPROVE is demoted to WAIT_INDEFINITELY — a silent timeout must never auto-execute a gated tool call (that is exactly what the gate exists to prevent).

  4. Absent config on every level leaves the default WAIT_INDEFINITELY.

The duration resolves down the same chain independently (rule → toolApprovalshitlConfig), so a rule that states only a policy still inherits a clock and actually fires.

Key rule: AUTO_APPROVE applies to a tool pause only when set explicitly on toolApprovals. Agent-level AUTO_APPROVE covers RULE pauses but is demoted for tool pauses. Save-time emits a WARN (not a 400) when an agent sets AUTO_APPROVE and a toolApprovals block without its own timeoutPolicy.

Resuming — per-call verdicts and amendments

A tool pause is resumed through the same POST /agents/{conversationId}/resume endpoint. The top-level verdict resolves the whole batch; the optional toolDecisions map (keyed by callId) lets a reviewer decide individual calls and amend an approved call's arguments before it runs. Calls not listed in toolDecisions inherit the top-level verdict.

Validation (all 400, applied before any state changes, so a bad body never consumes the pause):

  • toolDecisions is only valid for a TOOL_CALL pause.

  • An unknown callId (not in the pending batch) is rejected (the body names the valid ids).

  • Each listed call requires a verdict; per-call note1024 chars (top-level note ≤ 4096).

  • Top-level REJECTED combined with any per-call APPROVED is contradictory — set the top-level verdict to APPROVED to mix outcomes.

  • amendedArguments is only valid on an APPROVED call, must be a JSON object, is size-capped, and cannot amend a call whose arguments were truncated at pause time (approve/reject it as-is).

A REJECTED call is not executed; instead the LLM receives a structured rejection tool-result (with the reviewer's note) so it can produce a coherent tool-less answer — rejection feeds back into the model rather than aborting the turn.

pauseDetails for a tool pause

GET /agents/{id}/approval-status returns a TOOL_CALL pauseDetails object (computed at read time; see the pauseDetails shape reference above for the full JSON). Its calls[].arguments is always the redacted, size-capped value (argumentsRedacted) — the raw arguments never appear. executedUngatedCalls names ungated calls in the same batch that already ran (see decision 4). outcomeUnknown lists callIds with an EXECUTING journal entry — a prior approval that crashed mid-execution — and is empty in the common case.

Request pinning — approval binds to a request, not a tool name

For an http-sourced call, the tool name alone tells an approver little: it comes from the endpoint's operationId (or a generated slug) and says nothing about which resource is targeted or with what body. So at gate time each gated httpcall tool is resolvedIApiCallExecutor.resolve builds the method, URL, query, headers and body it would send, without sending it — and both a redacted preview and a SHA-256 fingerprint of that resolved request are persisted on the pause (PendingToolCall.requestPreview, .requestFingerprint). The preview is what an approver should actually look at, not the raw tool arguments.

GET .../approval-status surfaces this: each entry in pauseDetails.calls[] carries requestPinned (boolean) and, when the call could be resolved at all, requestPreview{method, uri, queryParams, headers, body, bodyTruncated}, all already redacted. The raw fingerprint itself is never exposed; it is an internal comparison value with no meaning to a human.

The two fields are independent, and a client must not infer one from the other. requestPinned says whether a fingerprint will be enforced before execution — not whether a preview exists. An http call carrying preRequest.propertyInstructions is previewed best-effort but deliberately left unpinnable, so it arrives with requestPinned: false and a non-null requestPreview: show it, but do not present it as guaranteed to be what runs. requestPreview: null means there was nothing to resolve (every non-http tool), which is not a resolution failure a caller should treat as an error.

On resume, an approved call is re-resolved and its fingerprint re-compared immediately before execution. A mismatch refuses the call — a synthetic {"status":"NOT_EXECUTED","reason":"the request changed after it was approved"} result, an alertable log marker (WARN hitl.tool.request_changed — tool + callId + reason only, never the request itself; not an audit-ledger entry, for the same reason as hitl.tool.outcome_unknown below), and the rest of the batch proceeds normally. This is what makes the approval bind to the request that runs, not to the name of the tool that was called.

Headers are fingerprinted redacted; the body is fingerprinted raw. For headers the fingerprint deliberately covers the redacted form: ApiCallExecutor resolves ${caller:token} into Authorization at execution time, and the approver is routinely a different person than whoever's turn raised the pause — fingerprinting the live header would mismatch on every cross-user approval (the normal case) and the check would have to be disabled. Redacting first makes the fingerprint answer what approval is actually about — what the request does — and leaves whose credentials carry it to authentication, which the fingerprint does not participate in.

The body has no such legitimate variance (${caller:token} is rejected outside headers, and a ${vault:…} reference resolves identically at gate time and at execution), so it is hashed as resolved and only the stored copy is redacted. Redacting before hashing would collapse two different credentials to one marker and therefore to one fingerprint, letting a swapped secret pass the pre-execution re-check unnoticed.

The fingerprint is never returned through the client API — but it is a SHA-256 digest, not encryption, and it is persisted on the pause record. Treat the stored value as sensitive internal data: for a predictable body it supports offline guessing, and equal digests reveal that two requests were identical. It is excluded from every client-facing projection for that reason, not merely because it is meaningless to a human.

Query parameters get the same treatment as the body — hashed as resolved, redacted for display — because ?api_key=… is a conventional way to pass a credential and the query string is shown to the approver too. A repeated parameter (?tag=a&tag=b) is canonicalised as one length-prefixed field per value, so a single value containing the display separator cannot impersonate two.

Body and query redaction are by value shape, not field name — a body is caller-defined JSON (or another format entirely) with no fixed key vocabulary to match on the way headers have. SecretRedactionFilter (the same filter behind argumentsRedacted) removes OpenAI/Anthropic-style keys, bearer tokens and vault references wherever they appear. A hand-rolled secret in a generically named field, matching none of those shapes, is not caught — the same limitation the redacted tool arguments already carry, and the reason a config write that must carry a credential belongs behind a vault reference rather than a literal.

A call can be unpinned, and that is a deliberate degrade, not a bug. A call is left unpinnable whenever execute() could legitimately build a request that resolve() did not — the guard is "never pin what cannot be honoured", so the set is defined by that property rather than by a list of features:

Unpinnable when

Why execute() can diverge from resolve()

The tool is not http (builtin/mcp/a2a)

There is no HTTP request on this side of the boundary to pin.

preRequest.propertyInstructions is set

Those write to conversation memory, so resolving them ahead of execution would apply them twice.

fireAndForget and preRequest.batchRequests

The batch expands at execution time into N requests, none of them the single one that was previewed.

postResponse.retryApiCallInstruction with maxRetries >= 1

buildRequest sits inside the retry loop, and each attempt re-renders templates against a memory that the previous attempt wrote to ({…Error}, {…HttpCode}, {responseObjectName}). Attempt 2 is a request nobody previewed.

The retry row is the easy one to trip over: RetryApiCallInstruction.maxRetries defaults to 3, so "postResponse": {"retryApiCallInstruction": {}} is by itself enough to unpin an otherwise-pinnable write — and a retry can fire on a 2xx response when responseValuePathMatchers matches, not only on retryOnHttpCodes. Read requestPinned per call; do not infer it from the endpoint.

An unpinned call (requestFingerprint == null) is approved on name and arguments alone, exactly as before pinning existed — nothing is ever refused on a comparison that was never sound. An amended call (amendedArguments set) is likewise never fingerprint-checked: the approver rewrote the request themselves, so the pin describes the request they replaced.

Three situations fail closed instead — refused, not silently allowed — because "cannot verify" is a different answer than "unchanged": the tool disappeared from the workflow between pause and resume, re-resolution throws, or the call's config gained any of the unpinnable properties above mid-pause (a pinned call becoming unpinnable — adding propertyInstructions, or a retryApiCallInstruction, while a human is deciding). Treating any of these as "unchanged" would make reconfiguring an agent while a human is deciding the way around the guard.

The execution journal (at-most-once)

Approved tool executions are protected by a write-ahead journal (IHitlToolJournalStore) so a human approval is executed at most once, across pod crashes and re-approvals:

  1. Before running an approved tool, tryClaim inserts an EXECUTING entry keyed by (conversationId, pauseEpoch, callId). The pauseEpoch is a per-pause UUID — providers may reuse tool-call ids across different pauses in one conversation, so the epoch keeps them distinct.

  2. After the tool returns, markExecuted stores the capped result.

  3. On resume, an EXECUTED entry replays its stored result — the tool is never re-run.

  4. An EXECUTING entry (a crash inside the tool) yields an honest EXECUTION_OUTCOME_UNKNOWN result fed to the LLM: "a previous execution attempt was interrupted; it may or may not have taken effect — verify externally before retrying." This is logged with an alertable marker (WARN hitl.tool.outcome_unknown, not an audit-ledger entry — the orchestrator has no audit collector in scope here) and surfaced in pauseDetails.outcomeUnknown.

Outcome-unknown contract: the framework never silently re-executes a tool that may have already run. A crash between "claimed" and "executed" is reported as genuinely unknown, not retried — the human (or a downstream check) decides what to do.

Slack

The Slack integration is tool-pause-aware. When hitlApprovalChannel is set, a TOOL_CALL pause posts an interactive approval card that additionally renders one context block per pending call — toolName plus its redacted, 300-char-truncated arguments (max 5 calls, then +N more) — so a reviewer sees what they are approving. The raw argument value is never accessed. The in-thread pause notice stays pause-reason-only. Slack buttons are all-or-nothing in v1 (Approve/Reject the whole batch); per-call verdicts and amendments require the REST/MCP resume body. See Slack Integration.

Group members

A group has no per-member human reviewer, so a member agent's gated tool call inside a group turn is resolved gracefully, not stranded: the framework issues a system:group REJECTED decision through the normal resume path (note: "tool approval is not available during group discussions in this version"), the member's LLM receives the rejection tool-result and produces a coherent tool-less contribution, and that becomes the turn's output. If the resume cannot complete in the member-turn budget (or the member re-pauses), it falls back to the existing member-pause handling (turn recorded SKIPPED, stranded pause auto-cancelled as system:group). The inGroupTurns: "INBOX" mode (routing member tool pauses to a human inbox instead) is reserved and rejected at save time in v1.

Frozen-transcript semantics

A tool pause serializes the exact in-flight langchain4j message list (the AiMessage + prior tool results of the current LLM loop) at pause time and persists it on the snapshot. On resume — even days later — the loop re-enters the same task index and replays that frozen transcript, then applies the verdicts and continues. The human therefore approves against, and the model resumes against, pause-time prompt state — not a transcript rebuilt from current memory. This is why the design persists the transcript rather than reconstructing it: a multi-iteration turn rebuilt from memory would lose intermediate tool results.

Rolling upgrade

The gate is guarded by a feature flag: eddi.hitl.tool.enabled (default true; injected into LlmTask). When false, the effective tool-approvals config resolves to null and the gate is inert.

  • Enabling on a cluster: complete the rollout of the new version to all pods before enabling gates in agent configs. A pod running the previous version cannot read a TOOL_CALL-paused conversation (it does not understand the AWAITING_HUMAN/TOOL_CALL state), so mixed-version pods should not be producing tool pauses.

  • Downgrading: set eddi.hitl.tool.enabled=false and drain existing tool pauses (resolve or cancel them) before rolling back to a version without this feature — otherwise those paused conversations are unreadable by the old code.


Timeout Policies

Policy
On timeout
Typical use case

WAIT_INDEFINITELY

Nothing is scheduled; waits forever

Critical decisions (compliance, safety)

AUTO_APPROVE

APPROVED decision (decidedBy: system:timeout), resumes

Progress monitoring, non-critical gates

AUTO_REJECT

REJECTED decision → regular: turn skipped; group: FAILED

Strict SLA gates

ABORT

Cancels the conversation/discussion

Safety-critical pipelines

Timeouts are one-shot schedules on EDDI's cluster-aware ScheduleFireExecutor — exactly one pod owns each claim, and they route through the same resume/cancel machinery as human decisions. Delivery is at-least-once (a lease-expired claim can be re-stolen after a pod crash), but that is safe: the timeout resolves through the resume/cancel state CAS, so a duplicate fire against an already-resolved conversation is a no-op.


Who May Decide

Approve/reject/cancel/status require, in order: the conversation owner, the eddi-admin role, or the dedicated eddi-approver role. Ownerless (legacy/internal) conversations fail closed for non-admin/non-approver callers. decidedBy is always set server-side from the authenticated identity — it cannot be spoofed via the request body. Approvers and admins see all entries in the pending-approvals listings; other users only their own.

Approver read scope: the approver role exists to decide pending approvals, not as a universal read grant. Approver-only callers (not owner, not admin) may read approval-status?detail=full — the full memory snapshot / transcript — only while the conversation is actually awaiting approval; outside a pause they get 403 and can use the summary view.

Auditing

Every HITL decision — human or automated — writes an immutable hitl.approval entry to the audit ledger (verdict, decider, automated flag, note), and the resumed pipeline leg records its task entries like any other turn. Cancelling a pending approval is a decision too: it writes an hitl.approval entry with verdict CANCELLED, the cancel mode, and the cancelling actor (decidedBy: the authenticated principal, or a system:* identifier such as system:timeout for ABORT policies and system:group for auto-cancelled member pauses). This is the EU AI Act human-oversight trail.

Crash Recovery

On startup, HitlCrashRecoveryObserver repairs HITL state; it never destroys a legitimate pause. The sweep runs on a background thread (application readiness does not block on it) and reads bounded projected summaries, not full conversation documents:

  • Paused conversations with a finite policy get their one-shot timeout schedule idempotently re-armed at the original due time (or shortly after startup if overdue). After creating the schedule the pause state is re-checked — if a resume/cancel landed in the window, the schedule is withdrawn.

  • WAIT_INDEFINITELY pauses are left untouched (and skipped without any further read).

  • Regular conversations stuck IN_PROGRESS with an intact HITL bookmark (pod died mid-resume) are restored to AWAITING_HUMAN so the approval can be re-issued.

Config: eddi.hitl.crash-recovery.enabled (default true), eddi.hitl.crash-recovery.recover-in-progress (default true; consider false in multi-pod deployments — see the class Javadoc for the rolling-restart caveat).

Operations

  • Metrics (/q/metrics): eddi_hitl_pause_count, eddi_hitl_resume_count, eddi_hitl_timeout_count, each tagged surface=regular|group; eddi_group_member_pause_skipped_count for auto-cancelled member pauses inside groups; eddi.operator.write.approval{decision=approved|rejected|timeout}, one per gated call the instant its verdict is resolved (timeout is a distinct bucket from approved/rejected — see Request pinning above; not operator-specific despite the name, it fires for any gated call regardless of which agent).

  • Operator canary/gate metrics (client-reported): the write canary and gate-installed check both run entirely client-side (there is no server-side notion of "the operator", just an agent with a particular hitlConfig), so the Manager reports outcomes via POST /administration/operator/{canary-result,gate-status} (eddi-admin only) purely so they show up on /q/metrics without a Manager tab open. This is a relay, not a verification — a report is trusted at face value. Produces eddi.operator.canary{outcome=pass|fail|unknown}, eddi.operator.canary.duration, and the gauge eddi.operator.gate.verified (1 only while every provisioned version last read back with a sound gate; defaults to 0, including on a deployment that has never activated an operator — that ambiguity is real and unresolved by this signal alone).

  • Undeploy: paused conversations do not count as active — an agent version with pending approvals can be undeployed. Resuming afterwards returns 409 agent not deployed and the pause is restored (redeploy, then retry). The idle-conversation cleanup sweep likewise spares AWAITING_HUMAN conversations — a pending approval is never silently force-ended by maintenance.

  • Cancel semantics (regular): cancels a paused conversation, or signals a turn executing on the same pod to stop at the next task boundary. CANCEL_IMMEDIATE currently degrades to graceful on the regular surface. Cancelling an idle conversation returns 409 (use endConversation).

  • Timeout schedules are not manually operable: HITL timeout schedules live in the schedule store but the schedule REST surface refuses to fire them manually (409, use /resume or /cancel — manual firing would bypass the approval authz), restricts update/delete/enable/disable to eddi-admin (403 otherwise, so an editor cannot disarm an ABORT/AUTO_REJECT safety timeout), and redacts them from non-admin listings.

  • Pause retention (optional): eddi.hitl.pending.max-age (ISO-8601 duration, default off) auto-cancels pauses older than the threshold — audited, schedule-disarmed, via the normal cancel path; eddi.hitl.pending.sweep-interval (default 6h) controls the sweep cadence. Under the default WAIT_INDEFINITELY policy, pauses otherwise accumulate until decided.

  • Mass-timeout drain: due schedules are claimed in configurable batches (eddi.schedule.poll-batch-size, default 100) and fired concurrently on virtual threads with per-fire error isolation.

Slack Integration

The Slack channel integration is HITL-aware end to end. Configuration lives in the Slack ChannelIntegrationConfiguration.platformConfig (both keys optional):

Key
Description

hitlApprovalChannel

Slack channel id that receives approval notifications when a conversation pauses.

hitlApproverUserIds

Comma-separated Slack user ids allowed to decide via buttons. Without this list, buttons are not rendered and interactive decisions are rejected — fail-closed.

Behavior:

  • In-thread pause notice — when a turn pauses, the thread receives the output so far plus "⏸️ awaiting human approval" with the pauseReason. Messages sent while paused get "Still awaiting approval" instead of a generic error (the input is not consumed).

  • Approval inbox message — if hitlApprovalChannel is set, an interactive Block Kit message is posted there (conversation, agent, reason, timeout deadline) with Approve/Reject buttons. For a RULE pause only the pause reason is shared — no conversation content (data minimization). For a TOOL_CALL pause the approval channel additionally renders one context block per pending call — toolName plus its redacted, size-capped arguments truncated to 300 chars for display (max 5 calls, then a +N more line) — because a reviewer cannot responsibly approve transfer_funds without seeing the amount. The raw (unredacted) argument value is never accessed; this relaxation applies to the approval channel only, not the in-thread notice (which stays pause-reason-only).

  • Deciding from Slack — buttons post to POST /integrations/slack/interactive (configure this as the Slack app's Interactivity Request URL; requests are signature-verified like the events webhook). The acting Slack user must be in hitlApproverUserIds; decisions are attributed decidedBy: slack:<userId> in the audit trail. Double-clicks and already-decided races update the message ("already resolved") without error spam. Group pauses work the same way (the button value carries group:<conversationId> and routes to the group approve machinery).

  • Continuation push — when the pause is resolved (human, Slack button, or system:timeout), the originating thread automatically receives the verdict and the resumed conversation's output — no polling. This works cross-pod: the resume fires an async CDI event (HitlResumeCompletedEvent) and the Slack observer resolves the thread from the persistent conversation mapping.

  • Group discussions — Slack-driven group discussions post pause/resume/cancel notices (and member_pause_skipped explanations) into their thread.

Delegated/managed conversations (agent-to-agent tools, MCP chat_managed): when a nested conversation pauses, the calling tool receives a structured PAUSED_FOR_APPROVAL result naming the conversation and reason instead of hanging — the delegated approval stays pending (it is not auto-cancelled; a reviewer decides, then the tool can be re-invoked).

MCP Surface

The HITL operations are also exposed over the MCP server (McpHitlTools), so an MCP client that drives a conversation (talk_to_agent / chat_managed) and receives a PAUSED_FOR_APPROVAL result can resolve the gate over the same transport instead of switching to REST. That paused payload names the tool to call next ("suggestNextTool": "resume_conversation").

MCP tool
Mirrors REST
Notes

list_pending_approvals

GET /agents/pending-approvals

Owner-scoped; includes RULE and TOOL_CALL pauses

get_approval_status

GET /agents/{id}/approval-status

Summary reports pauseType; detail=full returns the snapshot (incl. any pending tool-call batch)

resume_conversation

POST /agents/{id}/resume

A single verdict resolves both RULE and TOOL_CALL pauses

cancel_conversation

POST /agents/{id}/cancel

list_group_pending_approvals

GET /groups/{groupId}/conversations/pending-approvals

Owner-scoped

list_all_group_pending_approvals

GET /groups/pending-approvals

Cross-group inbox

get_group_approval_status

GET /groups/{groupId}/conversations/{gcId}/approval-status

Summary only; detail=full returns the whole conversation

approve_group_phase

POST /groups/{groupId}/conversations/{gcId}/approve

Optional taskApprovals JSON for TASK granularity; blocks and returns the resumed discussion (no SSE variant over MCP)

cancel_group_discussion

POST /groups/{groupId}/conversations/{gcId}/cancel

Authority is identical to REST — the shared HitlAccessGuard enforces owner / eddi-admin / eddi-approver per conversation (fail-closed on a missing descriptor), and the decision is attributed server-side to the authenticated caller, prefixed mcp: (e.g. mcp:alice, mirroring the system:timeout convention) — never taken from a tool argument. Errors are structured JSON (errorCodeNOT_FOUND | WRONG_STATE | FORBIDDEN | DISABLED | BAD_REQUEST | CONFLICT | INTERNAL) so a programmatic client can branch on the failure kind.

Kill-switch: set eddi.mcp.hitl.mutations.enabled=false to make MCP a read-only HITL surface — resume_conversation, cancel_conversation, approve_group_phase, and cancel_group_discussion then return a DISABLED error while the read-only list/status tools keep working. MCP is a transport, not a new authority: no tool lets an agent approve its own gate.

Known Limitations (v1)

  • Member-level HITL inside a group: a member agent's own PAUSE_CONVERSATION rule firing during a group turn is not supported — the turn is recorded as SKIPPED with an explanatory note, the member's stranded pause is auto-cancelled (audited as system:group), and a member_pause_skipped SSE event is emitted. Use group-level HITL (requiresApproval) instead.

  • Nested groups: requiresApproval inside a sub-group of a group-of-groups is not supported — the sub-pause is cancelled and the member turn is recorded as SKIPPED with an explanatory note.

  • Cross-pod cancel of an actively-running turn: the cooperative cancel flag is per-pod; the DB CAS covers paused states cluster-wide, and a running group leg re-checks the persisted state at every phase boundary, so a cross-pod cancel takes effect at the next boundary.

  • Tool-level HITL (pausing when the LLM invokes a gated tool) is implemented — see Tool-Level Approval Gating. Group-member tool pauses are auto-rejected gracefully (they have no reviewer); the inGroupTurns: "INBOX" mode is reserved (rejected at save time in v1).

  • Managed REST conversations (RestAgentManagement): a managed intent mapping stays pinned to a paused conversation until it is resumed or reset — a managed say against it answers 409 promptly, but the mapping is not automatically re-created while the approval is pending.

  • Upgrade notes:

    • DiscussionPhase.requiresApproval existed before this feature as an inert placeholder. Stored group configs that set it to true will begin pausing for approval after this upgrade.

    • PAUSE_CONVERSATION is a newly reserved action name. A stored behavior-rule config that already emitted an action with this exact name will begin pausing conversations after this upgrade — rename such actions before upgrading.

    • During a rolling upgrade, pods running the previous version cannot read conversations paused by new pods (unknown AWAITING_HUMAN state) — schedule pause-heavy traffic after the rollout completes.

Last updated

Was this helpful?