Scheduled Execution & Heartbeats
Overview
EDDI supports scheduled agent execution — agents can be triggered automatically on a timer without any user input. This enables proactive agents, background maintenance, periodic data processing, and memory consolidation.
Use Cases
Proactive Agents
Check for updates, send notifications, or perform monitoring at regular intervals
Dream Consolidation
Background memory maintenance — prune stale entries, detect contradictions, summarize facts
Data Pipelines
Periodically fetch data from external APIs and process it through the agent pipeline
Health Checks
Run diagnostic agents that verify system health and report anomalies
Report Generation
Generate daily/weekly summary reports through conversational agents
Concepts
Schedule
A Schedule defines when and how often an agent fires:
{
"agentId": "agent-123",
"agentVersion": 0,
"triggerType": "CRON",
"cronExpression": "0 2 * * *",
"conversationStrategy": "persistent",
"message": "Run maintenance cycle",
"userId": "system:scheduler",
"timeZone": "Europe/Vienna",
"enabled": true
}Trigger Types
CRON
Wall-clock aligned cron expression
new
0 2 * * * (daily at 2am)
HEARTBEAT
Fixed-interval, drift-proof
persistent
Every 300 seconds
A CRON schedule carrying oneTimeAt instead of cronExpression fires once at that instant rather than recurring. Exactly one of the two is required.
Conversation Strategies
persistent
Reuses the same conversation across all fires. Context accumulates.
Dream consolidation, ongoing monitoring, stateful agents
new
Creates a fresh conversation for each fire. Clean context each time.
Report generation, data pipelines, stateless tasks
Configuration
Creating a Schedule
Cron Expression Reference
EDDI uses standard 5-field cron expressions:
Common patterns:
0 2 * * *
Daily at 2:00 AM
*/30 * * * *
Every 30 minutes
0 */4 * * *
Every 4 hours
0 9 * * 1-5
Weekdays at 9:00 AM
0 0 1 * *
First day of each month at midnight
Heartbeat Configuration
For heartbeat triggers, use heartbeatIntervalSeconds instead of cronExpression:
Heartbeats are drift-proof — the next fire is the time this fire was due plus the interval, not the moment the turn happened to finish. A 40 s turn on a 60 s heartbeat still fires every 60 s. (lastFired + interval would be the drifting formula: lastFired is the completion instant. The one exception is a fire that overran a whole interval — anchoring on the due time would put the next fire in the past, which is a re-fire loop rather than catching up, so it is clamped to now + interval.)
Schedule Fields
agentId
string
required
Agent to trigger
agentVersion
int
0 (latest)
Agent version (0 = latest deployed)
triggerType
enum
CRON
CRON or HEARTBEAT
cronExpression
string
—
5-field cron (for CRON type)
heartbeatIntervalSeconds
long
—
Interval in seconds (for HEARTBEAT type)
conversationStrategy
string
varies
new or persistent
message
string
—
Message text sent to the agent on each fire
userId
string
system:scheduler
User identity for the fire
timeZone
string
UTC
IANA timezone (e.g., Europe/Vienna)
environment
string
production
Deployment environment
enabled
boolean
true
Whether the schedule is active
maxCostPerFire
double
-1 (unlimited)
Dollar ceiling per fire
oneTimeAt
string
—
ISO-8601 instant for a single fire. Mutually exclusive with cronExpression; exactly one of the two is required for a CRON trigger
metadata
object
—
Free-form markers read by the fire executor. {"dreamType": "dream_consolidation"} dispatches the fire to the Dream service — see Scheduling a Dream Cycle
Managing Schedules
POST
/schedulestore/schedules
Create a schedule
GET
/schedulestore/schedules
List schedules, newest first (optional ?agentId= filter; ?limit= default 500, max 1000; ?offset= default 0)
GET
/schedulestore/schedules/{id}
Get a specific schedule
PUT
/schedulestore/schedules/{id}
Update a schedule
DELETE
/schedulestore/schedules/{id}
Delete a schedule
POST
/schedulestore/schedules/{id}/enable
Enable a schedule
POST
/schedulestore/schedules/{id}/disable
Disable a schedule
POST
/schedulestore/schedules/{id}/fire
Manually trigger a fire immediately
Paging (wire change). The listing used to be one hard-capped page of 500 in whatever order the store returned; past that, the surplus schedules could not be found, disabled or deleted through the list at all. It is now ordered (newest
createdAtfirst, id breaking ties) and takeslimit/offset. A response holding exactlylimitentries may be truncated — ask for the next page to find out.
limit=0is now400, on all three listing endpoints. It used to be passed through to the store, where the two backends read it opposite ways: the MongoDB driver treatslimit(0)as no limit and dumped every row, while PostgreSQL'sLIMIT 0returned nothing. A client that sentlimit=0and got away with it must send a positive value.Firing a one-shot consumes it.
POST /{id}/fireruns the same state machine a polled fire does, so a successful manual fire of aoneTimeAtschedule disables it — it is the run, not a rehearsal. Re-arm it withPOST /{id}/enable.Firing a heartbeat manually consumes its next scheduled fire. Same reason: a successful fire re-arms the schedule from the fire it was due to make, so firing a daily heartbeat by hand in the morning moves the next one to a day after that due time — tonight's run is skipped, not brought forward.
A skipped manual fire does not. If the coordinator drops the turn because the conversation is busy or
AWAITING_HUMAN, nothing was delivered, so nothing is consumed: a due time still in the future is left exactly where it was and tonight's run happens as configured. Only a due time that has already passed is rolled forward to the next cadence.A manual fire is synchronous: the request holds open until the turn finishes or
eddi.schedule.fire-timeout(default 5 minutes) elapses, so a proxy or client with a shorter read timeout may give up before the fire log comes back. The fire itself continues, and the schedule stays claimed until it ends — a retry in the meantime answers409.
Admin Endpoints
GET
/schedulestore/schedules/{id}/fires
Read fire history, newest first (?limit= default 20, must be > 0, capped at 500)
GET
/schedulestore/schedules/admin/failed
List all failed/dead-lettered fires (?limit= default 50, must be > 0, capped at 500)
POST
/schedulestore/schedules/{id}/retry
Re-queue a dead-lettered schedule
POST
/schedulestore/schedules/{id}/dismiss
Reset dead-letter without immediate retry
Dream Consolidation
Dream Consolidation is a specialized schedule that performs background memory maintenance on an agent's persistent user memories. It's configured in the agent's UserMemoryConfig, not as a standalone schedule.
What It Does
Stale entry pruning — Removes outdated facts that are no longer relevant
Contradiction detection — Identifies conflicting memories (e.g., "user likes coffee" vs "user hates coffee") and logs them for review. Resolution is planned for a future version.
Fact summarization — Consolidates verbose entries into concise summaries
Configuration
Dream consolidation is configured in the agent configuration:
Scope: a dream cycle only touches memories the firing agent wrote (
sourceAgentId). SetcrossAgentMaintenance: trueto maintain the user's whole memory set across agents — without it, agent A'spruneStaleAfterDayswould delete agent B's memories and A's model endpoint would see B's private text.
maxSummarizationCallsis deprecated in favour ofmaxCostPerRun(a call count is a poor budget — consolidations differ wildly in cost). It is still honoured as a secondary backstop if a stored config sets it explicitly, so existing configurations keep their bound.
Cost Control
Dream cycles consume LLM tokens. Use maxCostPerRun (in the Agent Configuration) to set a dollar ceiling per run:
When the budget is exceeded, the agent stops processing. This prevents runaway costs on large memory stores.
Tip: Use a cheaper model (e.g.,
claude-sonnet-4-6orgpt-4o-mini) for dream consolidation — the task doesn't require top-tier reasoning.
Fire Logging
Every scheduled execution is logged. View fire history via the REST API:
What
costmeans depends on the fire path, and the two are not the same quantity. A conversation fire reports the tool spend of that turn — theToolCostTrackerdelta — so a schedule whose agent only talks to the model, with no tool calls, reports0.00however many tokens it used. A dream consolidation fire reports its own estimated LLM cost. Compare a fire log against others on the same path, and usemaxCostPerFire/maxCostPerRunrather than the logged number to bound spend.
Fire Logs and Erasure
A fire log carries the conversationId of the turn it started, and it is only findable by its scheduleId — so a fire log whose schedule has been deleted is personal data that nothing can reach again. Deleting a schedule therefore always deletes its fire logs, on the single-schedule path and on all three bulk paths (by agent, by name, and the GDPR erasure by user).
Two mechanisms keep that true even while the schedule is firing:
The write is conditional. A fire log is stored only if its schedule still exists at the moment of the write — on PostgreSQL an
INSERT … WHERE EXISTS (SELECT 1 FROM eddi_schedules WHERE id = ?), on MongoDB (which has no conditional insert) an insert that is verified against the schedule immediately afterwards and removed again if it has gone. A fire in flight when an erasure runs simply writes no log. The fire itself is unaffected; only its log is dropped.The delete sweeps twice. The cascade removes the logs that exist when it runs — on PostgreSQL in the same transaction as the schedule delete — and a second indexed pass runs after the schedules are gone. That remains the belt-and-braces for a log written by an instance that had not yet observed the delete.
The consequence for operators: a schedule deleted mid-fire may lose the fire log for that one attempt. That is deliberate — the alternative is an unreachable record of an erased user's conversation.
State Machine
Each schedule follows a state machine:
SKIPPED is a fire-log status only — it is never stored on the schedule itself. It means the coordinator dropped the scheduled turn without consuming the input because the conversation was already busy or AWAITING_HUMAN. That is the normal state of a persistent heartbeat while a human is chatting in its conversation, or while a previous fire waits on a HITL approval, so a skip is logged and counted but never enters the retry/backoff/dead-letter machine: failCount is left exactly as it was and the claim is released.
Re-arming is deliberately conservative, because a skip delivered nothing:
Due time already passed (every polled skip, and a manual fire of an overdue schedule) → advanced to the next cadence, on the same drift-proof anchor a successful fire uses.
Due time still in the future (only reachable through
POST /{id}/fire, which claims regardless ofnextFire) → left untouched. The pending fire is the next cadence, so advancing past it would silently cancel a scheduled delivery that nothing replaced.A one-shot whose moment has passed has no cadence to re-arm to, so it does go through retry/backoff — its single delivery genuinely never happened.
Cluster Awareness
The SchedulePollerService is cluster-aware — in multi-instance deployments, only one instance executes each scheduled fire. This is achieved via atomic claim operations (tryClaim), preventing duplicate execution when running EDDI behind a load balancer.
Two consequences are worth stating plainly, because they shape how a scheduled target must be written:
Claiming is per-lease compare-and-set, so exactly one instance wins a given poll. But delivery is at-least-once: if the winner dies mid-fire, the lease expires (
eddi.schedule.lease-timeout) and another instance re-claims the same fire. Scheduled targets must be idempotent.Instances identify themselves by
eddi.schedule.instance-id, which is auto-derived from the hostname when left empty. In an environment where hostnames are recycled or duplicated (some container schedulers), set it explicitly — two instances sharing an ID makes claim ownership ambiguous.
Deployment Configuration
Individual schedules are configuration documents; the poller that runs them is tuned deployment-wide in application.properties (or the matching environment variables — Quarkus maps eddi.schedule.poll-interval to EDDI_SCHEDULE_POLL_INTERVAL — every non-alphanumeric character becomes _).
eddi.schedule.enabled
true
Master switch. false stops all polling — schedules remain stored and simply never fire
eddi.schedule.poll-interval
15s
How often each instance looks for due schedules. This is the floor on firing punctuality: a schedule due at 12:00:00 fires somewhere in [12:00:00, 12:00:15)
eddi.schedule.poll-batch-size
100
Max schedules claimed per poll cycle. Claimed schedules dispatch concurrently on virtual threads; raise it to drain large bursts (e.g. many one-shot HITL approval timeouts expiring together)
eddi.schedule.lease-timeout
5m
How long a claimed schedule is considered owned before another instance may re-claim it. Set it comfortably above your longest fire, or a slow run gets executed twice
eddi.schedule.max-retries
5
Attempts before a fire is DEAD_LETTERED
eddi.schedule.backoff-base-seconds
15
Retry delay = base × multiplier^(attempt-1) seconds
eddi.schedule.backoff-multiplier
4
With the defaults: 15s, 60s, 4m, 16m, 64m
eddi.schedule.min-interval-seconds
60
Smallest cron interval a schedule may request. Guards against schedule bombing; a rejected create returns a message naming this property
eddi.schedule.instance-id
(hostname)
Identity used for cluster claim tracking
eddi.schedule.default-timezone
UTC
IANA zone applied to schedules that do not name one
eddi.schedule.fire-timeout
5m
How long one conversation fire may run before it is abandoned as failed. Keep it at or below lease-timeout — past the lease another instance may reclaim the schedule regardless
eddi.schedule.fire-log-retention
90d
Fire logs older than this are deleted by a periodic sweep. 0 keeps everything — note that a 60-second heartbeat alone writes ~525,600 rows a year
eddi.schedule.fire-log-prune-interval
1h
How often that sweep runs. The DELETE is by timestamp and therefore idempotent, so it needs no cluster claim
Observability
eddi.schedule.poll.count
Counter
Poller liveness. Flat means the poller is not running — check eddi.schedule.enabled
eddi.schedule.fire.count
Counter
Fires executed
eddi.schedule.fire.failed
Counter
Fires that raised. Compare against fire.count for a failure rate
eddi.schedule.fire.skipped
Counter
Fires dropped because the target conversation was busy or awaiting a human. Not failures and never dead-lettered, but a heartbeat that only ever skips is delivering nothing — compare against fire.count
eddi.schedule.fire.deadlettered
Counter
Fires that exhausted max-retries. Alert on any increase — these need manual retry or dismissal
eddi.schedule.fire.duration
Timer
If p99 approaches lease-timeout, double execution is imminent
eddi.schedule.claim.conflict
Counter
Instances racing for the same schedule. Normal and expected in a cluster; a sharp rise alongside falling fire.count suggests contention rather than work
eddi.schedule.firelog.pruned
Counter
Fire logs removed by the retention sweep. Flat while the table grows means retention is disabled (fire-log-retention=0) or the sweep is failing — check the logs
Best Practices
Start with longer intervals — Begin with hourly or daily schedules and increase frequency only if needed
Use
persistentstrategy for stateful work — Dream consolidation and monitoring agents benefit from accumulated contextSet cost ceilings — Always configure
maxCostPerFireormaxCostPerRunfor LLM-powered scheduled tasksMonitor fire logs — Check for recurring failures that might indicate configuration issues
Use cheap models for maintenance — Background tasks rarely need expensive frontier models
See Also
Managed Agents — Intent-based agent routing
User Memory — Persistent user memory (target of dream consolidation)
LLM Configuration — Agent configuration reference
Metrics — Monitoring scheduled execution performance
Last updated
Was this helpful?