Design a resilient system that executes tasks on a user‑defined schedule—similar to a cron service—and remains correct even when the scheduler itself or its workers crash. Your explanation must cover how a schedule is parsed, how the next run time is computed in a timezone‑aware manner, how individual task instances are created and dispatched, how workers claim and complete tasks, and how the system recovers from outages while respecting the desired handling of runs that were missed during downtime.
Example Scenario
A user creates a schedule to run a report every weekday at 9:00 AM US Eastern. During a network partition, the scheduler is offline from Tuesday 8:50 AM to Wednesday 10:00 AM. When it comes back, the system must decide what to do with the missed Tuesday 9:00 AM and Wednesday 9:00 AM windows, and it must not double‑schedule a run that was already created before the outage.
Design Requirements
Time‑based triggering
Schedules are defined with an explicit timezone (e.g., America/New_York). The design must address how daylight‑saving‑time transitions (spring‑forward gaps, fall‑back repeats) influence the next‑run calculation, and how editing a schedule affects already‑computed futures.
Occurrence identity and idempotency
Each logical scheduled moment must have a durable, unique identifier that persists across restarts, so that the scheduler can recognize a previously created occurrence and avoid duplicate dispatches. A worker might complete the external side‑effect and then crash before registering success; clarify how idempotent handling prevents double‑execution without promising exactly‑once delivery of external actions.
Misfire policy
When a scheduled time has passed without being acted on (e.g., after an outage), the system must implement a clean misfire strategy: skip, coalesce, or replay each missed occurrence individually. The choice must be explicit and configurable.
Concurrency and leasing
Workers must acquire a lease or claim on an occurrence before attempting it, and the design should describe how simultaneous attempts (including overlapping runs of the same scheduled job) are either permitted or prevented. Retry logic for transient failures and dead‑letter handling for permanent failures should be outlined.
High‑availability scheduling
The scheduling logic must be robust against multiple scheduler instances starting simultaneously, ensuring that no two instances create the same logical occurrence. This typically involves an atomic “grab next window and advance” operation, often backed by a durable store.
Operational visibility
The design should mention how to provide observability into backlog depth, pending occurrences, and how to safely pause or modify a schedule without corrupting existing work items.
Hint
A restarting scheduler may re‑examine the same time window and attempt to produce an occurrence that logically already exists. Assign each occurrence a stable identity and enforce a uniqueness constraint so that the second attempt is silently recognized as already‑done rather than creating a duplicate.
What a strong design will cover
Follow‑up scenarios to address
Goal
A durable cron‑like scheduler with timezone‑aware planning, configurable misfire policies, idempotent dispatch via worker leases, and reliable recovery from both scheduler and worker failures.