A distributed job scheduler receives instructions such as "run this task every day at 02:00 UTC" and ensures that the task is triggered at the intended time. On a single host, this is relatively simple. With multiple scheduler and worker machines sharing responsibility, the system must survive failures without losing a scheduled run or committing the same run twice.
Design a scheduler that supports immediate, one-time, and recurring jobs. Throughout the design, consider the following central problem: if multiple machines determine that one job is ready, how can the system create exactly one durable record for that scheduled occurrence and recover it safely when a failure happens?
2 seconds of its scheduled fire time.(job_id, scheduled_fire_time) must correspond to one durable slot and no more than one committed success.10k execution starts per second on a sustained basis, with 5-10x burst skew at minute or hour boundaries.| Metric | Assumed value | Used for |
|---|---|---|
| Peak execution starts | ~10k/s sustained | Worker pool, queue throughput, execution-state write rate |
| Cron burst skew | 5–10× at minute or hour boundaries | Shard sizing, smoothing, catch-up surge |
| Scheduler shard count | on the order of 100–1000 logical segments | Coordinated polling and horizontal scale |
| Normal dispatch lag target | ~2s | Sweep cadence, near-time queue horizon, autoscaling triggers |
| Worker timeout range | seconds to tens of minutes | Lease renewal, heartbeat interval, long-running policy |
| Execution history retention | 30–90 days hot in primary store | Indexing, partitioning, archive boundary |
Treat these figures as planning assumptions. They provide concrete inputs for reasoning about partitioning, queue depth, and recovery behavior in the sections that follow.
[!IMPORTANT] The central tension is that many scheduler nodes can independently discover that a job is due, while only one durable record for that scheduled occurrence should reach a final result. Keep separate the description of what should execute, the identity of the machine currently scheduling it, and the worker authorized to complete it.
Out of scope (below the line)