Requirements
- Schedule and run jobs of a given type. Each job has an SLA; if it exceeds the deadline, raise/record an error and let the user inspect logs.
- Support recurring jobs (e.g. a job that runs every minute).
- Guarantee at-most-once execution per scheduled run.
- Handle a broad range of failure cases gracefully — this is where most of the follow-up time goes.
Deep-dive directions interviewers push on:
-
What happens when a worker or a component in the pipeline dies mid-execution? How do you still honor at-most-once and continue retrying?
-
A recurring job that fails must still send a notification — design how the failure path reliably triggers the notification even when parts of the system are down.
-
Retry strategy and backoff; how retries interact with the at-most-once guarantee.
-
Explain deduplication and how to avoid duplicate execution.
-
Design data sharding and the database schema, including partition-key and sort-key choices.
-
Explain what happens when job status is not updated and how to prevent jobs from being lost.
-
Justify using a Redis queue and explain what happens if that queue fails.
-
Explain how updates are pushed or broadcast to subscribers.
Notes
- The round is open-book and the prompt is stable, so the front half carries little differentiation. Proactively and thoroughly cover reliability and fault tolerance early; candidates who leave it thin get walked into follow-up traps and run out of time.
- A common at-most-once implementation is conditional writes (compare-and-set on a claim/lease record) so two workers can't both execute the same run.
- A frequently added constraint is that the scheduler only handles a single job type — don't over-generalize the model.
Preparation
- Build a reference design end-to-end (queue + workers + a durable job/lease store + a notifier) and rehearse the at-most-once claim via conditional writes.
- Pre-script the failure-handling story: worker crash, lost notification on recurring-job failure, and retry/backoff — these are the graded follow-ups.
- Work the Hello Interview / ByteByteGo job-scheduler material, which maps directly onto this prompt.

