·Ankit Mehta·7 min read

How to Build an On-Call Schedule for the First Time

Informal incident handling works until it does not. Someone who built the service gets pinged in Slack. Whoever is online jumps in. The same two people take every painful night because they answer fastest. Then an alert lands while those people are traveling, or burned out, or already in another fire, and the room discovers that "we all kind of own this" means nobody is scheduled to.

If that sounds familiar, you are ready to formalize on-call. The first schedule will be imperfect. That is expected. The goal is a manageable burden with clear ownership, not a perfect roster on day one.

Signs informal handling is breaking

Formalize when you see more than one of these.

  • Alerts sit in a channel with no acknowledged owner.
  • The same engineers take most overnight work without a rotation.
  • New services ship with no named responder.
  • People hesitate to take leave because production might page "them" socially.
  • Post-incident reviews keep ending with "we were not sure who was on."

You do not need a fifty-person SRE org to justify a schedule. You need recurring pages and more than one person who can respond.

Who should be on-call, and who should not

On-call is for people who can do useful work under pressure with the access and docs they have. Seniority helps, but readiness matters more than title.

A primary should be able to

  • receive a phone or push page on a device they will hear
  • reach the runbook and the dashboards named in it
  • take basic mitigation steps for the services in scope
  • escalate without ego when they are stuck

Do not put someone on primary because the spreadsheet needed another name. Shadow first. Keep managers off the primary rotation unless they are truly in the technical response path. Managers are often better as a later escalation tier.

On-call is a burden. Say that out loud. Compensation, comp time, or load balancing across the quarter are how you keep the system human. Pretending it is a light privilege is how people quit the rotation by quitting the company.

Rotation design options

Weekly rotation. Most common for small teams. One primary for seven days. Enough time to build context, short enough that the end is visible. Start here for 5 to 15 engineers.

Follow-the-sun. Works when you have enough people across time zones that daytime coverage in each region replaces overnight pages for most. At fifteen-plus distributed engineers this can be excellent. At six people in one city it is fiction.

Bi-weekly. Longer stretches mean fewer handoffs and heavier individual load. Use later if weekly handoffs are the pain point and staffing can absorb two-week stints.

For a first schedule on a small team, pick weekly. Optimize after you have data from real pages.

Shadow on-call before primary

New engineers should observe a full rotation before they own one. They get the pages too, or sit with the primary, and practice acknowledgement, runbook use, and escalation without being the only name on the policy.

Shadowing is cheaper than a first solo night where the new primary freezes and the secondary does all the work anyway. Make shadow completion a gate, not a suggestion.

Override windows

People will have weddings, flights, and sick days. If overrides are awkward, people stay listed while unreachable.

Keep overrides simple.

  • Named substitute for a defined window
  • Secondary automatically covers if no substitute is set and the primary does not ack
  • No social punishment for taking a planned override

Publish how to request an override in the same place as the schedule. If the only path is "ask in Slack and hope," you do not have an override policy.

Escalation path

A schedule without escalation is a single point of failure with a calendar UI. Define at least

  1. Primary
  2. Secondary
  3. Engineering manager or service owner contact

Set acknowledgement windows and channels in the tool. Our escalation policy guide goes deeper on windows and tiers. For the first schedule, the minimum is enough that a missed primary page still reaches a human.

Pre-on-call checklist

Before someone is primary, all of this should be true.

  • Phone or push alerts tested with a real page
  • Access to production systems needed for the services in scope
  • Runbook links open and not pointing at retired hosts
  • They know who secondary is and how escalation fires
  • They know the severity definitions used for overnight pages
  • Handoff note from the previous primary is written, even if short

If any box is unchecked, they are not ready. Fix the box. Do not "start anyway."

Example. 6-person team, weekly rotation

Engineers. Alex, Blair, Casey, Drew, Eden, Fran.

Rules.

  • Weekly primary, Monday 11:00 to next Monday 11:00 local
  • Secondary is next week's primary, so context carries forward
  • New hires shadow one week before entering the primary pool
  • Overrides go to any other trained engineer, recorded on the schedule

| Week of | Primary | Secondary | | --- | --- | --- | | 6 Apr | Alex | Blair | | 13 Apr | Blair | Casey | | 20 Apr | Casey | Drew | | 27 Apr | Drew | Eden | | 4 May | Eden | Fran | | 11 May | Fran | Alex |

After six weeks, everyone has been primary once. Adjust order if one person owns a launch week that should not also carry primary.

First week review

Treat the first week as instrumentation for the process, not only for the product.

Track

  • pages fired vs pages that needed human action
  • time to acknowledge
  • whether secondary was pulled in
  • runbook gaps the primary hit
  • whether overrides were needed and whether they worked

Then change one thing. Alert thresholds, runbook steps, rotation order, or quiet hours for low severities. Do not rebuild the whole schedule because the first Tuesday was noisy.

Common first-schedule mistakes

Daily rotations. Too much handoff, too little context, constant context switching.

No secondary. The primary's dead phone becomes everyone's incident.

No override policy. Unreachable primaries stay on the roster.

Juniors on primary too early. Shadow exists for a reason.

Every service in one rotation with no scope. A six-person team on-call for thirty undocumented services will drown. Start with the services that already page, write runbooks for those, then expand.

Ignoring load. If one person took four overnight incidents and everyone else took zero, the schedule may be fine and the alert quality may not be. Fix noise before you add more bodies.

Handoffs that take five minutes

Weekly rotations fail quietly when handoffs are "you are up now" with no context. Require a short note at the boundary.

Include

  • open incidents or degraded components
  • alerts that are noisy right now
  • deploys or migrations landing this week
  • anything the secondary should know before the primary goes offline

Five minutes of writing saves forty minutes of rediscovery during the first page of the new shift. Store the note where the schedule lives, not in a private DM.

Make it boring, then improve it

A good first on-call schedule is a bit dull. Weekly cadence, clear primary and secondary, tested phones, written overrides, short handoffs. Excitement belongs in the product, not in the paging order.

When you want schedules and escalation in one workspace without paying per responder as the roster grows, Vigiles includes that on flat per-workspace pricing. The first version still comes from your team naming who is ready and which services are in scope. Iterate after the first real week of pages, not after the twentieth debate about the perfect roster.

Common questions

How do you build an on-call schedule for the first time?
Start with a weekly primary rotation among engineers who can follow a runbook, add a secondary, define overrides, and review the first week of pages before you optimize. Expect the first version to need changes.
Who should be on the on-call rotation?
People who can access production systems, follow the runbook, and escalate when stuck. Skip anyone who has not shadowed a rotation or does not yet have working phone alerts.
What is the best on-call rotation length for a small team?
Weekly is the usual starting point for a team of about 5 to 15. It balances context continuity against individual burden better than daily swaps or multi-week stretches for most first schedules.

Ready to try Vigiles?

Start monitoring your endpoints in under 2 minutes. Free forever for small projects.

Create Your Workspace Free