EngineeringAug 9, 2026·9 min read

How We Set a Realistic Uptime Target

A working note on setting an uptime target — what matters, what does not, and where these projects usually go sideways.

Muhammad Qitmeer
Muhammad Qitmeer
Co-Founder & CEO, Augere Labs
Share
A working note on setting an uptime target — what matters, what does not, and where these projects usually go sideways.

Setting an uptime target looks like a technical decision. In practice it is a scheduling and ownership decision wearing a technical costume. This post walks the order we actually use.

The problem underneath

The pattern is familiar. Someone raises it in a standup, a decision gets made in ten minutes, and nobody records why.

Six weeks later three people hold three different mental models. The rework costs more than the original choice ever did.

What this looks like in real projects

In projects like these, one version is local. A single workflow strains, everything else is fine, and two focused weeks clear it.

The other reads identically in a status update, but the strain is systemic. Treat that as local and you spend a quarter arriving back where you started.

Mistakes companies make

  • Choosing tools before the workflow is written down.
  • Scoping version one to cover every edge case.
  • Leaving the work unowned, then blaming the tool.
  • Skipping measurement, so nobody can prove it helped.
  • Treating launch day as the end of the cost.

The first and the last are the expensive ones.

How We Set a Realistic Uptime Target — setting an uptime target decision flow used by the Augere Labs team
How we frame setting an uptime target in the first week of a project.

The engineering view on setting an uptime target

From inside the codebase, setting an uptime target reduces to three questions. What happens when a step fails halfway. Who finds out. How you reverse it.

Design for partial failure before you need it. Step three fails after one and two succeeded, and that is the case people skip.

Give retries a ceiling and some jitter. A retry storm is an outage you built yourself.

How we approach it step by step

  1. Reproduce the pain with a real case, not a description of it.
  2. Write the target outcome as a single number.
  3. Pick the smallest change that could plausibly move that number.
  4. Build it with a rollback path.
  5. Release to one team or a slice of traffic.
  6. Review in two weeks, then widen, revise, or delete.

Deleting is a legitimate result. It happens less often than it should.

Practical guardrails

  • Instrument before optimising.
  • Cap spend and volume in code, not on the invoice.
  • Write down the decision, not only the outcome.
  • Keep one named owner with protected hours.
  • Set a review date ninety days out and keep it.

Trade-offs worth saying out loud

Speed against flexibility. Managed service against control. Cheap now against cheap later. None of it is free.

This trade-off usually appears when the second customer wants something the first one didn't. That is the moment to revisit setting an uptime target, not before.

Common misconceptions

“We need the best available option.” You need the one your team can operate at 2am. Rarely the same thing.

“We’ll do it properly later.” Sometimes true. Put a date on later or it never arrives.

“It’s a one-off.” Anything a customer touches becomes a product, support included.

Frequently asked questions

How do we know whether it worked?

Choose the number before you build — hours saved, error rate, response time, or conversion — then compare a two-week window either side.

Is it cheaper to buy a tool instead?

Often yes for the first version. Build when the workflow is a genuine differentiator or no tool fits the data you already hold.

Do we need to hire someone for this?

Not at the start. One named owner with a few protected hours a week, plus a small build team, is enough to prove value.

How long does setting an uptime target take to get right?

A narrow first version is usually four to six weeks. Anything quoted at three months with nothing shippable in between is a risk, not a plan.

When should we revisit the decision?

When a second customer asks for something the first never needed, or when volume changes by an order of magnitude.

Conclusion

The useful move on setting an uptime target is almost always the smaller one. Ship a narrow slice a real user can touch this month, measure it, then decide what earns the next four weeks.

Everything gets easier once something is live.

Related reading and next steps

Want a second opinion on setting an uptime target for your setup? Book a 30-minute call. If it is not worth building, we will say so.

FAQ

Frequently asked questions

How do we know whether it worked?+

Choose the number before you build — hours saved, error rate, response time, or conversion — then compare a two-week window either side.

Is it cheaper to buy a tool instead?+

Often yes for the first version. Build when the workflow is a genuine differentiator or no tool fits the data you already hold.

Do we need to hire someone for this?+

Not at the start. One named owner with a few protected hours a week, plus a small build team, is enough to prove value.

How long does setting an uptime target take to get right?+

A narrow first version is usually four to six weeks. Anything quoted at three months with nothing shippable in between is a risk, not a plan.

When should we revisit the decision?+

When a second customer asks for something the first never needed, or when volume changes by an order of magnitude.

Building something similar?

Let's talk in 30 minutes.

Book an intro
© 2026 Augere Labs