Cleaning a Messy Spreadsheet Before Automating Anything Around It
A working note on data cleanup before automation — what matters, what does not, and where projects usually go sideways.
The question behind data cleanup before automation is usually financial, not technical. Somebody wants to know what it costs to get this right and what it costs to get it wrong.
The problem underneath data cleanup before automation
Teams treat this as a tooling question. It is a workflow question wearing a tooling costume.
Swap the tool and the same friction shows up two months later with a different logo on it.
What this looks like in practice
In projects like these the shape repeats. Someone maps the current process, finds three painful steps, and discovers only one of them justifies real engineering.
A team we would typically advise starts with the step generating the most back-and-forth email. Not the most interesting one.
The first version covers the common case and a human handles the rest. That is the design, not a compromise.
Mistakes teams make with data cleanup before automation
- Treating launch as the finish line. Most of the cost arrives afterwards.
- No named owner. Unowned work drifts, then the technology takes the blame.
- Designing for the rare case. Build the common path first.
- Skipping measurement. If nobody can tell whether it worked, you will keep paying regardless.
- Picking the tool first. That is the last decision, not the first.
How we approach it technically
Start with the data model. Most bad decisions here are downstream of a schema that made an assumption nobody revisited.
Then the failure modes. Then the interface. Interfaces are cheap to change; schemas and contracts are not.
Alert on rate of change rather than fixed thresholds. Quiet degradation is the failure that costs customers without waking anyone.
A sequence that tends to work
- Write the outcome and the metric, one sentence each, agreed by whoever signs off.
- Map the process end to end, including the manual steps people are slightly embarrassed about.
- Pick the single highest-friction step and ignore the rest for now.
- Ship a narrow version behind a flag to a handful of real users.
- Watch it for two weeks against the number from step one.
- Expand only where the data says it pays.
Step three is where teams cheat. Keeping it honest turns a six-month project into a six-week one.
Practical guardrails
- Instrument before you optimise. Guessing at bottlenecks costs more than measuring them.
- Keep a rollback path for anything touching customer data.
- Document the decision, not just the result.
- Set a review date ninety days out.
- Cap spend and volume in code, not on the invoice.
The honest trade-offs
Going fast now usually means paying interest later. That is fine if you know the rate and have a date to refinance.
Going slow now to avoid rework only pays off if the requirements hold. Early on, they rarely do.
Things people believe that are not quite true
That more tooling reduces risk. Usually it moves the risk somewhere less visible.
That a rewrite resets the clock. It resets the bugs too, and you get a new set.
That the team will document it afterwards. They will not, unless it is part of the definition of done.
Frequently asked questions
How long does data cleanup before automation usually take?
A narrow first version is normally four to six weeks. Anything quoted at three months with no shippable slice in between is a risk, not a plan.
What is the most common mistake with data cleanup before automation?
Scoping too wide. Covering every case in version one delays feedback and inflates cost with no matching benefit.
Do we need a dedicated team for this?
Not at the start. One owner with a few hours a week plus a small build team is enough until the first version proves value.
How do we know whether it worked?
Pick the number before you build: hours saved, error rate, response time or conversion. Compare a two-week window before and after.
What should we do first?
Write one sentence describing the outcome of data cleanup before automation, then map the workflow it touches. Both take an afternoon and remove most of the guesswork.
Conclusion
The useful move on data cleanup before automation is almost always the smaller one. Ship a narrow slice a real user can touch this month, measure it, then decide what deserves the next four weeks.
Everything gets easier once something is live.
Related reading and next steps
- AI automations — how we run this kind of work.
- All Augere Labs services.
- More writing from the team.
Want a second opinion on data cleanup before automation for your setup? Book a 30-minute call. We will say plainly if it is not worth building.
FAQ
Frequently asked questions
How long does data cleanup before automation usually take?+
A narrow first version is normally four to six weeks. Anything quoted at three months with no shippable slice in between is a risk, not a plan.
What is the most common mistake with data cleanup before automation?+
Scoping too wide. Covering every case in version one delays feedback and inflates cost with no matching benefit.
Do we need a dedicated team for this?+
Not at the start. One owner with a few hours a week plus a small build team is enough until the first version proves value.
How do we know whether it worked?+
Pick the number before you build: hours saved, error rate, response time or conversion. Compare a two-week window before and after.
What should we do first?+
Write one sentence describing the outcome of data cleanup before automation, then map the workflow it touches. Both take an afternoon and remove most of the guesswork.
Building something similar?
Let's talk in 30 minutes.

