Software Architecture · Published 9/29/2026 · By Eric Njuguna
Scale the Workload Before You Scale the Architecture
A practical way to choose between a larger service, background work, caching, and distributed architecture using measured workload shape.
“Will it scale?” is a fair question with an unhelpful default answer: add another service, another queue, and another diagram.
Architecture should follow the shape of the workload. Start by learning what the system actually does: request volume, concurrency, payload size, read/write mix, latency expectations, burst patterns, and the cost of being briefly unavailable.
Measure the bottleneck before naming the fix
If responses slow down while the database is saturated, adding more web servers may make the database queue longer. If users wait on a report that runs for minutes, moving every endpoint to an event-driven architecture may be a wide detour around one long-running task.
Look at traces, service metrics, query timings, and representative load tests. Ask which resource reaches its limit first and how the user experiences it. “CPU is high” is a clue; “the report endpoint spends 82% of its time waiting on one query” points closer to a remedy.
Prefer the smallest change that buys room
For predictable, moderate traffic, a better-indexed query, a simpler data access pattern, or a larger service instance can be the most maintainable answer. If a task is slow but does not need to finish during the request, put that work in a durable background queue and let the user see its status.
If repeated reads dominate, caching may help—provided you can explain freshness, invalidation, and what happens when the cache is empty. If traffic arrives in uneven bursts, capacity limits and backpressure matter as much as the maximum number of instances.
Each option moves complexity. A queue adds retries and stuck-work handling. A cache adds staleness. Sharding adds coordination and operational overhead. Choose the move whose new failure modes the team is equipped to operate.
Make the decision reversible where you can
Write down the observation, the expected constraint, and the signal that would show the change worked. For example: “Move report generation to a background job because request timeouts interrupt valid work; track completion time and stuck-job count.” That gives the team a way to revisit the decision when the workload changes.
This is the same discipline used in dependable data platforms: observe first, make the failure mode legible, and scale the actual constraint. Architecture is not a trophy cabinet of technologies. It is a set of choices matched to real operating conditions.
For a review of a growing platform, see how I work with teams.