Insights
When Scaling Is Not the Solution
A short engagement where the assumed fix was horizontal scaling — and the real problem was elsewhere.
A recent engagement started with a familiar challenge: several users hitting an application at the same time, so the team assumed they needed to scale horizontally. That was the assumed direction. The useful role was independent advice on how to move forward without adding further performance risk. The organisation is not named.
01
The assumed direction
The team had already started to treat horizontal scaling as the next logical step. They had never operated the system in a scaled setup before, and the underlying architecture had not been validated for it. From a performance engineering perspective, that raised a concern that shows up often: scaling a system that is not properly understood often amplifies problems instead of solving them.
The first recommendation was therefore simple. Do not start with horizontal scaling. If scaling is needed at all, vertical scaling is usually the safer and more controlled step — especially when system behaviour is not yet fully understood. Even that, though, was not the real focus of the engagement. The useful work was to understand what the system was already doing under modest load.
02
What we actually found
After a short, focused discussion and initial analysis, it became clear that the issue was not concurrency or infrastructure capacity. The system was already showing performance issues at low load. There were underlying database-related bottlenecks in PostgreSQL, query behaviour was not well understood, and production observability was effectively missing — the team had only partial logging to work with. With that limited visibility, they were operating without a clear picture of what was actually happening inside the system.
Once they investigated more closely, the root cause surfaced in the data access path rather than in the number of concurrent users. Several queries were not well written, indexes were not properly utilised, and some database interactions were simply inefficient. Lightweight tooling helped them understand query execution and behaviour, which made it possible to identify problematic SQL patterns and bottlenecks with relatively low effort. The issue was not scale. It was inefficiency and lack of visibility.
03
From scaling to understanding
Instead of debating scaling strategies, we shifted the discussion towards understanding system behaviour. Practical next steps that could provide immediate insight included introducing basic tracing and database-level visibility, extracting usable performance-related information directly from PostgreSQL, and discussing options such as controlled benchmarking (for example with pgbench) for later testing. It quickly became evident that deeper analysis was needed at the query level before any infrastructure change would be justified.
That change of perspective matters because many teams reach for capacity when they lack evidence. Without answers to a few basic questions — how the system behaves under load, where time is actually spent, and what the bottleneck really is — scaling often becomes an expensive guess.
04
Embedding performance earlier
The conversation naturally moved from fixing the immediate issue to improving long-term capability. We discussed how performance engineering could be embedded earlier and more consistently into delivery: introducing performance-related quality gates, validating performance at single-user level early in development, applying lightweight profiling and analysis before scaling decisions, and making performance a shared responsibility across teams. The goal was not to add heavy process, but to enable better decisions earlier — so the next “we need to scale” conversation starts from evidence rather than assumption.
05
What a few hours changed
The entire engagement took only a few hours, yet the impact was significant. The team avoided premature and potentially costly scaling efforts, identified real technical bottlenecks, improved their understanding of system behaviour, and left with clear, actionable next steps.
This case highlights a pattern that appears frequently: many performance problems are not scaling problems. They are understanding problems. Performance engineering is not only about tools, load, or infrastructure. It is about making better decisions based on a clearer understanding of the system. In many cases, a short, focused advisory session is enough to unlock that understanding — and prevent much bigger problems later.
