Home
Storage

Storage Planning Isn’t Just Capacity

Storage infrastructure capacity-planning illustration
Pedro CoutoJul 5, 20263 min read
Share
23 reads

The first question is usually: “How many terabytes do you need?” It’s a valid question, but an incomplete one. In critical environments, storage also involves performance, availability, recovery, security, and operational predictability. Capacity is one part of the design — not the whole design.

Capacity is a variable

Before defining terabytes, there are better questions to ask:

  • What’s the workload, and what’s its I/O profile?
  • What latency is expected, and what growth is projected?
  • What’s the Recovery Point Objective (RPO) — how much data can you afford to lose?
  • What’s the Recovery Time Objective (RTO) — how long can you afford to be down?
  • What’s the backup window, and what’s the actual restore time?
  • What’s the impact of a controller going down? Of a site going down?
  • Does the environment maintain availability during a takeover, and does performance stay within what’s needed afterward?
  • Who can delete snapshots, and who can change policies?
  • Does replication protect the environment, or just replicate the problem?

These answers define the architecture. Everything else is sizing.

Availability needs to be designed

High availability means more than having two controllers. What matters is how the environment behaves during a failure. During a takeover, one controller assumes the services and workloads of the other. Service remains available, but processing, memory, cache, ports, and I/O demand are now concentrated on fewer resources.

The capacity constraint is straightforward. In a two node HA pair, if both controllers reach 70% CPU utilization during peak periods, the surviving controller cannot maintain the same performance after takeover. In simple terms, it would be expected to handle approximately 140% of a single controller’s normal workload.

When performance during takeover is a requirement, the environment should be sized so that one controller can support both workloads. A conservative starting point is to keep sustained peak utilization near 50% per controller. Alternatively, the design should clearly document which services may degrade, the expected level of degradation, and whether that impact remains within the agreed service level.

Workloads do not always combine linearly during takeover. Actual behavior depends on the controller model, protocol mix, I/O profile, cache efficiency, storage layout, and other resource constraints. For this reason, the 50% guideline should not be treated as a universal limit. The design must be validated using observed performance data, platform specifications, and, whenever possible, takeover testing.

The right question is not simply, “Does the environment have high availability?” It is, “Can the environment continue delivering the required service level during a takeover?”

That question changes how CPU, cache, ports, throughput, IOPS, latency, aggregates, SAN and NAS paths, and future growth must be evaluated.

When budget limitations make full takeover performance impractical, the requirement can be refined using real workload data. For example, performance monitoring can be filtered to business hours or other critical operating periods. This helps determine the resources required during the periods when service levels matter most and makes any accepted degradation explicit rather than accidental.

Technology doesn’t fix a bad premise

ONTAP, SAN, NAS, object storage, SnapMirror, MetroCluster, backup, cyber recovery, Kubernetes, VMware, and cloud are tools. Good tools help, but they don’t replace architecture. If the design doesn’t account for failure, operation, security, recovery, and growth, the environment is limited from the start,  even with good hardware

About the author

NetApp A-Team member focused on enterprise storage and AI infrastructure

Pedro Couto · Enterprise Infrastructure Architect

Pedro works at the intersection of NetApp ONTAP, hybrid cloud, data protection and high-performance AI platforms. Designing and delivering enterprise infrastructure for organizations where downtime is not an option. A recognized member of the NetApp A-Team and holder of the NetApp Subject Matter Expert Elite designation, he is one of the authors of many certifications of NetApp such as NCDA, NCIE-DP, NCIE-MetroCluster, Storage Engineer and more.

Comments

0

No comments yet. Start the conversation.

Up to 2,000 characters. Plain text only.