Cloud & Infrastructure
Updated
A one- to three-person IT team should not make one blanket decision about whether the organization is "on-premises" or "in the cloud." The useful decision unit is a workload: the application, its data, identities, integrations, recovery dependencies, and the people responsible for operating it.
This guide is not another list of generic cloud pros and cons. It is a repeatable placement review for a small team that must live with the result after the project ends. The output should be a defensible choice for each workload, a named owner, and a reversible transition plan.
Start with the workload, not the destination
Cloud can shift infrastructure work to a provider, but it does not remove operational responsibility. Someone still has to manage identity, configuration, data protection, cost, vendor escalation, monitoring, and recovery. On-premises systems can preserve local control or support a hard dependency, but they also require facilities, hardware lifecycle, patching, capacity, and replacement planning.
NIST defines several cloud service and deployment models, including public, private, and hybrid cloud. Those definitions are useful vocabulary, not a placement answer. A small IT team should first describe the business service and then decide which model can support it with an acceptable operating burden.
Build one fact sheet for every candidate workload
Do not begin with a vendor quote. Create a short fact sheet using evidence from the application owner, technical staff, contracts, and current monitoring.
| Decision area | Questions to answer | Evidence to retain |
|---|---|---|
| Business impact | Which work stops if this fails? How long can that interruption be tolerated? | Approved impact statement and recovery priorities |
| Data | What data is stored, where may it reside, and who may access it? | Data classification, retention needs, and access owners |
| Dependencies | Which identity, network, device, vendor, database, and integration services must work with it? | Current data-flow and dependency map |
| Demand | Is usage stable, seasonal, growing, or difficult to predict? | Capacity and performance history |
| Recovery | What must be restored first, and has the complete workflow been tested? | Recovery design and latest exercise record |
| Operations | Who patches, monitors, supports, approves changes, and handles incidents? | Responsibility matrix and escalation path |
| Economics | What are the full recurring, transition, connectivity, support, and exit costs? | Comparable multi-year cost model with assumptions |
Choose a treatment before choosing a platform
A migration is not always the right treatment. Microsoft's current Cloud Adoption Framework describes several choices, including retire, retain, rehost, replatform, refactor, rearchitect, replace, and rebuild. For a small team, this prevents two expensive mistakes: moving an obsolete workload and carrying an existing operational problem into a different hosting model.
- Retire: remove a workload that no longer has a valid owner or business need after dependencies and records obligations are checked.
- Retain: keep a stable workload where it is when a technical, contractual, latency, equipment, or risk constraint makes movement unjustified.
- Replace: use a supported SaaS service when the required business capability is standard and the integration and exit terms are acceptable.
- Rehost or replatform: move with limited application change when there is a clear operational benefit and the team can manage the target environment.
- Refactor, rearchitect, or rebuild: accept deeper change only when the business outcome warrants the skills, testing, and lifecycle burden.
A hybrid result is normal. It should be intentional, with documented identity, connectivity, monitoring, recovery, and support boundaries rather than an indefinite halfway state.
Apply the small-team capacity test
For each viable option, walk through a normal month, a staff absence, a major update, a provider outage, and a recovery event. Ask whether the team can perform or reliably oversee every required task. A technically sound target is still a poor choice if it creates an operating model the available staff cannot sustain.
- Can two people, not just the primary administrator, access the runbook and escalation contacts?
- Can the team explain the shared-responsibility boundary and show which controls remain theirs?
- Are alert triage, configuration review, access review, backup verification, and cost review assigned?
- Is specialist support available for the tasks the internal team cannot perform safely?
- Can the service still be administered if the usual administrator, network, or identity service is unavailable?
Use a decision matrix without hiding judgment
Rate each option against the same factors: business fit, dependency fit, recovery, security and data handling, staff capacity, vendor support, total cost, migration risk, and exit feasibility. Record the evidence and uncertainty behind each rating. A weighted score can organize discussion, but it should not overrule a known hard constraint.
Identify disqualifiers before scoring. Examples include an unsupported application version, an integration the vendor will not certify, an unavailable recovery path, unacceptable data-location terms, insufficient connectivity, or no named operational owner. Escalate unresolved assumptions instead of marking them as complete.
Move in a sequence the team can reverse
- Inventory and classify. Confirm the workload owner, users, data, integrations, licenses, and business impact.
- Baseline the current service. Capture performance, availability, support effort, recovery evidence, and actual cost before claiming improvement.
- Design the target operating model. Assign identity, security, monitoring, backup, cost, vendor, and incident responsibilities.
- Pilot a representative workflow. Test real authentication, device, printing, integration, performance, and support scenarios.
- Define cutover and rollback. State who may stop the change, what triggers rollback, how data is reconciled, and how users are informed.
- Prove recovery before retirement. Do not decommission the prior capability until the target workload and its dependent business process can be recovered.
- Review after stabilization. Compare the target with the original baseline, close temporary access, and update the decision record.
Keep a one-page placement decision record
The final record should name the workload and business owner, selected treatment and destination, alternatives rejected, hard constraints, assumptions, expected operating tasks, cost basis, recovery approach, security responsibilities, migration and rollback owners, review date, and exit path. This is more valuable to a small team than a long strategy deck that cannot guide the next support call.
Related planning guides
- Use the cloud migration checklist before approving a move.
- Build recurring cloud cost governance into the operating model.
- Review Microsoft 365 and Azure readiness for a care-dependent workload.
Primary sources
- NIST SP 800-145: The NIST Definition of Cloud Computing
- Microsoft Cloud Adoption Framework: develop a cloud adoption strategy
- Microsoft Cloud Adoption Framework: select a migration strategy
- CISA Cloud Security Technical Reference Architecture
Suggested next step
Talk with Cloud Core MSP if your team needs a workload inventory, placement workshop, or migration plan with explicit operational ownership and rollback criteria.