Choosing a Partner

Cloud Application Stack Management: What Should a Provider Own?

Anas MaqsoodCo-Founder & Chief Technology Officer8 min read

The short answer

A cloud application stack provider should operate the customer-controlled layers named in the agreement, not simply provision a cloud account. The scope should map ownership for identity, networking, runtime, data, delivery, observability, recovery, incidents and cost, with evidence for each layer and a clear boundary with both your team and the cloud platform.

What is cloud application stack management?

Cloud application stack management is the ongoing operation, security, change, observation and recovery of the customer-controlled layers that make a cloud application work. It is broader than creating cloud resources and narrower than an assumption that one supplier owns everything in production.

The boundary depends on the services you use. With virtual machines, your side usually retains more responsibility for operating systems and runtimes. A managed database or serverless platform transfers some lower-level work to the cloud provider, but your application code, configuration, identities, data choices, monitoring and recovery decisions still need owners. A mixed stack can have a different boundary for every service.

Which cloud application layers need a named owner?

LayerOwnership to defineEvidence to request
Architecture and governanceServices, environments, dependencies, risks and decision authorityCurrent architecture diagram, inventory and responsibility matrix
Cloud resourcesAccounts, regions, quotas, provisioning, policies, tagging and retirementControlled configuration or infrastructure-as-code history and drift findings
Identity and secretsHuman and workload access, privileged roles, keys and access removalRedacted permission matrix, access reviews and rotation records
Network and edgeDNS, certificates, ingress, egress, firewalls and public endpointsNetwork diagram, endpoint inventory and certificate or rule review
Compute and runtimeImages, patches, containers, orchestration, upgrades and capacitySupported-version inventory, maintenance records and upgrade plan
Data and recoveryStorage, encryption choices, retention, backups, restores and deletionData-store inventory, approved recovery objectives and dated restore test
Application deliveryBuilds, tests, configuration, migrations, releases and rollbackPipeline results, artefact identity, deployment log and rollback route
ObservabilityMetrics, logs, traces, dashboards, alerts and alert ownershipTelemetry map, dashboard, alert catalogue and recent alert history
Incidents and continuityTriage, escalation, communication, recovery and corrective workRunbooks, escalation matrix and recovery-exercise or incident evidence
Cost and capacityAllocation, budgets, anomalies, forecasts, lifecycle and rightsizingWorkload-level cost report, budget alerts and optimisation decisions
Cloud application stack management scope checklist

Not every engagement needs the provider to operate all ten layers. The important point is that no layer is left between teams. If your internal engineer owns data migrations while the provider owns deployments, the release process should say exactly where the hand-off occurs and who stops or reverses a failed change.

How should responsibility change with the hosting model?

Cloud platforms use a shared-responsibility model. The platform operates parts of the underlying cloud; the customer and any contracted operator retain responsibility for the workload-specific layers. Moving from infrastructure as a service to a more managed platform changes that division, but it does not remove the customer's side of it.

  • Infrastructure as a service: define guest operating-system patching, runtime support, application deployment, data configuration and customer-controlled network rules.
  • Platform or serverless services: the platform covers more of the host and runtime, while your side still needs owners for code, configuration, identities, data, telemetry and recovery choices.
  • Managed Kubernetes: specify control-plane, worker-node, add-on, policy, workload, upgrade, telemetry and backup ownership separately.
  • Software as a service: the vendor operates the product, but your organisation still controls users, access decisions, data use and tenant configuration.
  • Mixed stacks: assess every service on its own boundary instead of applying one label to the entire system.

This is why a provider's technology list is not enough. Two teams can both advertise AWS, Azure or Google Cloud while offering materially different operational coverage. The useful comparison is the work and evidence inside the boundary, not the logos outside it.

What evidence should a provider supply before you sign?

Ask for representative, redacted evidence rather than confidential access to another customer's environment. The goal is to see whether the operating process produces verifiable records, not to collect screenshots that disclose somebody else's systems.

  1. A responsibility matrix covering routine work, approvals, incidents and handover
  2. An example architecture and data-flow diagram with a resource inventory
  3. A redacted access model and privileged-access review
  4. A controlled configuration or infrastructure-as-code sample with change history
  5. A monitoring view, alert catalogue and escalation path
  6. Defined recovery objectives with a dated restore or recovery-test result
  7. A release record showing tests, artefact identity, deployment result and rollback route
  8. An incident runbook and a redacted example of tracked corrective work
  9. A workload-attributed cost view with budget or anomaly alerts
  10. An exit pack covering repositories, diagrams, runbooks, access transfer and supplier-access removal

Scale the evidence to the risk. A small internal prototype does not need the control pack of a regulated production system, but it still needs an owner for access, changes, data and recovery. The provider should be able to explain what is deliberately omitted and why.

Which questions expose an unclear managed-cloud offer?

  • Show us the ownership boundary for every layer, including what stays with our team and the cloud platform.
  • Which production actions can you take directly, and which require our approval?
  • What operating evidence will we receive, and how often?
  • Who owns alerts outside our normal working hours?
  • When was the recovery process last exercised, and what did the exercise cover?
  • How are emergency changes recorded, reviewed and reversed?
  • How will we see cost by workload and be warned of unusual spend?
  • What do we receive, and which supplier access is removed, when the engagement ends?

The answers should become part of the scope. A polished sales explanation does not resolve an ownership gap during an incident. A named role, a decision path and an artefact do.

What should a monthly cloud operations report contain?

A monthly operations report should show decisions and evidence, not a wall of unprioritised metrics. The exact contents depend on the workload, but the report should let a buyer connect changes, incidents, recovery readiness, access, capacity and spend to the people responsible for the next action.

Reporting areaWhat to includeDecision it should support
Changes and releasesMaterial releases, configuration changes, failed changes and rollback evidenceWhether the release process is controlled and where corrective work is needed
Incidents and alertsMaterial incidents, recurring alerts, impact, response and tracked corrective actionsWhich failure patterns need engineering work rather than more notifications
Recovery readinessBackup status, restore or recovery exercises, findings and unresolved gapsWhether the documented recovery path has current evidence behind it
Identity and accessPrivileged-access changes, removals, reviews and unresolved exceptionsWhether access still matches current roles and approved operational needs
Workload health and capacitySignals tied to application health, capacity constraints and meaningful trendsWhether the workload needs a configuration, architecture or capacity change
CostSpend by workload, budget exceptions, anomalies and approved optimisation decisionsWhich cost change has an owner and whether it reflects useful demand or waste
Known risks and maintenanceUpcoming upgrades, expiring dependencies, open risks and ownership decisionsWhat must be approved or scheduled before it becomes urgent
A practical cloud application stack management report

Microsoft's Azure Well-Architected guidance treats observability as a separate workload capability that collects metrics, logs, traces and events across infrastructure, application health, and build and release processes. Google Cloud's operational-excellence guidance likewise connects monitoring, capacity planning, incident management and change management. Those practices support a report that explains workload health and action, rather than one built around whichever charts are easiest to export.

What should the first cloud-stack engagement deliver?

Start with the ownership gap that creates the next decision, rather than asking a provider to take over the entire stack at once. Give each prospective provider the same architecture outline and ask for a bounded deliverable with an acceptance test. That makes the proposed work comparable and leaves your team with a usable artefact even if you choose a different long-term operator.

If the uncertainty is…Ask for…Accept it when…
No one can name who owns each production layerA responsibility map covering the application, platform, internal team and proposed operatorEvery material layer has an owner, an approver, an exclusion and a handover route
The release path depends on one person or undocumented accessA release and access review using the current pipeline and account rolesThe team can show who approves, deploys, reverses and removes access for a change
Backups exist but recovery is uncertainA recovery-readiness review of the workload's data and dependenciesThe team has agreed recovery objectives, named responders and evidence from a representative restore exercise
Choose a first deliverable from the ownership gap you need to resolve.

A review is not the same as ongoing cloud operation. Ask which findings the provider will implement, which remain with your team and what separate agreement would govern incident response or maintenance. The responsibility boundary should stay legible after the first deliverable is complete.

How can ApexStack scope the first engagement?

ApexStack's Product Blueprint starts from US$1,000 for a bounded planning and de-risking outcome. For a cloud application stack, that could be an architecture and ownership review, a deployment audit or a recovery-readiness assessment. It is not a production build or a blanket price for operating a complete cloud estate.

A Launch Sprint starts from US$2,500 and covers planning, UX direction, implementation, testing and deployment for one tightly scoped first release or core workflow. Authentication, billing, mobile applications, complex AI, multiple integrations, data migration, compliance work and extensive administration can increase the quote. Wider cloud operations are scoped only after the workload, access boundary, recovery needs and operational risks are understood. If you are comparing providers, send ApexStack the same architecture outline and acceptance criteria you send them; we can assess whether a bounded review or one implementation task is the right first step.

Sources

Frequently asked questions

Is cloud application stack management the same as DevOps?
They overlap, but the phrases are not interchangeable. DevOps describes ways of organising software delivery and operations. Cloud application stack management is the concrete operating scope across the customer-controlled application and cloud layers. A contract should define the work rather than rely on either label.
Does a cloud provider manage application security?
Only within its documented service boundary. Your organisation still needs ownership for customer-controlled identities, data, application code, configuration and other workload-specific controls. A contracted operator can take on some of that work, but the agreement must say which parts.
Should a managed-cloud service include backups?
Only if backups and recovery are explicitly in scope. Define which data is protected, retention, recovery objectives, who responds to a failure and how restores are tested. A successful backup job alone does not demonstrate that the application can be recovered.
How do I compare cloud app stack providers?
Give each provider the same architecture outline and ask for a layer-by-layer responsibility matrix, exclusions, approval path, evidence pack and exit process. Compare those boundaries alongside price and technology fit.

Where should ApexStack start with your cloud stack?

Send the current architecture, the team responsible for releases and incidents, and the one ownership gap that worries you most. ApexStack can scope a bounded review of that boundary, identify the evidence needed for a handover, and discuss an implementation path if a specific gap needs engineering work. A Product Blueprint starts from US$1,000 for planning and de-risking; a Launch Sprint starts from US$2,500 for one tightly scoped first release or core workflow. Wider operations, multiple integrations, data migration, compliance and extensive administration may increase the quote.