CloudWatch Cost Explosion: Why Startups Overpay for Observability
CloudWatch bills blow up on custom metrics and log ingestion, not dashboards. Here's the real cost math and when an open-source stack wins.

If your CloudWatch bill jumped from pocket change to hundreds of dollars a month, the culprit is almost never dashboards — it’s custom metrics, high-cardinality dimensions, and log ingestion. AWS charges roughly $0.30 per custom metric per month, $0.50 per 1,000 custom metric API calls via PutMetricData, and $0.50 per GB of logs ingested. A startup that emits a per-user or per-request dimension can manufacture tens of thousands of distinct metrics without noticing. For most small teams the honest answer is simple: keep CloudWatch for alarms and AWS-native signals, and move high-volume application metrics and verbose logs to an open-source stack only once your bill crosses a few hundred dollars a month.
Where the money actually goes #
CloudWatch pricing is deceptive because the line items that hurt are the ones you generate programmatically. Dashboards cost $3 each per month — annoying, not fatal. The real damage comes from three places. Custom metrics are billed per unique combination of namespace, metric name, and dimension values, so a single metric emitted with a customer_id dimension across 5,000 customers is 5,000 billable metrics, or about $1,500/month at $0.30 each. Log ingestion at $0.50/GB sounds cheap until a chatty service in debug mode ships 40 GB/day — that’s $600/month just to ingest, before storage. And Logs Insights queries bill $0.005 per GB scanned, so a dashboard that re-queries 100 GB every five minutes quietly adds up.
This is the same pattern we described in our anatomy of an AWS bill shock: the ambush line items are usage-metered, not capacity-metered, so they scale with how your code behaves rather than how big your servers are. Observability is the worst offender because it’s the one bill that grows fastest precisely when things are going well and traffic is climbing.
The open-source alternative, with real numbers #
The usual escape hatch is a self-hosted stack: Prometheus for metrics, Loki or OpenSearch for logs, Grafana for dashboards. The appeal is that you pay for compute and storage, not per-metric or per-GB-ingested. A Prometheus plus Grafana setup on a single t3.medium (~$30/month on-demand, closer to $12 on a reservation) can comfortably hold millions of time series that would cost thousands of dollars as CloudWatch custom metrics. Loki is cheaper still for logs because it indexes only labels, not full text, so you pay mostly for S3-backed object storage at roughly $0.023/GB-month instead of $0.50/GB ingested plus storage.
The catch is the word "self-hosted." You now own retention policies, upgrades, a Grafana you have to secure, and the operational burden of the thing that’s supposed to tell you when other things break. For a five-person team, a few hundred dollars of monthly savings can evaporate into a weekend of yak-shaving per quarter. That trade-off is the same judgment call we laid out in when to stop using AWS managed services and self-manage: managed is worth a premium until the premium outgrows the salary cost of operating the alternative.
A pragmatic middle path #
You don’t have to choose one stack for everything. The cheapest observability bill a startup can run usually mixes tools by signal type. Keep CloudWatch for what it does natively and cheaply — ECS, RDS, and ALB metrics arrive for free or near-free because they’re AWS-emitted standard metrics, and CloudWatch alarms are $0.10 each per month. Route your own application metrics to Prometheus so high-cardinality labels don’t translate into a per-dimension bill. And be ruthless about log volume before you decide which log store to pay for: dropping debug-level logs in production and sampling health-check lines often cuts ingestion by 60–80% regardless of where the logs land.
The single biggest lever, though, is cardinality discipline. Treat every metric dimension and every log label as a line item, because that’s literally what it becomes. One audit rule of thumb covers most of it:
- Any dimension whose cardinality grows with your user count — user IDs, session IDs, request IDs — belongs in logs or traces, never as a metric dimension.
How this plays out when Vylara provisions your infrastructure #
When you connect your repo through the Vylara GitHub App, the agent engine analyzes your code and detects the services your app actually needs — PostgreSQL, Redis, S3, and so on — then provisions them in your own AWS account on the first deploy. That means the observability picture is anchored to resources you own and control, not a black box. After launch, Vylara’s infrastructure chat can read your CloudWatch logs and metrics directly to diagnose issues; the log tail path keys off your log group and region, so when you ask why a service is returning 502s, the answer references the same log group you’re being billed for.
Because read operations like pulling recent logs and cost data happen through that chat without you wiring up a separate dashboard, you can inspect what’s driving volume before committing to a heavier open-source stack. And since everything lives in your account, migrating log destinations later — swapping a CloudWatch log group for a cheaper store — is a change you make to infrastructure you already own, with no vendor hosting in between. For a deeper view of how this fits a small team’s budget, see our guide for startups deploying to their own AWS.
The honest bottom line: CloudWatch is the right default at small scale, and the "explosion" is almost always a cardinality or log-volume bug rather than a pricing injustice. Fix the emission patterns first. If your bill still sits above $300–$500/month after that, a Prometheus-and-Loki stack on a $30 instance is a defensible move — just budget for the operational time it costs you to run it.
Review your cloud plan in Vylara, merge delivery changes as Git PRs, and deploy into your own AWS or Azure account when you’re ready.
Start freeFrequently asked questions
- Why did my CloudWatch bill suddenly spike?
- Almost always custom metrics or log ingestion, not dashboards. CloudWatch bills about $0.30 per unique custom metric per month and $0.50 per GB of logs ingested, so a metric emitted with a high-cardinality dimension (like user ID) or a service stuck in debug logging can add hundreds of dollars without any infrastructure change.
- Is Prometheus and Grafana actually cheaper than CloudWatch?
- For high-volume application metrics, yes — often dramatically. A single ~$30/month instance can hold millions of time series that would cost thousands as CloudWatch custom metrics. The savings shrink once you factor in the operational time to run, secure, and upgrade the stack yourself, so it mainly pays off above roughly $300–$500/month in CloudWatch spend.
- Does Vylara set up my observability stack?
- Vylara provisions the services your app needs in your own AWS account, and its infrastructure chat can read CloudWatch logs and metrics to help you troubleshoot, referencing the actual log group and region. It does not install a separate Prometheus or Grafana stack for you, but because everything runs in your account you're free to add or migrate log destinations yourself.



