Platform Strategy

Should You Optimize Datadog or Migrate? A Decision Framework

A structured way to compare optimization, hybrid architecture and migration based on cost, risk, capacity and time to value.

Optimization and migration solve different problems, and applying the wrong one to a given cause rarely produces the result the business expected.

Short answer

Whether to optimize Datadog, adopt a hybrid architecture, or migrate away from it depends on whether the cost driver is avoidable consumption or structural platform economics — and those require different evidence to diagnose. Optimization fixes governance gaps: unindexed-by-default logs, uncontrolled custom-metric cardinality, missing usage attribution. Migration addresses a case where the platform's pricing model no longer fits the workload even after governance is sound. Most organizations facing a large Datadog bill have not yet established which of these two situations they are in, and a large share of migration projects are actually attempts to solve a governance problem with an architecture project. The decision should follow a five-option ladder — keep as-is, optimize, hybrid, migrate selected workloads, full migration — ordered by reversibility and time to value, not treated as a binary choice between renewing and leaving.

The mistake this article exists to prevent

A Datadog renewal quote arrives, or a monthly bill crosses a threshold that gets executive attention, and the conversation collapses almost immediately into "should we migrate off Datadog." That framing skips the diagnostic step that determines whether migration would even help.

Migration is expensive, slow, and only partially reversible once dashboards, alerts and team habits have moved. Optimization is comparatively cheap, fast, and fully reversible. Using migration to solve a problem that optimization would have fixed is not a conservative choice — it is the more expensive and more disruptive option applied to a problem the cheaper option was built for. And the reverse mistake exists too: optimizing indefinitely against a workload whose growth trajectory makes the current platform structurally unaffordable regardless of governance quality, because nobody wants to have the harder migration conversation.

This article's job is to separate those two situations using evidence, not intuition.

Five options, not two

Treat the decision as a ladder, evaluated in order, rather than a single yes/no on migration.

1. Renew largely as-is. Appropriate when current commitments and forecast demand are already aligned and the products in use continue to earn their cost. This is the least common finding for organizations motivated enough to be reading this, but it is a legitimate outcome of the analysis, not a failure to find savings.

2. Optimize Datadog. Appropriate when the evidence shows avoidable consumption: logs indexed without exclusion filters, custom metrics carrying unbounded tag cardinality, non-production environments monitored at production intensity, no usage attribution by team or service. This is the default first move for most organizations, because it is reversible, fast, and does not require replacing anything the team already knows how to operate.

3. Hybrid architecture. Appropriate when different telemetry types or workloads have genuinely different economics — for example, integrated application observability staying on Datadog while high-volume, long-retention logs move to a lower-cost tier or a different backend, or while a security-specific dataset moves to a platform better suited to that workload. A hybrid architecture is not a compromise adopted to avoid a hard decision; it is the correct architecture when the workload itself is not homogeneous.

4. Migrate selected workloads. Appropriate when a specific, well-bounded telemetry type has structurally poor economics on the current platform but the rest of the estate does not. Moving that workload alone captures most of the benefit at a fraction of the risk and engineering cost of a full platform migration.

5. Full platform migration. Appropriate when the platform itself — not its configuration — is the problem across the majority of the estate, and the evidence supports a positive three-year outcome after migration engineering, dual-running, and retraining costs are included.

Each rung down this ladder costs more, takes longer, and is harder to reverse than the one above it. The evidence bar required to justify moving to the next rung should rise accordingly.

The evidence that actually distinguishes optimize from migrate

Avoidable consumption versus structural cost

Avoidable consumption is telemetry that provides little or no operational value relative to its cost: debug-level logs indexed by default, custom metrics carrying request IDs or customer IDs as tags, non-production environments monitored identically to production. This is a governance failure, and it is present in the overwhelming majority of Datadog estates that have grown organically without a cost review. The detailed mechanics of finding it are covered in Datadog Bill Too High? 12 Configuration Mistakes That Inflate Observability Costs; the log-specific version is in Datadog Logs Pricing: Ingestion, Indexing and Retention Explained; the metric-cardinality version is in Datadog Custom Metrics and Cardinality: How Costs Grow.

Structural cost is different: it is the cost of telemetry the organization genuinely needs, priced by a platform whose billing model no longer matches how that telemetry is generated. A platform engineering team running thousands of short-lived Kubernetes workloads, for example, may generate legitimate, high-value telemetry at a volume and cardinality that no amount of governance discipline reduces below a certain floor — see How Kubernetes Labels and Container Churn Increase Observability Costs for why dynamic infrastructure changes this floor specifically.

The test: after a genuine optimization pass — not a hypothetical one, an executed one — does the remaining cost still grow faster than the business it describes? If yes, the driver is structural. If the bill falls back in line with business growth once governance is applied, the driver was avoidable consumption, and migration would not have produced a better outcome than the optimization already did.

Contract timing

A renewal deadline creates commercial leverage. It does not create technical readiness, and it is a poor reason by itself to accelerate a migration decision that the workload evidence does not yet support. Where renewal timing is the actual pressure, How to Audit Datadog Costs Before Your Next Renewal covers how to build a renewal position from evidence rather than deadline pressure, including how to buy time for a proper migration evaluation without signing a long-term commitment the organization is not confident in.

Engineering capability and platform-team capacity

Migration away from an integrated SaaS platform generally increases the organization's operational surface, whatever the target. An organization without an established platform-engineering practice capable of running the target architecture is not comparing "expensive SaaS" against "cheaper self-managed platform." It is comparing a known cost against a project that will consume engineering capacity the organization does not currently have, on top of whatever the new platform ends up costing to run. This capability gap does not disqualify migration — but it belongs in the decision as a real cost, not an assumption that the team will simply absorb the work.

Feature and workflow dependencies

Every month an organization operates on a platform, it accumulates dependencies specific to that platform: dashboards, monitors, unified service tagging conventions, correlation between signals, and operational habits built around the product's particular workflow. These dependencies are real migration cost even when they never appear on an invoice, and they grow the longer the decision is deferred. This is not an argument to migrate early to avoid future cost — it is a reason to price the dependency honestly rather than assuming migration effort stays constant over time.

Expected growth trajectory

A workload growing 15% a year and a workload doubling every six months face structurally different decisions even with identical avoidable-consumption findings today. The growth trajectory determines how quickly a structural-cost finding, if one exists, will re-emerge after optimization — which is why the decision needs a forward demand view, not only a backward-looking usage review.

Decision matrix

Evidence pattern Most likely position on the ladder
Bill grew faster than usage, and no cost review has been performed Optimize — start here regardless of any other finding
Optimization executed, remaining cost still tracks business growth Renew as-is, or right-size the commitment
Optimization executed, one workload type remains structurally expensive Migrate that workload only, or adopt a hybrid architecture
Optimization executed, the majority of the estate remains structurally expensive relative to a credible target Evaluate full migration, with a genuine three-year TCO model — see Datadog vs OpenSearch: A Practical TCO Comparison
Renewal deadline is the primary pressure, workload evidence is incomplete Negotiate time or a shorter commitment; do not let the deadline substitute for evidence
Organization lacks platform-engineering capacity for the migration target Weight heavily toward optimize or hybrid; price the capability gap explicitly if migration is still pursued
Cost driver is concentrated in one telemetry type or one team Migrate selected workloads rather than the full estate

Why "do not migrate to solve a configuration problem" needs a caveat

The rule that opens most cost-optimization conversations — do not migrate to solve a configuration problem — is correct as a default, and it prevents the majority of premature migration decisions. It is not universally true.

There is a point at which continuing to optimize stops producing meaningful results, because the remaining cost is not a configuration artifact anymore. An organization that has executed real governance work — index and retention design, cardinality control, non-production right-sizing, usage attribution — and still finds the bill growing in proportion to genuine, necessary telemetry has exhausted what optimization can do. At that point, treating every renewal cycle as another optimization exercise is its own mistake: it consumes engineering time on diminishing returns while deferring a platform decision the evidence already supports.

The discipline this requires is honest measurement of what optimization actually achieved, not an assumption in either direction. Organizations that skip the optimization step and migrate immediately usually rediscover the same governance gaps on the new platform within a year. Organizations that never stop optimizing, on a workload whose economics no longer fit the platform, pay the structural cost indefinitely while calling it a work in progress.

Migration should not be used as a negotiation prop

A related failure mode deserves its own caution: raising migration as a stated intention primarily to extract a better renewal discount, without a genuine evaluation behind it. Vendors are generally capable of distinguishing a credible migration threat from a negotiating tactic, and a bluff that is called leaves the organization renewing anyway, from a weaker position, having spent the credibility a future genuine evaluation would need. If migration evidence exists, use it. If it does not yet exist, build it before the conversation, or negotiate on the actual evidence available — avoidable consumption, commitment right-sizing, contract structure — rather than an unbuilt migration case.

Time to value and reversibility as explicit decision inputs

Two properties of each option on the ladder deserve to be scored explicitly rather than left implicit.

Time to value — optimization changes typically show results within one to two billing cycles; a hybrid architecture takes longer because it requires standing up a second collection and storage path; a full migration is typically a multi-quarter program before the new platform is fully load-bearing. An organization under near-term budget pressure should weight this heavily, because a migration that pays back in year three does not help a renewal happening in ninety days.

Reversibility — optimization is fully reversible; a hybrid architecture is reversible per workload; a full migration becomes progressively less reversible as dashboards, alerts and team habits accumulate on the new platform. Reversibility matters most when the evidence supporting the decision is not yet strong — the less certain the evidence, the more that uncertainty should be reflected in choosing a more reversible option.

When this decision needs an outside review

An internal team is well positioned to execute optimization — the mechanics are documented and largely mechanical once the governance gaps are found. An internal team is less well positioned to make an unbiased optimize-versus-migrate call, because the people running the current platform have a natural stake in the outcome, and the people advocating for a new platform often have not yet had to operate one at the scale being discussed.

A vendor-neutral Observability Migration Assessment applies the ladder above to the organization's actual usage, contract and platform-team capacity, and states explicitly which rung the evidence supports — including the uncomfortable case where the honest answer is "the evidence does not yet exist to justify migration, and here is what would need to be true for it to."

FAQ

Start with optimization, because most Datadog bills contain avoidable consumption — unindexed-by-default logs, unbounded custom-metric cardinality, non-production environments monitored like production — that governance fixes quickly and reversibly. Migration is justified only when a genuine optimization pass has been executed and the remaining cost still grows faster than the business it describes, indicating a structural platform-economics problem rather than a configuration one.

A hybrid architecture keeps some telemetry on the current platform while moving other telemetry — often high-volume, long-retention logs or a specific security dataset — to a platform better suited to its economics. It makes sense when the workload is not homogeneous: different telemetry types genuinely have different cost profiles, and forcing all of them onto one platform is itself a source of avoidable cost.

Execute a real optimization pass — not a hypothetical one — covering log indexing and retention, custom-metric cardinality, non-production governance and usage attribution. If the remaining cost, after those changes, still tracks the growth of the underlying business or workload rather than growing faster than it, the driver is structural. If the cost falls back in line with business growth once governance is applied, it was avoidable consumption, and migration would not have produced a materially better outcome.

Rarely, but not never. If the workload's growth trajectory means the structural cost will clearly re-emerge shortly after optimization completes, and the evidence for that trajectory is strong, deferring a migration decision purely to finish an optimization exercise with diminishing returns can cost more in delay than it saves in due diligence. This is the exception, not the default, and it requires the growth evidence to be genuinely strong.

Only if the evaluation is real. Vendors can generally distinguish a credible migration case from a negotiating tactic, and an empty threat that gets tested leaves the organization renewing from a weaker position. If genuine migration evidence exists, it is legitimate context for a renewal conversation. If it does not exist yet, build it first, or negotiate on evidence that does exist — avoidable consumption, commitment sizing, contract structure.

At minimum: normalized historical usage by product and billing dimension, avoidable consumption identified and, where possible, already remediated, the growth trajectory of the underlying workload, the organization's platform-engineering capacity for the migration target under consideration, and contract timing. The decision should be built from that evidence, not from the size of the latest invoice alone.

Sources

  • FinOps Foundation — forecasting and rate-optimization guidance on combining historical usage with planned architecture and demand changes.
  • Datadog — Usage Attribution and Usage Metering documentation (evidence base for distinguishing avoidable from structural consumption).
  • AWS — total cost of ownership guidance for Amazon OpenSearch Service (operational management as a TCO component when migration is the option under evaluation).

This framework is a decision aid, not a substitute for a workload-specific evidence review. Validate usage, contract and platform-capacity assumptions against your organization's actual data before committing to any option on the ladder.

Not sure whether your Datadog problem is governance or architecture?

A vendor-neutral assessment applies this decision ladder to your actual usage, contract and platform-team capacity, and states plainly which option the evidence supports.