Most organizations considering a Splunk migration start with a question such as:
“What should we replace Splunk with?”
That question is usually too early.
A large Splunk estate is rarely just a place where logs are stored.
It may contain years of SPL searches, dashboards, alerts, macros, lookups, field extractions, accelerated data models, private applications, security detections, ITSI content, operational workflows and integrations that other systems depend on.
So the real question is not:
Which platform is cheaper than Splunk?
It is:
Which Splunk workloads should move, where should they move, and what cost and operational responsibility will move with them?
Datadog, Elastic and OpenSearch can all replace important parts of a Splunk estate.
They do not replace them in the same way.
Datadog moves more of the backend operating responsibility toward a SaaS provider.
Elastic offers several operating models and strong search capabilities, with assisted migration paths for selected Splunk Security content.
OpenSearch provides Apache 2.0 licensing and significant infrastructure control, particularly in AWS-oriented environments, but generally requires more ownership of the platform and more reconstruction of Splunk-specific content.
And sometimes the best decision is not to replace Splunk yet.
If an organization has years of deeply interconnected SPL, Enterprise Security content, private apps and operational processes, reducing the expensive parts of the Splunk estate first may produce a better three-year outcome than forcing a complete migration.
The decision is therefore not simply:
Splunk vs Datadog vs Elastic vs OpenSearch.
It is:
current workload → dependencies → target architecture → migration effort → operating model → three-year TCO → risk
That is the framework this article develops.
Executive summary
A credible Splunk migration decision requires five conclusions.
First, define what “Splunk” actually means inside your organization.
A Splunk estate may combine centralized logging, operational search, Enterprise Security, ITSI, Observability Cloud, custom applications and years of knowledge objects.
Those workloads do not necessarily belong on the same target platform.
Second, separate product choice from operating-model choice.
Choosing OpenSearch does not tell you whether it should be self-managed or consumed through Amazon OpenSearch Service.
Choosing Elastic does not tell you whether the target should be Elastic Cloud Hosted, Serverless, ECK or a self-managed deployment.
A self-managed platform cannot be compared fairly with Datadog SaaS until the engineering cost of operating it is included.
Third, normalize the economics before comparing prices.
Splunk, Datadog, Elastic and OpenSearch charge for different things.
An apparently cheaper unit price can simply mean that cost has moved from licensing into infrastructure, engineering, storage, SaaS consumption or operational risk.
Fourth, treat Splunk content migration as a separate engineering problem.
Moving telemetry is often easier than reproducing the behavior built around that telemetry.
SPL, macros, lookups, dashboards, detections, data models, alerts, RBAC and investigation workflows can dominate migration effort.
Fifth, compare migration against optimization.
The alternative to a full migration is not always “do nothing.”
An organization can reduce what reaches Splunk, route lower-value telemetry elsewhere, change retention, introduce OpenTelemetry or another neutral pipeline, migrate selected workloads and preserve the Splunk capabilities that remain difficult to replace.
The objective is not to find the platform with the lowest advertised price.
It is to determine where cost, responsibility and dependency should live after the migration.
What does a Splunk migration actually mean?
A Splunk migration is the controlled movement of selected data, searches, dashboards, alerts, security content and operational workflows from a Splunk estate to another platform or architecture.
The word selected matters.
A migration can mean:
- replacing Splunk completely,
- moving only observability workloads,
- moving operational logs while retaining Splunk Enterprise Security,
- moving historical data to cheaper storage,
- introducing another platform for new applications,
- or gradually shrinking Splunk behind a vendor-neutral telemetry pipeline.
These are materially different programs.
Consider a large bank as an illustrative example.
Splunk may simultaneously support application logs, infrastructure troubleshooting, fraud-related searches, SOC detections and audit retention.
Moving application logs to another observability backend does not automatically mean that the bank should migrate its security detections at the same time.
Now consider an industrial company.
Splunk might contain operational logs from IT systems, plant integrations, infrastructure alerts and years of custom dashboards used by support teams.
The optimal target for current production telemetry may be different from the optimal destination for historical data.
And in a large e-commerce company, Splunk might be heavily used by platform engineering while APM already lives on another platform.
In that case, the migration problem may be primarily about logs rather than the entire observability architecture.
These examples are deliberately different.
There is no industry-specific Splunk migration pattern.
The first task is always to identify what Splunk is doing in your organization today.
The first mistake: comparing products before defining the workload
It is tempting to build a spreadsheet with four columns:
| Platform | Cost | Features | Migration effort |
|---|---|---|---|
| Splunk | |||
| Datadog | |||
| Elastic | |||
| OpenSearch |
That spreadsheet is not useful yet.
Before comparing targets, separate the current Splunk estate into workload categories.
A useful initial decomposition is:
Operational logging
Application logs, infrastructure logs, troubleshooting and ad hoc search.
Observability
Logs connected with metrics, traces, services, alerts and incident investigation.
Security analytics
CIM mappings, correlation searches, risk objects, detections, threat intelligence and analyst workflows.
Search and analytics
Business or operational search use cases built around SPL.
Long-term retention
Data retained primarily for historical, regulatory or compliance reasons.
Custom Splunk capabilities
Private applications, custom commands, macros, lookups and workflows that have grown around the platform.
Only after that decomposition does the product comparison become meaningful.
The target for operational logs might be Datadog.
The SIEM target might remain Splunk temporarily.
Historical data might move to object storage.
Another organization might decide to move both logging and security analytics to Elastic.
A third might standardize telemetry around OpenTelemetry and OpenSearch while retaining a much smaller Splunk estate.
The migration architecture does not need to be symmetrical.
Two decisions are usually being confused
A Splunk replacement involves two separate decisions.
Decision 1: which platform fits the workload?
The shortlist might include:
- Datadog,
- Elastic,
- OpenSearch,
- an optimized Splunk estate,
- or a multi-platform architecture.
Decision 2: who operates the target?
The target might be:
- SaaS,
- a vendor-managed cluster,
- a serverless service,
- a self-managed virtual-machine deployment,
- or a self-managed Kubernetes platform.
These decisions must be separated because they change the economics dramatically.
Suppose an organization concludes:
“OpenSearch is significantly cheaper than Datadog.”
That statement tells us almost nothing without the operating model.
If the OpenSearch platform requires infrastructure capacity, Kubernetes operations, upgrades, snapshots, disaster recovery, performance tuning, incident response and several engineers to support it, those costs belong in the comparison.
Likewise:
“Datadog is expensive.”
may be true from a direct platform-spend perspective while still producing a competitive TCO if it removes a substantial amount of backend operational work.
The commercial invoice is only one part of the architecture.
Where each target is structurally strongest
There is no universal ranking.
A better starting point is to ask what the organization is trying to optimize.
| Target | Structurally strongest when | Main trade-off |
|---|---|---|
| Datadog | Integrated SaaS observability and low backend operational burden are priorities | Multiple usage dimensions require disciplined cost governance |
| Elastic | Search analytics, deployment flexibility and selected Splunk Security migration capabilities matter | Cost and operating responsibility vary significantly between deployment models |
| OpenSearch | Apache 2.0 licensing, AWS alignment or infrastructure control are strategic | More platform ownership and Splunk-content reconstruction may be required |
| Optimized Splunk | Existing SPL, Enterprise Security, private apps and knowledge objects create significant dependency | Licensing and workload costs may remain high |
| Multi-platform architecture | Logs, observability, security and archive workloads have different economics | Governance and cross-platform workflows become more complex |
These are starting positions.
They are not recommendations.
Datadog as a Splunk migration target
Datadog changes the operating model more than it simply changes the search engine.
With a traditional self-managed search platform, engineering teams think about things such as:
- cluster capacity,
- nodes,
- shards,
- storage,
- lifecycle,
- upgrades,
- recovery,
- and backend performance.
Datadog moves most of that responsibility behind a managed SaaS service.
The operational focus moves toward:
- agents,
- integrations,
- tagging,
- service ownership,
- sampling,
- log indexes,
- exclusion filters,
- archives,
- custom metrics,
- spans,
- usage commitments,
- and cost governance.
That distinction matters.
The migration is not:
Splunk cluster → Datadog cluster
It is closer to:
search-platform operations → SaaS consumption governance
For an organization whose platform team is already overloaded, that can be extremely valuable.
For an organization whose primary objective is minimizing variable SaaS consumption and that already has a strong internal search-platform team, it may be less attractive.
Where Datadog is most compelling
Datadog is usually a strong starting point when:
- observability is the dominant workload,
- SaaS simplicity is valuable,
- application and infrastructure telemetry should live in one operational experience,
- the organization wants to minimize backend cluster management,
- and it is prepared to govern several independent consumption dimensions.
The important economic warning
Datadog should not be evaluated using one price.
Different products can expose different billing dimensions.
A complete cost model may need to consider:
- infrastructure hosts,
- containers,
- APM,
- ingested logs,
- indexed events,
- retained logs,
- indexed spans,
- custom metrics,
- pipeline processing,
- security products,
- archives,
- rehydration,
- and support.
This means Datadog can simplify infrastructure operations while making usage governance more important.
The complexity has moved.
It has not disappeared.
Elastic as a Splunk migration target
“Elastic” is not one operating model.
That distinction is essential.
An organization can use:
- Elastic Stack self-managed,
- Elastic Cloud on Kubernetes,
- Elastic Cloud Hosted,
- or Elastic Cloud Serverless.
They share an Elasticsearch foundation.
They do not have identical economics or responsibilities.
A self-managed Elastic deployment gives the organization significant control.
It also leaves the organization responsible for much more of the platform lifecycle.
Elastic Cloud Hosted removes part of that responsibility.
Serverless changes the consumption model again.
So a statement such as:
“We compared Splunk with Elastic.”
is incomplete.
The real comparison needs to state which Elastic deployment model was evaluated.
Where Elastic is most compelling
Elastic is a strong candidate when the organization values:
- search and analytics capabilities,
- deployment flexibility,
- the Elasticsearch ecosystem,
- the ability to choose between managed and self-managed operating models,
- or assisted migration for selected Splunk Security assets.
Elastic also provides a clearer first-party assisted path than the other targets in this comparison for some Splunk Security content.
That does not make a Splunk migration automatic.
It changes one part of the migration equation.
OpenSearch as a Splunk migration target
OpenSearch also needs to be split into several operating models:
- self-managed OpenSearch,
- provisioned Amazon OpenSearch Service,
- Amazon OpenSearch Serverless.
These are not economically equivalent.
The OpenSearch project is distributed under Apache License 2.0.
That can matter strategically for organizations that want broad software-use rights, infrastructure control or alignment with AWS.
But one misconception needs to disappear immediately:
No software-license fee does not mean no platform cost.
A self-managed OpenSearch platform still requires:
- compute,
- storage,
- networking,
- observability,
- backups,
- upgrades,
- security,
- engineering,
- support,
- incident response,
- and capacity management.
Apache 2.0 changes the licensing model.
It does not operate the system.
Where OpenSearch is most compelling
OpenSearch becomes particularly interesting when:
- software licensing freedom matters,
- AWS integration is strategically important,
- the organization already has strong search-platform engineering capability,
- infrastructure control is preferred over SaaS abstraction,
- or the workload economics favor infrastructure-based operation.
The trade-off is that more responsibility generally stays with the customer.
Staying on Splunk is also an architectural option
Migration discussions often create an artificial binary choice:
migrate or accept the current situation.
There is a third option:
change the Splunk estate without replacing all of Splunk.
For example, an organization might:
- filter low-value data before indexing,
- remove duplicate telemetry,
- reduce excessive retention,
- route archive-oriented data elsewhere,
- migrate selected operational workloads,
- introduce OpenTelemetry or another neutral pipeline,
- optimize expensive searches,
- review acceleration,
- renegotiate the commercial model,
- and preserve the Splunk workloads whose migration risk remains high.
This can be particularly rational when the estate contains significant:
- SPL,
- Enterprise Security content,
- private applications,
- macros,
- lookups,
- accelerated data models,
- or custom operational workflows.
A partial exit is not a failed migration.
It can be the target architecture.
The invoice is not the TCO
The pricing models of these platforms are not directly comparable.
Splunk may charge based on indexed-data volume or workload capacity depending on the product and agreement.
Datadog uses several product-specific consumption dimensions.
Elastic Cloud Hosted is largely capacity-oriented, while Elastic Cloud Serverless introduces consumption-based dimensions.
Self-managed Elastic combines infrastructure, subscription and engineering cost.
OpenSearch software has no software-license fee, while managed AWS services charge for the infrastructure or serverless capacity consumed.
These are fundamentally different economic models.
That is why comparing advertised units directly often produces misleading conclusions.
A useful three-year TCO model
A defensible comparison can start with:
Three-year TCO = platform charges + cloud infrastructure + support + implementation + coexistence + migration engineering + ongoing platform labor + risk contingency − negotiated discounts and committed-use benefits
Each term matters.
Ignore platform labor and self-managed platforms look artificially cheap
Infrastructure is not the entire cost of Elastic or OpenSearch.
Someone must operate the platform.
Ignore SaaS consumption dimensions and managed platforms look artificially predictable
A SaaS platform can have several cost drivers growing independently.
Ignore coexistence and every migration looks artificially cheap
For part of the migration, both platforms may run simultaneously.
That period has a cost.
Ignore migration engineering and “moving logs” appears to be the migration
It is not.
The expensive work may sit in schema redesign, SPL conversion, dashboard validation, detections, RBAC and operational workflows.
Build the workload baseline before calculating anything
Every target should be evaluated against the same workload.
At minimum, the migration baseline should understand:
- logs, metrics and traces by volume,
- peak versus average ingestion,
- enrichment expansion,
- host and container counts,
- metric-series behavior,
- span volume,
- retention by data class,
- search concurrency,
- scheduled searches,
- active users,
- security workloads,
- compliance requirements,
- availability requirements,
- and recovery objectives.
Without that baseline, pricing comparisons are mostly estimates of different things.
Pricing models by platform
| Platform/model | Main pricing dimension | Main forecasting risk |
|---|---|---|
| Splunk ingest pricing | Indexed-data volume | Daily indexed volume growth |
| Splunk workload pricing | Applicable workload capacity | Search, indexing and workload behavior |
| Datadog | Multiple product-specific usage dimensions | Several usage dimensions can grow independently |
| Elastic Cloud Hosted | Provisioned capacity, transfer and snapshots | Required capacity and topology |
| Elastic Cloud Serverless | Consumption such as ingest, retention and egress | Billed consumption growth |
| Self-managed Elastic | Infrastructure, subscription and staffing | Platform labor and capacity headroom |
| Self-managed OpenSearch | Infrastructure, staffing and support | Engineering and reliability effort |
| Amazon OpenSearch Service | Managed infrastructure and storage | Shards, nodes and retained capacity |
| OpenSearch Serverless | Serverless capacity and managed storage | Workload behavior and deployment mode |
The purpose of this table is not to determine which vendor is cheapest.
It is to show why a single “price per GB” cannot normalize the decision.
Cost does not disappear. It changes location.
This is one of the most important ideas in any Splunk migration.
Consider the same operational logging workload across three target models.
With a SaaS target, more cost may appear directly in the vendor invoice.
With a managed search service, part of the cost moves into cloud infrastructure.
With a self-managed platform, the direct software bill may become smaller while engineering and operational responsibility increase.
The architecture can therefore move cost between:
software licensing
SaaS consumption
cloud infrastructure
storage
data transfer
engineering labor
support
operational risk
The correct comparison is the combined system.
Not the invoice that is easiest to see.
Moving the data is often not the hardest part
This is where many Splunk migration estimates fail.
A Splunk estate may contain:
- indexes,
- sourcetypes,
- dashboards,
- saved searches,
- reports,
- alerts,
- macros,
- lookups,
- field extractions,
- event types,
- tags,
- workflow actions,
- custom commands,
- accelerated data models,
- summary indexes,
- security detections,
- risk objects,
- and integrations.
Data can be copied.
Behavior has to be reproduced.
Consider a hypothetical enterprise operations team.
They have a dashboard that appears to contain ten charts.
From a migration spreadsheet, that may look like:
1 dashboard to migrate.
But the dashboard may rely on several SPL searches.
Those searches may rely on:
- macros,
- lookup tables,
- search-time field extraction,
- aliases,
- calculated fields,
- accelerated data models,
- and specific timestamp semantics.
The actual migration unit is not the dashboard.
It is the dependency graph behind the dashboard.
That is why counting dashboards and alerts is not enough.
Schema migration is often harder than dashboard migration
A dashboard can render successfully on the target platform and still be wrong.
Suppose, as an illustrative example, that a company migrates an operational dashboard from Splunk to another platform.
The visualization shows CPU alerts, application errors and transaction failures.
Everything appears visually correct.
But Splunk originally created several fields through search-time extraction and enriched others through lookups.
If the target schema produces different fields or different event classifications, the new dashboard can return different results while looking almost identical.
A successful rendering proves almost nothing.
The migration must validate:
- fields,
- filters,
- aggregations,
- timestamps,
- source completeness,
- query semantics,
- and returned results.
Splunk-specific schemas and conventions may also need to be mapped to target conventions.
Elastic commonly uses ECS.
OpenSearch environments may use SS4O or organization-specific schemas.
Datadog relies heavily on standard attributes, tagging and service identity.
The target therefore needs an explicit schema for concepts such as:
- timestamp,
- service,
- environment,
- host,
- container,
- user,
- asset,
- severity,
- trace,
- span,
- network fields,
- security categories,
- ownership,
- and data classification.
This work should happen before hundreds of dashboards are mechanically recreated.
OpenTelemetry can reduce one layer of lock-in
OpenTelemetry is useful during a Splunk migration because it can separate instrumentation and telemetry transport from the final backend.
It can help an organization:
- standardize application instrumentation,
- route telemetry to more than one platform,
- process telemetry before storage,
- sample data,
- enrich metadata,
- and support coexistence during migration.
Conceptually:
Applications / infrastructure
↓
OpenTelemetry
↓
Processing / routing
↙ ↘
Splunk Target
This can be extremely useful during a phased migration.
New applications do not necessarily need to become tightly coupled to the old platform while the migration is still running.
But OpenTelemetry solves only part of the problem.
It does not automatically make portable:
- dashboards,
- alerts,
- monitors,
- detection rules,
- RBAC,
- saved searches,
- query languages,
- investigation workflows,
- or vendor-specific enrichment.
OpenTelemetry can reduce instrumentation and transport lock-in.
It does not eliminate content, schema or operating-model lock-in.
What can actually be automated?
No target should be assumed to provide a universal, lossless converter for an enterprise Splunk estate.
That does not mean all migration work is manual.
It means the scope of automation needs to be understood precisely.
Splunk to Elastic
Elastic provides an assisted migration path for selected Splunk Security assets.
Elastic Automatic Migration can help translate supported Splunk rules and dashboards into Elastic Security.
That can remove meaningful manual work.
It does not mean:
“Splunk Enterprise Security can be automatically migrated to Elastic.”
The complete migration still includes:
- data pipelines,
- target schemas,
- identities,
- actions,
- investigations,
- operational processes,
- and validation.
Elastic therefore receives an advantage for supported migration tooling.
It should not receive an assumption of complete functional equivalence.
Splunk to Datadog
Datadog provides several useful migration building blocks, including APIs, observability pipelines, dual-routing patterns, Splunk-forwarder integration and archive/backfill capabilities.
A partner Splunk-to-Datadog migration service can also convert selected dashboards, queries and monitors.
Again, the distinction matters.
That is useful tooling.
It is not evidence that an entire Splunk estate can be converted automatically.
For organizations planning coexistence, Datadog Observability Pipelines can also be used to split log traffic between destinations.
That can support a gradual migration rather than an immediate cutover.
Splunk to OpenSearch
A Splunk-to-OpenSearch migration normally requires a different expectation.
AWS Migration Assistant for Amazon OpenSearch Service can support migrations from Elasticsearch, OpenSearch and Apache Solr into supported OpenSearch targets.
That is useful for the migrations it supports.
It does not translate Splunk-specific assets such as:
- SPL,
- macros,
- lookups,
- dashboards,
- Enterprise Security content,
- or knowledge objects.
A Splunk-to-OpenSearch program should therefore assume more custom content reconstruction or translation.
The lower software licensing constraint should not be confused with lower migration complexity.
They are separate dimensions.
Build the dependency inventory before selecting the winner
Before committing to Datadog, Elastic, OpenSearch or even a full migration, inventory the Splunk estate.
The purpose is not bureaucratic documentation.
The purpose is to identify what the organization is actually dependent on.
Data inventory
Understand:
- indexes,
- sourcetypes,
- ownership,
- average and peak volume,
- retention,
- legal hold,
- parsing,
- index-time transformation,
- data quality,
- duplicates,
- and late-event behavior.
Content inventory
Count and classify:
- dashboards,
- searches,
- alerts,
- reports,
- macros,
- lookups,
- field extractions,
- event types,
- tags,
- workflow actions,
- and custom commands.
Do not stop at object counts.
Capture dependencies.
Performance inventory
Identify:
- accelerated data models,
- summary indexing,
- scheduled-search concurrency,
- real-time searches,
- expensive ad hoc searches,
- and search-head pressure.
Some Splunk constructs exist because of the architecture of the current platform.
They should not always be reproduced literally on the target.
Some should be redesigned.
Security inventory
If security is in scope, inventory:
- CIM mappings,
- correlation searches,
- risk objects,
- suppression,
- notable-event workflows,
- threat intelligence,
- response actions,
- investigation processes,
- and audit requirements.
This work deserves its own migration stream.
Platform inventory
Finally, understand:
- forwarders,
- HEC integrations,
- applications,
- add-ons,
- private apps,
- authentication,
- RBAC,
- service accounts,
- APIs,
- support,
- and disaster recovery.
If this inventory does not exist, target selection is premature.
A realistic Splunk migration sequence
A migration should progressively reduce uncertainty.
That is very different from selecting a product and starting to copy dashboards.
Phase 1 — Discovery and dependency mapping
Create the inventory.
Identify owners.
Build the dependency graph.
Then classify existing content into categories such as:
- retire,
- temporarily retain in Splunk,
- migrate automatically,
- migrate with human correction,
- redesign,
- or replace with target-native functionality.
This classification can eliminate large amounts of unnecessary migration work.
There is no reason to reproduce an obsolete dashboard on the new platform simply because it exists on the old one.
Phase 2 — Define the target architecture
Before migrating content, decide:
- target products,
- operating model,
- schema,
- tenancy,
- RBAC,
- retention,
- availability,
- recovery,
- support,
- and cost guardrails.
This is also where the organization decides whether it is building one target or several.
Phase 3 — Run a representative pilot
Do not select only easy workloads.
That creates a successful demonstration rather than a useful migration test.
A representative pilot should include different types of difficulty.
For example:
- a high-volume data source,
- a complex SPL search,
- an operational alert,
- an important dashboard,
- a security detection if security is in scope,
- a macro or lookup dependency,
- historical search,
- and retention requirements.
The objective is to discover the hard parts before the migration is scaled.
Phase 4 — Migrate enough content to test equivalence
Translate enough searches, dashboards and alerts to test the target architecture.
This is also where schema problems become visible.
Do not postpone metadata and schema work until after dual ingestion starts.
By then, incorrect conventions may already have been propagated across the target.
Phase 5 — Dual routing or dual ingestion
Selected telemetry can now be sent to both systems.
Conceptually:
Telemetry
↓
Processing / routing layer
↓
┌───────────────┐
│ │
Splunk New platform
This creates a period during which the two platforms can be compared.
But dual running is not free.
Budget for:
- duplicated ingestion,
- processing,
- transfer,
- target storage,
- source storage,
- monitoring,
- and reconciliation.
Migration business cases often underestimate this phase because the future-state architecture is modeled while the transition-state architecture is ignored.
The organization pays for the transition too.
Phase 6 — Decide what historical data actually needs to move
Historical migration can become one of the least economically justified parts of a program.
Do not assume every retained byte must be copied into the new searchable platform.
Ask:
- Does it need to remain searchable?
- For how long?
- For which users?
- For which compliance obligation?
- Can it remain in immutable archive storage?
- Can only the required searchable window be migrated?
The cheapest historical migration is often the data that does not need to migrate.
Phase 7 — Migrate in waves
Group workloads around useful boundaries.
For example:
- business service,
- application family,
- data domain,
- security use case,
- engineering team,
- regulatory boundary,
- or technical dependency.
This makes ownership and rollback clearer than an arbitrary migration of thousands of independent objects.
Phase 8 — Perform acceptance testing
Migration acceptance should measure behavior, not appearance.
Validate:
- event completeness,
- timestamps,
- fields,
- query-result variance,
- alert behavior,
- dashboard latency,
- detection latency,
- RBAC,
- retention,
- recovery,
- and cost.
A dashboard appearing on-screen is not an acceptance criterion.
Phase 9 — Controlled cutover
Define explicitly:
- who owns the cutover,
- the change window,
- success criteria,
- stop conditions,
- escalation,
- and stakeholder communication.
“Zero downtime” should also be defined precisely.
Does zero downtime mean:
- telemetry is never lost?
- dashboards remain available?
- searches remain available?
- alerts continue evaluating?
- security detections continue operating?
Without a definition, the phrase has little architectural value.
Phase 10 — Keep a rollback window
Do not decommission Splunk immediately after the first successful target deployment.
Keep the source available until:
- required data has been reconciled,
- alerts are stable,
- users have accepted the target,
- incident workflows have been tested,
- audit requirements are satisfied,
- and retention and contractual obligations are understood.
Decommissioning is part of the migration.
It is not simply procurement cancelling a renewal.
Splunk Enterprise Security should usually be treated separately
A Splunk Enterprise Security migration is not just a logging-platform migration.
Security content can include:
- CIM mappings,
- data models,
- correlation searches,
- risk-based alerting,
- suppression,
- notable-event workflows,
- threat intelligence,
- response actions,
- analyst roles,
- investigation processes,
- and audit evidence.
Moving those capabilities changes how a security team works.
The organization should therefore explicitly decide whether to:
- migrate security and observability together,
- keep Splunk ES while migrating operational logging,
- select another SIEM,
- redesign detections around a new schema,
- or operate a multi-platform transition.
For many enterprises, separating these decisions reduces risk.
The observability team does not need to wait for a complete SIEM transformation before changing the economics of operational logging.
And the SOC should not be forced into a security-platform migration solely because the platform team wants to move logs.
What operations look like after the migration
Migration architecture should be judged by the system that exists after the project team leaves.
That future operating model differs considerably by target.
Datadog
Most search-cluster operations disappear from the customer's responsibility.
The ongoing work shifts toward:
- tagging,
- agents,
- integrations,
- indexes,
- exclusions,
- retention,
- archives,
- spans,
- custom metrics,
- product usage,
- and contractual commitments.
The operating model becomes primarily SaaS governance.
That is lower infrastructure responsibility.
It is not zero governance.
Elastic Cloud
Elastic Cloud reduces infrastructure operations.
Teams still need to govern:
- data models,
- mappings,
- integrations,
- lifecycle policies,
- query efficiency,
- access,
- application-level upgrade implications,
- and cost.
The managed service operates significant parts of the platform.
It does not make architecture decisions for the organization.
Self-managed Elastic or ECK
More responsibility moves back to the internal platform organization.
That includes areas such as:
- VM or Kubernetes capacity,
- upgrades,
- snapshots,
- recovery,
- resource tuning,
- certificates,
- scaling,
- support procedures,
- and operator lifecycle.
This is why self-managed Elastic should not be modeled simply as:
Elastic license + cloud instances.
The platform team is part of the TCO.
Amazon OpenSearch Service
Managed OpenSearch reduces node-level administration.
The customer still makes important decisions around:
- domain or collection design,
- mappings,
- shard strategy,
- capacity,
- networking,
- access policies,
- KMS,
- upgrades,
- recovery,
- and cost controls.
Managed does not mean architecture-free.
Self-managed OpenSearch
This keeps the broadest operational surface inside the organization.
The platform team may own:
- infrastructure,
- orchestration,
- plugins,
- security,
- upgrades,
- backup,
- recovery,
- performance,
- support,
- and incident response.
Again:
Apache 2.0 changes software freedom.
It does not operate the platform.
A practical decision matrix
The decision can now be expressed more clearly.
| Dominant requirement | Likely starting point | Why |
|---|---|---|
| Minimize backend operational burden | Datadog | Integrated SaaS operating model |
| Assisted conversion of selected Splunk Security content | Elastic | Documented migration assistance for supported content |
| Maximum deployment flexibility | Elastic | Hosted, Serverless, ECK and self-managed options |
| Apache 2.0 software | OpenSearch | Open-source project licensing |
| AWS-oriented managed search | Amazon OpenSearch Service | Integration with the AWS operating model |
| Very large SPL/custom-app dependency | Optimize or phase Splunk exit | Migration dependency may dominate economics |
| Different economics for logs, APM and SIEM | Multi-platform architecture | Different workloads can have different optimal backends |
| Splunk renewal approaching but migration not ready | Optimize Splunk + introduce pipeline | Reduces immediate pressure without forcing an unsafe cutover |
This matrix should not determine the final architecture.
It should determine what deserves deeper evaluation.
Use a weighted scorecard instead of arguing about products
Product debates often become ideological.
One team prefers SaaS.
Another prefers open source.
Another wants to stay entirely on AWS.
Another is focused on security migration.
A weighted decision model forces these priorities into the same framework.
A reasonable starting structure is:
| Evaluation dimension | Example weight |
|---|---|
| Functional fit | 20% |
| Three-year TCO | 20% |
| Migration complexity | 15% |
| Operating-model fit | 15% |
| Security and compliance | 10% |
| Portability and future exit cost | 10% |
| Support and ecosystem | 5% |
| Strategic roadmap | 5% |
The weights should change with context.
A highly regulated organization may increase security and compliance.
A company with a small platform team may increase operating-model fit.
An organization six months from a major Splunk renewal may increase migration complexity and time-to-value.
The scorecard does not remove judgment.
It makes the judgment explicit.
Should you optimize Splunk before migrating?
Often, yes.
Migration should not prevent immediate optimization.
Before approving a backend replacement, inspect the current estate for:
- unused data,
- verbose sources,
- duplicated events,
- debug logs,
- excessive retention,
- low-value indexes,
- inefficient searches,
- real-time searches,
- acceleration costs,
- routing opportunities,
- archive strategy,
- and unused contractual entitlement.
Some of these changes create savings whether the organization eventually stays or leaves.
They can also reduce migration scope.
If 20 data sources are discovered to have little operational value, there is little benefit in paying to migrate them perfectly to another platform.
Optimization can therefore be part of migration preparation.
The most useful intermediate architecture is often a neutral pipeline
A telemetry pipeline can change the economics before the final backend decision is complete.
Instead of:
Sources
↓
Splunk
the organization can gradually move toward:
Sources
↓
Telemetry pipeline
↓
Processing / filtering / routing
↓
┌────────────┬────────────┬─────────────┐
│ │ │ │
Splunk Datadog Elastic OpenSearch
Not every organization needs four destinations.
That is not the point.
The architectural benefit is that collection and routing become less tightly coupled to the backend.
The organization can then:
- filter before expensive indexing,
- route different workloads differently,
- dual-send during migration,
- migrate applications in waves,
- and reduce the blast radius of future platform changes.
OpenTelemetry can play an important role in such an architecture.
So can vendor-specific or independent pipeline technologies.
The objective is optionality.
Common Splunk migration misconceptions
“OpenTelemetry eliminates vendor lock-in.”
No.
It can significantly reduce instrumentation and transport dependency.
It does not standardize dashboards, queries, alert definitions, RBAC, security detections or investigation workflows.
It removes one layer of lock-in.
Not all of them.
“Open source is automatically cheaper.”
No.
Open-source software can reduce or remove software-license cost.
Infrastructure, engineering, support and operational risk remain.
For an organization with an experienced internal platform team, the economics may be excellent.
For an organization that needs to create that capability from scratch, they may not be.
“Datadog log cost is simply the ingestion price.”
No.
A complete Datadog model may involve multiple product and usage dimensions.
The correct calculation needs the actual workload and proposed commercial agreement.
“Elastic is one deployment model.”
No.
Hosted, Serverless, ECK and self-managed deployments move cost and responsibility differently.
A serious comparison must name the actual target architecture.
“OpenSearch is a drop-in replacement for Splunk.”
No.
OpenSearch can provide the search and analytics foundation for many Splunk workloads.
Splunk-specific content such as SPL, macros, dashboards, knowledge objects and security workflows still needs translation or redesign.
“Serverless means we no longer need capacity or cost planning.”
Serverless removes part of the infrastructure planning problem.
It does not remove consumption planning.
Ingest, storage, query activity, egress, workload behavior and service limits still affect economics.
“If the migrated dashboard works, the migration works.”
No.
A dashboard can render and return the wrong data.
Acceptance must validate the result, not simply the visualization.
What should a company actually compare?
At this stage, the architecture review should answer a relatively small number of questions.
1. What are we actually using Splunk for?
Not what licenses were purchased.
What workloads depend on it?
2. Which workloads are candidates for migration?
Observability?
Logging?
Security?
Archive?
All of them?
3. Which dependencies make those workloads difficult to move?
SPL?
Apps?
Data models?
Detections?
Lookups?
Operational workflows?
4. Which target operating model fits the organization?
SaaS?
Managed service?
Self-managed?
Hybrid?
5. Where will cost move?
Licensing?
Cloud?
SaaS consumption?
Engineering?
Support?
6. What does the migration itself cost?
Discovery?
Translation?
Dual running?
Training?
Testing?
Historical data?
7. What would happen if we did not fully migrate?
Could optimization or a partial exit produce a better result?
Those questions lead to an architecture decision.
Product demos do not.
FAQ
What is the best alternative to Splunk?
There is no universal best alternative.
Datadog is generally strongest when managed SaaS observability and lower backend operational responsibility dominate the decision.
Elastic is a strong candidate when search capabilities, deployment flexibility or selected Splunk Security migration tooling matter.
OpenSearch is particularly attractive when Apache 2.0 licensing, AWS integration or infrastructure control are strategic priorities.
For estates with significant Splunk-specific dependencies, optimizing or partially exiting Splunk can also be the best near-term architecture.
Is Datadog cheaper than Splunk?
It can be.
It can also be more expensive for a particular workload.
The answer depends on actual hosts, containers, logs, metrics, traces, indexing, retention, custom metrics, security products and commercial terms.
Datadog's lower backend operational burden should also be included in the comparison.
The only useful method is to replay the real workload against the proposed contract.
Is Elastic cheaper than Splunk?
Potentially.
But “Elastic” must first be defined.
Elastic Cloud Hosted, Elastic Cloud Serverless, ECK and self-managed Elastic have different cost and operating models.
Infrastructure, subscription entitlements, support and platform engineering should all be included in TCO.
Is OpenSearch cheaper than Splunk?
OpenSearch removes the OpenSearch software-license cost.
It does not remove infrastructure, cloud-service charges, staffing, support or migration engineering.
The economics are strongest when the organization has the right engineering capability or when AWS and infrastructure-level control align well with the workload.
Can SPL dashboards and alerts be migrated automatically?
Some assets can be assisted or partially translated.
No target should be assumed to provide complete, lossless conversion of an enterprise Splunk estate.
Macros, lookups, custom commands, data models, field semantics and other dependencies generally require human validation.
Does OpenTelemetry make Splunk migration easier?
It can make instrumentation, routing and coexistence significantly easier.
It does not migrate existing SPL, dashboards, alerts, security content or RBAC.
Use it to reduce future backend dependency, not as a universal migration converter.
How long does an enterprise Splunk migration take?
It depends much more on scope and dependency than on raw data volume.
A relatively simple logging estate can be migrated much faster than a multi-business-unit environment containing Enterprise Security, private applications and thousands of interconnected knowledge objects.
The latter can become a multi-quarter transformation.
What are the most underestimated migration costs?
Usually:
- discovery,
- dependency mapping,
- schema normalization,
- SPL and content translation,
- coexistence,
- historical backfill,
- acceptance testing,
- training,
- support-model changes,
- and ongoing operation of the target platform.
Should Splunk Enterprise Security be migrated separately?
Often, yes.
A SIEM migration involves security schemas, detections, risk models, investigations, analyst processes and audit requirements.
Those are materially different from migrating application logs.
Separating the decisions can reduce migration risk.
Should we migrate before our Splunk renewal?
A renewal deadline creates commercial leverage.
It does not create technical readiness.
Start early enough to compare:
- Splunk optimization,
- partial migration,
- full migration,
- architectural coexistence,
- and contract renegotiation.
A forced cutover caused by procurement timing is not a migration strategy.
When a vendor-neutral Splunk architecture review is justified
A vendor-neutral review is most useful before the target has already been selected.
Once an organization has internally decided that “we are moving everything to platform X,” the assessment often becomes validation of that choice rather than architecture analysis.
A proper review should instead normalize four different things:
the current Splunk estate
the workload
the target architectures
the economics
That is particularly valuable when:
- annual Splunk spend is material,
- renewal is approaching,
- teams disagree on the target,
- Enterprise Security or ITSI is involved,
- the estate contains large amounts of SPL or saved content,
- ownership is unclear,
- no reliable workload inventory exists,
- SaaS is being compared with self-managed infrastructure,
- or vendor proposals use incompatible pricing models.
The deliverable should not simply be:
“We recommend Elastic.”
or:
“OpenSearch is 40% cheaper.”
It should explain why.
A useful assessment should produce:
- current-state inventory,
- dependency map,
- workload and cost baseline,
- target shortlist,
- licensing analysis,
- three-year TCO scenarios,
- target architecture,
- migration waves,
- coexistence design,
- acceptance criteria,
- rollback strategy,
- and an executive recommendation.
Who should review a Splunk architecture before migration?
The reviewer needs to understand both sides of the migration.
Understanding only Splunk is insufficient.
Understanding only Datadog, Elastic or OpenSearch is also insufficient.
The assessment needs to connect:
current Splunk architecture
→ commercial model
→ telemetry pipelines
→ content dependencies
→ target architectures
→ migration sequence
→ three-year economics
A reviewer focused entirely on the destination platform can easily underestimate the source-side dependency.
That is often where the migration complexity lives.
What should a Splunk migration consultant deliver?
A credible Splunk migration engagement should leave the organization with a decision model and an executable architecture.
That normally includes:
- current-state inventory,
- dependency mapping,
- normalized workload baseline,
- target options,
- three-year TCO,
- migration waves,
- coexistence architecture,
- acceptance tests,
- rollback criteria,
- and a final recommendation.
A platform demonstration is useful during evaluation.
It is not a migration assessment.
A price comparison is useful during procurement.
It is not a migration architecture.
Can Splunk cost be reduced without replacing Splunk?
Yes.
An organization may be able to reduce cost through:
- pre-ingestion filtering,
- duplicate reduction,
- retention changes,
- search optimization,
- acceleration review,
- lower-cost archive,
- routing changes,
- premium-product review,
- workload redistribution,
- and contract renegotiation.
These options should be quantified before a full replacement is approved.
Sometimes they are temporary measures that fund the migration.
Sometimes they change the business case enough that a complete migration is no longer necessary.
Both outcomes are valid.
From “Which vendor?” to “Which architecture?”
The wrong way to begin a Splunk migration is:
Splunk is expensive. Which platform should replace it?
The better sequence is:
What does Splunk do today?
↓
Which workloads should move?
↓
What dependencies do those workloads have?
↓
Which target architectures can replace them?
↓
What responsibility moves with each architecture?
↓
What does each scenario cost over three years?
↓
What migration sequence reduces risk?
Only then does the vendor decision become meaningful.
Datadog, Elastic and OpenSearch can all form credible parts of a post-Splunk architecture.
They optimize for different things.
Datadog transfers more backend responsibility to a managed SaaS platform and makes usage governance central to the operating model.
Elastic provides broad deployment flexibility and the clearest documented assisted path among these targets for selected Splunk Security assets.
OpenSearch provides Apache 2.0 software freedom and strong infrastructure control, particularly in AWS-oriented architectures, while generally placing more platform responsibility on the organization.
Splunk itself may remain part of the target architecture when existing content makes immediate replacement uneconomic.
And different workloads may legitimately end up on different platforms.
The goal is not to remove the Splunk logo from the architecture diagram.
The goal is to build an architecture in which the three-year cost, operational responsibility and migration risk are understood and intentional.
Vendor-Neutral Observability Migration Assessment
SAB Consulting helps organizations evaluate Splunk migrations across Datadog, Elastic, OpenSearch and hybrid architectures without beginning from a predetermined target vendor.
A Vendor-Neutral Observability Migration Assessment can examine:
- the current Splunk architecture,
- workload and usage baselines,
- SPL and knowledge-object dependencies,
- security and observability boundaries,
- target-platform options,
- licensing and operating-model implications,
- three-year TCO,
- telemetry-pipeline architecture,
- migration sequencing,
- coexistence,
- acceptance criteria,
- and rollback strategy.
The result may recommend a full migration.
It may recommend a phased migration.
It may recommend moving only selected workloads.
Or it may conclude that optimizing Splunk first produces the best near-term outcome.
That is the point of doing the architecture and economic analysis before selecting the destination.
Samuel Abouelfateh — Founder, SAB Consulting
Splunk Migration · Observability Architecture · Observability FinOps · Elasticsearch · OpenSearch · Kubernetes · OpenTelemetry
Selected official references
- Splunk Platform pricing models
- Splunk workload pricing
- Elastic Automatic Migration documentation
- Datadog Observability Pipelines documentation
- Datadog Splunk forwarder source
- AWS Migration Assistant for Amazon OpenSearch Service
- OpenSearch project and Apache 2.0 licensing
Prices, packaging, feature entitlements, SLAs and migration capabilities change frequently. Validate estimates against current official documentation and the commercial proposals supplied to your organization.