# The Best Data Streaming Platforms in 2026: Kafka, Confluent, Kinesis, and Pub/Sub

> Kafka, Confluent, Kinesis, Pub/Sub, Event Hubs and Redpanda compared on scaling, retention and real pricing. See which one fits your workload.

Source: https://www.erathos.com/en/blog/best-data-streaming-platforms-2026
Em português: https://www.erathos.com/blog/best-data-streaming-platforms-2026
Published: 2026-09-29
Category: Tool Guides

![The Best Data Streaming Platforms](https://cms-media.erathos.com/The Best Data Streaming Platforms.png)

A data streaming platform moves events from the systems that create them to the systems that need them, as they happen. In 2026 the choice comes down to Apache Kafka (run by you or by a vendor), the three cloud-native brokers (Amazon Kinesis, Google Pub/Sub, Azure Event Hubs), and two Kafka alternatives (Redpanda and Pulsar via StreamNative).

This guide compares them on the numbers that matter when you sign a contract: how they split load, how long they keep data, what a unit of capacity costs, and how they bill you when traffic is low. Every price and limit below comes from the vendor's own pricing or docs page, linked where it appears. It ends with a question I rarely see addressed: whether you need a streaming broker at all to keep a data warehouse fresh.

## What is a data streaming platform?

A data streaming platform is a service that receives a continuous flow of events, stores them in order for a set time, and lets many programs read them independently. The job breaks down into [capturing, storing, processing, and routing streams of events](https://kafka.apache.org/documentation/).

That definition covers three separate jobs:

Job

What it does

Examples

Event broker

Receives, stores, and hands out events

Kafka, Kinesis Data Streams, Pub/Sub, Event Hubs, Redpanda, Pulsar

Stream processing

Runs joins, filters, and aggregations on events in motion

Kafka Streams, Confluent Cloud for Apache Flink, ksqlDB

Change data capture into a warehouse

Reads database changes and loads them into a warehouse such as BigQuery

Kafka Connect, Erathos

This article is about the first job, the broker.

## How do Kafka partitions, Kinesis shards, and Pub/Sub topics scale?

Kafka and Kinesis split a topic into fixed pieces (partitions or shards) and you scale by adding pieces. Pub/Sub has no pieces to manage: a topic [autoscales according to demand](https://docs.cloud.google.com/pubsub/docs/migrating-from-kafka-to-pubsub) and ordering is an opt-in feature.

In Kafka, a topic is spread over partitions on different brokers. Events with the same key [always land in the same partition](https://kafka.apache.org/documentation/), and any consumer reading that partition gets the events in the exact order they were written. Ordering is therefore per partition, never across the whole topic. A consumer group splits partitions among its members, and a second consumer group reads the same partitions independently.

![Kafka topic partitions consumer group](https://cms-media.erathos.com/Kafka topic partitions consumer group.png)

Kinesis uses the same idea under the name shard. One shard handles up to [1 MB/s or 1,000 records per second of writes, and 2 MB/s or 2,000 records per second of reads](https://docs.aws.amazon.com/streams/latest/dev/service-sizes-and-limits.html). You pick a partition key per record, and records with the same key go to the same shard. In on-demand mode AWS manages shards for you: a new stream starts at 4 MB/s write and 8 MB/s read, and scales up to 10 GB/s write in some regions or 200 MB/s in the rest.

Pub/Sub [does not have partitions](https://docs.cloud.google.com/pubsub/docs/migrating-from-kafka-to-pubsub). Publishers write to a topic, each subscription gets its own copy of the stream, and Google adds capacity automatically. If you turn on ordering keys, messages [sent in the same region with the same key](https://docs.cloud.google.com/pubsub/docs/subscription-properties) arrive in publish order. Without a key there is no order guarantee.

Kafka

Kinesis Data Streams

Pub/Sub

Unit of scale

Partition

Shard (provisioned) or automatic (on-demand)

None, topic autoscales

Ordering scope

Within a partition

Within a shard

Within an ordering key, same region

Who adds capacity

You

You (provisioned) or AWS (on-demand)

Google

Max message size

You set it per topic or broker

10 MiB

10 MB

## Apache Kafka: when running it yourself makes sense

Apache Kafka is the open-source broker the rest of this list is measured against, and the current release is [4.3.1, published June 25, 2026](https://kafka.apache.org/downloads). It is licensed under [Apache License 2.0](https://github.com/apache/kafka/blob/trunk/LICENSE), so the software costs nothing. You pay for servers, disks, and the people who run them.

Two changes define Kafka in 2026. Since 4.0, Kafka [runs entirely without ZooKeeper](https://kafka.apache.org/blog/2025/03/18/apache-kafka-4.0.0-release-announcement/). The brokers store their own metadata using a built-in consensus protocol called KRaft, so there is one system to operate instead of two. The same release moved brokers to Java 17 and made the new consumer group protocol (KIP-848) generally available, which speeds up rebalances when consumers join or leave.

Retention is time-based by default. A topic keeps data for [7 days](https://kafka.apache.org/41/configuration/topic-configs/) and has no size limit unless you set one. Tiered storage, which moves older segments to object storage such as S3, [is off by default](https://kafka.apache.org/41/operations/tiered-storage/) and Kafka does not ship a storage plugin for it. You bring your own implementation of the RemoteStorageManager interface, and compacted topics are not supported. Managed vendors sell this as a finished feature; open-source Kafka does not.

Self-managed Kafka fits when you already run Kubernetes or VMs at scale, need full control over networking and versions, and can staff on-call for a stateful system. Google's own estimate for a three-broker, three-zone cluster on Compute Engine at 10 MiB/s sustained is [about $0.9K per month](https://cloud.google.com/managed-service-for-apache-kafka/pricing) in infrastructure, before anyone's salary.

## Confluent Cloud: managed Kafka with two billing models

Confluent Cloud is fully managed Kafka sold as five cluster types, and the important split is between elastic clusters billed in eCKUs and Dedicated clusters billed in fixed CKUs. A CKU (Confluent Unit for Kafka) is a block of reserved throughput; an eCKU is the elastic version that bills only what you use each hour.

Cluster type

Intended use

eCKU or CKU range

Per unit: ingress / egress / partitions

Basic

Development and testing

1 to 50 eCKU

5 MBps / 15 MBps / 30

Standard

Production on public networking

1 to 10 eCKU

25 MBps / 75 MBps / 250

Enterprise

Production on private networking

1 to 32 eCKU

60 MBps / 180 MBps / 3,000

Dedicated

High throughput, manual scaling

Fixed CKUs

60 MBps / 180 MBps / 4,500

Freight

Cost-optimized, relaxed latency

2 to 152 eCKU

60 MBps / 180 MBps / 3,000

Source: [Confluent cluster types](https://docs.confluent.io/cloud/current/clusters/cluster-types.html). Standard and Enterprise need 2 eCKUs for the [99.99% SLA](https://docs.confluent.io/cloud/current/clusters/cluster-types.html) (service level agreement, the uptime promise) and get 99.9% at 1 eCKU. Freight does not support idempotent producers or transactions, so it fits log shipping better than payment events.

Confluent bills on [cluster capacity, data in, data out, storage, cluster linking, connectors, ksqlDB, Flink SQL, Tableflow, and audit logs](https://docs.confluent.io/cloud/current/billing/billing-dimensions.html). At idle, a Dedicated cluster charges the full CKU cost when nothing is flowing because capacity is reserved. An eCKU cluster charges nothing at zero consumption. Flink SQL charges only for the minutes queries are running, and Tableflow (which writes topics out as Iceberg tables) is billed by topic-hour and GB processed.

New Basic, Standard, Enterprise, and Freight accounts get [$400 in free credits](https://docs.confluent.io/cloud/current/clusters/cluster-types.html). Confluent's public docs describe the billing units but do not publish a flat dollar table for CKUs and eCKUs, so the exact rate comes from their pricing calculator for your cloud and region.

## Amazon Kinesis Data Streams and MSK: two ways to stream on AWS

On AWS you choose between Kinesis Data Streams, a proprietary broker with its own API, and Amazon MSK, which is managed Apache Kafka. Kinesis is simpler to operate; MSK keeps your code portable.

Kinesis Data Streams has [three billing modes](https://aws.amazon.com/kinesis/data-streams/pricing/): On-demand Standard, On-demand Advantage, and Provisioned. The US-East example rates on the pricing page:

Mode

Data in

Data out

Fixed charge

On-demand Standard

$0.08 per GB

$0.040 per GB

$0.040 per stream-hour

On-demand Advantage

$0.032 per GB

$0.016 per GB

25 MB/s in and 25 MB/s out account minimum

Provisioned

$0.014 per million PUT payload units (1 KB each)

Included in shard

$0.015 per shard-hour

Worked example at those rates, writing 1,000 GB a month and reading it once. On-demand Standard: 1,000 × $0.08 = $80 in, 1,000 × $0.040 = $40 out, 730 × $0.040 = $29.20 in stream-hours, so about $149. Provisioned with 5 shards (enough for a 5 MB/s peak): 5 × $0.015 × 730 = $54.75 in shard-hours, plus about 1 billion PUT payload units × $0.014 per million = $14, so about $69. Provisioned wins on a steady load, on-demand wins when peaks are unpredictable and you would otherwise over-provision shards.

Retention starts at [24 hours](https://aws.amazon.com/kinesis/data-streams/pricing/) and can be extended to 7 days or up to 365 days. Extended retention costs $0.020 per shard-hour on provisioned streams. Records are rounded up to 1 KB for data-in billing, so many tiny events cost more per byte than a few large ones.

Amazon MSK is billed per broker and per GB of storage. The [MSK pricing page](https://aws.amazon.com/msk/pricing/) example is three kafka.m5.large brokers at $0.21 per hour plus $0.10 per GB-month of storage, coming to $620.33 a month. MSK Serverless in the same example lists $0.75 per cluster-hour, $0.0015 per partition-hour, $0.10 per GB in, $0.05 per GB out, and $0.10 per GB-month stored. Serverless still has a fixed cluster-hour charge, so it is never free at idle.

## Google Cloud Pub/Sub: the simplest to operate

Pub/Sub is a fully managed message service with no clusters, brokers, or partitions to size, and it bills on data volume. After the first 10 GiB of throughput each month, message delivery costs [$40 per TiB in all regions](https://cloud.google.com/pubsub/pricing). Subscriptions that write straight into BigQuery or Cloud Storage cost $50 per TiB, and retained storage costs $0.27 per GiB-month.

Messages are limited to [10 MB](https://docs.cloud.google.com/pubsub/docs/migrating-from-kafka-to-pubsub). A subscription retains unacknowledged messages for [7 days by default](https://docs.cloud.google.com/pubsub/docs/subscription-properties), configurable from 10 minutes to 31 days. The acknowledgement deadline defaults to 10 seconds with a 600-second maximum, and delivery is at-least-once, so subscribers must handle repeats.

Two features change that picture if you turn them on. Ordering keys give in-order delivery for messages with the same key in the same region. Exactly-once delivery [works only on pull subscriptions and within one region](https://cloud.google.com/pubsub/docs/exactly-once-delivery). With it enabled, a message acknowledged successfully is never redelivered. Combining ordering with exactly-once limits throughput to thousands of messages per second per client and raises end-to-end latency, so it fits payment ledgers more than clickstreams.

Pub/Sub Lite, the older reserved-capacity variant, [is deprecated and shut down as of March 18, 2026](https://cloud.google.com/pubsub/pricing). It should not appear in any new design.

Google also sells Kafka on GCP as Managed Service for Apache Kafka. It bills [$0.09 per hour per vCPU with 4 GiB of RAM](https://cloud.google.com/managed-service-for-apache-kafka/pricing), plus $0.000232877 per GiB-hour of local storage and $0.01 per GiB of inter-zone transfer. Google's own estimate puts a 10 MiB/s cluster at about $1.1K a month versus $0.9K for self-run Kafka on Compute Engine, and $11K versus $9.1K at 100 MiB/s. Committed use discounts of [20% for one year and 40% for three years](https://cloud.google.com/managed-service-for-apache-kafka/pricing) apply to the compute portion.

## Azure Event Hubs: Azure's ingestion service with a Kafka endpoint

Azure Event Hubs is a managed ingestion service that also accepts Kafka clients, so existing producers and consumers can [connect without code changes](https://azure.microsoft.com/en-us/pricing/details/event-hubs/) on Standard tier and above. Capacity is sold in throughput units (TUs), where one TU gives 1 MB/s of ingress and 2 MB/s of egress.

Tier

Capacity unit

Price (Central US, as displayed)

Ingress

Kafka endpoint

Max retention

Basic

Throughput Unit

$0.015 per hour

$0.028 per million events

No

1 day

Standard

Throughput Unit

$0.03 per hour

$0.028 per million events

Yes

7 days

Premium

Processing Unit

$1.233 per hour

Included

Yes

90 days

Dedicated

Capacity Unit

$6.849 per hour

Included

Yes

90 days

Source: [Azure Event Hubs pricing](https://azure.microsoft.com/en-us/pricing/details/event-hubs/). Microsoft notes the displayed figures are estimates and vary by agreement, region, and currency. Capture, which writes events to Azure Storage automatically, costs $73 per month per TU on Standard and is included in Premium and Dedicated.

![Azure Event Hubs](https://cms-media.erathos.com/Azure Event Hubs.png)

_Azure Event Hubs tier pricing as displayed on the Azure pricing page_

A single Standard TU running all month is 730 × $0.03 = $21.90 plus ingress events. That makes Event Hubs one of the cheapest entry points on this list for a small Kafka-protocol workload already running in Azure. Basic tier has no Kafka endpoint and only one day of retention, so the Kafka path starts at Standard.

## Redpanda and StreamNative: Kafka alternatives that keep the Kafka API

Redpanda and StreamNative both accept Kafka clients while replacing the engine behind them. Redpanda is a single binary written in C++ that [needs no ZooKeeper](https://www.redpanda.com/data-streaming/bring-your-own-cloud-byoc), and StreamNative runs Apache Pulsar with a Kafka-compatible layer.

Redpanda's source is under the [Business Source License 1.1](https://github.com/redpanda-data/redpanda/blob/dev/licenses/bsl.md). You can run it in production for your own workloads, but you cannot offer it to others as a streaming or queuing service. Four years after a version ships, that version converts to Apache 2.0. Redpanda Cloud comes in three shapes:

Redpanda Cloud

Tenancy

Runs in

Max write / read

Max partitions

[SLA](https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview/)

Serverless

Multi-tenant

Redpanda's AWS or GCP account

100 MB/s / 300 MB/s

5,000

99.9%

Dedicated

Single-tenant

Redpanda's AWS, Azure, or GCP account

400 MB/s / 800 MB/s

45,600

99.99%, multi-AZ

BYOC

Single-tenant

Your own AWS, Azure, or GCP account

2 GB/s / 4 GB/s

112,500

99.99%, multi-AZ

Source: [Redpanda Cloud overview](https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview/). BYOC means bring your own cloud: the data plane runs in your account and Redpanda manages it remotely. Serverless pricing is [usage-based on ingress, egress, storage, and partitions](https://www.redpanda.com/data-streaming/serverless), and free trials get $100 in the first 30 days.

StreamNative publishes the clearest price list of any vendor here. Serverless is billed per ETU (elastic throughput unit) at [$0.10 per hour, plus $0.13 per GB in, $0.04 per GB out, and $0.09 per GB-month stored](https://docs.streamnative.io/cloud/billing/billing-overview). One ETU covers 5 MBps in and 15 MBps out, and Serverless tops out at 20 ETUs. Dedicated Kafka, in public preview, reserves RTUs (reserved throughput units) at $0.75 per hour, with 25 MBps in and 75 MBps out per RTU.

StreamNative's Pulsar splits the broker from storage. The [broker is stateless](https://pulsar.apache.org/docs/next/concepts-architecture-overview/) and messages persist in a separate Apache BookKeeper cluster. Pulsar ships tiered storage with drivers for [S3, Google Cloud Storage, and filesystem](https://pulsar.apache.org/docs/next/cookbooks-tiered-storage/), and you can set an automatic offload threshold per namespace. Kafka's tiered storage needs a plugin you supply, unlike Pulsar's built-in drivers.

## How do capacity-based and pay-per-throughput pricing differ?

Capacity-based pricing charges for the units you configure (shards, CKUs, TUs, RTUs) every hour they exist, whether or not data flows. Pay-per-throughput pricing charges for GB written, GB read, and GB stored, so a quiet month costs close to nothing.

Billing model

Products

Idle cost

Fits

Reserved capacity

Kinesis Provisioned, Confluent Dedicated CKU, Event Hubs TU/PU/CU, StreamNative Dedicated RTU, MSK Provisioned brokers

Full unit price

Steady, predictable throughput

Usage-based

Kinesis On-demand, Pub/Sub, Confluent eCKU clusters, Redpanda Serverless, StreamNative Serverless ETU

Zero or a small floor

Bursty or new workloads

Self-managed

Apache Kafka, Pulsar

Servers keep running

Teams with platform engineers and predictable scale

![capacity-based streaming bill](https://cms-media.erathos.com/capacity-based streaming bill.png)

The two models cross over at some utilization. In the Kinesis example above, provisioned came to $69 and on-demand to $149 for the same 1,000 GB, because the load kept 5 shards reasonably busy. Run those same 5 shards at 10% utilization and provisioned still costs $54.75 in shard-hours, while on-demand drops with the traffic. The same split shows up in Confluent Cloud, where a Dedicated cluster at idle pays the [full CKU cost while an eCKU cluster pays nothing at zero consumption](https://docs.confluent.io/cloud/current/billing/billing-dimensions.html).

Usage-based plans can still have a minimum. Kinesis On-demand Advantage needs a 25 MB/s in and 25 MB/s out account minimum to get its lower per-GB rate, and MSK Serverless charges $0.75 per cluster-hour before any data moves.

## What about benchmarks?

Vendor benchmarks for Kafka versus Redpanda point in opposite directions, and the most useful independent reading on them concludes that the numbers only hold for the exact setup that produced them. Jack Vanlightly, who discloses that he works at Confluent, ran both systems on [identical i3en.6xlarge hardware](https://jack-vanlightly.com/blog/2023/5/15/kafka-vs-redpanda-performance-do-the-claims-add-up) and found that changing producer count, retention state, TLS, keys, or test duration flipped the results.

Vanlightly's conclusion applies directly to a purchasing decision. Benchmarks are only useful when you run them yourself, on your own workload. A short proof of concept with your real message sizes and key distribution tells you more than any published chart, including the ones on this page.

## Do you need a streaming platform to keep your warehouse fresh?

For analytics, change data capture (CDC) usually does the job instead of a streaming broker. CDC reads a database's own change log and copies inserts, updates, and deletes to the warehouse. How often that copy runs is a separate decision from how the changes are captured.

The distinction matters because the reliability argument for streaming is often really an argument for CDC. A cursor-based batch sync (select rows where updated\_at is greater than the last run) cannot see a row that was inserted and deleted between runs, and it misses intermediate states. A [CDC pipeline that runs once an hour still captures every change](https://www.erathos.com/en/blog/cursor-based-sync-vs-change-data-capture?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026), because it reads the log rather than taking a snapshot.

This is how Erathos moves data. For PostgreSQL, it reads the write-ahead log (WAL, the file where PostgreSQL records every change before applying it) through a logical replication slot, and the connector [supports CDC for both overwrite and append modes](https://www.erathos.com/en/blog/cursor-based-sync-vs-change-data-capture?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026). For MySQL, it [reads the binary log](https://www.erathos.com/en/blog/mysql-cdc-to-bigquery?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026), which needs ROW format and FULL row images on the source.

Extracted data is [staged in a cloud bucket](https://docs.erathos.com/platform/how-we-move-data) (S3, GCS, or Azure Blob Storage), loaded into the destination with a bulk COPY or equivalent, and the temporary files are deleted after the load finishes. Sync modes cover [Full Refresh, Partial Overwrite, Partial Append, and Partial Versioned](https://docs.erathos.com/platform/connections/sync-types), where the partial modes need a primary key and a date or datetime cursor.

![erathos cdc sync into the warehouse](https://cms-media.erathos.com/erathos cdc sync into the warehouse.png)

A streaming broker fits when applications, not dashboards, consume the events: fraud checks, inventory updates, notifications, or anything that must react in under a second. When the consumer is a warehouse table that refreshes hourly, CDC into the warehouse does the same job with one fewer system to run. For the broader pipeline design, see [how to build and manage data pipelines](https://www.erathos.com/en/blog/how-to-build-and-manage-data-pipeline?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026).

## FAQ

### Is Apache Kafka free?

Yes. Kafka is licensed under [Apache License 2.0](https://github.com/apache/kafka/blob/trunk/LICENSE), which permits commercial use, modification, and distribution at no charge. The cost is infrastructure and operations: Google estimates a three-broker cluster at 10 MiB/s at [about $0.9K a month](https://cloud.google.com/managed-service-for-apache-kafka/pricing) in Compute Engine spend, and managed Kafka from Confluent, AWS, or Google adds a margin on top of that in exchange for running it.

### Kafka vs Kinesis: which should I pick on AWS?

Kinesis Data Streams is the better fit when your producers and consumers are AWS services and you want no clusters to manage. MSK is the better fit when you need the Kafka API or want to keep your code portable. Kinesis shards are fixed at [1 MB/s in and 2 MB/s out](https://docs.aws.amazon.com/streams/latest/dev/service-sizes-and-limits.html) and retention defaults to 24 hours. MSK is real Kafka, starting around [$620 a month](https://aws.amazon.com/msk/pricing/) for three small brokers in the AWS example.

### Pub/Sub vs Kafka: when is Pub/Sub the simpler choice?

Pub/Sub is simpler when you do not need partitions, transactions, or replay beyond 31 days, and you want a bill that scales with data at [$40 per TiB](https://cloud.google.com/pubsub/pricing). Kafka is the better fit when you need per-key ordering across a high-volume topic, transactional writes across topics, or long retention with tiered storage. On GCP, Managed Service for Apache Kafka gives you Kafka without leaving Google's billing.

### Is Azure Event Hubs a Kafka replacement?

For producers and consumers that use the Kafka protocol, yes on Standard tier and above. One throughput unit gives [1 MB/s in and 2 MB/s out](https://azure.microsoft.com/en-us/pricing/details/event-hubs/), Standard retains data up to 7 days, and Premium or Dedicated extend that to 90 days. The Basic tier has no Kafka endpoint.

### Should I choose Redpanda or Pulsar instead of Kafka?

Redpanda fits when you want the Kafka API with a single C++ binary to operate and are comfortable with the [BSL 1.1 license](https://github.com/redpanda-data/redpanda/blob/dev/licenses/bsl.md), which blocks reselling it as a service. Pulsar fits when you need [stateless brokers](https://pulsar.apache.org/docs/next/concepts-architecture-overview/) that scale separately from storage, or built-in tiered storage to S3 and GCS. Both should be tested against your own workload before you commit, since published benchmarks vary with test setup.

### Do I need streaming, or is batch enough?

Batch is enough when the consumer is a warehouse or dashboard that refreshes on a schedule, as long as the capture method is CDC rather than a cursor query. CDC reads the database log, so an [hourly run still records every insert, update, and delete](https://www.erathos.com/en/blog/cursor-based-sync-vs-change-data-capture?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026). Streaming is needed when an application must react to each event within seconds.

## Try Erathos free for 14 days

If your goal is a warehouse that reflects your production databases, Erathos runs log-based CDC from PostgreSQL and MySQL into your warehouse without a broker for you to operate. [Start a free 14-day trial](https://app.erathos.com/signup?utm_source=blog&utm_medium=organic&utm_content=bydefault&utm_campaign=best-data-streaming-platforms-2026) and connect your first source.
