Transforming How We Run Kafka at Honeycomb
Honeycomb moved six Kafka clusters off Confluent and ZooKeeper to KRaft on EKS. They rehearsed full forward and backward rollbacks first; runs fell from five hours to two, and seven teams became one.

The migration was born in a December 2025 incident that emptied a cluster and, in doing so, showed that Retriever could reset its offsets and switch clusters, which became the move Honeycomb repeated for every cluster. Each run followed the same script: producers cut over first, then consumers, with a window of downtime between them accepted in advance, and a set of OpenTelemetry dashboards built for the occasion so the team could watch the queues drain. The number that moved with practice was the run itself, from four to five hours to two to three, and the last migration ran without its lead in the building. Broker replacement used to cost 48 to 72 hours; the post says only that it is faster now. Parsons warns your results will differ. The bill he does name is the downtime window, chosen rather than suffered.
Comments
Loading comments…