Your Galera cluster’s hardest problem was never the replication library

Share this Post:

Ask anyone who runs a multi-node MySQL Galera Cluster in production about their worst night. Almost nobody says write-set replication was wrong. They say something closer to this:

“All queries piling up… no error messages in logs.”

Twenty minutes of downtime, ended only when a human picked a node to kill. (codership-team)

Or the flow-control stall where writes time out and wsrep_flow_control_sent still reads 0, so there’s nothing to alert on and nothing to write in the post-mortem. (codership-team)

You have your own version of these. You didn’t need mine.

And here’s the thing about all of them: the root cause is usually not Galera. It’s flashcache, or a shadowed wsrep_provider_options line, or DNS resolving to the wrong interface, or a one-way firewall path. One team diagnosed a Kubernetes crash loop by reading the operator’s source code.

So your MTTR isn’t dominated by repair. It’s dominated by diagnosis. The cluster is not hard to fix once you know what’s wrong, but it’s hard to know what’s wrong at 3am with writes stalled and clean logs.

That’s what changed on 30 September 2026.

What the EOL actually takes away

MariaDB plc acquired Codership in May 2025 and set the end of support for all current MySQL Galera Cluster versions at 30 September 2026. No development, no maintenance, no binary releases after that.

Your cluster will behave on 1 October exactly as it did the day before. Nothing new breaks. What disappears is the two things that used to bound those incidents: someone to escalate to, and a release that eventually contains the fix.

MariaDB Enterprise’s recommended way out is migrating to MariaDB Galera Cluster. That’s a different database: distinct system table structure, a data dictionary that differs fundamentally from MySQL’s, and user privileges that MariaDB’s own migration notes say you have to systematically recreate. You built this HA layer on open source MySQL because you wanted control. Changing database vendor to keep a support contract is the opposite of control.

Two paths that don’t ask you to:

Keep the cluster, buy back the escalation path

If diagnosis is the expensive part, what shortens it is a senior engineer who has seen your failure mode before. Not a new binary.

Percona MySQL Support covers your existing MySQL Galera Cluster as it is. No migration, no re-platforming, no privilege rebuild. Senior MySQL engineers 24x7x365 on follow-the-sun, consultative support as well as incident triage (wsrep tuning, DDL strategy, write distribution, quorum design), and contractual SLAs down to 15-minute initial response for Severity 1 incidents on Premium tier. Details in the support datasheet.

Or move to PXC, where the cluster layer still gets fixed

Percona XtraDB Cluster is Percona Server for MySQL plus the Galera library – our own fork, in a public repo, compatible with MySQL Community Edition.

The difference from the MariaDB path is mechanical: your schema, system tables, and user accounts carry over as they are. Same wsrep behavior, same ecosystem, same tools, same connectors. The ten-step walkthrough is public and free to follow on your own.

The reason to pick PXC isn’t that it’s another Galera build. It’s that the failure modes at the top of this post are what we ship fixes for: IST and SST behavior, node eviction, flow control, DDL under concurrent writes. Public release notes, Jira IDs attached, every quarter:

That cadence goes back to 2012, across 5.5, 5.6, 5.7, 8.0 and now 8.4, with the 8.0 line still shipping alongside. In the last 90 days, nearly 4,000 PXC host instances reported telemetry outside Kubernetes, and the PXC Operator runs across thousands more Kubernetes deployments. The engineers you’d escalate to are in this cluster layer every week.- See the recent work on gcache inspection and cross-site replication in the Operator.

If you’d rather not migrate alone, our migration service includes a test environment on the target cluster with a wsrep behavior comparison against your current one: you watch it handle the same failures before you commit, plus a documented rollback plan, scheduled cutover with off-hours support, and a health audit a few weeks after go-live.

What’s not an option

Running unmaintained. Not because of the date since your cluster won’t notice the date. Because the next time all the queries pile up and the logs say nothing, there’s no ticket to open and no release to wait for.

Support buys time. Migration ends the exposure. Doing both, in that order, is a perfectly good plan. Either way your clustering layer stays open source, stays on MySQL, and stays supported.

Talk to us about which one fits.

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Oldest
Newest Most Voted

Far
Enough.

Said no pioneer ever.
MySQL, PostgreSQL, InnoDB, MariaDB, MongoDB and Kubernetes are trademarks for their respective owners.
© 2026 Percona All Rights Reserved