Learn OpenStack · deploy/reference/known-issues

Known Issues — 2025.1 (Epoxy)

Partial upgrade with --limit corrupts cluster state

Symptom: After running kolla-ansible upgrade --limit <host>, services on non-limited hosts fail to communicate with upgraded hosts. API errors, RPC timeouts, and DB schema mismatches appear across the cluster.

Cause: OpenStack upgrades involve DB schema migrations and RPC version negotiations. Running upgrade on a subset of hosts leaves the cluster in a split-version state where migrated schemas and RPC versions are incompatible with non-upgraded services.

Workaround: Never use --limit with kolla-ansible upgrade. Always upgrade the entire cluster in one operation. If you've already done a partial upgrade, complete the upgrade on all remaining hosts immediately.

Affected versions: All kolla-ansible versions Status: By design — not a bug, but a known operational constraint


MariaDB Galera cluster fails to bootstrap after power loss

Symptom: After a full power loss (all controllers down simultaneously), MariaDB containers fail to start. Logs show WSREP: failed to open channel and the galera cluster cannot form.

Cause: No node has safe_to_bootstrap: 1 in its grastate.dat because all nodes were shut down uncleanly.

Workaround: 1. On each controller, check /var/lib/docker/volumes/mariadb/_data/grastate.dat 2. Find the node with the highest seqno value 3. On that node, set safe_to_bootstrap: 1 in grastate.dat 4. Run kolla-ansible mariadb-recovery

Affected versions: All kolla-ansible versions with MariaDB Galera Status: Open — inherent to Galera's crash recovery design


RabbitMQ cluster partition after network split

Symptom: After a network partition between controllers, RabbitMQ enters a split-brain state. OpenStack services log ConnectionResetError or AMQP connection lost repeatedly. Some services work, others don't depending on which RabbitMQ node they're connected to.

Cause: RabbitMQ's partition handling policy. Kolla-ansible defaults to pause_minority which stops the minority side, but if the partition is symmetric (2-node cluster) or transient, recovery may not be automatic.

Workaround: 1. Identify the partition: docker exec rabbitmq rabbitmqctl cluster_status 2. If partitioned: restart RabbitMQ on the minority nodes: docker restart rabbitmq on the affected hosts 3. If that fails: stop RabbitMQ on all nodes, then start them one at a time starting with the node that has the most recent data 4. Run kolla-ansible reconfigure -t rabbitmq to ensure consistent state

Affected versions: All kolla-ansible versions Status: Open — inherent to distributed messaging systems


Nova live migration fails with "Unsafe migration" error

Symptom: openstack server migrate --live fails with Migration pre-check error: Unsafe migration: Migration may not be safe or similar libvirt errors about CPU model incompatibility.

Cause: Compute nodes have different CPU models or features. Libvirt refuses live migration when the destination CPU doesn't support all features the guest is using.

Workaround: Set a common CPU model in config/nova.conf:

[libvirt]
cpu_mode = custom
cpu_model = Cascadelake-Server  # or another model common to all compute nodes

Then run kolla-ansible reconfigure -t nova.

Affected versions: All versions, any heterogeneous compute environment Status: Open — requires operator configuration


Horizon session timeout causes 500 errors instead of redirect to login

Symptom: After session timeout, Horizon shows a 500 Internal Server Error instead of redirecting to the login page. Browser console shows Django middleware errors.

Cause: Memcached session backend loses the session data, and Horizon's error handling doesn't gracefully handle the missing session.

Workaround: Clear browser cookies and reload. To reduce frequency, increase session timeout in config/horizon/local_settings:

SESSION_TIMEOUT = 86400  # 24 hours instead of default 1800

Then run kolla-ansible reconfigure -t horizon.

Affected versions: 2025.1.0 Status: Fixed in 2025.1.1