# On a 3 node broker cluster, one node is always down

**URL:** <https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993>\
**Category:** Containers\
**Created:** [3 November 2023 04:19 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993 "2023-11-03T04:19:27Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Udayendu](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/udayendu/32/3314_2.png) [@Udayendu](https://forum.confluent.io/u/Udayendu)\
**Post date:** [3 November 2023 04:19 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/1 "2023-11-03T04:19:27Z")

</div>

Using docker compose, I am trying to deploy a 3 node kafka cluster. Randomly 2 nodes are working fine and 3rd node is always getting the following error:

```auto
[2023-11-03 04:12:59,865] INFO [broker-1-to-controller-forwarding-channel-manager]: Starting (kafka.server.BrokerToControllerRequestThread)
[2023-11-03 04:12:59,959] INFO [MetadataLoader id=1] initializeNewPublishers: the loader is still catching up because we still don't know the high water mark yet. (org.apache.kafka.image.loader.MetadataLoader)
[2023-11-03 04:13:00,026] INFO [RaftManager id=1] Registered the listener org.apache.kafka.image.loader.MetadataLoader@351741309 (org.apache.kafka.raft.KafkaRaftClient)
[2023-11-03 04:13:00,061] INFO [MetadataLoader id=1] initializeNewPublishers: the loader is still catching up because we still don't know the high water mark yet. (org.apache.kafka.image.loader.MetadataLoader)
[2023-11-03 04:13:00,162] INFO [MetadataLoader id=1] initializeNewPublishers: the loader is still catching up because we still don't know the high water mark yet. (org.apache.kafka.image.loader.MetadataLoader)
[2023-11-03 04:13:00,218] ERROR Encountered fatal fault: Unexpected error in raft IO thread (org.apache.kafka.server.fault.ProcessTerminatingFaultHandler)
java.lang.IllegalStateException: Received request or response with leader OptionalInt[1] and epoch 18 which is inconsistent with current leader OptionalInt.empty and epoch 0
	at org.apache.kafka.raft.KafkaRaftClient.maybeTransition(KafkaRaftClient.java:1513)
	at org.apache.kafka.raft.KafkaRaftClient.maybeHandleCommonResponse(KafkaRaftClient.java:1473)
	at org.apache.kafka.raft.KafkaRaftClient.handleFetchResponse(KafkaRaftClient.java:1071)
	at org.apache.kafka.raft.KafkaRaftClient.handleResponse(KafkaRaftClient.java:1550)
	at org.apache.kafka.raft.KafkaRaftClient.handleInboundMessage(KafkaRaftClient.java:1676)
	at org.apache.kafka.raft.KafkaRaftClient.poll(KafkaRaftClient.java:2251)
	at kafka.raft.KafkaRaftManager$RaftIoThread.doWork(RaftManager.scala:64)
	at org.apache.kafka.server.util.ShutdownableThread.run(ShutdownableThread.java:127)

```

I have 3 dedicated controller nodes running separately. Here is the broker config that I am using for node1 which is down currently:

```auto
---
version: '2'
services:

  broker:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-broker-1
    container_name: kafka-broker-1
    ports:
      - "9092:9092"
      - "29092:29092"
    environment:
      KAFKA_NODE_ID: 1
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT'
      KAFKA_ADVERTISED_LISTENERS: 'PLAINTEXT://vskafka-broker-1:29092'
      KAFKA_PROCESS_ROLES: 'broker'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'PLAINTEXT://vskafka-broker-1:9092'
      KAFKA_INTER_BROKER_LISTENER_NAME: 'PLAINTEXT'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-broker-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

```

This node can reach to the mentioned controllers without any issue using telnet. And also can talk to the other broker nodes.

---

<div class="post-metadata">

**Author:** ![mmuehlbeyer](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/mmuehlbeyer/32/1088_2.png) [@mmuehlbeyer](https://forum.confluent.io/u/mmuehlbeyer)\
**Post date:** [3 November 2023 08:12 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/2 "2023-11-03T08:12:58Z")

</div>

hey @Udayendu

hmm looks strange  
could you share your complete docker-compose?

---

<div class="post-metadata">

**Author:** ![Udayendu](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/udayendu/32/3314_2.png) [@Udayendu](https://forum.confluent.io/u/Udayendu)\
**Post date:** [3 November 2023 08:21 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/3 "2023-11-03T08:21:10Z")

</div>

Hi @mmuehlbeyer

I have 3 controller nodes and 3 broker nodes.  
I am using the same cluster ID for all the 6 nodes. All the nodes have their own docker-compose file.

I am doing the following steps:

- deployed all the controller nodes first
- then started the deployment of broker nodes. some time randomly node2 is not working and some time node1. Out of 3 nodes only two nodes are showing as working and 3rd is not.
- for broker node1, controller quorum voters are controller node2 and 3. And like wise its repeated for all the 3 broker nodes.

Here is the config that I am using for now to configure brokers:

```auto
---
version: '2'
services:

  broker:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-broker-1
    container_name: kafka-broker-1
    ports:
      - "9092:9092"
      - "29092:29092"
    environment:
      KAFKA_NODE_ID: 1
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT'
      KAFKA_ADVERTISED_LISTENERS: 'PLAINTEXT://vskafka-broker-1:29092'
      KAFKA_PROCESS_ROLES: 'broker'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'PLAINTEXT://vskafka-broker-1:9092'
      KAFKA_INTER_BROKER_LISTENER_NAME: 'PLAINTEXT'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-broker-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'  

---
version: '2'
services:

  broker:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-broker-2
    container_name: kafka-broker-2
    ports:
      - "9092:9092"
      - "29092:29092"
    environment:
      KAFKA_NODE_ID: 2
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT'
      KAFKA_ADVERTISED_LISTENERS: 'PLAINTEXT://vskafka-broker-2:29092'
      KAFKA_PROCESS_ROLES: 'broker'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'PLAINTEXT://vskafka-broker-2:9092'
      KAFKA_INTER_BROKER_LISTENER_NAME: 'PLAINTEXT'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-broker-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

---
version: '2'
services:

  broker:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-broker-3
    container_name: kafka-broker-3
    ports:
      - "9092:9092"
      - "29092:29092"
    environment:
      KAFKA_NODE_ID: 3
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT'
      KAFKA_ADVERTISED_LISTENERS: 'PLAINTEXT://vskafka-broker-3:29092'
      KAFKA_PROCESS_ROLES: 'broker'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093'
      KAFKA_LISTENERS: 'PLAINTEXT://vskafka-broker-3:9092'
      KAFKA_INTER_BROKER_LISTENER_NAME: 'PLAINTEXT'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-broker-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

```

And my controller configs are:

```auto
---
version: '2'
services:

  controller:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-controller-1
    container_name: kafka-controller-1
    ports:
      - "9093:9093"
    environment:
      KAFKA_NODE_ID: 1
      KAFKA_PROCESS_ROLES: 'controller'
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'CONTROLLER://vskafka-controller-1:9093'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-controller-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

---
version: '2'
services:

  controller:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-controller-2
    container_name: kafka-controller-2
    ports:
      - "9093:9093"
    environment:
      KAFKA_NODE_ID: 2
      KAFKA_PROCESS_ROLES: 'controller'
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'CONTROLLER://vskafka-controller-2:9093'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-controller-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

---
version: '2'
services:

  controller:
    image: confluentinc/cp-kafka:7.5.1
    hostname: vskafka-controller-3
    container_name: kafka-controller-3
    ports:
      - "9093:9093"
    environment:
      KAFKA_NODE_ID: 3
      KAFKA_PROCESS_ROLES: 'controller'
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: 'CONTROLLER:PLAINTEXT'
      KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'
      KAFKA_LISTENERS: 'CONTROLLER://vskafka-controller-3:9093'
      KAFKA_CONTROLLER_LISTENER_NAMES: 'CONTROLLER'
      KAFKA_LOG_DIRS: '/tmp/kraft-controller-logs'
      CLUSTER_ID: 'mX-qLvc-T2y2OPeJ3AMRXg'

```

Let me know if you need further info from my side.

---

<div class="post-metadata">

**Author:** ![mmuehlbeyer](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/mmuehlbeyer/32/1088_2.png) [@mmuehlbeyer](https://forum.confluent.io/u/mmuehlbeyer)\
**Post date:** [3 November 2023 08:36 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/4 "2023-11-03T08:36:23Z")

</div>

ok understood  
so 3 separate nodes with a controller and a broker on each of the nodes correct?

---

<div class="post-metadata">

**Author:** ![mmuehlbeyer](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/mmuehlbeyer/32/1088_2.png) [@mmuehlbeyer](https://forum.confluent.io/u/mmuehlbeyer)\
**Post date:** [3 November 2023 08:42 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/5 "2023-11-03T08:42:41Z")

</div>

and why did you not specify all 3 nodes in  
` KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093'`

---

<div class="post-metadata">

**Author:** ![Udayendu](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/udayendu/32/3314_2.png) [@Udayendu](https://forum.confluent.io/u/Udayendu)\
**Post date:** [3 November 2023 08:55 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/6 "2023-11-03T08:55:09Z")

</div>

yes, all are on separate nodes.

---

<div class="post-metadata">

**Author:** ![Udayendu](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/udayendu/32/3314_2.png) [@Udayendu](https://forum.confluent.io/u/Udayendu)\
**Post date:** [3 November 2023 08:58 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/7 "2023-11-03T08:58:03Z")

</div>

When I am specifying all the 3 controller nodes, its complaining with the following error:

```auto
# docker logs kafka-broker-1
===> User
uid=1000(appuser) gid=1000(appuser) groups=1000(appuser)
===> Configuring ...
Running in KRaft mode...
===> Running preflight checks ...
===> Check if /var/lib/kafka/data is writable ...
===> Running in KRaft mode, skipping Zookeeper health check...
===> Using provided cluster id mX-qLvc-T2y2OPeJ3AMRXg ...
Exception in thread "main" java.lang.IllegalArgumentException: requirement failed: If process.roles contains just the 'broker' role, the node id 1 must not be included in the set of voters controller.quorum.voters=Set(1, 2, 3) at scala.Predef$.require(Predef.scala:337) at kafka.server.KafkaConfig.validateValues(KafkaConfig.scala:2246) at kafka.server.KafkaConfig.<init>(KafkaConfig.scala:2160) at kafka.server.KafkaConfig.<init>(KafkaConfig.scala:1568) at kafka.tools.StorageTool$.$anonfun$main$1(StorageTool.scala:50) at scala.Option.flatMap(Option.scala:283) at kafka.tools.StorageTool$.main(StorageTool.scala:50) at kafka.tools.StorageTool.main(StorageTool.scala)

```

Let say if I am deploying broker1, the I cant add it for controller 1. Same logic is applicable for all the remaining two nodes as well.

---

<div class="post-metadata">

**Author:** ![Udayendu](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.confluent.io/udayendu/32/3314_2.png) [@Udayendu](https://forum.confluent.io/u/Udayendu)\
**Post date:** [5 November 2023 15:05 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/8 "2023-11-05T15:05:28Z")

</div>

To fix this issue, I changed the ID of all brokers to 101, 102 and 103 and for all controllers 1, 2 and 3. Then added all the controllers to the quorum list for brokers and started the services like:

```auto
KAFKA_CONTROLLER_QUORUM_VOTERS: '1@vskafka-controller-1:9093,2@vskafka-controller-2:9093,3@vskafka-controller-3:9093'

```

Its started working well now.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex019/uploads/confluentcommunity/original/1X/c49438c90c9df282e9996fdf6971be890c71b65a.svg) [@system](https://forum.confluent.io/u/system)\
**Post date:** [12 November 2023 15:05 UTC](https://forum.confluent.io/t/on-a-3-node-broker-cluster-one-node-is-always-down/8993/9 "2023-11-12T15:05:35Z")

</div>

This topic was automatically closed 7 days after the last reply. New replies are no longer allowed.
