Start and stop Hyper-Converged Database (HCD)
Use these commands to start and stop Hyper-Converged Database (HCD) nodes in a cluster.
When starting a multi-node cluster, start seed nodes first.
|
Many HCD configuration changes require a rolling restart. This means that you stop and start each node one at a time, rather than stopping the entire cluster simultaneously. Rolling restarts allow the cluster to continue servicing requests while you apply configuration changes. |
Start and stop nodes with Mission Control
|
Nodes managed by Mission Control must be started and stopped using Mission Control. Don’t start or stop them manually. The operator attempts to reconcile the actual state of the nodes with the desired state defined in Mission Control; manually stopped or started nodes will be reverted to the desired state by the operator. |
For clusters managed by Mission Control, use Mission Control to start and stop nodes.
Start HCD as a standalone process (tarball)
When installed from a binary tarball, HCD runs as a standalone process.
-
Run
bin/hcd cassandrafrom your HCD installation directory:INSTALL_DIRECTORY/bin/hcd cassandra -
To start multiple nodes, repeat the previous command on each node.
-
Check node and cluster state to verify that the node is running.
Start HCD as a service (package)
When installed from an RHEL or Debian package, HCD runs as a service.
Package installations include start and stop scripts for the HCD service.
Binary tarballs don’t include these scripts.
-
Start HCD on a node:
sudo service hcd start -
To start multiple nodes, repeat the previous command on each node.
-
Check node and cluster state to verify that the node is running.
Stop a node
The commands to stop HCD on a node depend on your installation method:
- Package installations
-
-
Run
nodetool drainto flush commit logs to disk before stopping the node:nodetool drainIf you disabled durable writes (not recommended), you must run drain the node to prevent data loss.
When durable writes are enabled (the default), draining the node is beneficial when restarting nodes. Because the commit logs are not replayed, the startup process is faster.
-
Stop the HCD service:
sudo service hcd stopFor RHEL package installations with systemd, run
systemctl stopwith the full node ID, including thehcd-prefix. The default node ID ishcd.systemctl stop FULL_NODE_ID
-
- Tarball installations
-
Stop the HCD process with
cassandra-stop. You don’t need to runnodetool drainbecausecassandra-stopautomatically drains the node before stopping it. If necessary, runcassandra-stopwithsudo.INSTALL_DIRECTORY/bin/hcd cassandra-stop
Check node and cluster state
After you start or stop nodes, run nodetool status to check the state of the cluster:
- Package installations
-
nodetool status - Tarball installations
-
INSTALL_DIRECTORY/bin/nodetool status
The output lists all nodes in the cluster, separated by datacenter. The first column reports the state and status of the node. The contents of the other columns depends on your cluster configuration, such as the use of virtual nodes (vnodes). For example:
- With vnodes
-
With vnodes, the
Tokenscolumn indicates the number of virtual nodes assigned to each physical node:Datacenter: DC1 =============== Status=Up/Down |/ State=Normal/Leaving/Joining/Moving -- Address Load Tokens Owns (effective) Host ID Rack UN 10.0.0.1 97.87 KiB 16 100.0% e4b7eb22-31d0-43ed-95e0-97915efcf58d RAC1The
Ownscolumn reports the percentage of data in the cluster that the node is responsible for. A?indicates that the ownership information is unavailable or undetermined, such as when bootstrapping a new node or rebalancing the cluster. - Without vnodes
-
Without vnodes, the
Tokencolumn lists the specific token value assigned to each physical node:Datacenter: DC1 ===================== Status=Up/Down |/ State=Normal/Leaving/Joining/Moving -- Address Load Owns Host ID Token Rack UN 172.16.222.136 103.24 KB ? 3c1d0657-0990-4f78-a3c0-3e0c37fc3a06 1647352612226902707 rack1
When starting or stopping a node, the following status and state combinations are expected from healthy nodes.
If a node reports an unexpected status or fails to appear in the nodetool status output, see Troubleshoot issues with starting and stopping nodes.
| Operation | Status/State | Description |
|---|---|---|
Start a node |
|
The node is running and joined the cluster. |
Start a node |
|
When bootstrapping a node, the node temporarily reports as Run |
Stop a node |
|
The node is stopped and remains in the cluster. Stopped nodes don’t leave the cluster unless other operations occur, such as decommissioning the node. This status is uniquely reported when you run |
Stop a node |
|
The node is stopped and remains in the cluster. Stopped nodes don’t leave the cluster unless other operations occur, such as decommissioning the node. This status is reported for stopped nodes when you run |
Troubleshoot issues with starting and stopping nodes
The following errors might occur when starting or stopping nodes:
- Timed out while waiting for HCD to start
-
This error can indicate a genuine timeout, but it can also occur when the user running the service doesn’t have permission to read the required files. For more information, see HCD cannot start with YAML parsing error.
- Node missing from
nodetool statusoutput, ornodetool statusreports only the current node in a multi-node cluster -
This issue indicates that the node failed to join the cluster (ring). Common causes of this issue include:
Cause Description Resolution Seed node started after non-seed nodes
When starting a new cluster or performing a full cluster restart, you must start the seed nodes first. The seed nodes describe the cluster topology and help new nodes discover and join the cluster.
Stop all nodes, and then start the seed nodes before starting the non-seed nodes.
Mismatched versions or software
Unless you are in the process of upgrading, all nodes in a cluster must have the same version of HCD installed in the same way (tarball, Debian package, or RHEL package). Nodes might fail to join the ring if they are running different HCD versions, different HCD distributions, or other Cassandra-based database distributions.
Compare the node’s HCD version to the other nodes in the cluster, and then take action accordingly to align the versions.
Misconfiguration
Various configuration issues can prevent a node from joining the cluster, such as:
-
Incorrect or incomplete seed node topology
-
Mismatched cluster names
-
Non-seed node cannot reach seed node
-
Incorrect or mismatched hostnames, IP addresses, or port numbers
-
Intentional node isolation is enabled, such as
write_surveymode
Investigate the cluster and node logs, configuration files, and startup options (JVM properties).
Network connectivity
Outages, timeouts, or firewall rules can block communication between nodes.
Investigate network configurations and logs.
-
cassandra-stopfails due to missing Java process ID (PID)-
This error is exclusive to tarball installations;
cassandra-stopisn’t available to package installations. To resolve this error:-
Get the HCD PID:
ps auwx | grep hcd -
Pass the PID to
cassandra-stop:bin/hcd cassandra-stop -p PID
-