Data replication
Hyper-Converged Database (HCD) can replicate data to multiple nodes to make the cluster more reliable and resilient. Replication is set in the keyspace definition, and consists of a replication factor and replication strategy. All tables in a keyspace inherit the replication strategy and replication factor from the keyspace. Data that has different replication requirements within the same datacenter must be stored in different keyspaces.
Replication factor
The replication factor determines the number of copies (replicas) of data that are stored in the cluster. Production databases typically have a replication factor of 3 or more. The replication factor should not exceed the total number of nodes.
|
Never use a replication factor of 1 in production or any sensitive environment. If the node containing the row goes down, the row cannot be retrieved. |
Replication strategy
The replication strategy determines whether replicas are placed in a single datacenter or across multiple datacenters. The following replication strategies are available:
- NetworkTopologyStrategy
-
Recommended for all deployments because it is easier to scale to multiple datacenters and racks, even if you don’t initially require these features.
NetworkTopologyStrategy places replicas in the same datacenter by walking the ring clockwise until it reaches the first node in another rack. It also attempts to place replicas on distinct racks because nodes in the same rack (or similar physical grouping) can fail at the same time due to power, cooling, or network issues. This strategy requires that you specify how many replicas you want in each datacenter per keyspace. When deciding how many replicas to configure in each datacenter, be mindful of cross-datacenter latency on the read path and available nodes in the event of a failure. Typically, multi-datacenter clusters are configured as follows:
-
Two replicas in each datacenter: This configuration tolerates the failure of a single node per replication group and still allows local reads at a consistency level of
ONE. -
Three replicas in each datacenter: This configuration tolerates either the failure of one node per replication group at a strong consistency level of
LOCAL_QUORUMor multiple node failures per datacenter using consistency levelONE. -
Asymmetrical replication groupings are also possible. For example, you can have three replicas in one datacenter to serve real-time application requests and use a single replica elsewhere for running analytics.
NetworkTopologyStrategy requires a network-aware snitch to correctly place replicas across racks and datacenters. It cannot use the default snitch.
-
- SimpleStrategy
-
Use SimpleStrategy only for development clusters that will always have one datacenter and one rack. In all other cases, including all production clusters, use NetworkTopologyStrategy.
SimpleStrategy places the first replica on a node that is selected by the partitioner. Additional replicas are placed on subsequent nodes clockwise around the ring without considering topology (rack or datacenter location).
Replication map examples
The following examples use the NetworkTopologyStrategy.
- Single datacenter
-
Create a keyspace with 3 replicas in one datacenter:
CREATE KEYSPACE user_profiles WITH replication = { 'class': 'NetworkTopologyStrategy', 'datacenter1': 3 }; - Multiple datacenters
-
Create a keyspace with 3 replicas in two datacenters (named
eastandwest):CREATE KEYSPACE user_profiles WITH replication = { 'class': 'NetworkTopologyStrategy', 'east': 3, 'west': 3 };