Restoring a snapshot into a new cluster
Suppose you want to copy a snapshot of SSTable data files from a three node DataStax Enterprise (DSE) cluster with vnodes enabled (128 tokens) and recover it on another newly created three node cluster (128 tokens). The token ranges will not match, because the token ranges cannot be exactly the same in the new cluster. You need to specify the tokens for the new cluster that were used in the old cluster.
This procedure assumes you are familiar with restoring a snapshot and configuring and initializing a cluster.
-
To recover the snapshot on the new cluster:
-
From the old cluster, retrieve the list of tokens associated with each node’s IP:
nodetool ring | grep -w <ip_address_of_node> | awk '{print $NF ","}' | xargs -
In the
cassandra.yamlfile for each node in the new cluster, add the list of tokens you obtained in the previous step to the initial_token parameter using the same num_tokens setting as in the old cluster.If nodes are assigned to racks, make sure the token allocation and rack assignments in the new cluster are identical to those of the old.
-
Make any other necessary changes in the new cluster’s
cassandra.yamland property files so that the new nodes match the old cluster settings. Make sure the seed nodes are set for the new cluster. -
Clear the system table data from each new node:
sudo rm -rf /var/lib/cassandra/data/system/*This allows the new nodes to use the initial tokens defined in the
cassandra.yamlwhen they restart. -
Start each node using the specified list of token ranges in new cluster’s
cassandra.yaml:initial_token: -9211270970129494930, -9138351317258731895, -8980763462514965928, ... -
Recreate the schema in the new cluster.
All schemas from the old cluster must be reproduced in the new cluster.
-
Don’t use
nodetool refresh, and don’t manually copy files into the data directory of a running node. These operations are unsafe because, while a node is running, the files in the data directory can be silently overwritten by new SSTables with the same filename resulting from memtable flushes or compaction. -
After stopping the node, restore the SSTables from the snapshots to the corresponding directory paths on the new cluster, replacing the UUIDs in the snapshot paths with the new cluster’s UUIDs.
The paths on both clusters are the same except for UUIDs in directory names, such as
/var/lib/cassandra/data/keyspace/table-UUID/snapshots/snapshot_name. Your target paths must use the new UUIDs.SSTable restoration is required to populate the new cluster with data before starting it.