Ec2MultiRegionSnitch
Use the Ec2MultiRegionSnitch for deployments on Amazon EC2 where the cluster spans multiple regions.
Use the Ec2MultiRegionSnitch for deployments on Amazon EC2 where the cluster spans multiple regions.
You must configure settings in both the cassandra.yaml file and the property file (cassandra-rackdc.properties) used by the Ec2MultiRegionSnitch.
Configuring cassandra.yaml for cross-region communication
The Ec2MultiRegionSnitch uses public IP designated in the broadcast_address to allow cross-region connectivity. Configure each node as follows:
- In the cassandra.yaml, set the listen_address to the private
IP address of the node, and the broadcast_address to the public IP address of the node.
This allows Cassandra nodes in one EC2 region to bind to nodes in another region, thus enabling multiple data center support. For intra-region traffic, Cassandra switches to the private IP after establishing a connection.
- Set the addresses of the seed nodes in the cassandra.yaml file to
that of the public IP. Private IP are not routable between networks. For
example:
seeds: 50.34.16.33, 60.247.70.52
To find the public IP address, from each of the seed nodes in EC2:
$ curl http://instance-data/latest/meta-data/public-ipv4
Note: Do not make all nodes seeds, see Internode communications (gossip). - Be sure that the storage_port or ssl_storage_port is open on the public IP firewall.
Configuring the snitch for cross-region communication
In EC2 deployments, the region name is treated as the data center name and availability zones are treated as racks within a data center. For example, if a node is in the us-east-1 region, us-east is the data center name and 1 is the rack location. (Racks are important for distributing replicas, but not for data center naming.)
For each node, specify its data center in the cassandra-rackdc.properties. The dc_suffix option defines the data centers used by the snitch. Any other lines are ignored.
In the example below, there are two cassandra data centers and each data center is named for its workload. The data center naming convention in this example is based on the workload. You can use other conventions, such as DC1, DC2 or 100, 200. (Data center names are case-sensitive.)
Region: us-east | Region: us-west |
---|---|
Node and data center:
This results in four us-east data
centers:
|
Node and data center:
This results in four us-west data
centers:
|
Keyspace strategy options
When defining your keyspace strategy options, use the EC2 region name, such as ``us-east``, as your data center name.
Package installations | /etc/cassandra/cassandra-rackdc.properties |
Tarball installations | install_location/conf/cassandra-rackdc.properties |
Windows installations | C:\Program Files\DataStax Community\apache-cassandra\conf\cassandra-rackdc.properties |
Package installations | /etc/cassandra/cassandra.yaml |
Tarball installations | install_location/resources/cassandra/conf/cassandra.yaml |
Windows installations | C:\Program Files\DataStax Community\apache-cassandra\conf\cassandra.yaml |