Start Apache Spark
Before you start Apache Spark™, configure RPC for the DseClientTool object.
RPC permission for the DseClientTool object is required to run Apache Spark because the DseClientTool object is called implicitly by the Spark launcher.
Prerequisites
By default DSEFS is required to execute Spark applications.
DSEFS should not be disabled when you enable Apache Spark on a DataStax Enterprise (DSE) node.
If there is a strong reason not to use DSEFS as the default file system, reconfigure Apache Spark to use a different file system.
For example to use a local file system set the following properties in spark-daemon-defaults.conf:
spark.hadoop.fs.defaultFS=file:///
spark.hadoop.hive.metastore.warehouse.dir=file:///tmp/warehouse
Start a Spark cluster
How you start Apache Spark depends on the installation and if you want to run in Spark mode or SearchAnalytics mode.
|
Test SearchAnalytics mode on a non-production cluster before enabling it in production. |
- Package installations
-
To start the Spark trackers on a cluster of analytics nodes, set
SPARK_ENABLEDto1in the/etc/default/dsefile. Then, when you start DSE as a service, the node is launched as a Spark node.To start the node in combined SearchAnalytics mode, set
SPARK_ENABLED=1andSEARCH_ENABLED=1in the/etc/default/dsefile. SearchAnalytics mode also requires that you setcql_solr_query_paging: driverindse.yaml. - Tarball installations
-
To start the Spark trackers on a cluster of analytics nodes, use the
-koption:INSTALL_DIRECTORY/bin/dse cassandra -kNodes started with
-kare automatically assigned to the default Analytics datacenter if you do not configure a datacenter in the snitch property file.To start the node in combined SearchAnalytics mode, use the
-kand-soptions. SearchAnalytics mode also requires that you setcql_solr_query_paging: driverindse.yaml.INSTALL_DIRECTORY/bin/dse cassandra -k -s
When you start a node with the Spark option, a node is designated as a Spark master.
To identify master nodes, run dsetool ring, and then find the nodes with Analytics(SM) in the Workload column:
Address DC Rack Workload Graph Status State Load Owns Token Health [0,1]
0
10.200.175.149 Analytics rack1 Analytics(SM) no Up Normal 185 KiB ? -9223372036854775808 0.90
10.200.175.148 Analytics rack1 Analytics(SW) no Up Normal 194.5 KiB ? 0 0.90
NOTE: you must specify a keyspace to get ownership information.
Launch Apache Spark
After starting a Spark node, use dse Spark commands to launch Apache Spark.
The directory in which you run dse Spark commands must be writable by the current user.
- Commands
-
-
dse spark: Launch the Apache Spark interactive shell. -
dse spark-submit: Launch applications on a cluster in the same way that you would use the Apache Sparkspark-submitcommand. With this interface, you can use Spark cluster managers without separate configurations for each application.
-
- Configuration
-
You can use Apache Cassandra specific properties to configure DSE Spark startup and operation.
Apache Spark binds to the
listen_addressthat is specified incassandra.yaml. - Authentication
-
Internal authentication is supported.
DataStax recommends using the environment variables
DSE_USERNAMEandDSE_PASSWORDto prevent these credentials from appearing in the Spark logs or the process list on the Spark Web UI.These environment variables are supported for all
dse sparkanddse client-toolcommands.Set these environment variables in your Bash
.profileor.bash_profilefile:export DSE_USERNAME=user export DSE_PASSWORD=secretFor other ways to provide authentication credentials, see Provide credentials from DSE tools.
Specify Spark URLs
You do not need to specify the Spark master address when starting Spark jobs with DSE. If you connect to any Spark node in a datacenter, DSE will automatically discover the master address and connect the client to the master.
Specify the URL for any Spark node using the following format:
dse://[<Spark node address>[:<port number>]]?[<parameter name>=<parameter value>;]<...>
By default the URL is dse://?, which is equivalent to dse://localhost:9042.
Any parameters you set in the URL will override the configuration read from DSE’s Apache Spark configuration settings.
You can specify the work pool in which the application will be run by adding the workpool=<work pool name> as a URL parameter.
For example, dse://1.1.1.1:123?workpool=workpool2.
Valid parameters are CassandraConnectorConf settings without the spark.cassandra. prefix.
For example, you can set the spark.cassandra.connection.local_dc option to dc2 by specifying dse://?connection.local_dc=dc2.
Or to specify multiple spark.cassandra.connection.host addresses for high-availability if the specified connection point is down: dse://1.1.1.1:123?connection.host=1.1.2.2,1.1.3.3.
If the connection.host parameter is specified, the host provided in the standard URL is prepended to the list of hosts set in connection.host.
If the port is specified in the standard URL, it overrides the port number set in the connection.port parameter.
Connection options when using dse spark-submit are retrieved in the following order: from the master URL, then the Apache Cassandra Spark Connector options, then the DSE configuration files.
Detect Spark application failures
DSE has a failure detector for Spark applications, which detects whether a running Spark application is dead or alive. If the application has failed, the application will be removed from the DSE Spark Resource Manager.
The failure detector works by keeping an open TCP connection from a DSE Spark node to the Spark driver in the application.
No data is exchanged, but regular TCP connection keep-alive control messages are sent and received.
When the connection is interrupted, the failure detector will attempt to reacquire the connection every 1 second for the duration of the appReconnectionTimeoutSeconds timeout value (5 seconds by default).
If it fails to reacquire the connection during that time, the application is removed.
A custom timeout value is specified by adding appReconnectionTimeoutSeconds=<value> in the master URI when submitting the application.
For example to set the timeout value to 10 seconds:
dse spark --master dse://?appReconnectionTimeoutSeconds=10