Running the Spark MLlib demo application
The Spark MLlib demo application demonstrates how to run machine-learning analytic jobs using Spark and DataStax Enterprise. The demo solves the classic iris flower classification problem, using the iris flower dataset. The application will use the iris flower dataset to build a Naive Bayes classifier that will recognize a flower based on four feature measurements.
We strongly recommend that you install the BLAS library on your machines before running Spark MLlib jobs.
The BLAS library is not distributed with DSE due to licensing restrictions, but improves MLlib performance significantly.
-
Install Gradle.
-
In a terminal, change to the
/demos/spark-mlibdirectory of your DSE installation.The default location of the
demosdirectory depends on the type of installation:-
Package installations:
/usr/share/dse/demosYou might need to install the demos packages, if you didn’t install them during your initial installation. For more information, see Install DataStax Enterprise (DSE) 5.1 on Debian-based systems with APT and Install DataStax Enterprise (DSE) 5.1 on RHEL-based systems with Yum.
-
Tarball installations:
INSTALL_DIRECTORY/demos
README.mdfiles are provided with your DSE installation and in thedemosdirectories with instructions on building and running the demo applications. -
-
Build the application using the
gradlebuild tool.gradle -
Use
spark-submitto submit the application JAR.The Spark MLlib demo application reads the Spark demo
directory/spark-mllib/iris.csvfile on each node. This file must be accessible in the same location on each node. If some nodes do not have the same local file path, set up a shared network location accessible to all the nodes in the cluster.To run the application where each node has access to the same local location of
iris.csv.dse spark-submit NaiveBayesDemo.jarTo specify a shared location of
iris.csv:dse spark-submit NaiveBayesDemo.jar /mnt/shared/iris.csv