Get started with the Data API
As an alternative to the Cassandra Query Language (CQL) and Cassandra drivers, you can use the Data API to programmatically interact with DataStax Enterprise (DSE).
You can interact with the Data API through direct HTTP requests or through the clients.
This quickstart explains how to install a client, connect to a database with the client, and then issue Data API commands to create a DSE keyspace, load data and autogenerate embeddings with $vectorize, and then run a similarity search.
Prerequisites
You need a running DSE cluster with the Data API. For exploration, consider running a DSE Docker container with the Data API.
Generate credentials
You need credentials to authorize Data API commands:
-
Get database user credentials.
The client uses these credentials to generate an authentication token.
In an isolated test environment, you can use the default superuser credentials (
cassandra/cassandra).When you install DSE, it creates a
cassandrasuperuser role in the database, and DSE runs as this user. Don’t use the defaultcassandrarole in production because it is a security risk. Instead, create a new superuser role for running DSE. -
Get your database’s Data API endpoint, which is the node’s hostname or IP address and port number.
The default port number is
8181. For example, if you are running a DSE Docker container, the API endpoint ishttp://localhost:8181. -
Optional: If you want to use the
vectorizeservice to automatically generate embeddings from text strings, create an OpenAI account and generate an OpenAI API key.If you plan to provide your own embeddings generated prior to loading data, you don’t need this API key.
-
Recommended: Set environment variables for your credentials.
For example:
export DB_USERNAME=DB_USERNAME # Your database username export DB_PASSWORD=DB_PASSWORD # Your database password export DB_API_ENDPOINT=DB_API_ENDPOINT # Your database API endpoint export OPENAI_API_KEY=API_KEY # Your OpenAI API key to generate embeddingsIf you don’t want to use environment variables, you can provide your credentials and API endpoint directly in the client code.
Install a client
Install the Data API client library for the language and package manager you’re using:
- Python
-
AstraPy is the official Python client for the Data API. For more information, see the Python client reference.
Install the Python client with pip:
-
Verify that pip is version 23.0 or later:
pip --version -
Upgrade pip if needed:
python -m pip install --upgrade pip -
Install the
astrapypackage.Python version 3.9 to 3.14 is required.
pip install astrapy
-
- TypeScript
-
astra-db-tsis the official TypeScript client for the Data API. For more information, see the TypeScript client reference.To install the TypeScript client:
-
Verify your Node version.
Node.js version 16.20.2 or later is required.
node --version -
Use npm or Yarn to install the TypeScript client:
With npm:
npm install @datastax/astra-db-tsOr with Yarn 2.0 or later:
yarn add @datastax/astra-db-ts
-
- Java
-
astra-db-javais the official Java client for the Data API. For more information, see the Java client reference.Use Maven or Gradle to install the Java client:
-
Install Java 11 or later.
-
Install Apache Maven™ 3.9 or later or Gradle.
-
Create a project that includes the Java client:
With Maven, add the Java client to your
pom.xmlfile:<dependencies> <dependency> <groupId>com.datastax.astra</groupId> <artifactId>astra-db-java</artifactId> <version>1.0.0</version> </dependency> </dependencies>Or with Gradle, add the Java client to your
build.gradlefile:dependencies { implementation 'com.datastax.astra:astra-db-java:1.0.0' }
-
Instantiate a client object
When you create apps using the Data API clients, your main entry point is the DataAPIClient object.
The client only needs a token because it isn’t specific to a database.
For DSE databases, the token is created using UsernamePasswordTokenProvider.
Initialize the Python client
To avoid a namespace collision, don’t name your Python files astrapy.py.
To follow along with this quickstart, create a file named quickstart.py in your project.
import os
from astrapy import DataAPIClient
from astrapy.constants import Environment
from astrapy.authentication import UsernamePasswordTokenProvider
from astrapy.constants import VectorMetric
from astrapy.ids import UUID
from astrapy.exceptions import InsertManyException
from astrapy.info import CollectionVectorServiceOptions
# Database settings
DB_USERNAME = "cassandra"
DB_PASSWORD = "cassandra"
DB_API_ENDPOINT = "http://localhost:8181"
DB_KEYSPACE = "my_keyspace"
DB_COLLECTION = "vector_test"
# Database settings if you exported them as environment variables
# DB_USERNAME = os.environ.get("DB_USERNAME")
# DB_PASSWORD = os.environ.get("DB_PASSWORD")
# DB_API_ENDPOINT = os.environ.get("DB_API_ENDPOINT")
# Embedding provider settings
EMBEDDING_PROVIDER = "openai";
EMBEDDING_MODEL_NAME = "text-embedding-3-small";
EMBEDDING_DIMENSIONS = 1024
EMBEDDING_API_KEY = os.environ.get("EMBEDDING_API_KEY");
# Build a token
tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD)
# Initialize the client and get a "Database" object
client = DataAPIClient(environment=Environment.DSE)
database = client.get_database(DB_API_ENDPOINT, token=tp)
database.get_database_admin().create_keyspace(DB_KEYSPACE, update_db_keyspace=True)
The specification for initializing the Python client is as follows. For more information, see the AstraPy API reference.
- Syntax
-
# Build a token tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD) # Initialize the client and get a "Database" object client = DataAPIClient(token=tp, environment=Environment.DB_ENVIRONMENT) - Parameters
-
Name Type Summary username
strA username for token creation. Example:
cassandra.password
strA password for token creation. Example:
cassandra.environment
strThe environment to use with
Environment. Example:DSE. - Returns
-
DataAPIClient: An instance of the client class
Initialize the TypeScript client
To follow along with this quickstart, create a file named quickstart.ts in your project:
import { DataAPIClient, UsernamePasswordTokenProvider, VectorDoc, UUID } from '@datastax/astra-db-ts';
// Database settings
const DB_USERNAME = "cassandra";
const DB_PASSWORD = "cassandra";
const DB_API_ENDPOINT = "http://localhost:8181";
const DB_ENVIRONMENT = "dse";
const DB_KEYSPACE = "cycling";
// Database settings if you exported them as environment variables
// const DB_USERNAME = process.env.DB_USERNAME;
// const DB_PASSWORD = process.env.DB_PASSWORD;
// const DB_API_ENDPOINT = process.env.DB_API_ENDPOINT;
// OpenAI settings
const OPEN_AI_PROVIDER = "openai";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY
const MODEL_NAME = "text-embedding-3-small";
// Build a token in the required format
const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD);
// Initialize the client and get a "Db" object
const client = new DataAPIClient({ environment: DB_ENVIRONMENT });
const db = client.db(DB_API_ENDPOINT, { token: tp });
const dbAdmin = db.admin({ environment: DB_ENVIRONMENT });
The specification for initializing the TypeScript client is as follows. For more information, see the astra-db-ts API Reference.
- Syntax
-
// Build a token in the required format const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD); // Initialize the client and get a "Db" object const client = new DataAPIClient({ environment: DB_ENVIRONMENT }); - General parameters
-
Name Type Summary username
strA username for token creation. Example:
cassandra.password
strA password for token creation. Example:
cassandra.options?
The options to use for the client, including defaults.
DataAPIClientOptions-
Name Type Summary Options related to the API requests the client makes.
The
DataAPIHttpOptionstype is a discriminated union on theclientfield. There are four available behaviors for theclientfield:+
-
httpOptionsnot set: Usefetch-h2if available or fall back tofetch. -
client: 'default'or unset: Usefetch-h2if available or throw an error. -
client: 'fetch': Only use the nativefetchAPI. -
client: 'custom': Pass a custom Fetcher implementation to the client.
+
fetch-h2is generally available by default on node runtimes only. On other runtimes, you might need to use the nativefetchAPI or, if your code is minified, pass in thefetch-h2module manually.+ For more information, see the astra-db-ts README and the DataAPIHttpOptions reference.
Allows default options for when spawning a Db instance.
Allows default options for when spawning some Admin instance.
-
- Monitoring/logging
-
For information on setting up commands monitoring, see the astra-db-ts README.
- Returns
-
DataAPIClient: An instance of the client class
Initialize the Java client
To follow along with this quickstart, create a file named Quickstart.java in the /src/main/java/com/example/ directory of your project.
The following code is incomplete.
You will add more code in the next steps.
import com.datastax.astra.client.Collection;
import com.datastax.astra.client.DataAPIClient;
import com.datastax.astra.client.Database;
import com.datastax.astra.client.admin.DataAPIDatabaseAdmin;
import com.datastax.astra.client.model.CollectionOptions;
import com.datastax.astra.client.model.CommandOptions;
import com.datastax.astra.client.model.Document;
import com.datastax.astra.client.model.FindOneOptions;
import com.datastax.astra.client.model.KeyspaceOptions;
import com.datastax.astra.client.model.SimilarityMetric;
import com.datastax.astra.internal.auth.UsernamePasswordTokenProvider;
import java.util.Optional;
import static com.datastax.astra.client.DataAPIClients.DEFAULT_ENDPOINT_LOCAL;
import static com.datastax.astra.client.DataAPIOptions.DataAPIDestination.HCD;
import static com.datastax.astra.client.DataAPIOptions.builder;
import static com.datastax.astra.client.model.Filters.eq;
public class QuickStartDSE69 {
public static void main(String[] args) {
// Database Settings
String cassandraUserName = "cassandra";
String cassandraPassword = "cassandra";
String dataApiUrl = DEFAULT_ENDPOINT_LOCAL; // http://localhost:8181
String databaseEnvironment = "DSE" // DSE, HCD, or ASTRA
String keyspaceName = "ks1";
String collectionName = "lyrics";
// Database settings if you export them as environment variables
// String cassandraUserName = System.getenv("DB_USERNAME");
// String cassandraPassword = System.getenv("DB_PASSWORD");
// String dataApiUrl = System.getenv("DB_API_ENDPOINT");
// OpenAI Embeddings
String openAiProvider = "openai";
String openAiKey = System.getenv("OPENAI_API_KEY"); // Need to export OPENAI_API_KEY
String openAiModel = "text-embedding-3-small";
int openAiEmbeddingDimension = 1536;
// Build a token in the form of Cassandra:base64(username):base64(password)
String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString();
System.out.println("1/7 - Creating Token: " + token);
// Initialize the client
DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build());
System.out.println("2/7 - Connected to Data API");
The specification for initializing the Java client is as follows:
- Syntax
-
// Build a token in the form of Cassandra:base64(username):base64(password) String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString(); System.out.println("1/7 - Creating Token: " + token); // Initialize the client DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build()); - Parameters
-
Name Type Summary username
strA username for token creation. Example:
cassandra.password
strA password for token creation. Example:
cassandra.options
A class wrapping the advanced configuration of the client such as as
HttpClientsettings (timeouts…). - Returns
-
DataAPIClient: An instance of the client class
Create a keyspace
Create a keyspace or use an existing one:
- Python
-
database.get_database_admin().create_keyspace(DB_KEYSPACE) - TypeScript
-
(async () => { await dbAdmin.createKeyspace(DB_KEYSPACE); console.log(await dbAdmin.listKeyspaces()); })(); - Java
-
// Create a default keyspace ((DataAPIDatabaseAdmin) client .getDatabase(dataApiUrl) .getDatabaseAdmin()).createKeyspace(keyspaceName, KeyspaceOptions.simpleStrategy(1)); System.out.println("3/7 - Keyspace '" + keyspaceName + "'created "); Database db = client.getDatabase(dataApiUrl, keyspaceName); System.out.println("4/7 - Connected to Database");
Create a collection
Create a collection in your keyspace.
Choose dimensions that match your vector data and pick an appropriate similarity metric: cosine (default), dot_product, or euclidean.
Add the following code after your client initialization and keyspace code.
- Python
-
# Create a collection. The default similarity metric is cosine. If you're not # sure what dimension to set, use whatever dimension vector your embeddings # model produces. collection = database.create_collection( DB_COLLECTION, dimension=EMBEDDING_DIMENSIONS, metric=VectorMetric.COSINE, service={ "provider": EMBEDDING_PROVIDER, "modelName": EMBEDDING_MODEL_NAME, }, embedding_api_key=EMBEDDING_API_KEY, keyspace=DB_KEYSPACE, check_exists=False, ) print(f"* Collection: {collection.full_name}\n") - TypeScript
-
// Schema for the collection (VectorDoc adds the $vector field) interface Idea extends VectorDoc { idea: string, } (async function () { // Create a typed, vector-enabled collection. The default metric is cosine. // If you're not sure what dimension to set, use whatever dimension vector // your embeddings model produces. const collection = await db.createCollection<Idea>('vector_test', { keyspace: DB_KEYSPACE, vector: { service: { provider: OPEN_AI_PROVIDER, modelName: MODEL_NAME }, dimension: 5, metric: 'cosine', }, embeddingApiKey: OPENAI_API_KEY, checkExists: false }); console.log(* Created collection ${collection.keyspace}.${collection.collectionName}); - Java
-
// Create a collection Collection<Document> collectionLyrics = db.createCollection(collectionName, CollectionOptions.builder() .vectorDimension(5) .vectorSimilarity(SimilarityMetric.COSINE) .build(), System.out.println("5/7 - Collection created");
Load vector data
Insert a few documents into the collection.
Two methods are available for inserting vector data:
-
The
$vectorizemethod generates embeddings using a specified embedding service. -
The
$vectormethod is used when you already have embeddings.
These methods are mutually exclusive, but you can use both in the same database.
For example, you can use $vectorize as your primary embedding generation method, and then use $vector to insert some bespoke embeddings as needed.
Add the code for your preferred method to your quickstart script.
Use the $vectorize method
The $vectorize Data API method supports various embedding providers.
The following examples use OpenAI as the embedding provider.
To find and configure other providers, see your client’s reference.
You need an API key for the provider.
- Python
-
# Insert documents into the collection. # (UUIDs here are version 7.) documents = [ { "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"), "text": "Chat bot integrated sneakers that talk to you", "$vectorize": "Wild! How can they do that?" }, { "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"), "text": "An AI quilt to help you sleep forever", "$vectorize": "Sleep like a baby soft and cuddly" }, { "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"), "text": "A deep learning display that controls your mood", "$vectorize": "I do not want my mood controlled!" }, ] try: insertion_result = collection.insert_many(documents) print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n") except InsertManyException: print("* Documents found on DB already. Let's move on.\n") - TypeScript
-
// Insert documents into the collection (using UUIDv7s) const documents = [ { _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'), text: 'ChatGPT integrated sneakers that talk to you', $vectorize: 'Wild! How can they do that?', }, { _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'), text: 'An AI quilt to help you sleep forever', $vectorize: 'Sleep like a baby soft and cuddly', }, { _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'), text: 'A deep learning display that controls your mood', $vectorize: 'I do not want my mood controlled!', }, ]; try { const inserted = await collection.insertMany(documents); console.log(`* Inserted ${inserted.insertedCount} items.`); } catch (e) { console.log('* Documents found on DB already. Let\'s move on!'); } - Java
-
// Insert some documents collectionLyrics.insertMany( new Document(1).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("A lovestruck Romeo sings the streets a serenade"), new Document(2).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Says something like, You and me babe, how about it?"), new Document(4).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Juliet says,Hey, it's Romeo, you nearly gimme a heart attack"), new Document(5).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("He's underneath the window"), new Document(6).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("She's singing, Hey la, my boyfriend's back"), new Document(7).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("You shouldn't come around here singing up at people like that"), new Document(8).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Anyway, what you gonna do about it?")); System.out.println("6/7 - Collection populated");
Use the $vector method
The $vector method can be used if you generate embeddings before loading data.
- Python
-
# Insert documents into the collection. # (UUIDs here are version 7.) documents = [ { "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"), "$vectorize": "Chat bot integrated sneakers that talk to you", }, { "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"), "$vectorize": "An AI quilt to help you sleep forever", }, { "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"), "$vectorize": "A deep learning display that controls your mood", }, ] try: insertion_result = collection.insert_many(documents) print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n") except InsertManyException: print("* Documents found on DB already. Let's move on.\n") - TypeScript
-
// Insert documents into the collection (using UUIDv7s) const documents = [ { _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'), text: 'ChatGPT integrated sneakers that talk to you', $vector: [0.25, 0.25, 0.25, 0.25, 0.45], }, { _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'), text: 'An AI quilt to help you sleep forever', $vector: [0.10, 0.15, 0.25, 0.25, 0.15], }, { _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'), text: 'A deep learning display that controls your mood', $vector: 'I do not want my mood controlled!', }, ]; try { const inserted = await collection.insertMany(documents); console.log(`* Inserted ${inserted.insertedCount} items.`); } catch (e) { console.log('* Documents found on DB already. Let\'s move on!'); } - Java
-
// Insert some documents collection.insertMany( new Document("1") .append("text", "ChatGPT integrated sneakers that talk to you") .vector(new float[]{0.1f, 0.15f, 0.3f, 0.12f, 0.05f}), new Document("2") .append("text", "An AI quilt to help you sleep forever") .vector(new float[]{0.45f, 0.09f, 0.01f, 0.2f, 0.11f}), new Document("3") .append("text", "A deep learning display that controls your mood") .vector(new float[]{0.1f, 0.05f, 0.08f, 0.3f, 0.6f})); System.out.println("6/7 - Collection populated");
Run a similarity search
A vector search (similarity search) finds documents that are close to a specific vector embedding.
The output is a list of documents sorted by similarity.
For this example, the calculation uses cosine similarity.
The similarity metric and dimensions are set when you create a collection.
The following examples show a complete script using an existing keyspace, $vectorize, and a similarity search.
Try running your quickstart script on a test database.
Then, try modifying the script to insert more documents or run different commands.
- Python
-
import os from astrapy import DataAPIClient from astrapy.constants import Environment from astrapy.authentication import UsernamePasswordTokenProvider from astrapy.constants import VectorMetric from astrapy.ids import UUID from astrapy.exceptions import InsertManyException from astrapy.info import CollectionVectorServiceOptions # Database settings DB_USERNAME = "cassandra" DB_PASSWORD = "cassandra" DB_API_ENDPOINT = "http://localhost:8181" DB_KEYSPACE = "my_keyspace" DB_COLLECTION = "vector_test" # Database settings if you exported them as environment variables # DB_USERNAME = os.environ.get("DB_USERNAME") # DB_PASSWORD = os.environ.get("DB_PASSWORD") # DB_API_ENDPOINT = os.environ.get("DB_API_ENDPOINT") # Embedding provider settings EMBEDDING_PROVIDER = "openai"; EMBEDDING_MODEL_NAME = "text-embedding-3-small"; EMBEDDING_DIMENSIONS = 1024 EMBEDDING_API_KEY = os.environ.get("EMBEDDING_API_KEY"); # Build a token tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD) # Initialize the client and get a "Database" object client = DataAPIClient(environment=Environment.DSE) database = client.get_database(DB_API_ENDPOINT, token=tp) database.get_database_admin().create_keyspace(DB_KEYSPACE, update_db_keyspace=True) # Create a collection. The default similarity metric is cosine. If you're not # sure what dimension to set, use whatever dimension vector your embeddings # model produces. collection = database.create_collection( DB_COLLECTION, dimension=EMBEDDING_DIMENSIONS, metric=VectorMetric.COSINE, service={ "provider": EMBEDDING_PROVIDER, "modelName": EMBEDDING_MODEL_NAME, }, embedding_api_key=EMBEDDING_API_KEY, keyspace=DB_KEYSPACE, check_exists=False, ) print(f"* Collection: {collection.full_name}\n") # Insert documents into the collection. # (UUIDs here are version 7.) documents = [ { "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"), "$vectorize": "Chat bot integrated sneakers that talk to you", }, { "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"), "$vectorize": "An AI quilt to help you sleep forever", }, { "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"), "$vectorize": "A deep learning display that controls your mood", }, ] try: insertion_result = collection.insert_many(documents) print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n") except InsertManyException: print("* Documents found on DB already. Let's move on.\n") # Perform a similarity search query = [0.15, 0.1, 0.1, 0.35, 0.55] results = collection.find( sort={"$vector": query}, limit=10, ) print("Vector search results:") for document in results: print(" ", document)Run the script with
python quickstart.pyor the name of your script file. - TypeScript
-
import { DataAPIClient, UsernamePasswordTokenProvider, VectorDoc, UUID } from '@datastax/astra-db-ts'; // Database settings const DB_USERNAME = "cassandra"; const DB_PASSWORD = "cassandra"; const DB_API_ENDPOINT = "http://localhost:8181"; const DB_ENVIRONMENT = "dse"; const DB_KEYSPACE = "cycling"; // Database settings if you exported them as environment variables // const DB_USERNAME = process.env.DB_USERNAME; // const DB_PASSWORD = process.env.DB_PASSWORD; // const DB_API_ENDPOINT = process.env.DB_API_ENDPOINT; // OpenAI settings const OPEN_AI_PROVIDER = "openai"; const OPENAI_API_KEY = process.env.OPENAI_API_KEY const MODEL_NAME = "text-embedding-3-small"; // Build a token in the required format const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD); // Initialize the client and get a "Db" object const client = new DataAPIClient({ environment: DB_ENVIRONMENT }); const db = client.db(DB_API_ENDPOINT, { token: tp }); const dbAdmin = db.admin({ environment: DB_ENVIRONMENT }); // Schema for the collection (VectorDoc adds the $vector field) interface Idea extends VectorDoc { idea: string, } (async function () { // Create a typed, vector-enabled collection. The default metric is cosine. // If you're not sure what dimension to set, use whatever dimension vector // your embeddings model produces. const collection = await db.createCollection<Idea>('vector_test', { keyspace: DB_KEYSPACE, vector: { service: { provider: OPEN_AI_PROVIDER, modelName: MODEL_NAME }, dimension: 5, metric: 'cosine', }, embeddingApiKey: OPENAI_API_KEY, checkExists: false }); console.log(`* Created collection ${collection.keyspace}.${collection.collectionName}`); // Insert documents into the collection (using UUIDv7s) const documents = [ { _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'), text: 'ChatGPT integrated sneakers that talk to you', $vector: [0.25, 0.25, 0.25, 0.25, 0.45], }, { _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'), text: 'An AI quilt to help you sleep forever', $vector: [0.10, 0.15, 0.25, 0.25, 0.15], }, { _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'), text: 'A deep learning display that controls your mood', $vector: 'I do not want my mood controlled!', }, ]; try { const inserted = await collection.insertMany(documents); console.log(`* Inserted ${inserted.insertedCount} items.`); } catch (e) { console.log('* Documents found on DB already. Let\'s move on!'); } // Perform a similarity search const cursor = await collection.find({}, { vector: [0.15, 0.1, 0.1, 0.35, 0.55], limit: 10, includeSimilarity: true, }); console.log('* Search results:') for await (const doc of cursor) { console.log(' ', doc.text, doc.$similarity); } // Cleanup (if desired) //await db.dropCollection('vector_test'); //console.log('* Collection dropped.'); // Close the client await client.close(); })();Run the script with npm or Yarn and the name of your script file, such as
npx tsx quickstart.tsoryarn dlx tsx quickstart.ts. - Java
-
import com.datastax.astra.client.Collection; import com.datastax.astra.client.DataAPIClient; import com.datastax.astra.client.Database; import com.datastax.astra.client.admin.DataAPIDatabaseAdmin; import com.datastax.astra.client.model.CollectionOptions; import com.datastax.astra.client.model.CommandOptions; import com.datastax.astra.client.model.Document; import com.datastax.astra.client.model.FindOneOptions; import com.datastax.astra.client.model.KeyspaceOptions; import com.datastax.astra.client.model.SimilarityMetric; import com.datastax.astra.internal.auth.UsernamePasswordTokenProvider; import java.util.Optional; import static com.datastax.astra.client.DataAPIClients.DEFAULT_ENDPOINT_LOCAL; import static com.datastax.astra.client.DataAPIOptions.DataAPIDestination.HCD; import static com.datastax.astra.client.DataAPIOptions.builder; import static com.datastax.astra.client.model.Filters.eq; public class QuickStartDSE69 { public static void main(String[] args) { // Database Settings String cassandraUserName = "cassandra"; String cassandraPassword = "cassandra"; String dataApiUrl = DEFAULT_ENDPOINT_LOCAL; // http://localhost:8181 String databaseEnvironment = "DSE" // DSE, HCD, or ASTRA String keyspaceName = "ks1"; String collectionName = "lyrics"; // Database settings if you export them as environment variables // String cassandraUserName = System.getenv("DB_USERNAME"); // String cassandraPassword = System.getenv("DB_PASSWORD"); // String dataApiUrl = System.getenv("DB_API_ENDPOINT"); // OpenAI Embeddings String openAiProvider = "openai"; String openAiKey = System.getenv("OPENAI_API_KEY"); // Need to export OPENAI_API_KEY String openAiModel = "text-embedding-3-small"; int openAiEmbeddingDimension = 1536; // Build a token in the form of Cassandra:base64(username):base64(password) String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString(); System.out.println("1/7 - Creating Token: " + token); // Initialize the client DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build()); System.out.println("2/7 - Connected to Data API"); // Create a collection Collection<Document> collectionLyrics = db.createCollection(collectionName, CollectionOptions.builder() .vectorDimension(5) .vectorSimilarity(SimilarityMetric.COSINE) .build(), System.out.println("5/7 - Collection created"); // Insert some documents collection.insertMany( new Document("1") .append("text", "ChatGPT integrated sneakers that talk to you") .vector(new float[]{0.1f, 0.15f, 0.3f, 0.12f, 0.05f}), new Document("2") .append("text", "An AI quilt to help you sleep forever") .vector(new float[]{0.45f, 0.09f, 0.01f, 0.2f, 0.11f}), new Document("3") .append("text", "A deep learning display that controls your mood") .vector(new float[]{0.1f, 0.05f, 0.08f, 0.3f, 0.6f})); System.out.println("6/7 - Collection populated"); FindIterable<Document> resultsSet = collection.find( new float[]{0.15f, 0.1f, 0.1f, 0.35f, 0.55f}, 10 ); resultsSet.forEach(System.out::println); // Cleanup (if desired) //collection.drop(); //System.out.println("Deleted the collection"); } }Build and run your Java project. For example, with Maven:
mvn clean compile export OPENAI_API_KEY=<your-api-key> mvn exec:java -Dexec.mainClass="com.example.QuickStartDSE"Or with Gradle:
gradle build gradle run
Data API limits
The following limits apply to Data API usage with DSE.
| Entity | Limit | Notes |
|---|---|---|
Property naming conventions |
User-defined property names must:
|
System-defined operator and property names are prefixed by |
Supported data types |
|
See your client’s reference for information about working with dates, UUIDs and ObjectIDs. |
Number of collections per keyspace |
Five |
Up to five collections in a DSE keyspace. |
Page size |
20 |
A page may contain up to 20 documents. After that per-page maximum is reached, you can load any additional documents on the next page via the |
Sort page size |
100 |
Document page size for sorting; implemented as separate from page size because sort operations need more rows per page. |
Maximum property name |
100 |
Maximum of 100 characters in a property name. |
Maximum path length |
1,000 |
Maximum of 1,000 characters in a path name; total for all segments, including any dots (.) between properties in a path. |
String property maximum bytes |
8,000 |
Maximum of 8,000 UTF-8 bytes for |
Number property maximum characters |
100 |
Maximum of 100 characters for |
Maximum elements per array |
1,000 |
Maximum number of elements in an array. This limit applies to indexed properties only. This limit is ignored for non-indexed properties. |
Maximum dimensions in vector-enabled collection |
4,096 |
Maximum size of dimensions you can define for a vector-enabled collection. |
Maximum number of properties per JSON object |
1,000 |
Maximum number of properties for a JSON object. This limit applies to indexed properties only. This limit is ignored for non-indexed properties. A given JSON object may have nested objects, also known as sub-documents. This maximum total count of 1,000 refers to all the indexed properties in the main document, plus a count of 1 for each sub-document (if any). |
Maximum number of properties per JSON document |
2,000 |
Maximum number of properties allowed in a single JSON document is 2,000. This limit includes intermediate properties as well as leaf properties.
For example, the document |
Maximum document size in characters |
4 million |
Maximum size of each document in a collection is 4 million characters. |
Maximum inserted batch size in characters |
20 million |
Maximum size of an entire batch of documents submitted via an |
Maximum number of documents deleted per transaction |
20 |
Maximum number of documents that can be deleted in each transaction. |
Maximum number of documents updated per transaction |
20 |
Maximum number of documents that can be updated in each transaction. |
Maximum number of documents inserted per transaction |
20 |
Maximum number of documents that can be inserted in each transaction when using |
Maximum size |
100 |
Maximum size of an |
Maximum number of documents returned with each vector search |
1,000 |
Maximum number of documents returned with each vector search. |
|
If your request is valid but the command exceeds a limit, the Data API responds with It is also possible to receive a response containing both data and errors. Always inspect the response for error messages. For example, if you exceed the per-transaction limit of 20 documents in an
|
Data API operators
Data API supports the following logical and update operators that you can use in filters. For usage examples, see Documents reference.
| Operator type | Name | Purpose |
|---|---|---|
Logical query |
|
Joins query clauses with a logical |
Logical query |
|
Joins query clauses with a logical |
Logical query |
|
Returns documents that do not match the conditions of the filter clause. |
Range query |
|
Matches documents where the given property is greater than the specified value. |
Range query |
|
Matches documents where the given property is greater than or equal to the specified value. |
Range query |
|
Matches documents where the given property is less than the specified value. |
Range query |
|
Matches documents where the given property is less than or equal to the specified value. |
Comparison query |
|
Matches documents where the value of a property equals the specified value. This is the default when you do not specify an operator. |
Comparison query |
|
Matches documents where the value of a property does not equal the specified value. |
Comparison query |
|
Matches any of the values specified in the array. |
Comparison query |
|
Matches any of the values that are NOT IN the array. |
Element query |
|
Matches documents that have the specified property. |
Array query |
|
Matches arrays that contain all elements in the specified array. |
Array query |
|
Selects documents where the array has the specified number of elements. |
Property update |
|
Used in an update operation.
In the following example, the
|
Property update |
|
Increments the value of the property by the specified amount. |
Property update |
|
Updates the property only if the specified value is less than the existing property value. |
Property update |
|
Updates the property only if the specified value is greater than the existing property value. |
Property update |
|
Multiply the value of a property in the document. For example:
|
Property update |
|
Renames the specified property in each matching document. |
Property update |
|
Sets the value of a property in each matching document. |
Property update |
|
Set the value of a property in the document if an upsert is performed. For example:
|
Property update |
|
Removes the specified property from each matching document. |
Array update |
|
Adds elements to the array only if they do not already exist in the set. |
Array update |
|
Removes the first or last item of the array, depending on the value of the operator ( |
Array update |
|
Adds or appends data to the end of the property value. Or, if the value is not yet an array:
|
Array update |
|
An array update that modifies the |
Array update |
|
An array update that modifies the |
Next steps
Learn more about Data API commands: