Get started with the Data API

As an alternative to the Cassandra Query Language (CQL) and Cassandra drivers, you can use the Data API to programmatically interact with DataStax Enterprise (DSE).

You can interact with the Data API through direct HTTP requests or through the clients.

This quickstart explains how to install a client, connect to a database with the client, and then issue Data API commands to create a DSE keyspace, load data and autogenerate embeddings with $vectorize, and then run a similarity search.

Prerequisites

You need a running DSE cluster with the Data API. For exploration, consider running a DSE Docker container with the Data API.

Generate credentials

You need credentials to authorize Data API commands:

  1. Get database user credentials.

    The client uses these credentials to generate an authentication token.

    In an isolated test environment, you can use the default superuser credentials (cassandra/cassandra).

    When you install DSE, it creates a cassandra superuser role in the database, and DSE runs as this user. Don’t use the default cassandra role in production because it is a security risk. Instead, create a new superuser role for running DSE.

  2. Get your database’s Data API endpoint, which is the node’s hostname or IP address and port number.

    The default port number is 8181. For example, if you are running a DSE Docker container, the API endpoint is http://localhost:8181.

  3. Optional: If you want to use the vectorize service to automatically generate embeddings from text strings, create an OpenAI account and generate an OpenAI API key.

    If you plan to provide your own embeddings generated prior to loading data, you don’t need this API key.

  4. Recommended: Set environment variables for your credentials.

    For example:

    export DB_USERNAME=DB_USERNAME # Your database username
    export DB_PASSWORD=DB_PASSWORD # Your database password
    export DB_API_ENDPOINT=DB_API_ENDPOINT # Your database API endpoint
    export OPENAI_API_KEY=API_KEY # Your OpenAI API key to generate embeddings

    If you don’t want to use environment variables, you can provide your credentials and API endpoint directly in the client code.

Install a client

Install the Data API client library for the language and package manager you’re using:

Python

AstraPy is the official Python client for the Data API. For more information, see the Python client reference.

Install the Python client with pip:

  1. Verify that pip is version 23.0 or later:

    pip --version
  2. Upgrade pip if needed:

    python -m pip install --upgrade pip
  3. Install the astrapy package.

    Python version 3.9 to 3.14 is required.

    pip install astrapy
TypeScript

astra-db-ts is the official TypeScript client for the Data API. For more information, see the TypeScript client reference.

To install the TypeScript client:

  1. Verify your Node version.

    Node.js version 16.20.2 or later is required.

    node --version
  2. Use npm or Yarn to install the TypeScript client:

    With npm:

    npm install @datastax/astra-db-ts

    Or with Yarn 2.0 or later:

    yarn add @datastax/astra-db-ts
Java

astra-db-java is the official Java client for the Data API. For more information, see the Java client reference.

Use Maven or Gradle to install the Java client:

  1. Install Java 11 or later.

  2. Install Apache Maven™ 3.9 or later or Gradle.

  3. Create a project that includes the Java client:

    With Maven, add the Java client to your pom.xml file:

      <dependencies>
        <dependency>
          <groupId>com.datastax.astra</groupId>
          <artifactId>astra-db-java</artifactId>
          <version>1.0.0</version>
        </dependency>
      </dependencies>

    Or with Gradle, add the Java client to your build.gradle file:

    dependencies {
        implementation 'com.datastax.astra:astra-db-java:1.0.0'
    }

Instantiate a client object

When you create apps using the Data API clients, your main entry point is the DataAPIClient object. The client only needs a token because it isn’t specific to a database. For DSE databases, the token is created using UsernamePasswordTokenProvider.

Initialize the Python client

To avoid a namespace collision, don’t name your Python files astrapy.py. To follow along with this quickstart, create a file named quickstart.py in your project.

import os

from astrapy import DataAPIClient
from astrapy.constants import Environment
from astrapy.authentication import UsernamePasswordTokenProvider
from astrapy.constants import VectorMetric
from astrapy.ids import UUID
from astrapy.exceptions import InsertManyException
from astrapy.info import CollectionVectorServiceOptions

# Database settings
DB_USERNAME = "cassandra"
DB_PASSWORD = "cassandra"
DB_API_ENDPOINT = "http://localhost:8181"
DB_KEYSPACE = "my_keyspace"
DB_COLLECTION = "vector_test"

# Database settings if you exported them as environment variables
# DB_USERNAME = os.environ.get("DB_USERNAME")
# DB_PASSWORD = os.environ.get("DB_PASSWORD")
# DB_API_ENDPOINT = os.environ.get("DB_API_ENDPOINT")

# Embedding provider settings
EMBEDDING_PROVIDER = "openai";
EMBEDDING_MODEL_NAME = "text-embedding-3-small";
EMBEDDING_DIMENSIONS = 1024
EMBEDDING_API_KEY = os.environ.get("EMBEDDING_API_KEY");

# Build a token
tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD)

# Initialize the client and get a "Database" object
client = DataAPIClient(environment=Environment.DSE)
database = client.get_database(DB_API_ENDPOINT, token=tp)
database.get_database_admin().create_keyspace(DB_KEYSPACE, update_db_keyspace=True)

The specification for initializing the Python client is as follows. For more information, see the AstraPy API reference.

Syntax
# Build a token
tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD)

# Initialize the client and get a "Database" object
client = DataAPIClient(token=tp, environment=Environment.DB_ENVIRONMENT)
Parameters
Name Type Summary

username

str

A username for token creation. Example: cassandra.

password

str

A password for token creation. Example: cassandra.

environment

str

The environment to use with Environment. Example: DSE.

Returns

DataAPIClient: An instance of the client class

Initialize the TypeScript client

To follow along with this quickstart, create a file named quickstart.ts in your project:

import { DataAPIClient, UsernamePasswordTokenProvider, VectorDoc, UUID } from '@datastax/astra-db-ts';

// Database settings
const DB_USERNAME = "cassandra";
const DB_PASSWORD = "cassandra";
const DB_API_ENDPOINT = "http://localhost:8181";
const DB_ENVIRONMENT = "dse";
const DB_KEYSPACE = "cycling";

// Database settings if you exported them as environment variables
// const DB_USERNAME = process.env.DB_USERNAME;
// const DB_PASSWORD = process.env.DB_PASSWORD;
// const DB_API_ENDPOINT = process.env.DB_API_ENDPOINT;

// OpenAI settings
const OPEN_AI_PROVIDER = "openai";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY
const MODEL_NAME = "text-embedding-3-small";

// Build a token in the required format
const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD);

// Initialize the client and get a "Db" object
const client = new DataAPIClient({ environment: DB_ENVIRONMENT });
const db = client.db(DB_API_ENDPOINT, { token: tp });
const dbAdmin = db.admin({ environment: DB_ENVIRONMENT });

The specification for initializing the TypeScript client is as follows. For more information, see the astra-db-ts API Reference.

Syntax
// Build a token in the required format
const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD);

// Initialize the client and get a "Db" object
const client = new DataAPIClient({ environment: DB_ENVIRONMENT });
General parameters
Name Type Summary

username

str

A username for token creation. Example: cassandra.

password

str

A password for token creation. Example: cassandra.

options?

DataAPIClientOptions

The options to use for the client, including defaults.

DataAPIClientOptions
Name Type Summary

httpOptions?

DataAPIHttpOptions

Options related to the API requests the client makes.

The DataAPIHttpOptions type is a discriminated union on the client field. There are four available behaviors for the client field:

+

  • httpOptions not set: Use fetch-h2 if available or fall back to fetch.

  • client: 'default' or unset: Use fetch-h2 if available or throw an error.

  • client: 'fetch': Only use the native fetch API.

  • client: 'custom': Pass a custom Fetcher implementation to the client.

+ fetch-h2 is generally available by default on node runtimes only. On other runtimes, you might need to use the native fetch API or, if your code is minified, pass in the fetch-h2 module manually.

+ For more information, see the astra-db-ts README and the DataAPIHttpOptions reference.

dbOptions?

DbSpawnOptions

Allows default options for when spawning a Db instance.

adminOptions?

AdminSpawnOptions

Allows default options for when spawning some Admin instance.

Monitoring/logging

For information on setting up commands monitoring, see the astra-db-ts README.

Returns

DataAPIClient: An instance of the client class

Initialize the Java client

To follow along with this quickstart, create a file named Quickstart.java in the /src/main/java/com/example/ directory of your project. The following code is incomplete. You will add more code in the next steps.

import com.datastax.astra.client.Collection;
import com.datastax.astra.client.DataAPIClient;
import com.datastax.astra.client.Database;
import com.datastax.astra.client.admin.DataAPIDatabaseAdmin;
import com.datastax.astra.client.model.CollectionOptions;
import com.datastax.astra.client.model.CommandOptions;
import com.datastax.astra.client.model.Document;
import com.datastax.astra.client.model.FindOneOptions;
import com.datastax.astra.client.model.KeyspaceOptions;
import com.datastax.astra.client.model.SimilarityMetric;
import com.datastax.astra.internal.auth.UsernamePasswordTokenProvider;

import java.util.Optional;

import static com.datastax.astra.client.DataAPIClients.DEFAULT_ENDPOINT_LOCAL;
import static com.datastax.astra.client.DataAPIOptions.DataAPIDestination.HCD;
import static com.datastax.astra.client.DataAPIOptions.builder;
import static com.datastax.astra.client.model.Filters.eq;

public class QuickStartDSE69 {

    public static void main(String[] args) {

        // Database Settings
        String cassandraUserName     = "cassandra";
        String cassandraPassword     = "cassandra";
        String dataApiUrl            = DEFAULT_ENDPOINT_LOCAL;  // http://localhost:8181
        String databaseEnvironment   = "DSE" // DSE, HCD, or ASTRA
        String keyspaceName          = "ks1";
        String collectionName        = "lyrics";

        // Database settings if you export them as environment variables
        // String cassandraUserName            = System.getenv("DB_USERNAME");
        // String cassandraPassword            = System.getenv("DB_PASSWORD");
        // String dataApiUrl                   = System.getenv("DB_API_ENDPOINT");

        // OpenAI Embeddings
        String openAiProvider        = "openai";
        String openAiKey             = System.getenv("OPENAI_API_KEY"); // Need to export OPENAI_API_KEY
        String openAiModel           = "text-embedding-3-small";
        int openAiEmbeddingDimension = 1536;

        // Build a token in the form of Cassandra:base64(username):base64(password)
        String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString();
        System.out.println("1/7 - Creating Token: " + token);

        // Initialize the client
        DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build());
        System.out.println("2/7 - Connected to Data API");

The specification for initializing the Java client is as follows:

Syntax
// Build a token in the form of Cassandra:base64(username):base64(password)
String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString();
System.out.println("1/7 - Creating Token: " + token);

// Initialize the client
DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build());
Parameters
Name Type Summary

username

str

A username for token creation. Example: cassandra.

password

str

A password for token creation. Example: cassandra.

options

DataAPIOptions

A class wrapping the advanced configuration of the client such as as HttpClient settings (timeouts…​).

Returns

DataAPIClient: An instance of the client class

Create a keyspace

Create a keyspace or use an existing one:

Python
database.get_database_admin().create_keyspace(DB_KEYSPACE)
TypeScript
(async () => {
  await dbAdmin.createKeyspace(DB_KEYSPACE);
  console.log(await dbAdmin.listKeyspaces());
})();
Java
        // Create a default keyspace
        ((DataAPIDatabaseAdmin) client
                .getDatabase(dataApiUrl)
                .getDatabaseAdmin()).createKeyspace(keyspaceName, KeyspaceOptions.simpleStrategy(1));
        System.out.println("3/7 - Keyspace '" + keyspaceName + "'created ");

        Database db = client.getDatabase(dataApiUrl, keyspaceName);
        System.out.println("4/7 - Connected to Database");

Create a collection

Create a collection in your keyspace.

Choose dimensions that match your vector data and pick an appropriate similarity metric: cosine (default), dot_product, or euclidean.

Add the following code after your client initialization and keyspace code.

Python
# Create a collection. The default similarity metric is cosine. If you're not
# sure what dimension to set, use whatever dimension vector your embeddings
# model produces.
collection = database.create_collection(
    DB_COLLECTION,
    dimension=EMBEDDING_DIMENSIONS,
    metric=VectorMetric.COSINE,
    service={
        "provider": EMBEDDING_PROVIDER,
        "modelName": EMBEDDING_MODEL_NAME,
    },
    embedding_api_key=EMBEDDING_API_KEY,
    keyspace=DB_KEYSPACE,
    check_exists=False,
)
print(f"* Collection: {collection.full_name}\n")
TypeScript
// Schema for the collection (VectorDoc adds the $vector field)
interface Idea extends VectorDoc {
  idea: string,
}

(async function () {
  // Create a typed, vector-enabled collection. The default metric is cosine.
  // If you're not sure what dimension to set, use whatever dimension vector
  // your embeddings model produces.
  const collection = await db.createCollection<Idea>('vector_test', {
    keyspace: DB_KEYSPACE,
    vector: {
      service: {
        provider: OPEN_AI_PROVIDER,
        modelName: MODEL_NAME
      },
      dimension: 5,
      metric: 'cosine',
    },
    embeddingApiKey: OPENAI_API_KEY,
    checkExists: false
  });
  console.log(* Created collection ${collection.keyspace}.${collection.collectionName});
Java
        // Create a collection
        Collection<Document> collectionLyrics =  db.createCollection(collectionName, CollectionOptions.builder()
        .vectorDimension(5)
        .vectorSimilarity(SimilarityMetric.COSINE)
        .build(),
        System.out.println("5/7 - Collection created");

Load vector data

Insert a few documents into the collection.

Two methods are available for inserting vector data:

  • The $vectorize method generates embeddings using a specified embedding service.

  • The $vector method is used when you already have embeddings.

These methods are mutually exclusive, but you can use both in the same database. For example, you can use $vectorize as your primary embedding generation method, and then use $vector to insert some bespoke embeddings as needed.

Add the code for your preferred method to your quickstart script.

Use the $vectorize method

The $vectorize Data API method supports various embedding providers. The following examples use OpenAI as the embedding provider. To find and configure other providers, see your client’s reference. You need an API key for the provider.

Python
# Insert documents into the collection.
# (UUIDs here are version 7.)
documents = [
    {
        "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"),
        "text": "Chat bot integrated sneakers that talk to you",
         "$vectorize": "Wild! How can they do that?"
    },
    {
        "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"),
        "text": "An AI quilt to help you sleep forever",
         "$vectorize": "Sleep like a baby soft and cuddly"
    },
    {
        "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"),
        "text": "A deep learning display that controls your mood",
         "$vectorize": "I do not want my mood controlled!"
    },
]
try:
    insertion_result = collection.insert_many(documents)
    print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n")
except InsertManyException:
    print("* Documents found on DB already. Let's move on.\n")
TypeScript
  // Insert documents into the collection (using UUIDv7s)
  const documents = [
    {
      _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'),
      text: 'ChatGPT integrated sneakers that talk to you',
      $vectorize: 'Wild! How can they do that?',
    },
    {
      _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'),
      text: 'An AI quilt to help you sleep forever',
      $vectorize: 'Sleep like a baby soft and cuddly',
    },
    {
      _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'),
      text: 'A deep learning display that controls your mood',
      $vectorize: 'I do not want my mood controlled!',
    },
  ];

  try {
    const inserted = await collection.insertMany(documents);
    console.log(`* Inserted ${inserted.insertedCount} items.`);
  } catch (e) {
    console.log('* Documents found on DB already. Let\'s move on!');
  }
Java
    // Insert some documents
    collectionLyrics.insertMany(
        new Document(1).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("A lovestruck Romeo sings the streets a serenade"),
        new Document(2).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Says something like, You and me babe, how about it?"),
        new Document(4).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Juliet says,Hey, it's Romeo, you nearly gimme a heart attack"),
        new Document(5).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("He's underneath the window"),
        new Document(6).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("She's singing, Hey la, my boyfriend's back"),
        new Document(7).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("You shouldn't come around here singing up at people like that"),
        new Document(8).append("band", "Dire Straits").append("song", "Romeo And Juliet").vectorize("Anyway, what you gonna do about it?"));
    System.out.println("6/7 - Collection populated");

Use the $vector method

The $vector method can be used if you generate embeddings before loading data.

Python
# Insert documents into the collection.
# (UUIDs here are version 7.)
documents = [
    {
        "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"),
         "$vectorize": "Chat bot integrated sneakers that talk to you",
    },
    {
        "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"),
         "$vectorize": "An AI quilt to help you sleep forever",
    },
    {
        "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"),
         "$vectorize": "A deep learning display that controls your mood",
    },
]
try:
    insertion_result = collection.insert_many(documents)
    print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n")
except InsertManyException:
    print("* Documents found on DB already. Let's move on.\n")
TypeScript
  // Insert documents into the collection (using UUIDv7s)
  const documents = [
    {
      _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'),
      text: 'ChatGPT integrated sneakers that talk to you',
      $vector: [0.25, 0.25, 0.25, 0.25, 0.45],
    },
    {
      _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'),
      text: 'An AI quilt to help you sleep forever',
      $vector: [0.10, 0.15, 0.25, 0.25, 0.15],
    },
    {
      _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'),
      text: 'A deep learning display that controls your mood',
      $vector: 'I do not want my mood controlled!',
    },
  ];

  try {
    const inserted = await collection.insertMany(documents);
    console.log(`* Inserted ${inserted.insertedCount} items.`);
  } catch (e) {
    console.log('* Documents found on DB already. Let\'s move on!');
  }
Java
    // Insert some documents
    collection.insertMany(
        new Document("1")
                .append("text", "ChatGPT integrated sneakers that talk to you")
                .vector(new float[]{0.1f, 0.15f, 0.3f, 0.12f, 0.05f}),
        new Document("2")
                .append("text", "An AI quilt to help you sleep forever")
                .vector(new float[]{0.45f, 0.09f, 0.01f, 0.2f, 0.11f}),
        new Document("3")
                .append("text", "A deep learning display that controls your mood")
                .vector(new float[]{0.1f, 0.05f, 0.08f, 0.3f, 0.6f}));
    System.out.println("6/7 - Collection populated");

A vector search (similarity search) finds documents that are close to a specific vector embedding.

The output is a list of documents sorted by similarity. For this example, the calculation uses cosine similarity. The similarity metric and dimensions are set when you create a collection.

The following examples show a complete script using an existing keyspace, $vectorize, and a similarity search. Try running your quickstart script on a test database. Then, try modifying the script to insert more documents or run different commands.

Python
import os

from astrapy import DataAPIClient
from astrapy.constants import Environment
from astrapy.authentication import UsernamePasswordTokenProvider
from astrapy.constants import VectorMetric
from astrapy.ids import UUID
from astrapy.exceptions import InsertManyException
from astrapy.info import CollectionVectorServiceOptions

# Database settings
DB_USERNAME = "cassandra"
DB_PASSWORD = "cassandra"
DB_API_ENDPOINT = "http://localhost:8181"
DB_KEYSPACE = "my_keyspace"
DB_COLLECTION = "vector_test"

# Database settings if you exported them as environment variables
# DB_USERNAME = os.environ.get("DB_USERNAME")
# DB_PASSWORD = os.environ.get("DB_PASSWORD")
# DB_API_ENDPOINT = os.environ.get("DB_API_ENDPOINT")

# Embedding provider settings
EMBEDDING_PROVIDER = "openai";
EMBEDDING_MODEL_NAME = "text-embedding-3-small";
EMBEDDING_DIMENSIONS = 1024
EMBEDDING_API_KEY = os.environ.get("EMBEDDING_API_KEY");

# Build a token
tp = UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD)

# Initialize the client and get a "Database" object
client = DataAPIClient(environment=Environment.DSE)
database = client.get_database(DB_API_ENDPOINT, token=tp)
database.get_database_admin().create_keyspace(DB_KEYSPACE, update_db_keyspace=True)

# Create a collection. The default similarity metric is cosine. If you're not
# sure what dimension to set, use whatever dimension vector your embeddings
# model produces.
collection = database.create_collection(
    DB_COLLECTION,
    dimension=EMBEDDING_DIMENSIONS,
    metric=VectorMetric.COSINE,
    service={
        "provider": EMBEDDING_PROVIDER,
        "modelName": EMBEDDING_MODEL_NAME,
    },
    embedding_api_key=EMBEDDING_API_KEY,
    keyspace=DB_KEYSPACE,
    check_exists=False,
)
print(f"* Collection: {collection.full_name}\n")

# Insert documents into the collection.
# (UUIDs here are version 7.)
documents = [
    {
        "_id": UUID("018e65c9-df45-7913-89f8-175f28bd7f74"),
         "$vectorize": "Chat bot integrated sneakers that talk to you",
    },
    {
        "_id": UUID("018e65c9-e1b7-7048-a593-db452be1e4c2"),
         "$vectorize": "An AI quilt to help you sleep forever",
    },
    {
        "_id": UUID("018e65c9-e33d-749b-9386-e848739582f0"),
         "$vectorize": "A deep learning display that controls your mood",
    },
]
try:
    insertion_result = collection.insert_many(documents)
    print(f"* Inserted {len(insertion_result.inserted_ids)} items.\n")
except InsertManyException:
    print("* Documents found on DB already. Let's move on.\n")

# Perform a similarity search
query = [0.15, 0.1, 0.1, 0.35, 0.55]
results = collection.find(
    sort={"$vector": query},
    limit=10,
)
print("Vector search results:")
for document in results:
    print("    ", document)

Run the script with python quickstart.py or the name of your script file.

TypeScript
import { DataAPIClient, UsernamePasswordTokenProvider, VectorDoc, UUID } from '@datastax/astra-db-ts';

// Database settings
const DB_USERNAME = "cassandra";
const DB_PASSWORD = "cassandra";
const DB_API_ENDPOINT = "http://localhost:8181";
const DB_ENVIRONMENT = "dse";
const DB_KEYSPACE = "cycling";

// Database settings if you exported them as environment variables
// const DB_USERNAME = process.env.DB_USERNAME;
// const DB_PASSWORD = process.env.DB_PASSWORD;
// const DB_API_ENDPOINT = process.env.DB_API_ENDPOINT;

// OpenAI settings
const OPEN_AI_PROVIDER = "openai";
const OPENAI_API_KEY = process.env.OPENAI_API_KEY
const MODEL_NAME = "text-embedding-3-small";

// Build a token in the required format
const tp = new UsernamePasswordTokenProvider(DB_USERNAME, DB_PASSWORD);

// Initialize the client and get a "Db" object
const client = new DataAPIClient({ environment: DB_ENVIRONMENT });
const db = client.db(DB_API_ENDPOINT, { token: tp });
const dbAdmin = db.admin({ environment: DB_ENVIRONMENT });

// Schema for the collection (VectorDoc adds the $vector field)
interface Idea extends VectorDoc {
  idea: string,
}

(async function () {
  // Create a typed, vector-enabled collection. The default metric is cosine.
  // If you're not sure what dimension to set, use whatever dimension vector
  // your embeddings model produces.
  const collection = await db.createCollection<Idea>('vector_test', {
    keyspace: DB_KEYSPACE,
    vector: {
      service: {
        provider: OPEN_AI_PROVIDER,
        modelName: MODEL_NAME
      },
      dimension: 5,
      metric: 'cosine',
    },
    embeddingApiKey: OPENAI_API_KEY,
    checkExists: false
  });
  console.log(`* Created collection ${collection.keyspace}.${collection.collectionName}`);

  // Insert documents into the collection (using UUIDv7s)
  const documents = [
    {
      _id: new UUID('018e65c9-df45-7913-89f8-175f28bd7f74'),
      text: 'ChatGPT integrated sneakers that talk to you',
      $vector: [0.25, 0.25, 0.25, 0.25, 0.45],
    },
    {
      _id: new UUID('018e65c9-e1b7-7048-a593-db452be1e4c2'),
      text: 'An AI quilt to help you sleep forever',
      $vector: [0.10, 0.15, 0.25, 0.25, 0.15],
    },
    {
      _id: new UUID('018e65c9-e33d-749b-9386-e848739582f0'),
      text: 'A deep learning display that controls your mood',
      $vector: 'I do not want my mood controlled!',
    },
  ];

  try {
    const inserted = await collection.insertMany(documents);
    console.log(`* Inserted ${inserted.insertedCount} items.`);
  } catch (e) {
    console.log('* Documents found on DB already. Let\'s move on!');
  }

  // Perform a similarity search
  const cursor = await collection.find({}, {
    vector: [0.15, 0.1, 0.1, 0.35, 0.55],
    limit: 10,
    includeSimilarity: true,
  });

  console.log('* Search results:')
  for await (const doc of cursor) {
    console.log('  ', doc.text, doc.$similarity);
  }

  // Cleanup (if desired)
  //await db.dropCollection('vector_test');
  //console.log('* Collection dropped.');

  // Close the client
  await client.close();

})();

Run the script with npm or Yarn and the name of your script file, such as npx tsx quickstart.ts or yarn dlx tsx quickstart.ts.

Java
import com.datastax.astra.client.Collection;
import com.datastax.astra.client.DataAPIClient;
import com.datastax.astra.client.Database;
import com.datastax.astra.client.admin.DataAPIDatabaseAdmin;
import com.datastax.astra.client.model.CollectionOptions;
import com.datastax.astra.client.model.CommandOptions;
import com.datastax.astra.client.model.Document;
import com.datastax.astra.client.model.FindOneOptions;
import com.datastax.astra.client.model.KeyspaceOptions;
import com.datastax.astra.client.model.SimilarityMetric;
import com.datastax.astra.internal.auth.UsernamePasswordTokenProvider;

import java.util.Optional;

import static com.datastax.astra.client.DataAPIClients.DEFAULT_ENDPOINT_LOCAL;
import static com.datastax.astra.client.DataAPIOptions.DataAPIDestination.HCD;
import static com.datastax.astra.client.DataAPIOptions.builder;
import static com.datastax.astra.client.model.Filters.eq;

public class QuickStartDSE69 {

    public static void main(String[] args) {

        // Database Settings
        String cassandraUserName     = "cassandra";
        String cassandraPassword     = "cassandra";
        String dataApiUrl            = DEFAULT_ENDPOINT_LOCAL;  // http://localhost:8181
        String databaseEnvironment   = "DSE" // DSE, HCD, or ASTRA
        String keyspaceName          = "ks1";
        String collectionName        = "lyrics";

        // Database settings if you export them as environment variables
        // String cassandraUserName            = System.getenv("DB_USERNAME");
        // String cassandraPassword            = System.getenv("DB_PASSWORD");
        // String dataApiUrl                   = System.getenv("DB_API_ENDPOINT");

        // OpenAI Embeddings
        String openAiProvider        = "openai";
        String openAiKey             = System.getenv("OPENAI_API_KEY"); // Need to export OPENAI_API_KEY
        String openAiModel           = "text-embedding-3-small";
        int openAiEmbeddingDimension = 1536;

        // Build a token in the form of Cassandra:base64(username):base64(password)
        String token = new UsernamePasswordTokenProvider(cassandraUserName, cassandraPassword).getTokenAsString();
        System.out.println("1/7 - Creating Token: " + token);

        // Initialize the client
        DataAPIClient client = new DataAPIClient(token, builder().withDestination(databaseEnvironment).build());
        System.out.println("2/7 - Connected to Data API");

        // Create a collection
        Collection<Document> collectionLyrics =  db.createCollection(collectionName, CollectionOptions.builder()
        .vectorDimension(5)
        .vectorSimilarity(SimilarityMetric.COSINE)
        .build(),
        System.out.println("5/7 - Collection created");

    // Insert some documents
    collection.insertMany(
        new Document("1")
                .append("text", "ChatGPT integrated sneakers that talk to you")
                .vector(new float[]{0.1f, 0.15f, 0.3f, 0.12f, 0.05f}),
        new Document("2")
                .append("text", "An AI quilt to help you sleep forever")
                .vector(new float[]{0.45f, 0.09f, 0.01f, 0.2f, 0.11f}),
        new Document("3")
                .append("text", "A deep learning display that controls your mood")
                .vector(new float[]{0.1f, 0.05f, 0.08f, 0.3f, 0.6f}));
    System.out.println("6/7 - Collection populated");

FindIterable<Document> resultsSet = collection.find(
    new float[]{0.15f, 0.1f, 0.1f, 0.35f, 0.55f},
    10
);
resultsSet.forEach(System.out::println);
// Cleanup (if desired)
//collection.drop();
//System.out.println("Deleted the collection");

    }
}

Build and run your Java project. For example, with Maven:

mvn clean compile
export OPENAI_API_KEY=<your-api-key>
mvn exec:java -Dexec.mainClass="com.example.QuickStartDSE"

Or with Gradle:

gradle build
gradle run

Data API limits

The following limits apply to Data API usage with DSE.

Entity Limit Notes

Property naming conventions

User-defined property names must:

  • Start and end with a letter or an underscore.

  • Contain only letters, numbers, and underscores.

  • No more than 48 characters in length.

  • Use the reserved property name _id only to set a document’s unique ID.

System-defined operator and property names are prefixed by $, such as $exists, $and, $or, and $vector.

Supported data types

  • String

  • Number

  • Object (JSON object)

  • Array

  • Boolean

  • Vector ($vector)

  • Date ($date)

  • Null

  • UUID ($uuid)

  • ObjectId ($objectId)

See your client’s reference for information about working with dates, UUIDs and ObjectIDs.

Number of collections per keyspace

Five

Up to five collections in a DSE keyspace.

Page size

20

A page may contain up to 20 documents. After that per-page maximum is reached, you can load any additional documents on the next page via the nextPageState generated ID found in a JSON API command’s response.

Sort page size

100

Document page size for sorting; implemented as separate from page size because sort operations need more rows per page.

Maximum property name

100

Maximum of 100 characters in a property name.

Maximum path length

1,000

Maximum of 1,000 characters in a path name; total for all segments, including any dots (.) between properties in a path.

String property maximum bytes

8,000

Maximum of 8,000 UTF-8 bytes for string length in an indexed property.

Number property maximum characters

100

Maximum of 100 characters for number length in a property.

Maximum elements per array

1,000

Maximum number of elements in an array. This limit applies to indexed properties only. This limit is ignored for non-indexed properties.

Maximum dimensions in vector-enabled collection

4,096

Maximum size of dimensions you can define for a vector-enabled collection.

Maximum number of properties per JSON object

1,000

Maximum number of properties for a JSON object. This limit applies to indexed properties only. This limit is ignored for non-indexed properties.

A given JSON object may have nested objects, also known as sub-documents. This maximum total count of 1,000 refers to all the indexed properties in the main document, plus a count of 1 for each sub-document (if any).

Maximum number of properties per JSON document

2,000

Maximum number of properties allowed in a single JSON document is 2,000. This limit includes intermediate properties as well as leaf properties. For example, the document { "root": { "branch": { "leaf": 42 } } } has three properties: root, root.branch, and root.branch.leaf.

Maximum document size in characters

4 million

Maximum size of each document in a collection is 4 million characters.

Maximum inserted batch size in characters

20 million

Maximum size of an entire batch of documents submitted via an insertMany or updateMany command is 20 million characters.

Maximum number of documents deleted per transaction

20

Maximum number of documents that can be deleted in each transaction.

Maximum number of documents updated per transaction

20

Maximum number of documents that can be updated in each transaction.

Maximum number of documents inserted per transaction

20

Maximum number of documents that can be inserted in each transaction when using insertMany.

Maximum size _id values array via $in

100

Maximum size of an _id values array that can be sent via the $in operator.

Maximum number of documents returned with each vector search

1,000

Maximum number of documents returned with each vector search.

If your request is valid but the command exceeds a limit, the Data API responds with HTTP 200 OK and an error message.

It is also possible to receive a response containing both data and errors. Always inspect the response for error messages.

For example, if you exceed the per-transaction limit of 20 documents in an updateMany command, the Data API response contains the following message:

{
  "errors": [
    {
      "message": "Request invalid: field 'command.documents' value \"[...]\" not valid. Problem: amount of documents to update is over the max limit (21 vs 20).",
      "errorCode": "COMMAND_FIELD_INVALID"
    }
  ]
}

Data API operators

Data API supports the following logical and update operators that you can use in filters. For usage examples, see Documents reference.

Operator type Name Purpose

Logical query

$and

Joins query clauses with a logical AND, returning the documents that match the conditions of both clauses.

Logical query

$or

Joins query clauses with a logical OR, returning the documents that match the conditions of either clause.

Logical query

$not

Returns documents that do not match the conditions of the filter clause.

Range query

$gt

Matches documents where the given property is greater than the specified value.

Range query

$gte

Matches documents where the given property is greater than or equal to the specified value.

Range query

$lt

Matches documents where the given property is less than the specified value.

Range query

$lte

Matches documents where the given property is less than or equal to the specified value.

Comparison query

$eq

Matches documents where the value of a property equals the specified value. This is the default when you do not specify an operator.

Comparison query

$ne

Matches documents where the value of a property does not equal the specified value.

Comparison query

$in

Matches any of the values specified in the array.

Comparison query

$nin

Matches any of the values that are NOT IN the array.

Element query

$exists

Matches documents that have the specified property.

Array query

$all

Matches arrays that contain all elements in the specified array.

Array query

$size

Selects documents where the array has the specified number of elements.

Property update

$currentDate

Used in an update operation. In the following example, the createdAt property is updated to use the current date:

{
  "findOneAndUpdate": {
    "filter" : {"_id" : "doc1"},
    "update" : {
      "$currentDate": {
        "createdAt": true
        }
      }
    }
}

Property update

$inc

Increments the value of the property by the specified amount.

Property update

$min

Updates the property only if the specified value is less than the existing property value.

Property update

$max

Updates the property only if the specified value is greater than the existing property value.

Property update

$mul

Multiply the value of a property in the document. For example:

{
    "findOneAndUpdate": {
        "filter": {
            "_id": "upsert-id"
        },
        "update": {
            "$currentDate": {
                "field": true
            },
            "$mul": {
                "min_col": 5.2
            }
        },
        "options": {
            "returnDocument": "after"
        }
    }
}

Property update

$rename

Renames the specified property in each matching document.

Property update

$set

Sets the value of a property in each matching document.

Property update

$setOnInsert

Set the value of a property in the document if an upsert is performed. For example:

{
    "findOneAndUpdate": {
        "filter": {
            "_id": "upsert-id"
        },
        "update": {
            "$currentDate": {
                "field": true
            },
            "$setOnInsert": {
                "customer.name": "James B."
            }
        },
        "options": {
            "returnDocument": "after"
        }
    }
}

Property update

$unset

Removes the specified property from each matching document.

Array update

$addToSet

Adds elements to the array only if they do not already exist in the set.

Array update

$pop

Removes the first or last item of the array, depending on the value of the operator (-1 to remove the first item; 1 to remove the last item).

Array update

$push

Adds or appends data to the end of the property value. Or, if the value is not yet an array:

  • If the property has no value, creates a one-element array (containing the item given).

  • If the property has a non-array value, creates a two-element array, with the old value as the first entry, and the specified item as the second entry.

Array update

$each

An array update that modifies the $push and $addToSet operators to append multiple items for array updates.

Array update

$position

An array update that modifies the $push operator to specify the position in the array to add elements.

Was this helpful?

Give Feedback

How can we improve the documentation?

© Copyright IBM Corporation 2026 | Privacy policy | Terms of use |  Manage Privacy Choices

Apache, Apache Cassandra, Cassandra, Apache Tomcat, Tomcat, Apache Lucene, Apache Solr, Apache Hadoop, Hadoop, Apache Pulsar, Pulsar, Apache Spark, Spark, Apache TinkerPop, TinkerPop, Apache Kafka and Kafka are either registered trademarks or trademarks of the Apache Software Foundation or its subsidiaries in Canada, the United States and/or other countries. Kubernetes is the registered trademark of the Linux Foundation.

General Inquiries: Contact IBM