Integrate Griptape with Astra DB Serverless

Griptape can use the vector capabilities of Astra DB Serverless with the dedicated Astra DB Vector Store Driver.

The following tutorial creates a Python script to integrate with Griptape.

Prerequisites

This guide requires the following:

Connect to your database

Import libraries and connect to the database:

  1. Create a .env file in your Python project directory, and then set the following environment variables:

    APPLICATION_TOKEN="APPLICATION_TOKEN"
    API_ENDPOINT="API_ENDPOINT"
    KEYSPACE_NAME="default_keyspace"
    GRIPTAPE_COLLECTION_NAME="griptape_integration"
    OPENAI_API_KEY="API_KEY"

    Replace the placeholders with the credentials from the Prerequisites.

    KEYSPACE_NAME must be set to the keyspace associated with your griptape_integration collection. The default keyspace for collections in Serverless (vector) databases is default_keyspace.

  2. Create a Python file for your integration script.

    To avoid a namespace collision, don’t name the file griptape.py. To follow along with this tutorial in a local script, name the file integrate.py.

  3. In your Python file, import dependencies:

    import os
    from dotenv import load_dotenv
    
    from griptape.drivers import (
        AstraDbVectorStoreDriver,
        OpenAiChatPromptDriver,
        OpenAiEmbeddingDriver,
    )
    from griptape.engines.rag import RagEngine
    from griptape.engines.rag.modules import (
        PromptResponseRagModule,
        VectorStoreRetrievalRagModule,
    )
    from griptape.engines.rag.stages import ResponseRagStage, RetrievalRagStage
    from griptape.loaders import WebLoader
    from griptape.structures import Agent
    from griptape.tools import RagTool
  4. Load the environment variables:

    load_dotenv()
    
    APPLICATION_TOKEN = os.environ["APPLICATION_TOKEN"]
    API_ENDPOINT = os.environ["API_ENDPOINT"]
    KEYSPACE_NAME = os.environ.get("KEYSPACE_NAME")
    GRIPTAPE_COLLECTION_NAME = os.environ["GRIPTAPE_COLLECTION_NAME"]

Initialize the vector store and RAG engine

  1. Initialize the vector store driver and pass it to the RAG Engine, which is a Griptape component that drives RAG pipelines.

    When working with the Griptape astradb_vector_store_driver, the Griptape namespace is a label for entries in a vector store (within an Astra DB collection). The astra_db_namespace attribute is your Astra DB keyspace.

    namespace = "datastax_blog"
    
    vector_store_driver = AstraDbVectorStoreDriver(
        embedding_driver=OpenAiEmbeddingDriver(),
        api_endpoint=API_ENDPOINT,
        token=APPLICATION_TOKEN,
        collection_name=GRIPTAPE_COLLECTION_NAME,
        astra_db_namespace=KEYSPACE_NAME,
    )
    
    engine = RagEngine(
        retrieval_stage=RetrievalRagStage(
            retrieval_modules=[
                VectorStoreRetrievalRagModule(
                    vector_store_driver=vector_store_driver,
                    query_params={
                        "count": 2,
                        "namespace": namespace,
                    },
                )
            ]
        ),
        response_stage=ResponseRagStage(
            response_modules=[
                PromptResponseRagModule(
                    prompt_driver=OpenAiChatPromptDriver(model="gpt-4o"),
                ),
            ],
        ),
    )
  2. Ingest a web page into the vector store:

    input_blogpost = (
        "www.datastax.com/blog/indexing-all-of-wikipedia-on-a-laptop"
    )
    
    vector_store_driver.upsert_text_artifacts(
        {namespace: WebLoader(max_tokens=256).load(input_blogpost)}
    )
  3. Wrap the RAG Engine in a RAGClient, and then pass it to a Griptape agent as a tool:

    rag_tool = RagTool(
        description="A DataStax blog post",
        rag_engine=engine,
    )
    agent = Agent(tools=[rag_tool])
  4. Run a RAG-powered question-and-answer process based on the ingested content:

    agent.run(
        "what engine did DataStax develop to index such an amount of data on a "
        "laptop? Please summarize its main features."
    )
    
    answer = agent.output_task.output.value
    
    print(answer)

Run the code

Run the script to test the integration:

python integrate.py

The following sample response is truncated for clarity:

[08/21/24 00:47:28] INFO     ToolkitTask 09ca0fb83cd24f2590155aab415af651
                             Input: what engine did DataStax develop to index such an amount of data on a laptop?
                             Please summarize its main features.
[08/21/24 00:47:29] INFO     Subtask bec12f6a6d72467bbf2ab059f5f5a59a
                             Actions: [
                               {
                                 "tag": "call_JuuW9d9xQxf6M5DRcGDGNhqW",
                                 "name": "RagTool",
                                 "path": "search",
                                 "input": {
                                   "values": {
                                     "query": "DataStax engine to index large amounts of data on a laptop"
                                   }
                                 }
                               }
                             ]
[08/21/24 00:47:31] INFO     Subtask bec12f6a6d72467bbf2ab059f5f5a59a
                             Response: DataStax Astra DB uses the [...]
[08/21/24 00:47:33] INFO     ToolkitTask 09ca0fb83cd24f2590155aab415af651
                             Output: DataStax developed the [...]

DataStax developed the **JVector library** to index large amounts of data
on a laptop. Here are its main features:

1. **Support for Larger-than-Memory Datasets**: JVector can handle datasets that
exceed the available memory by using compressed vectors.
2. **Efficient Construction-Related Searches**: It performs searches efficiently
even with compressed data.
3. **Memory Optimization**: The edge lists fit in memory, while the uncompressed
vectors do not, optimizing the use of available memory resources.

This approach makes it feasible to index large datasets, such as Wikipedia,
on a laptop.

Was this helpful?

Give Feedback

How can we improve the documentation?

© Copyright IBM Corporation 2026 | Privacy policy | Terms of use |  Manage Privacy Choices

Apache, Apache Cassandra, Cassandra, Apache Tomcat, Tomcat, Apache Lucene, Apache Solr, Apache Hadoop, Hadoop, Apache Pulsar, Pulsar, Apache Spark, Spark, Apache TinkerPop, TinkerPop, Apache Kafka and Kafka are either registered trademarks or trademarks of the Apache Software Foundation or its subsidiaries in Canada, the United States and/or other countries. Kubernetes is the registered trademark of the Linux Foundation.

General Inquiries: Contact IBM