Integrate Glean with Astra DB Serverless

With the Glean platform integration, you can use information from Astra DB Serverless databases as a data source for your Glean searches.

To do this, you use the Data API to push data from Astra DB through the Glean Indexing API to a custom data source in Glean.

Use this integration if you need to query your Astra DB data alongside your other Glean data sources. This is recommended for use cases where your Astra DB data contains human-readable information that you want users to find in Glean search results. This integration is best for non-vector data, such as non-vector CSV data, that you cannot ingest through other Glean data sources.

For a complete script example, see the mini-demo-astradb-glean GitHub repository.

Prerequisites

This integration requires the following:

Create a Python project

This tutorial script uses the Data API Python client. If you use a different Data API client or HTTP, you must modify the example project and script accordingly.

  1. Prepare a virtual environment:

    python3 -m venv my_virtual_env
  2. Activate the virtual environment and install the required packages:

    source my_virtual_env/bin/activate
    # on Windows run: my_virtual_env\Scripts\activate
    
    pip install \
        "astrapy>=2.0,<3.0" \
        "https://app.glean.com/meta/indexing_api_client.zip"

Set environment variables

  1. Create a .env file in a Python project directory.

  2. Set Astra DB environment variables:

    APPLICATION_TOKEN=APPLICATION_TOKEN
    API_ENDPOINT=API_ENDPOINT
    ASTRA_DB_COLLECTION_NAME=COLLECTION_NAME
    KEYSPACE_NAME=KEYSPACE_NAME

    Replace the following:

    • APPLICATION_TOKEN: An Astra application token with the Database Administrator role.

    • API_ENDPOINT: Your database’s API endpoint.

    • COLLECTION_NAME: The name of a collection in your database where you will store the data you want to push to Glean. This can be an existing collection or a collection that the script will create at runtime.

      This tutorial uses a vector-enabled collection in a Serverless (vector) database. Both vector and non-vector data can be saved in vector-enabled collections.

    • KEYSPACE_NAME: The keyspace where your collection is (or will be) located. If unspecified, the script uses the Data API default value of default_keyspace.

  3. Set Glean environment variables:

    GLEAN_CUSTOMER=GLEAN_CUSTOMER_NAME
    GLEAN_DATASOURCE_NAME=GLEAN_DATASOURCE_NAME
    GLEAN_API_TOKEN=GLEAN_INDEXING_API_TOKEN

    Replace the following:

    • GLEAN_CUSTOMER_NAME: Your Glean customer name from your Glean API endpoint. The endpoint format is https://GLEAN_CUSTOMER_NAME-be.glean.com/api/index/v1.

    • GLEAN_DATASOURCE_NAME: The name of a custom data source in Glean for indexing your Astra DB collection data. This can be an existing data source or a data source that the script will create at runtime.

    • GLEAN_INDEXING_API_TOKEN: A Glean Indexing API token.

Create a Glean indexing script

  1. Create a Python file in your Python project directory.

    In the next steps, you will add code to this file to create your Glean indexing script. To follow along with this tutorial, name the file astra-glean-import-job.py.

  2. Import dependencies:

    import os
    import requests
    
    from astrapy import DataAPIClient
    from colorama import Fore, Style
    from dotenv import load_dotenv
    
    import glean_indexing_api_client as indexing_api
    from glean_indexing_api_client.api import datasources_api, documents_api
    from glean_indexing_api_client.model.custom_datasource_config import (
        CustomDatasourceConfig,
    )
    from glean_indexing_api_client.model.object_definition import ObjectDefinition
    from glean_indexing_api_client.model.index_document_request import IndexDocumentRequest
    from glean_indexing_api_client.model.document_definition import DocumentDefinition
    from glean_indexing_api_client.model.content_definition import ContentDefinition
    from glean_indexing_api_client.model.document_permissions_definition import (
        DocumentPermissionsDefinition,
    )
  3. Load environment variables:

    # Load environment variables from .env
    load_dotenv()
    
    APPLICATION_TOKEN = os.environ["APPLICATION_TOKEN"]
    API_ENDPOINT = os.environ["API_ENDPOINT"]
    ASTRA_DB_COLLECTION_NAME = os.environ["ASTRA_DB_COLLECTION_NAME"]
    KEYSPACE_NAME = os.getenv("KEYSPACE_NAME")
    
    GLEAN_API_TOKEN = os.environ["GLEAN_API_TOKEN"]
    GLEAN_CUSTOMER = os.environ["GLEAN_CUSTOMER"]
    GLEAN_DATASOURCE_NAME = os.environ["GLEAN_DATASOURCE_NAME"]
    
    
    print(f"{Fore.GREEN}============================={Style.RESET_ALL}")
    print(f"{Fore.GREEN} ASTRADB - GLEAN INTEGRATION {Style.RESET_ALL}")
    print(f"{Fore.GREEN}============================={Style.RESET_ALL}\n")
  4. Initialize the Data API client and connect to your database:

    # Initialize Astra DB client
    client = DataAPIClient(callers=[("glean", "1.0")])
    database = client.get_database(
        API_ENDPOINT,
        token=APPLICATION_TOKEN,
        keyspace=KEYSPACE_NAME,
    )
    print(
        f"{Fore.CYAN}[ OK ] - Credentials are OK, your database name is "
        f"{Style.RESET_ALL}{database.name()}{Fore.CYAN}."
    )
  5. Create a collection in your database:

    # Create collection
    source_collection = database.create_collection(ASTRA_DB_COLLECTION_NAME)
    print(
        f"{Fore.CYAN}[ OK ] - Collection {Style.RESET_ALL}{source_collection.name}"
        f"{Fore.CYAN} is ready{Style.RESET_ALL}{Fore.CYAN}."
    )

    Or, use an existing collection:

    # Get an existing collection
    source_collection = database.get_collection(ASTRA_DB_COLLECTION_NAME)
    print(
        f"{Fore.CYAN}[ OK ] - Collection {Style.RESET_ALL}{source_collection.name}"
        f"{Fore.CYAN} is ready{Style.RESET_ALL}{Fore.CYAN}."
    )

    To use get_collection, the collection must have the exact name set in ASTRA_DB_COLLECTION_NAME, and it must exist in the specified KEYSPACE_NAME or the default keyspace. If the collection doesn’t exist in your database, create the collection before running the script.

  6. Load a sample dataset into your collection.

    If you already loaded data into your collection, don’t include this code in your script.

    Glean can index vector and non-vector data. If your collection contains vector data, be aware that Glean does not differentiate embeddings from other data. In the same way that Glean processes your other data sources, Glean indexes all collection data, including embeddings, as text data.

    # Load philosophers dataset
    print(f"{Fore.CYAN}[INFO] - Downloading data.{Style.RESET_ALL}")
    philo_dataset = requests.get(
        "https://raw.githubusercontent.com/"
        "datastaxdevs/mini-datasets/refs/heads/main/datasets/"
        "philosopher-quotes.json"
    ).json()
    print(f"{Fore.CYAN}[ OK ] - Dataset loaded in memory.{Style.RESET_ALL}")
    print(f"{Fore.CYAN}[INFO] - Sample record: {Style.RESET_ALL}{philo_dataset[16]}")
    
    
    def load_to_astra_db(data_to_insert, collection):
        """Load all of the provided data into a collection."""
        def split_tags(t):
            return [tag for tag in (t or "").split(";") if tag]
    
        documents_to_insert = [
            {
                **item,
                **{"_id": index, "tags": split_tags(item["tags"])},
            }
            for index, item in enumerate(data_to_insert)
        ]
        collection.insert_many(documents_to_insert)
    
    # Insert documents into Astra DB
    philo_count = len(philo_dataset)
    print(
        f"{Fore.CYAN}[INFO] - Inserting {philo_count} documents into Astra DB..."
        f"{Style.RESET_ALL}"
    )
    load_to_astra_db(philo_dataset, source_collection)
    print(f"{Fore.CYAN}[ OK ] - Insertion finished.{Style.RESET_ALL}")
  7. Initialize the Glean API client:

    # Setup Glean API
    GLEAN_API_ENDPOINT = f"https://{GLEAN_CUSTOMER}-be.glean.com/api/index/v1"
    print(
        f"{Fore.CYAN}[INFO] - Glean API setup, endpoint is:"
        f"{Style.RESET_ALL} {GLEAN_API_ENDPOINT}"
    )
    
    # Initialize Glean client
    configuration = indexing_api.Configuration(
        host=GLEAN_API_ENDPOINT, access_token=GLEAN_API_TOKEN
    )
    api_client = indexing_api.ApiClient(configuration)
    datasource_api = datasources_api.DatasourcesApi(api_client)
    print(f"{Fore.CYAN}[ OK ] - Glean client initialized{Style.RESET_ALL}")
  8. Create and register a Glean data source for your Astra DB database.

    Use the following code to create the data source at runtime. This code uses the Glean Indexing API /adddatasource endpoint.

    Omit this code if you want to use an existing data source or create a data source before running the script. Your data source name must match the GLEAN_DATASOURCE_NAME environment variable.

    # Create and register data source in Glean
    datasource_config = CustomDatasourceConfig(
        name=GLEAN_DATASOURCE_NAME,
        display_name="Astra DB Collection Data Source",
        datasource_category="PUBLISHED_CONTENT",
        url_regex=f"^{API_ENDPOINT}",
        object_definitions=[
            ObjectDefinition(doc_category="PUBLISHED_CONTENT", name="AstraVectorEntry")
        ],
    )
    
    try:
        datasource_api.adddatasource_post(datasource_config)
        print(
            f"{Fore.GREEN}[ OK ] - Data source has been created!"
            f"{Style.RESET_ALL}{Fore.GREEN}."
        )
    except indexing_api.ApiException as e:
        print(
            f"{Fore.RED}[ ERROR ] - Error creating data source: "
            f"{e}{Style.RESET_ALL}{Fore.GREEN}."
        )
  9. Index documents from your collection into Glean:

    def index_astra_db_document_into_glean(astra_document):
        """Index one Astra DB document into Glean."""
        document_id = str(astra_document["_id"])
        title = f"{astra_document['author']} quote_{astra_document['_id']}"
        body_text = astra_document["quote"]
        datasource_name = GLEAN_DATASOURCE_NAME
        request = IndexDocumentRequest(
            document=DocumentDefinition(
                datasource=datasource_name,
                title=title,
                id=document_id,
                view_url=API_ENDPOINT,
                body=ContentDefinition(mime_type="text/plain", text_content=body_text),
                permissions=DocumentPermissionsDefinition(allow_anonymous_access=True),
            )
        )
        documents_api_client = documents_api.DocumentsApi(api_client)
        try:
            documents_api_client.indexdocument_post(request)
        except indexing_api.ApiException as e:
            print(f"{Fore.RED}Error indexing document {document_id}: {e}{Style.RESET_ALL}")
    
    
    def index_documents_to_glean(collection):
        """Index all documents from an Astra DB collection to Glean."""
        total_docs = collection.count_documents({}, upper_bound=1000)
        print(
            f"{Fore.CYAN}[INFO] - Indexing {total_docs} "
            f"documents into Glean...{Style.RESET_ALL}"
        )
        for doc in collection.find():
            try:
                index_astra_db_document_into_glean(doc)
            except Exception as error:
                print(
                    f"{Fore.RED}Error indexing document "
                    f"{doc['_id']}: {error}{Style.RESET_ALL}"
                )
        print(f"{Fore.CYAN}[ OK ] - Indexing finished.{Style.RESET_ALL}")
    
    
    # Use the function to index documents into Glean
    index_documents_to_glean(source_collection)
    
    print(f"{Fore.GREEN}Import job completed successfully!{Style.RESET_ALL}")

Test the integration

  1. Run the script to test the integration:

    python3 astra-glean-import-job.py
  2. Try searching Glean for the content indexed from your collection.

    If the results reflect your collection data, the integration succeeded.

    If there are no results or no evidence of the expected response, check the following:

    • The script ran without error.

    • The data source category is accurate in the data source configuration.

    • Indexing is complete, and the data source is populated in Glean. For more information, see Debugging Glean data sources.

    • Your Astra DB collection contains only text data. Glean can index text data only. If your database contains only non-text data types, the content might not be fully supported in your Glean searches.

Next steps

Try these next steps to expand this integration:

  • Use a cron job to automatically run the script at regular intervals.

  • Extend the existing script or create additional scripts to pull from other databases and collections.

Was this helpful?

Give Feedback

How can we improve the documentation?

© Copyright IBM Corporation 2026 | Privacy policy | Terms of use |  Manage Privacy Choices

Apache, Apache Cassandra, Cassandra, Apache Tomcat, Tomcat, Apache Lucene, Apache Solr, Apache Hadoop, Hadoop, Apache Pulsar, Pulsar, Apache Spark, Spark, Apache TinkerPop, TinkerPop, Apache Kafka and Kafka are either registered trademarks or trademarks of the Apache Software Foundation or its subsidiaries in Canada, the United States and/or other countries. Kubernetes is the registered trademark of the Linux Foundation.

General Inquiries: Contact IBM