docs: astra component update (#6720)

* starter-project-update

* update-component-add-vectorize

* update-quickstart

* style-cleanup

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

* split-large-steps-add-admonition

* dimensions-not-required-for-astra-vectorize

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

* fix-numbering

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

---------

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>
This commit is contained in:
Mendon Kissling 2025-02-24 12:02:03 -05:00 • committed by GitHub
commit c18f65cac9
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
4 changed files with 83 additions and 38 deletions

View file

@ -37,31 +37,46 @@ For more information, see the [DataStax documentation](https://docs.datastax.com
| Name | Display Name | Info | | Name | Display Name | Info |
|------|--------------|------| |------|--------------|------|
| collection_name | Collection Name | The name of the collection within Astra DB where the vectors will be stored (required) | | token | Astra DB Application Token | The authentication token for accessing Astra DB. |
| token | Astra DB Application Token | Authentication token for accessing Astra DB (required) | | environment | Environment | The environment for the Astra DB API Endpoint. For example, `dev` or `prod`. |
| api_endpoint | API Endpoint | API endpoint URL for the Astra DB service (required) | | database_name | Database | The database name for the Astra DB instance. |
| search_input | Search Input | Query string for similarity search | | api_endpoint | Astra DB API Endpoint | The API endpoint for the Astra DB instance. This supersedes the database selection. |
| ingest_data | Ingest Data | Data to be ingested into the vector store | | collection_name | Collection | The name of the collection within Astra DB where the vectors are stored. |
| namespace | Namespace | Optional namespace within Astra DB to use for the collection | | keyspace | Keyspace | An optional keyspace within Astra DB to use for the collection. |
| embedding_choice | Embedding Model or Astra Vectorize | Determines whether to use an Embedding Model or Astra Vectorize for the collection | | embedding_choice | Embedding Model or Astra Vectorize | Choose an embedding model or use Astra vectorize. |
| embedding | Embedding Model | Allows an embedding model configuration (when using Embedding Model) | | embedding_model | Embedding Model | Specify the embedding model. Not required for Astra vectorize collections. |
| provider | Vectorize Provider | Provider for Astra Vectorize (when using Astra Vectorize) | | number_of_results | Number of Search Results | The number of search results to return (default: `4`). |
| metric | Metric | Optional distance metric for vector comparisons | | search_type | Search Type | The search type to use. The options are `Similarity`, `Similarity with score threshold`, and `MMR (Max Marginal Relevance)`. |
| batch_size | Batch Size | Optional number of data to process in a single batch | | search_score_threshold | Search Score Threshold | The minimum similarity score threshold for search results when using the `Similarity with score threshold` option. |
| setup_mode | Setup Mode | Configuration mode for setting up the vector store (options: "Sync", "Async", "Off", default: "Sync") | | advanced_search_filter | Search Metadata Filter | An optional dictionary of filters to apply to the search query. |
| pre_delete_collection | Pre Delete Collection | Boolean flag to determine whether to delete the collection before creating a new one | | autodetect_collection | Autodetect Collection | A boolean flag to determine whether to autodetect the collection. |
| number_of_results | Number of Results | Number of results to return in similarity search (default: 4) | | content_field | Content Field | A field to use as the text content field for the vector store. |
| search_type | Search Type | Search type to use (options: "Similarity", "Similarity with score threshold", "MMR (Max Marginal Relevance)") | | deletion_field | Deletion Based On Field | When provided, documents in the target collection with metadata field values matching the input metadata field value are deleted before new data is loaded. |
| search_score_threshold | Search Score Threshold | Minimum similarity score threshold for search results | | ignore_invalid_documents | Ignore Invalid Documents | A boolean flag to determine whether to ignore invalid documents at runtime. |
| search_filter | Search Metadata Filter | Optional dictionary of filters to apply to the search query | | astradb_vectorstore_kwargs | AstraDBVectorStore Parameters | An optional dictionary of additional parameters for the AstraDBVectorStore. |
### Outputs ### Outputs
| Name | Display Name | Info | | Name | Display Name | Info |
|------|--------------|------| |------|--------------|------|
| vector_store | Vector Store | Astra DB vector store instance configured with the specified parameters. | | vector_store | Vector Store | Astra DB vector store instance configured with the specified parameters. |
| search_results | Search Results | The results of the similarity search as a list of `Data` objects. | | search_results | Search Results | The results of the similarity search as a list of [Data](/concepts-objects#data-object) objects. |
### Generate embeddings
The **Astra DB Vector Store** component offers two methods for generating embeddings.
1. **Embedding Model**: Use your own embedding model by connecting an [Embeddings](/components-embedding-models) component in Langflow.
2. **Astra Vectorize**: Use Astra DB's built-in embedding generation service. When creating a new collection, choose the embeddings provider and models, including NVIDIA's `NV-Embed-QA` model hosted by Datastax.
:::important
The embedding model selection is made when creating a new collection and cannot be changed later.
:::
For an example of using the **Astra DB Vector Store** component with an embedding model, see the [Vector Store RAG starter project](/starter-projects-vector-store-rag).
For more information, see the [Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/embedding-generation.html).
## AstraDB Graph vector store ## AstraDB Graph vector store

View file

@ -11,8 +11,8 @@ Get to know Langflow by building an OpenAI-powered chatbot application. After yo
* [An OpenAI API key](https://platform.openai.com/) * [An OpenAI API key](https://platform.openai.com/)
* [An Astra DB vector database](https://docs.datastax.com/en/astra-db-serverless/get-started/quickstart.html) with: * [An Astra DB vector database](https://docs.datastax.com/en/astra-db-serverless/get-started/quickstart.html) with:
* An AstraDB application token * An Astra DB application token scoped to read and write to the database
* [A collection in Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection) * A collection created in [Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection) or a new collection created in the **Astra DB** component
## Open Langflow and start a new project ## Open Langflow and start a new project
@ -31,7 +31,7 @@ Continue to [Run the basic prompting flow](#run-basic-prompting-flow).
The Basic Prompting flow will look like this when it's completed: The Basic Prompting flow will look like this when it's completed:
![](/img/starter-flow-basic-prompting.png) ![Completed basic prompting flow](/img/starter-flow-basic-prompting.png)
To build the **Basic Prompting** flow, follow these steps: To build the **Basic Prompting** flow, follow these steps:
@ -46,7 +46,7 @@ The [OpenAI](components-models#openai) model component sends the user input and
You should now have a flow that looks like this: You should now have a flow that looks like this:
![](/img/quickstart-basic-prompt-no-connections.png) ![Basic prompting flow with no connections](/img/quickstart-basic-prompt-no-connections.png)
With no connections between them, the components won't interact with each other. With no connections between them, the components won't interact with each other.
You want data to flow from **Chat Input** to **Chat Output** through the connections between the components. You want data to flow from **Chat Input** to **Chat Output** through the connections between the components.
@ -111,7 +111,7 @@ If you don't want to create a blank flow, click **New Flow**, and then select **
Adding vector RAG to the basic prompting flow will look like this when completed: Adding vector RAG to the basic prompting flow will look like this when completed:
![](/img/quickstart-add-document-ingestion.png) ![Add document ingestion to the basic prompting flow](/img/quickstart-add-document-ingestion.png)
To build the flow, follow these steps: To build the flow, follow these steps:
@ -120,24 +120,39 @@ To build the flow, follow these steps:
The [Astra DB vector store](/components-vector-stores#astra-db-vector-store) component connects to your **Astra DB** database. The [Astra DB vector store](/components-vector-stores#astra-db-vector-store) component connects to your **Astra DB** database.
3. Click **Data**, select the **File** component, and then drag it to the canvas. 3. Click **Data**, select the **File** component, and then drag it to the canvas.
The [File](/components-data#file) component loads files from your local machine. The [File](/components-data#file) component loads files from your local machine.
3. Click **Processing**, select the **Split Text** component, and then drag it to the canvas. 4. Click **Processing**, select the **Split Text** component, and then drag it to the canvas.
The [Split Text](/components-processing#split-text) component splits the loaded text into smaller chunks. The [Split Text](/components-processing#split-text) component splits the loaded text into smaller chunks.
4. Click **Processing**, select the **Parse Data** component, and then drag it to the canvas. 5. Click **Processing**, select the **Parse Data** component, and then drag it to the canvas.
The [Data to Message](/components-processing#data-to-message) component converts the data from the **Astra DB** component into plain text. The [Data to Message](/components-processing#data-to-message) component converts the data from the **Astra DB** component into plain text.
5. Click **Embeddings**, select the **OpenAI Embeddings** component, and then drag it to the canvas. 6. Click **Embeddings**, select the **OpenAI Embeddings** component, and then drag it to the canvas.
The [OpenAI Embeddings](/components-embedding-models#openai-embeddings) component generates embeddings for the user's input, which are compared to the vector data in the database. The [OpenAI Embeddings](/components-embedding-models#openai-embeddings) component generates embeddings for the user's input, which are compared to the vector data in the database.
6. Connect the new components into the existing flow, so your flow looks like this: 7. Connect the new components into the existing flow, so your flow looks like this:
![](/img/quickstart-add-document-ingestion.png) ![Add document ingestion to the basic prompting flow](/img/quickstart-add-document-ingestion.png)
8. Configure the **Astra DB** component. 8. Configure the **Astra DB** component.
1. In the **Astra DB Application Token** field, add your **Astra DB** application token. 1. In the **Astra DB Application Token** field, add your **Astra DB** application token.
The component connects to your database and populates the menus with existing databases and collections. The component connects to your database and populates the menus with existing databases and collections.
2. Select your **Database**. 2. Select your **Database**.
If you don't have a collection, select **New database**.
Complete the **Name**, **Cloud provider**, and **Region** fields, and then click **Create**. **Database creation takes a few minutes**.
3. Select your **Collection**. Collections are created in your [Astra DB deployment](https://astra.datastax.com) for storing vector data. 3. Select your **Collection**. Collections are created in your [Astra DB deployment](https://astra.datastax.com) for storing vector data.
If you don't have a collection, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection). :::info
4. Select **Embedding Model** to bring your own embeddings model, which is the connected **OpenAI Embeddings** component. If you select a collection embedded with NVIDIA through Astra's vectorize service, the **Embedding Model** port is removed, because you have already generated embeddings for this collection with the NVIDIA `NV-Embed-QA` model. The component fetches the data from the collection, and uses the same embeddings for queries.
The **Dimensions** value must match the dimensions of your collection. This value can be found in your **Collection** in your [Astra DB deployment](https://astra.datastax.com). :::
9. If you don't have a collection, create a new one within the component.
1. Select **New collection**.
2. Complete the **Name**, **Embedding generation method**, **Embedding model**, and **Dimensions** fields, and then click **Create**.
Your choice for the **Embedding generation method** and **Embedding model** depends on whether you want to use embeddings generated by a provider through Astra's vectorize service, or generated by a component in Langflow.
* To use embeddings generated by a provider through Astra's vectorize service, select the model from the **Embedding generation method** dropdown menu, and then select the model from the **Embedding model** dropdown menu.
* To use embeddings generated by a component in Langflow, select **Bring your own** for both the **Embedding generation method** and **Embedding model** fields. In this starter project, the option for the embeddings method and model is the **OpenAI Embeddings** component connected to the **Astra DB** component.
* The **Dimensions** value must match the dimensions of your collection. This field is **not required** if you use embeddings generated through Astra's vectorize service. You can find this value in the **Collection** in your [Astra DB deployment](https://astra.datastax.com).
For more information, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/embedding-generation.html).
If you used Langflow's **Global Variables** feature, the RAG application flow components are already configured with the necessary credentials. If you used Langflow's **Global Variables** feature, the RAG application flow components are already configured with the necessary credentials.

View file

@ -22,7 +22,7 @@ This opens a starter flow with the necessary components to run an agentic applic
## Simple Agent flow ## Simple Agent flow
<img src="/img/starter-flow-simple-agent.png" alt="Starter flow simple agent" width="75%"/> ![Simple agent starter flow](/img/starter-flow-simple-agent.png)
The **Simple Agent** flow consists of these components: The **Simple Agent** flow consists of these components:

View file

@ -20,9 +20,9 @@ We've chosen [Astra DB](https://astra.datastax.com/signup?utm_source=langflow-p
## Prerequisites ## Prerequisites
* [An OpenAI API key](https://platform.openai.com/) * [An OpenAI API key](https://platform.openai.com/)
* [An Astra DB vector database](https://docs.datastax.com/en/astra-db-serverless/get-started/quickstart.html) with: * [An Astra DB vector database](https://docs.datastax.com/en/astra-db-serverless/get-started/quickstart.html) with the following:
* An Astra DB application token * An Astra DB application token scoped to read and write to the database
* [A collection in Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection) * A collection created in [Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection) or a new collection created in the **Astra DB** component
## Open Langflow and start a new project ## Open Langflow and start a new project
@ -60,10 +60,25 @@ The **Retriever Flow** (top of the screen) embeds the user's queries into vecto
1. In the **Astra DB Application Token** field, add your **Astra DB** application token. 1. In the **Astra DB Application Token** field, add your **Astra DB** application token.
The component connects to your database and populates the menus with existing databases and collections. The component connects to your database and populates the menus with existing databases and collections.
2. Select your **Database**. 2. Select your **Database**.
If you don't have a collection, select **New database**.
Complete the **Name**, **Cloud provider**, and **Region** fields, and then click **Create**. **Database creation takes a few minutes**.
3. Select your **Collection**. Collections are created in your [Astra DB deployment](https://astra.datastax.com) for storing vector data. 3. Select your **Collection**. Collections are created in your [Astra DB deployment](https://astra.datastax.com) for storing vector data.
If you don't have a collection, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection). :::info
4. Select **Embedding Model** to bring your own embeddings model, which is the connected **OpenAI Embeddings** component. If you select a collection embedded with Nvidia through Astra's vectorize service, the **Embedding Model** port is removed, because you have already generated embeddings for this collection with the Nvidia `NV-Embed-QA` model. The component fetches the data from the collection, and uses the same embeddings for queries.
The **Dimensions** value must match the dimensions of your collection. You can find this value in the **Collection** in your [Astra DB deployment](https://astra.datastax.com). :::
3. If you don't have a collection, create a new one within the component.
1. Select **New collection**.
2. Complete the **Name**, **Embedding generation method**, **Embedding model**, and **Dimensions** fields, and then click **Create**.
Your choice for the **Embedding generation method** and **Embedding model** depends on whether you want to use embeddings generated by a provider through Astra's vectorize service, or generated by a component in Langflow.
* To use embeddings generated by a provider through Astra's vectorize service, select the model from the **Embedding generation method** dropdown menu, and then select the model from the **Embedding model** dropdown menu.
* To use embeddings generated by a component in Langflow, select **Bring your own** for both the **Embedding generation method** and **Embedding model** fields. In this starter project, the option for the embeddings method and model is the **OpenAI Embeddings** component connected to the **Astra DB** component.
* The **Dimensions** value must match the dimensions of your collection. This field is **not required** if you use embeddings generated through Astra's vectorize service. You can find this value in the **Collection** in your [Astra DB deployment](https://astra.datastax.com).
For more information, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/embedding-generation.html).
If you used Langflow's **Global Variables** feature, the RAG application flow components are already configured with the necessary credentials. If you used Langflow's **Global Variables** feature, the RAG application flow components are already configured with the necessary credentials.