docs: astra component update (#6720)

* starter-project-update

* update-component-add-vectorize

* update-quickstart

* style-cleanup

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

* split-large-steps-add-admonition

* dimensions-not-required-for-astra-vectorize

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

* fix-numbering

* Apply suggestions from code review

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>

---------

Co-authored-by: KimberlyFields <46325568+KimberlyFields@users.noreply.github.com>
This commit is contained in:
Mendon Kissling 2025-02-24 12:02:03 -05:00 • committed by GitHub
commit c18f65cac9
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
4 changed files with 83 additions and 38 deletions

View file

@ -11,8 +11,8 @@ Get to know Langflow by building an OpenAI-powered chatbot application. After yo
* [An OpenAI API key](https://platform.openai.com/)
* [An Astra DB vector database](https://docs.datastax.com/en/astra-db-serverless/get-started/quickstart.html) with:
* An AstraDB application token
* [A collection in Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection)
* An Astra DB application token scoped to read and write to the database
* A collection created in [Astra](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection) or a new collection created in the **Astra DB** component
## Open Langflow and start a new project
@ -31,7 +31,7 @@ Continue to [Run the basic prompting flow](#run-basic-prompting-flow).
The Basic Prompting flow will look like this when it's completed:
![](/img/starter-flow-basic-prompting.png)
![Completed basic prompting flow](/img/starter-flow-basic-prompting.png)
To build the **Basic Prompting** flow, follow these steps:
@ -46,7 +46,7 @@ The [OpenAI](components-models#openai) model component sends the user input and
You should now have a flow that looks like this:
![](/img/quickstart-basic-prompt-no-connections.png)
![Basic prompting flow with no connections](/img/quickstart-basic-prompt-no-connections.png)
With no connections between them, the components won't interact with each other.
You want data to flow from **Chat Input** to **Chat Output** through the connections between the components.
@ -111,7 +111,7 @@ If you don't want to create a blank flow, click **New Flow**, and then select **
Adding vector RAG to the basic prompting flow will look like this when completed:
![](/img/quickstart-add-document-ingestion.png)
![Add document ingestion to the basic prompting flow](/img/quickstart-add-document-ingestion.png)
To build the flow, follow these steps:
@ -120,24 +120,39 @@ To build the flow, follow these steps:
The [Astra DB vector store](/components-vector-stores#astra-db-vector-store) component connects to your **Astra DB** database.
3. Click **Data**, select the **File** component, and then drag it to the canvas.
The [File](/components-data#file) component loads files from your local machine.
3. Click **Processing**, select the **Split Text** component, and then drag it to the canvas.
4. Click **Processing**, select the **Split Text** component, and then drag it to the canvas.
The [Split Text](/components-processing#split-text) component splits the loaded text into smaller chunks.
4. Click **Processing**, select the **Parse Data** component, and then drag it to the canvas.
5. Click **Processing**, select the **Parse Data** component, and then drag it to the canvas.
The [Data to Message](/components-processing#data-to-message) component converts the data from the **Astra DB** component into plain text.
5. Click **Embeddings**, select the **OpenAI Embeddings** component, and then drag it to the canvas.
6. Click **Embeddings**, select the **OpenAI Embeddings** component, and then drag it to the canvas.
The [OpenAI Embeddings](/components-embedding-models#openai-embeddings) component generates embeddings for the user's input, which are compared to the vector data in the database.
6. Connect the new components into the existing flow, so your flow looks like this:
7. Connect the new components into the existing flow, so your flow looks like this:
![](/img/quickstart-add-document-ingestion.png)
![Add document ingestion to the basic prompting flow](/img/quickstart-add-document-ingestion.png)
8. Configure the **Astra DB** component.
1. In the **Astra DB Application Token** field, add your **Astra DB** application token.
The component connects to your database and populates the menus with existing databases and collections.
2. Select your **Database**.
If you don't have a collection, select **New database**.
Complete the **Name**, **Cloud provider**, and **Region** fields, and then click **Create**. **Database creation takes a few minutes**.
3. Select your **Collection**. Collections are created in your [Astra DB deployment](https://astra.datastax.com) for storing vector data.
If you don't have a collection, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/manage-collections.html#create-collection).
4. Select **Embedding Model** to bring your own embeddings model, which is the connected **OpenAI Embeddings** component.
The **Dimensions** value must match the dimensions of your collection. This value can be found in your **Collection** in your [Astra DB deployment](https://astra.datastax.com).
:::info
If you select a collection embedded with NVIDIA through Astra's vectorize service, the **Embedding Model** port is removed, because you have already generated embeddings for this collection with the NVIDIA `NV-Embed-QA` model. The component fetches the data from the collection, and uses the same embeddings for queries.
:::
9. If you don't have a collection, create a new one within the component.
1. Select **New collection**.
2. Complete the **Name**, **Embedding generation method**, **Embedding model**, and **Dimensions** fields, and then click **Create**.
Your choice for the **Embedding generation method** and **Embedding model** depends on whether you want to use embeddings generated by a provider through Astra's vectorize service, or generated by a component in Langflow.
* To use embeddings generated by a provider through Astra's vectorize service, select the model from the **Embedding generation method** dropdown menu, and then select the model from the **Embedding model** dropdown menu.
* To use embeddings generated by a component in Langflow, select **Bring your own** for both the **Embedding generation method** and **Embedding model** fields. In this starter project, the option for the embeddings method and model is the **OpenAI Embeddings** component connected to the **Astra DB** component.
* The **Dimensions** value must match the dimensions of your collection. This field is **not required** if you use embeddings generated through Astra's vectorize service. You can find this value in the **Collection** in your [Astra DB deployment](https://astra.datastax.com).
For more information, see the [DataStax Astra DB Serverless documentation](https://docs.datastax.com/en/astra-db-serverless/databases/embedding-generation.html).
If you used Langflow's **Global Variables** feature, the RAG application flow components are already configured with the necessary credentials.