feat: Add ParseDataFrameComponent for DataFrame-to-text conversion with tests (#5594)

* add dataframe outputs to vector stores, directory, url, split text

* [autofix.ci] apply automated fixes

* [autofix.ci] apply automated fixes (attempt 2/3)

* add parse dataframe

* [autofix.ci] apply automated fixes

* Refactor: Update DataFrame handling in components

- Added import of DataFrame in directory and url components.
- Renamed variable 'df' to 'dataframe' in ParseDataFrameComponent for clarity.
- Updated method _clean_args and parse_data to use 'dataframe' instead of 'df' for consistency.

These changes enhance code readability and maintainability by standardizing the terminology used for DataFrame objects.

* [autofix.ci] apply automated fixes

* remove parse dataframe

* feat: add parse dataframe component

* [autofix.ci] apply automated fixes

* Refactor: Remove duplicate as_dataframe method in LCVectorStoreComponent

This commit eliminates the redundant as_dataframe method in the LCVectorStoreComponent class, streamlining the code and improving maintainability. The method was previously defined twice, and this change enhances clarity by ensuring only one implementation exists.

* [autofix.ci] apply automated fixes

* Refactor: Standardize DataFrame variable naming in ParseDataFrameComponent

This commit renames the variable 'df' to 'dataframe' in the ParseDataFrameComponent class to improve clarity and consistency. The changes are reflected in the _clean_args and parse_data methods, enhancing code readability and maintainability.

* test: add unit tests for ParseDataFrameComponent

This commit introduces a comprehensive suite of unit tests for the ParseDataFrameComponent, covering various scenarios including successful parsing with default and custom templates, handling of empty dataframes, invalid template keys, and performance on large dataframes. The tests ensure that the component behaves correctly with different data types and separators, and validate its functionality in both synchronous and asynchronous contexts. These additions enhance the reliability and maintainability of the component.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org>
Co-authored-by: Edwin Jose <edwin.jose@datastax.com>
This commit is contained in:
Rodrigo Nader 2025-01-17 20:41:42 -03:00 • committed by GitHub
commit 8d902e6c74
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 210 additions and 0 deletions

View file

@ -24,6 +24,7 @@ __all__ = [
"MergeDataComponent",
"MessageToDataComponent",
"ParseDataComponent",
"ParseDataFrameComponent",
"ParseJSONDataComponent",
"SelectDataComponent",
"SplitTextComponent",

View file

@ -0,0 +1,67 @@
from langflow.custom import Component
from langflow.io import DataFrameInput, MultilineInput, Output, StrInput
from langflow.schema.message import Message
class ParseDataFrameComponent(Component):
display_name = "Parse DataFrame"
description = (
"Convert a DataFrame into plain text following a specified template. "
"Each column in the DataFrame is treated as a possible template key, e.g. {col_name}."
)
icon = "braces"
name = "ParseDataFrame"
inputs = [
DataFrameInput(name="df", display_name="DataFrame", info="The DataFrame to convert to text rows."),
MultilineInput(
name="template",
display_name="Template",
info=(
"The template for formatting each row. "
"Use placeholders matching column names in the DataFrame, for example '{col1}', '{col2}'."
),
value="{text}",
),
StrInput(
name="sep",
display_name="Separator",
advanced=True,
value="\n",
info="String that joins all row texts when building the single Text output.",
),
]
outputs = [
Output(
display_name="Text",
name="text",
info="All rows combined into a single text, each row formatted by the template and separated by `sep`.",
method="parse_data",
),
]
def _clean_args(self):
dataframe = self.df
template = self.template or "{text}"
sep = self.sep or "\n"
return dataframe, template, sep
def parse_data(self) -> Message:
"""Converts each row of the DataFrame into a formatted string using the template.
then joins them with `sep`. Returns a single combined string as a Message.
"""
dataframe, template, sep = self._clean_args()
lines = []
# For each row in the DataFrame, build a dict and format
for _, row in dataframe.iterrows():
row_dict = row.to_dict()
text_line = template.format(**row_dict) # e.g. template="{text}", row_dict={"text": "Hello"}
lines.append(text_line)
# Join all lines with the provided separator
result_string = sep.join(lines)
self.status = result_string # store in self.status for UI logs
return Message(text=result_string)