feat: truncate parsed uploads to prevent database and frontend blocking caused by excessively large files (#3914)

* 📝 (constants.ts): increase maxSizeFilesInBytes constant value from 10MB to 100MB to allow larger file uploads

* 🐛 (inputFileComponent): fix bug in setting the maximum file size alert message to display the correct file size limit of 100 bytes instead of 10 bytes

* 📝 (schemas.py): Add a new field_serializer method to serialize data in VertexBuildResponse class
📝 (schemas.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries
📝 (model.py): Add a new field_serializer method to serialize outputs in TransactionBase class
📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries
📝 (model.py): Add a new field_serializer method to serialize data and artifacts in VertexBuildBase class
📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries

* 🐛 (schemas.py): fix truncation length of text fields to 10 characters instead of 99999
🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999
🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999
🐛 (index.tsx): truncate resultMessage to 99999 characters and add message if text is too long

* 🔧 (switchOutputView/index.tsx): Use useMemo to memoize resultMessage transformations for performance optimization

* 🐛 (model.py): Fix typo in the path for 'base_retriever' data field
🐛 (model.py): Fix typo in the path for 'base_retriever' data field
🐛 (model.py): Fix typo in the path for 'base_retriever' data field
🐛 (model.py): Fix typo in the path for 'base_retriever' data field
🐛 (index.tsx): Fix logic to correctly handle resultMessageMemoized when it is an object

* 📝 (model.py): refactor truncate_text function to truncate_long_strings for better clarity and consistency
📝 (model.py): update serialize_outputs and serialize_artifacts functions to use truncate_long_strings for string truncation
📝 (model.py): introduce MAX_TEXT_LENGTH constant for defining the maximum length of text to truncate in the models

* 📝 (schemas.py): refactor serialize_data method in VertexBuildResponse class to use a new helper function truncate_long_strings for better code readability and maintainability

* 🔧 (schemas.py): Move the `truncate_long_strings` function to a separate module to improve code organization and reusability
🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error
🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error

* 📝 (util.py): add function truncate_long_strings to recursively truncate long strings in dictionaries and lists to prevent exceeding the maximum text length.

* 📝 (constants.py): add constant MAX_TEXT_LENGTH with value 99999 for defining maximum text length allowed in the application

* 📝 (model.py): update import path for truncate_long_strings function to match new location in util module

* ✨ (test_truncate_long_strings_on_objects.py): Add unit tests for the function truncate_long_strings to ensure correct behavior when truncating long strings in various data structures
🐛 (switchOutputView/index.tsx): Fix truncation logic to correctly truncate long strings by adding ellipsis at the end instead of displaying additional text about truncation.

* [autofix.ci] apply automated fixes

* ✨ (test_truncate_long_strings_on_objects.py): Update import path for truncate_long_strings function
📝 (test_truncate_long_strings_on_objects.py): Add additional tests for handling negative, zero, and small max_length values in truncate_long_strings function

* ♻️ (schemas.py): refactor import statement to use the updated module name util_strings instead of util for better clarity and consistency.

* 📝 (model.py): Update import path for util_strings module to fix module import error
📝 (util.py): Remove redundant code for truncating long strings and move it to a separate util_strings module for better organization and separation of concerns.

* 📝 (schemas.py): refactor serialize_data method to handle both BaseModel and non-BaseModel data inputs in VertexBuildResponse class

* 📝 (util_strings.py): Update util_strings.py to improve string truncation function for dictionaries and lists
🔧 (test_truncate_long_strings_on_objects.py): Update test cases for string truncation function to cover additional scenarios and edge cases

* Update src/backend/base/langflow/utils/util_strings.py

Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org>

* 📝 (vite.config.mts): update environment variable MAX_FILE_SIZE to be defined in vite config for frontend to use in the application.

* 📝 (constants.ts): update maxSizeFilesInBytes constant to use process.env.MAX_FILE_SIZE environment variable for configurable file size limit
📝 (constants.ts): add MAX_TEXT_LENGTH constant with a value of 99999 for maximum text length limit

* 📝 (switchOutputView/index.tsx): import MAX_TEXT_LENGTH constant from shared constants file to improve code organization and reusability

* ✨ (langflow/__main__.py): add support for defining maximum file size for upload in MB to improve user experience and prevent large file uploads

* 🐛 (files.py): add validation to check if uploaded file size exceeds the maximum allowed size before processing it

* ✨ (schemas.py): add max_file_size_upload field to ConfigResponse schema to handle maximum file size allowed for upload

* 🔧 (vite.config.mts): remove MAX_FILE_SIZE environment variable configuration as it is no longer needed

* ✨ (base.py): introduce max_file_size_upload setting to limit the file size for uploads in MB

* 🐛 (util.py): add support for setting max_file_size_upload in update_settings function to allow configuring maximum file size for uploads

* 📝 (inputFileComponent/index.tsx): add support for retrieving max file size upload from utility store to improve code modularity and reusability
🐛 (inputFileComponent/index.tsx): fix error handling logic to display error message when uploading a file fails

* 📝 (constants.ts): remove maxSizeFilesInBytes constant as it is no longer used and update MAX_TEXT_LENGTH constant to a higher value

* ✨ (use-get-config.ts): add functionality to set max file size upload value from config response

* ✨ (utilityStore.ts): introduce maxFileSizeUpload property and setMaxFileSizeUpload function to handle maximum file size upload in bytes

* ✨ (frontend): introduce maxFileSizeUpload property and setMaxFileSizeUpload method to handle maximum file size upload functionality in the UtilityStoreType

* ♻️ (util_strings.py): refactor truncate_long_strings function to improve code readability and consistency by removing unnecessary whitespace and aligning assignment operators.

* 🐛 (files.py): fix formatting issue in the raise statement to improve code readability and maintain consistency

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org>
This commit is contained in:
Cristhian Zanforlin Lousa 2024-09-27 12:44:05 -03:00 • committed by GitHub
commit 948b150946
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
21 changed files with 414 additions and 21 deletions

View file

@ -138,6 +138,11 @@ def run(
help="Defines the number of retries for the health check.",
envvar="LANGFLOW_HEALTH_CHECK_MAX_RETRIES",
),
max_file_size_upload: int = typer.Option(
100,
help="Defines the maximum file size for the upload in MB.",
envvar="LANGFLOW_MAX_FILE_SIZE_UPLOAD",
),
):
"""
Run Langflow.
@ -158,6 +163,7 @@ def run(
auto_saving=auto_saving,
auto_saving_interval=auto_saving_interval,
health_check_max_retries=health_check_max_retries,
max_file_size_upload=max_file_size_upload,
)
# create path object if path is provided
static_files_dir: Path | None = Path(path) if path else None

View file

@ -43,6 +43,12 @@ async def upload_file(
storage_service: StorageService = Depends(get_storage_service),
):
try:
max_file_size_upload = get_storage_service().settings_service.settings.max_file_size_upload
if file.size > max_file_size_upload * 1024 * 1024:
raise HTTPException(
status_code=413, detail=f"File size is larger than the maximum file size {max_file_size_upload}MB."
)
flow_id_str = str(flow_id)
file_content = await file.read()
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")

View file

@ -16,6 +16,7 @@ from langflow.services.database.models.base import orjson_dumps
from langflow.services.database.models.flow import FlowCreate, FlowRead
from langflow.services.database.models.user import UserRead
from langflow.services.tracing.schema import Log
from langflow.utils.util_strings import truncate_long_strings
class BuildStatus(Enum):
@ -281,6 +282,12 @@ class VertexBuildResponse(BaseModel):
timestamp: datetime | None = Field(default_factory=lambda: datetime.now(timezone.utc))
"""Timestamp of the build."""
@field_serializer("data")
def serialize_data(self, data: ResultDataResponse) -> dict:
data_dict = data.model_dump() if isinstance(data, BaseModel) else data
truncated_data = truncate_long_strings(data_dict)
return truncated_data
class VerticesBuiltResponse(BaseModel):
vertices: list[VertexBuildResponse]
@ -341,3 +348,4 @@ class ConfigResponse(BaseModel):
auto_saving: bool
auto_saving_interval: int
health_check_max_retries: int
max_file_size_upload: int

View file

@ -2,12 +2,14 @@ from datetime import datetime, timezone
from typing import TYPE_CHECKING
from uuid import UUID, uuid4
from pydantic import field_validator
from pydantic import field_serializer, field_validator
from sqlmodel import JSON, Column, Field, Relationship, SQLModel
if TYPE_CHECKING:
from langflow.services.database.models.flow.model import Flow
from langflow.utils.util_strings import truncate_long_strings
class TransactionBase(SQLModel):
timestamp: datetime = Field(default_factory=lambda: datetime.now(timezone.utc))
@ -32,6 +34,11 @@ class TransactionBase(SQLModel):
value = UUID(value)
return value
@field_serializer("outputs")
def serialize_outputs(self, data) -> dict:
truncated_data = truncate_long_strings(data)
return truncated_data
class TransactionTable(TransactionBase, table=True): # type: ignore
__tablename__ = "transaction"

View file

@ -8,6 +8,8 @@ from sqlmodel import JSON, Column, Field, Relationship, SQLModel
if TYPE_CHECKING:
from langflow.services.database.models.flow.model import Flow
from langflow.utils.util_strings import truncate_long_strings
class VertexBuildBase(SQLModel):
timestamp: datetime = Field(default_factory=lambda: datetime.now(timezone.utc))
@ -38,6 +40,16 @@ class VertexBuildBase(SQLModel):
value = value.replace(tzinfo=timezone.utc)
return value
@field_serializer("data")
def serialize_data(self, data: dict) -> dict:
truncated_data = truncate_long_strings(data)
return truncated_data
@field_serializer("artifacts")
def serialize_artifacts(self, data) -> dict:
truncated_data = truncate_long_strings(data)
return truncated_data
class VertexBuildTable(VertexBuildBase, table=True): # type: ignore
__tablename__ = "vertex_build"

View file

@ -153,6 +153,8 @@ class Settings(BaseSettings):
"""The interval in ms at which Langflow will auto save flows."""
health_check_max_retries: int = 5
"""The maximum number of retries for the health check."""
max_file_size_upload: int = 100
"""The maximum file size for the upload in MB."""
@field_validator("dev")
@classmethod

View file

@ -183,3 +183,5 @@ MESSAGE_SENDER_AI = "Machine"
MESSAGE_SENDER_USER = "User"
MESSAGE_SENDER_NAME_AI = "AI"
MESSAGE_SENDER_NAME_USER = "User"
MAX_TEXT_LENGTH = 99999

View file

@ -431,6 +431,7 @@ def update_settings(
auto_saving: bool = True,
auto_saving_interval: int = 1000,
health_check_max_retries: int = 5,
max_file_size_upload: int = 100,
):
"""Update the settings from a config file."""
from langflow.services.utils import initialize_settings_service
@ -463,6 +464,9 @@ def update_settings(
if health_check_max_retries is not None:
logger.debug(f"Setting health_check_max_retries to {health_check_max_retries}")
settings_service.settings.update_settings(health_check_max_retries=health_check_max_retries)
if max_file_size_upload is not None:
logger.debug(f"Setting max_file_size_upload to {max_file_size_upload}")
settings_service.settings.update_settings(max_file_size_upload=max_file_size_upload)
def is_class_method(func, cls):

View file

@ -0,0 +1,28 @@
from langflow.utils import constants
def truncate_long_strings(data, max_length=None):
"""
Recursively traverse the dictionary or list and truncate strings longer than max_length.
"""
if max_length is None:
max_length = constants.MAX_TEXT_LENGTH
if max_length < 0 or not isinstance(data, dict | list):
return data
if isinstance(data, dict):
for key, value in data.items():
if isinstance(value, str) and len(value) > max_length:
data[key] = value[:max_length] + "..."
elif isinstance(value, (dict | list)):
truncate_long_strings(value, max_length)
elif isinstance(data, list):
for index, item in enumerate(data):
if isinstance(item, str) and len(item) > max_length:
data[index] = item[:max_length] + "..."
elif isinstance(item, (dict | list)):
truncate_long_strings(item, max_length)
return data

View file

@ -0,0 +1,98 @@
from langflow.utils.util_strings import truncate_long_strings
from langflow.utils.constants import MAX_TEXT_LENGTH
import pytest
@pytest.mark.parametrize(
"input_data, max_length, expected",
[
# Test case 1: Simple string truncation
({"key": "a" * 100}, 10, {"key": "a" * 10 + "..."}),
# Test case 2: Nested dictionary
({"outer": {"inner": "b" * 100}}, 5, {"outer": {"inner": "b" * 5 + "..."}}),
# Test case 3: List of strings
(["short", "a" * 100, "also short"], 7, ["short", "a" * 7 + "...", "also sh" + "..."]),
# Test case 4: Mixed nested structure
(
{"key1": ["a" * 100, {"nested": "b" * 100}], "key2": "c" * 100},
8,
{"key1": ["a" * 8 + "...", {"nested": "b" * 8 + "..."}], "key2": "c" * 8 + "..."},
),
# Test case 5: Empty structures
({}, 10, {}),
([], 10, []),
# Test case 6: Strings at exact max_length
({"exact": "a" * 10}, 10, {"exact": "a" * 10}),
# Test case 7: Non-string values
({"num": 12345, "bool": True, "none": None}, 5, {"num": 12345, "bool": True, "none": None}),
# Test case 8: Unicode characters
({"unicode": "こんにちは世界"}, 3, {"unicode": "こんに..."}),
# Test case 9: Very large structure
(
{"key" + str(i): "value" * i for i in range(1000)},
10,
{"key" + str(i): ("value" * i)[:10] + "..." if len("value" * i) > 10 else "value" * i for i in range(1000)},
),
],
)
def test_truncate_long_strings(input_data, max_length, expected):
result = truncate_long_strings(input_data, max_length)
assert result == expected
def test_truncate_long_strings_default_max_length():
long_string = "a" * (MAX_TEXT_LENGTH + 1)
input_data = {"key": long_string}
result = truncate_long_strings(input_data)
assert len(result["key"]) == MAX_TEXT_LENGTH + 3 # +3 for the "..."
def test_truncate_long_strings_no_modification():
input_data = {"short": "short string", "nested": {"also_short": "another short string"}}
result = truncate_long_strings(input_data, 100)
assert result == input_data
# Test for type preservation
def test_truncate_long_strings_type_preservation():
input_data = {"str": "a" * 100, "list": ["b" * 100], "dict": {"nested": "c" * 100}}
result = truncate_long_strings(input_data, 10)
assert isinstance(result, dict)
assert isinstance(result["str"], str)
assert isinstance(result["list"], list)
assert isinstance(result["dict"], dict)
# Test for in-place modification
def test_truncate_long_strings_in_place_modification():
input_data = {"key": "a" * 100}
result = truncate_long_strings(input_data, 10)
assert result is input_data # Check if the same object is returned
# Test for invalid input
def test_truncate_long_strings_invalid_input():
input_string = "not a dict or list"
result = truncate_long_strings(input_string, 10)
assert result == input_string # The function should return the input unchanged
# Updated test for negative max_length
def test_truncate_long_strings_negative_max_length():
input_data = {"key": "value"}
result = truncate_long_strings(input_data, -1)
assert result == input_data # Assuming the function ignores negative max_length
# Additional test for zero max_length
def test_truncate_long_strings_zero_max_length():
input_data = {"key": "value"}
result = truncate_long_strings(input_data, 0)
assert result == {"key": "..."} # Assuming the function truncates to just "..."
# Test for very small positive max_length
def test_truncate_long_strings_small_max_length():
input_data = {"key": "value"}
result = truncate_long_strings(input_data, 1)
assert result == {"key": "v..."} # Assuming the function keeps at least one character