feat: truncate parsed uploads to prevent database and frontend blocking caused by excessively large files (#3914)
* 📝 (constants.ts): increase maxSizeFilesInBytes constant value from 10MB to 100MB to allow larger file uploads * 🐛 (inputFileComponent): fix bug in setting the maximum file size alert message to display the correct file size limit of 100 bytes instead of 10 bytes * 📝 (schemas.py): Add a new field_serializer method to serialize data in VertexBuildResponse class 📝 (schemas.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries 📝 (model.py): Add a new field_serializer method to serialize outputs in TransactionBase class 📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries 📝 (model.py): Add a new field_serializer method to serialize data and artifacts in VertexBuildBase class 📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries * 🐛 (schemas.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (index.tsx): truncate resultMessage to 99999 characters and add message if text is too long * 🔧 (switchOutputView/index.tsx): Use useMemo to memoize resultMessage transformations for performance optimization * 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (index.tsx): Fix logic to correctly handle resultMessageMemoized when it is an object * 📝 (model.py): refactor truncate_text function to truncate_long_strings for better clarity and consistency 📝 (model.py): update serialize_outputs and serialize_artifacts functions to use truncate_long_strings for string truncation 📝 (model.py): introduce MAX_TEXT_LENGTH constant for defining the maximum length of text to truncate in the models * 📝 (schemas.py): refactor serialize_data method in VertexBuildResponse class to use a new helper function truncate_long_strings for better code readability and maintainability * 🔧 (schemas.py): Move the `truncate_long_strings` function to a separate module to improve code organization and reusability 🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error 🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error * 📝 (util.py): add function truncate_long_strings to recursively truncate long strings in dictionaries and lists to prevent exceeding the maximum text length. * 📝 (constants.py): add constant MAX_TEXT_LENGTH with value 99999 for defining maximum text length allowed in the application * 📝 (model.py): update import path for truncate_long_strings function to match new location in util module * ✨ (test_truncate_long_strings_on_objects.py): Add unit tests for the function truncate_long_strings to ensure correct behavior when truncating long strings in various data structures 🐛 (switchOutputView/index.tsx): Fix truncation logic to correctly truncate long strings by adding ellipsis at the end instead of displaying additional text about truncation. * [autofix.ci] apply automated fixes * ✨ (test_truncate_long_strings_on_objects.py): Update import path for truncate_long_strings function 📝 (test_truncate_long_strings_on_objects.py): Add additional tests for handling negative, zero, and small max_length values in truncate_long_strings function * ♻️ (schemas.py): refactor import statement to use the updated module name util_strings instead of util for better clarity and consistency. * 📝 (model.py): Update import path for util_strings module to fix module import error 📝 (util.py): Remove redundant code for truncating long strings and move it to a separate util_strings module for better organization and separation of concerns. * 📝 (schemas.py): refactor serialize_data method to handle both BaseModel and non-BaseModel data inputs in VertexBuildResponse class * 📝 (util_strings.py): Update util_strings.py to improve string truncation function for dictionaries and lists 🔧 (test_truncate_long_strings_on_objects.py): Update test cases for string truncation function to cover additional scenarios and edge cases * Update src/backend/base/langflow/utils/util_strings.py Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org> * 📝 (vite.config.mts): update environment variable MAX_FILE_SIZE to be defined in vite config for frontend to use in the application. * 📝 (constants.ts): update maxSizeFilesInBytes constant to use process.env.MAX_FILE_SIZE environment variable for configurable file size limit 📝 (constants.ts): add MAX_TEXT_LENGTH constant with a value of 99999 for maximum text length limit * 📝 (switchOutputView/index.tsx): import MAX_TEXT_LENGTH constant from shared constants file to improve code organization and reusability * ✨ (langflow/__main__.py): add support for defining maximum file size for upload in MB to improve user experience and prevent large file uploads * 🐛 (files.py): add validation to check if uploaded file size exceeds the maximum allowed size before processing it * ✨ (schemas.py): add max_file_size_upload field to ConfigResponse schema to handle maximum file size allowed for upload * 🔧 (vite.config.mts): remove MAX_FILE_SIZE environment variable configuration as it is no longer needed * ✨ (base.py): introduce max_file_size_upload setting to limit the file size for uploads in MB * 🐛 (util.py): add support for setting max_file_size_upload in update_settings function to allow configuring maximum file size for uploads * 📝 (inputFileComponent/index.tsx): add support for retrieving max file size upload from utility store to improve code modularity and reusability 🐛 (inputFileComponent/index.tsx): fix error handling logic to display error message when uploading a file fails * 📝 (constants.ts): remove maxSizeFilesInBytes constant as it is no longer used and update MAX_TEXT_LENGTH constant to a higher value * ✨ (use-get-config.ts): add functionality to set max file size upload value from config response * ✨ (utilityStore.ts): introduce maxFileSizeUpload property and setMaxFileSizeUpload function to handle maximum file size upload in bytes * ✨ (frontend): introduce maxFileSizeUpload property and setMaxFileSizeUpload method to handle maximum file size upload functionality in the UtilityStoreType * ♻️ (util_strings.py): refactor truncate_long_strings function to improve code readability and consistency by removing unnecessary whitespace and aligning assignment operators. * 🐛 (files.py): fix formatting issue in the raise statement to improve code readability and maintain consistency --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com> Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org>
This commit is contained in:
parent
b34a7c7f02
commit
948b150946
21 changed files with 414 additions and 21 deletions
|
|
@ -138,6 +138,11 @@ def run(
|
|||
help="Defines the number of retries for the health check.",
|
||||
envvar="LANGFLOW_HEALTH_CHECK_MAX_RETRIES",
|
||||
),
|
||||
max_file_size_upload: int = typer.Option(
|
||||
100,
|
||||
help="Defines the maximum file size for the upload in MB.",
|
||||
envvar="LANGFLOW_MAX_FILE_SIZE_UPLOAD",
|
||||
),
|
||||
):
|
||||
"""
|
||||
Run Langflow.
|
||||
|
|
@ -158,6 +163,7 @@ def run(
|
|||
auto_saving=auto_saving,
|
||||
auto_saving_interval=auto_saving_interval,
|
||||
health_check_max_retries=health_check_max_retries,
|
||||
max_file_size_upload=max_file_size_upload,
|
||||
)
|
||||
# create path object if path is provided
|
||||
static_files_dir: Path | None = Path(path) if path else None
|
||||
|
|
|
|||
|
|
@ -43,6 +43,12 @@ async def upload_file(
|
|||
storage_service: StorageService = Depends(get_storage_service),
|
||||
):
|
||||
try:
|
||||
max_file_size_upload = get_storage_service().settings_service.settings.max_file_size_upload
|
||||
if file.size > max_file_size_upload * 1024 * 1024:
|
||||
raise HTTPException(
|
||||
status_code=413, detail=f"File size is larger than the maximum file size {max_file_size_upload}MB."
|
||||
)
|
||||
|
||||
flow_id_str = str(flow_id)
|
||||
file_content = await file.read()
|
||||
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
|
||||
|
|
|
|||
|
|
@ -16,6 +16,7 @@ from langflow.services.database.models.base import orjson_dumps
|
|||
from langflow.services.database.models.flow import FlowCreate, FlowRead
|
||||
from langflow.services.database.models.user import UserRead
|
||||
from langflow.services.tracing.schema import Log
|
||||
from langflow.utils.util_strings import truncate_long_strings
|
||||
|
||||
|
||||
class BuildStatus(Enum):
|
||||
|
|
@ -281,6 +282,12 @@ class VertexBuildResponse(BaseModel):
|
|||
timestamp: datetime | None = Field(default_factory=lambda: datetime.now(timezone.utc))
|
||||
"""Timestamp of the build."""
|
||||
|
||||
@field_serializer("data")
|
||||
def serialize_data(self, data: ResultDataResponse) -> dict:
|
||||
data_dict = data.model_dump() if isinstance(data, BaseModel) else data
|
||||
truncated_data = truncate_long_strings(data_dict)
|
||||
return truncated_data
|
||||
|
||||
|
||||
class VerticesBuiltResponse(BaseModel):
|
||||
vertices: list[VertexBuildResponse]
|
||||
|
|
@ -341,3 +348,4 @@ class ConfigResponse(BaseModel):
|
|||
auto_saving: bool
|
||||
auto_saving_interval: int
|
||||
health_check_max_retries: int
|
||||
max_file_size_upload: int
|
||||
|
|
|
|||
|
|
@ -2,12 +2,14 @@ from datetime import datetime, timezone
|
|||
from typing import TYPE_CHECKING
|
||||
from uuid import UUID, uuid4
|
||||
|
||||
from pydantic import field_validator
|
||||
from pydantic import field_serializer, field_validator
|
||||
from sqlmodel import JSON, Column, Field, Relationship, SQLModel
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from langflow.services.database.models.flow.model import Flow
|
||||
|
||||
from langflow.utils.util_strings import truncate_long_strings
|
||||
|
||||
|
||||
class TransactionBase(SQLModel):
|
||||
timestamp: datetime = Field(default_factory=lambda: datetime.now(timezone.utc))
|
||||
|
|
@ -32,6 +34,11 @@ class TransactionBase(SQLModel):
|
|||
value = UUID(value)
|
||||
return value
|
||||
|
||||
@field_serializer("outputs")
|
||||
def serialize_outputs(self, data) -> dict:
|
||||
truncated_data = truncate_long_strings(data)
|
||||
return truncated_data
|
||||
|
||||
|
||||
class TransactionTable(TransactionBase, table=True): # type: ignore
|
||||
__tablename__ = "transaction"
|
||||
|
|
|
|||
|
|
@ -8,6 +8,8 @@ from sqlmodel import JSON, Column, Field, Relationship, SQLModel
|
|||
if TYPE_CHECKING:
|
||||
from langflow.services.database.models.flow.model import Flow
|
||||
|
||||
from langflow.utils.util_strings import truncate_long_strings
|
||||
|
||||
|
||||
class VertexBuildBase(SQLModel):
|
||||
timestamp: datetime = Field(default_factory=lambda: datetime.now(timezone.utc))
|
||||
|
|
@ -38,6 +40,16 @@ class VertexBuildBase(SQLModel):
|
|||
value = value.replace(tzinfo=timezone.utc)
|
||||
return value
|
||||
|
||||
@field_serializer("data")
|
||||
def serialize_data(self, data: dict) -> dict:
|
||||
truncated_data = truncate_long_strings(data)
|
||||
return truncated_data
|
||||
|
||||
@field_serializer("artifacts")
|
||||
def serialize_artifacts(self, data) -> dict:
|
||||
truncated_data = truncate_long_strings(data)
|
||||
return truncated_data
|
||||
|
||||
|
||||
class VertexBuildTable(VertexBuildBase, table=True): # type: ignore
|
||||
__tablename__ = "vertex_build"
|
||||
|
|
|
|||
|
|
@ -153,6 +153,8 @@ class Settings(BaseSettings):
|
|||
"""The interval in ms at which Langflow will auto save flows."""
|
||||
health_check_max_retries: int = 5
|
||||
"""The maximum number of retries for the health check."""
|
||||
max_file_size_upload: int = 100
|
||||
"""The maximum file size for the upload in MB."""
|
||||
|
||||
@field_validator("dev")
|
||||
@classmethod
|
||||
|
|
|
|||
|
|
@ -183,3 +183,5 @@ MESSAGE_SENDER_AI = "Machine"
|
|||
MESSAGE_SENDER_USER = "User"
|
||||
MESSAGE_SENDER_NAME_AI = "AI"
|
||||
MESSAGE_SENDER_NAME_USER = "User"
|
||||
|
||||
MAX_TEXT_LENGTH = 99999
|
||||
|
|
|
|||
|
|
@ -431,6 +431,7 @@ def update_settings(
|
|||
auto_saving: bool = True,
|
||||
auto_saving_interval: int = 1000,
|
||||
health_check_max_retries: int = 5,
|
||||
max_file_size_upload: int = 100,
|
||||
):
|
||||
"""Update the settings from a config file."""
|
||||
from langflow.services.utils import initialize_settings_service
|
||||
|
|
@ -463,6 +464,9 @@ def update_settings(
|
|||
if health_check_max_retries is not None:
|
||||
logger.debug(f"Setting health_check_max_retries to {health_check_max_retries}")
|
||||
settings_service.settings.update_settings(health_check_max_retries=health_check_max_retries)
|
||||
if max_file_size_upload is not None:
|
||||
logger.debug(f"Setting max_file_size_upload to {max_file_size_upload}")
|
||||
settings_service.settings.update_settings(max_file_size_upload=max_file_size_upload)
|
||||
|
||||
|
||||
def is_class_method(func, cls):
|
||||
|
|
|
|||
28
src/backend/base/langflow/utils/util_strings.py
Normal file
28
src/backend/base/langflow/utils/util_strings.py
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
from langflow.utils import constants
|
||||
|
||||
|
||||
def truncate_long_strings(data, max_length=None):
|
||||
"""
|
||||
Recursively traverse the dictionary or list and truncate strings longer than max_length.
|
||||
"""
|
||||
|
||||
if max_length is None:
|
||||
max_length = constants.MAX_TEXT_LENGTH
|
||||
|
||||
if max_length < 0 or not isinstance(data, dict | list):
|
||||
return data
|
||||
|
||||
if isinstance(data, dict):
|
||||
for key, value in data.items():
|
||||
if isinstance(value, str) and len(value) > max_length:
|
||||
data[key] = value[:max_length] + "..."
|
||||
elif isinstance(value, (dict | list)):
|
||||
truncate_long_strings(value, max_length)
|
||||
elif isinstance(data, list):
|
||||
for index, item in enumerate(data):
|
||||
if isinstance(item, str) and len(item) > max_length:
|
||||
data[index] = item[:max_length] + "..."
|
||||
elif isinstance(item, (dict | list)):
|
||||
truncate_long_strings(item, max_length)
|
||||
|
||||
return data
|
||||
|
|
@ -0,0 +1,98 @@
|
|||
from langflow.utils.util_strings import truncate_long_strings
|
||||
from langflow.utils.constants import MAX_TEXT_LENGTH
|
||||
import pytest
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"input_data, max_length, expected",
|
||||
[
|
||||
# Test case 1: Simple string truncation
|
||||
({"key": "a" * 100}, 10, {"key": "a" * 10 + "..."}),
|
||||
# Test case 2: Nested dictionary
|
||||
({"outer": {"inner": "b" * 100}}, 5, {"outer": {"inner": "b" * 5 + "..."}}),
|
||||
# Test case 3: List of strings
|
||||
(["short", "a" * 100, "also short"], 7, ["short", "a" * 7 + "...", "also sh" + "..."]),
|
||||
# Test case 4: Mixed nested structure
|
||||
(
|
||||
{"key1": ["a" * 100, {"nested": "b" * 100}], "key2": "c" * 100},
|
||||
8,
|
||||
{"key1": ["a" * 8 + "...", {"nested": "b" * 8 + "..."}], "key2": "c" * 8 + "..."},
|
||||
),
|
||||
# Test case 5: Empty structures
|
||||
({}, 10, {}),
|
||||
([], 10, []),
|
||||
# Test case 6: Strings at exact max_length
|
||||
({"exact": "a" * 10}, 10, {"exact": "a" * 10}),
|
||||
# Test case 7: Non-string values
|
||||
({"num": 12345, "bool": True, "none": None}, 5, {"num": 12345, "bool": True, "none": None}),
|
||||
# Test case 8: Unicode characters
|
||||
({"unicode": "こんにちは世界"}, 3, {"unicode": "こんに..."}),
|
||||
# Test case 9: Very large structure
|
||||
(
|
||||
{"key" + str(i): "value" * i for i in range(1000)},
|
||||
10,
|
||||
{"key" + str(i): ("value" * i)[:10] + "..." if len("value" * i) > 10 else "value" * i for i in range(1000)},
|
||||
),
|
||||
],
|
||||
)
|
||||
def test_truncate_long_strings(input_data, max_length, expected):
|
||||
result = truncate_long_strings(input_data, max_length)
|
||||
assert result == expected
|
||||
|
||||
|
||||
def test_truncate_long_strings_default_max_length():
|
||||
long_string = "a" * (MAX_TEXT_LENGTH + 1)
|
||||
input_data = {"key": long_string}
|
||||
result = truncate_long_strings(input_data)
|
||||
assert len(result["key"]) == MAX_TEXT_LENGTH + 3 # +3 for the "..."
|
||||
|
||||
|
||||
def test_truncate_long_strings_no_modification():
|
||||
input_data = {"short": "short string", "nested": {"also_short": "another short string"}}
|
||||
result = truncate_long_strings(input_data, 100)
|
||||
assert result == input_data
|
||||
|
||||
|
||||
# Test for type preservation
|
||||
def test_truncate_long_strings_type_preservation():
|
||||
input_data = {"str": "a" * 100, "list": ["b" * 100], "dict": {"nested": "c" * 100}}
|
||||
result = truncate_long_strings(input_data, 10)
|
||||
assert isinstance(result, dict)
|
||||
assert isinstance(result["str"], str)
|
||||
assert isinstance(result["list"], list)
|
||||
assert isinstance(result["dict"], dict)
|
||||
|
||||
|
||||
# Test for in-place modification
|
||||
def test_truncate_long_strings_in_place_modification():
|
||||
input_data = {"key": "a" * 100}
|
||||
result = truncate_long_strings(input_data, 10)
|
||||
assert result is input_data # Check if the same object is returned
|
||||
|
||||
|
||||
# Test for invalid input
|
||||
def test_truncate_long_strings_invalid_input():
|
||||
input_string = "not a dict or list"
|
||||
result = truncate_long_strings(input_string, 10)
|
||||
assert result == input_string # The function should return the input unchanged
|
||||
|
||||
|
||||
# Updated test for negative max_length
|
||||
def test_truncate_long_strings_negative_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, -1)
|
||||
assert result == input_data # Assuming the function ignores negative max_length
|
||||
|
||||
|
||||
# Additional test for zero max_length
|
||||
def test_truncate_long_strings_zero_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, 0)
|
||||
assert result == {"key": "..."} # Assuming the function truncates to just "..."
|
||||
|
||||
|
||||
# Test for very small positive max_length
|
||||
def test_truncate_long_strings_small_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, 1)
|
||||
assert result == {"key": "v..."} # Assuming the function keeps at least one character
|
||||
Loading…
Add table
Add a link
Reference in a new issue