feat: truncate parsed uploads to prevent database and frontend blocking caused by excessively large files (#3914)
* 📝 (constants.ts): increase maxSizeFilesInBytes constant value from 10MB to 100MB to allow larger file uploads * 🐛 (inputFileComponent): fix bug in setting the maximum file size alert message to display the correct file size limit of 100 bytes instead of 10 bytes * 📝 (schemas.py): Add a new field_serializer method to serialize data in VertexBuildResponse class 📝 (schemas.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries 📝 (model.py): Add a new field_serializer method to serialize outputs in TransactionBase class 📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries 📝 (model.py): Add a new field_serializer method to serialize data and artifacts in VertexBuildBase class 📝 (model.py): Add a new truncate_text helper function to safely truncate text in nested dictionaries * 🐛 (schemas.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (model.py): fix truncation length of text fields to 10 characters instead of 99999 🐛 (index.tsx): truncate resultMessage to 99999 characters and add message if text is too long * 🔧 (switchOutputView/index.tsx): Use useMemo to memoize resultMessage transformations for performance optimization * 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (model.py): Fix typo in the path for 'base_retriever' data field 🐛 (index.tsx): Fix logic to correctly handle resultMessageMemoized when it is an object * 📝 (model.py): refactor truncate_text function to truncate_long_strings for better clarity and consistency 📝 (model.py): update serialize_outputs and serialize_artifacts functions to use truncate_long_strings for string truncation 📝 (model.py): introduce MAX_TEXT_LENGTH constant for defining the maximum length of text to truncate in the models * 📝 (schemas.py): refactor serialize_data method in VertexBuildResponse class to use a new helper function truncate_long_strings for better code readability and maintainability * 🔧 (schemas.py): Move the `truncate_long_strings` function to a separate module to improve code organization and reusability 🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error 🔧 (model.py): Import the `truncate_long_strings` function from the correct module to fix the reference error * 📝 (util.py): add function truncate_long_strings to recursively truncate long strings in dictionaries and lists to prevent exceeding the maximum text length. * 📝 (constants.py): add constant MAX_TEXT_LENGTH with value 99999 for defining maximum text length allowed in the application * 📝 (model.py): update import path for truncate_long_strings function to match new location in util module * ✨ (test_truncate_long_strings_on_objects.py): Add unit tests for the function truncate_long_strings to ensure correct behavior when truncating long strings in various data structures 🐛 (switchOutputView/index.tsx): Fix truncation logic to correctly truncate long strings by adding ellipsis at the end instead of displaying additional text about truncation. * [autofix.ci] apply automated fixes * ✨ (test_truncate_long_strings_on_objects.py): Update import path for truncate_long_strings function 📝 (test_truncate_long_strings_on_objects.py): Add additional tests for handling negative, zero, and small max_length values in truncate_long_strings function * ♻️ (schemas.py): refactor import statement to use the updated module name util_strings instead of util for better clarity and consistency. * 📝 (model.py): Update import path for util_strings module to fix module import error 📝 (util.py): Remove redundant code for truncating long strings and move it to a separate util_strings module for better organization and separation of concerns. * 📝 (schemas.py): refactor serialize_data method to handle both BaseModel and non-BaseModel data inputs in VertexBuildResponse class * 📝 (util_strings.py): Update util_strings.py to improve string truncation function for dictionaries and lists 🔧 (test_truncate_long_strings_on_objects.py): Update test cases for string truncation function to cover additional scenarios and edge cases * Update src/backend/base/langflow/utils/util_strings.py Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org> * 📝 (vite.config.mts): update environment variable MAX_FILE_SIZE to be defined in vite config for frontend to use in the application. * 📝 (constants.ts): update maxSizeFilesInBytes constant to use process.env.MAX_FILE_SIZE environment variable for configurable file size limit 📝 (constants.ts): add MAX_TEXT_LENGTH constant with a value of 99999 for maximum text length limit * 📝 (switchOutputView/index.tsx): import MAX_TEXT_LENGTH constant from shared constants file to improve code organization and reusability * ✨ (langflow/__main__.py): add support for defining maximum file size for upload in MB to improve user experience and prevent large file uploads * 🐛 (files.py): add validation to check if uploaded file size exceeds the maximum allowed size before processing it * ✨ (schemas.py): add max_file_size_upload field to ConfigResponse schema to handle maximum file size allowed for upload * 🔧 (vite.config.mts): remove MAX_FILE_SIZE environment variable configuration as it is no longer needed * ✨ (base.py): introduce max_file_size_upload setting to limit the file size for uploads in MB * 🐛 (util.py): add support for setting max_file_size_upload in update_settings function to allow configuring maximum file size for uploads * 📝 (inputFileComponent/index.tsx): add support for retrieving max file size upload from utility store to improve code modularity and reusability 🐛 (inputFileComponent/index.tsx): fix error handling logic to display error message when uploading a file fails * 📝 (constants.ts): remove maxSizeFilesInBytes constant as it is no longer used and update MAX_TEXT_LENGTH constant to a higher value * ✨ (use-get-config.ts): add functionality to set max file size upload value from config response * ✨ (utilityStore.ts): introduce maxFileSizeUpload property and setMaxFileSizeUpload function to handle maximum file size upload in bytes * ✨ (frontend): introduce maxFileSizeUpload property and setMaxFileSizeUpload method to handle maximum file size upload functionality in the UtilityStoreType * ♻️ (util_strings.py): refactor truncate_long_strings function to improve code readability and consistency by removing unnecessary whitespace and aligning assignment operators. * 🐛 (files.py): fix formatting issue in the raise statement to improve code readability and maintain consistency --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com> Co-authored-by: Gabriel Luiz Freitas Almeida <gabriel@langflow.org>
This commit is contained in:
parent
b34a7c7f02
commit
948b150946
21 changed files with 414 additions and 21 deletions
|
|
@ -0,0 +1,98 @@
|
|||
from langflow.utils.util_strings import truncate_long_strings
|
||||
from langflow.utils.constants import MAX_TEXT_LENGTH
|
||||
import pytest
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"input_data, max_length, expected",
|
||||
[
|
||||
# Test case 1: Simple string truncation
|
||||
({"key": "a" * 100}, 10, {"key": "a" * 10 + "..."}),
|
||||
# Test case 2: Nested dictionary
|
||||
({"outer": {"inner": "b" * 100}}, 5, {"outer": {"inner": "b" * 5 + "..."}}),
|
||||
# Test case 3: List of strings
|
||||
(["short", "a" * 100, "also short"], 7, ["short", "a" * 7 + "...", "also sh" + "..."]),
|
||||
# Test case 4: Mixed nested structure
|
||||
(
|
||||
{"key1": ["a" * 100, {"nested": "b" * 100}], "key2": "c" * 100},
|
||||
8,
|
||||
{"key1": ["a" * 8 + "...", {"nested": "b" * 8 + "..."}], "key2": "c" * 8 + "..."},
|
||||
),
|
||||
# Test case 5: Empty structures
|
||||
({}, 10, {}),
|
||||
([], 10, []),
|
||||
# Test case 6: Strings at exact max_length
|
||||
({"exact": "a" * 10}, 10, {"exact": "a" * 10}),
|
||||
# Test case 7: Non-string values
|
||||
({"num": 12345, "bool": True, "none": None}, 5, {"num": 12345, "bool": True, "none": None}),
|
||||
# Test case 8: Unicode characters
|
||||
({"unicode": "こんにちは世界"}, 3, {"unicode": "こんに..."}),
|
||||
# Test case 9: Very large structure
|
||||
(
|
||||
{"key" + str(i): "value" * i for i in range(1000)},
|
||||
10,
|
||||
{"key" + str(i): ("value" * i)[:10] + "..." if len("value" * i) > 10 else "value" * i for i in range(1000)},
|
||||
),
|
||||
],
|
||||
)
|
||||
def test_truncate_long_strings(input_data, max_length, expected):
|
||||
result = truncate_long_strings(input_data, max_length)
|
||||
assert result == expected
|
||||
|
||||
|
||||
def test_truncate_long_strings_default_max_length():
|
||||
long_string = "a" * (MAX_TEXT_LENGTH + 1)
|
||||
input_data = {"key": long_string}
|
||||
result = truncate_long_strings(input_data)
|
||||
assert len(result["key"]) == MAX_TEXT_LENGTH + 3 # +3 for the "..."
|
||||
|
||||
|
||||
def test_truncate_long_strings_no_modification():
|
||||
input_data = {"short": "short string", "nested": {"also_short": "another short string"}}
|
||||
result = truncate_long_strings(input_data, 100)
|
||||
assert result == input_data
|
||||
|
||||
|
||||
# Test for type preservation
|
||||
def test_truncate_long_strings_type_preservation():
|
||||
input_data = {"str": "a" * 100, "list": ["b" * 100], "dict": {"nested": "c" * 100}}
|
||||
result = truncate_long_strings(input_data, 10)
|
||||
assert isinstance(result, dict)
|
||||
assert isinstance(result["str"], str)
|
||||
assert isinstance(result["list"], list)
|
||||
assert isinstance(result["dict"], dict)
|
||||
|
||||
|
||||
# Test for in-place modification
|
||||
def test_truncate_long_strings_in_place_modification():
|
||||
input_data = {"key": "a" * 100}
|
||||
result = truncate_long_strings(input_data, 10)
|
||||
assert result is input_data # Check if the same object is returned
|
||||
|
||||
|
||||
# Test for invalid input
|
||||
def test_truncate_long_strings_invalid_input():
|
||||
input_string = "not a dict or list"
|
||||
result = truncate_long_strings(input_string, 10)
|
||||
assert result == input_string # The function should return the input unchanged
|
||||
|
||||
|
||||
# Updated test for negative max_length
|
||||
def test_truncate_long_strings_negative_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, -1)
|
||||
assert result == input_data # Assuming the function ignores negative max_length
|
||||
|
||||
|
||||
# Additional test for zero max_length
|
||||
def test_truncate_long_strings_zero_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, 0)
|
||||
assert result == {"key": "..."} # Assuming the function truncates to just "..."
|
||||
|
||||
|
||||
# Test for very small positive max_length
|
||||
def test_truncate_long_strings_small_max_length():
|
||||
input_data = {"key": "value"}
|
||||
result = truncate_long_strings(input_data, 1)
|
||||
assert result == {"key": "v..."} # Assuming the function keeps at least one character
|
||||
Loading…
Add table
Add a link
Reference in a new issue