chore: Enhance Locust load testing and optimize database settings (#6265)

* feat: Enhance Locust load testing for Langflow run endpoint

Refactor locustfile to provide more robust and configurable load testing:
- Add dynamic configuration via environment variables
- Improve error handling and logging
- Implement realistic flow run simulation
- Add connection and timeout handling
- Support API key authentication
- Enhance stats tracking and error reporting

* fix: Improve transaction logging error handling and performance

- Add `no_autoflush` context to prevent unnecessary database operations
- Change transaction logging error from exception to error level logging
- Simplify error handling in log_transaction function

* chore: Add Locust to development dependencies

Update project dependencies by adding Locust (version 2.32.9) to the development requirements, supporting load testing capabilities

* feat: Optimize database connection settings for improved performance and scalability

- Increase default pool_size from 10 to 20 for better connection handling
- Adjust max_overflow to 40 to support higher concurrent connections
- Extend db_connect_timeout from 20 to 30 seconds
- Add pool_recycle and echo settings to db_connection_settings
- Enhance documentation for database connection settings, highlighting SQLite limitations

* feat: Add Locust load testing configuration to Makefile

- Introduce comprehensive Locust load testing target with configurable parameters
- Support flexible testing scenarios with customizable users, spawn rate, and host
- Enable headless and interactive testing modes
- Add environment variable support for API key, flow ID, and other testing parameters
- Provide sensible default values for load testing configuration

* refactor: Remove unused retry configuration in Locust load testing

- Remove RETRY_DELAY and MAX_RETRIES environment variables
- Simplify FlowRunUser configuration by eliminating unused retry settings
- Maintain existing wait time configuration for load testing

* feat: Enforce FLOW_ID requirement for Locust load testing

- Add mandatory validation for FLOW_ID environment variable
- Raise a clear ValueError if FLOW_ID is not provided
- Remove default flow ID to ensure explicit configuration
- Improve load testing configuration robustness

* feat: Add configurable request timeout for Locust load testing

- Introduce `locust_request_timeout` parameter in Makefile
- Update locustfile to use configurable request timeout from environment variable
- Set dynamic connection and network timeout based on REQUEST_TIMEOUT
- Improve request handling with flexible timeout configuration

* revert change to database connection retry
This commit is contained in:
Gabriel Luiz Freitas Almeida 2025-02-17 11:26:36 -03:00 • committed by GitHub
commit e1fb90074c
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 536 additions and 125 deletions

View file

@ -135,10 +135,11 @@ async def log_transaction(
flow_id=flow_id if isinstance(flow_id, UUID) else UUID(flow_id),
)
async with session_getter(get_db_service()) as session:
inserted = await crud_log_transaction(session, transaction)
logger.debug(f"Logged transaction: {inserted.id}")
with session.no_autoflush:
inserted = await crud_log_transaction(session, transaction)
logger.debug(f"Logged transaction: {inserted.id}")
except Exception: # noqa: BLE001
logger.exception("Error logging transaction")
logger.error("Error logging transaction")
async def log_vertex_build(

View file

@ -76,14 +76,13 @@ class Settings(BaseSettings):
`postgresql+psycopg` respectively)."""
database_connection_retry: bool = False
"""If True, Langflow will retry to connect to the database if it fails."""
pool_size: int = 10
"""DEPRECATED: Use db_connection_settings['pool_size'] instead.
The number of connections to keep open in the connection pool. If not provided, the default is 10."""
max_overflow: int = 20
"""DEPRECATED: Use db_connection_settings['max_overflow'] instead.
The number of connections to allow that can be opened beyond the pool size.
If not provided, the default is 20."""
db_connect_timeout: int = 20
pool_size: int = 20
"""The number of connections to keep open in the connection pool.
For high load scenarios, this should be increased based on expected concurrent users."""
max_overflow: int = 30
"""The number of connections to allow that can be opened beyond the pool size.
Should be 2x the pool_size for optimal performance under load."""
db_connect_timeout: int = 30
"""The number of seconds to wait before giving up on a lock to released or establishing a connection to the
database."""
@ -92,12 +91,27 @@ class Settings(BaseSettings):
"""SQLite pragmas to use when connecting to the database."""
db_connection_settings: dict | None = {
"pool_size": 10,
"max_overflow": 20,
"pool_timeout": 30,
"pool_pre_ping": True,
"pool_size": 20, # Match the pool_size above
"max_overflow": 30, # Match the max_overflow above
"pool_timeout": 30, # Seconds to wait for a connection from pool
"pool_pre_ping": True, # Check connection validity before using
"pool_recycle": 1800, # Recycle connections after 30 minutes
"echo": False, # Set to True for debugging only
}
"""Common database connection settings."""
"""Database connection settings optimized for high load scenarios.
Note: These settings are most effective with PostgreSQL. For SQLite:
- Reduce pool_size and max_overflow if experiencing lock contention
- SQLite has limited concurrent write capability even with WAL mode
- Best for read-heavy or moderate write workloads
Settings:
- pool_size: Number of connections to maintain (increase for higher concurrency)
- max_overflow: Additional connections allowed beyond pool_size
- pool_timeout: Seconds to wait for an available connection
- pool_pre_ping: Validates connections before use to prevent stale connections
- pool_recycle: Seconds before connections are recycled (prevents timeouts)
- echo: Enable SQL query logging (development only)
"""
# cache configuration
cache_type: Literal["async", "redis", "memory", "disk"] = "async"