I have set up a 2-node memberlist cluster using VMs, each configured with 8 vCPUs and 8GB RAM.
Trace data is generated via the otel-ctlcommand-line tool, deployed across 15 VMs. Two infinite loop scripts run on each VM, designed to produce approximately 1,000 traces simultaneously per second at the same time.
Testing reveals the system currently handles only ~300 traces per second. Traces exceeding this limit are lost (likely discarded), despite low resource utilization (CPU: 5%, RAM: 25%).
Objective: Optimize the configuration to achieve a throughput of 1,000 traces per second.
Architecture: otel-ctl → Nginx → Alloy → Nginx → Tempo
My Tempo version is 2.8.
tempo.yml:
target: scalable-single-binary
server:
http_listen_port: 3200
grpc_listen_port: 9095
grpc_server_max_recv_msg_size: 107374182400
grpc_server_max_send_msg_size: 107374182400
http_server_read_timeout: 777s
http_server_write_timeout: 600s
log_level: debug
distributor:
ring:
kvstore:
store: memberlist
instance_addr: ${HOSTNAME}
instance_port: 9095
receivers:
otlp:
protocols:
grpc:
endpoint: "0.0.0.0:4317"
http:
endpoint: "0.0.0.0:4318"
retry_after_on_resource_exhausted: 5s
ingester:
lifecycler:
ring:
kvstore:
store: memberlist
replication_factor: 2
heartbeat_timeout: 1m
address: ${HOSTNAME}
port: 9095
heartbeat_period: 2s
heartbeat_timeout: 3s
observe_period: 5s
join_after: 5s
final_sleep: 5s
min_ready_duration: 3s
flush_all_on_shutdown: true
flush_check_period: 60s
complete_block_timeout: 600s
max_block_bytes: 10485760
trace_idle_period: 60s
max_block_duration: 600s
query_frontend:
search:
default_result_limit: 100
max_duration: 720h
duration_slo: 5s
throughput_bytes_slo: 1073741824
metadata_slo:
duration_slo: 5s
throughput_bytes_slo: 1073741824
trace_by_id:
duration_slo: 5s
metrics:
max_duration: 720h
concurrent_jobs: 100
multi_tenant_queries_enabled: false
metrics_generator:
ring:
kvstore:
store: memberlist
instance_addr: ${HOSTNAME}
instance_port: 9095
processor:
service_graphs:
enable_client_server_prefix: true
enable_messaging_system_latency_histogram: true
peer_attributes: []
enable_virtual_node_label: true
histogram_buckets: [0.5, 0.9, 1, 2, 3, 4, 5, 10,30, 60]
span_metrics:
intrinsic_dimensions:
status_message: true
enable_target_info: true
histogram_buckets: [0.5, 0.9, 1, 2, 3, 4, 5, 10,30, 60]
local_blocks:
block:
version: vParquet3
filter_server_spans: true
flush_to_storage: true
storage:
path: /export/tempo_data/metrics_storage
remote_write:
- url: http://<IP3>:9090/api/v1/write
send_exemplars: true
traces_storage:
path: /export/tempo_data/metrics_traces_storage
version: vParquet3
querier:
frontend_worker:
frontend_address: ${HOSTNAME}:9095
grpc_client_config:
max_recv_msg_size: 10737418240
max_send_msg_size: 10737418240
search:
query_timeout: 600s
max_concurrent_queries: 30
compactor:
ring:
kvstore:
store: memberlist
instance_addr: ${HOSTNAME}
instance_port: 9095
compaction:
compaction_cycle: 120s
compaction_window: 600s
block_retention: 10m
compacted_block_retention: 1h
retention_concurrency: 10
v2_out_buffer_bytes: 20971520
storage:
trace:
backend: azure
azure:
storage_account_name: xxxxx
container_name: tempo
blocklist_poll: 5m
blocklist_poll_tolerate_consecutive_errors: 10
blocklist_poll_tolerate_tenant_failures: 10
wal:
path: /export/tempo_data/storage_wal
version: vParquet3
block:
version: vParquet3
memberlist:
randomize_node_name: false
retransmit_factor: 2
gossip_nodes: 2
message_history_buffer_bytes: 1000
gossip_to_dead_nodes_time: 3s
pull_push_interval: 3s
leave_timeout: 3s
bind_addr: ["0.0.0.0"]
bind_port: 7946
abort_if_cluster_join_fails: false
join_members:
- <IP1>:7946
- <IP2>:7946
overrides:
defaults:
ingestion:
rate_strategy: global
metrics_generator:
processors:
- local-blocks
- service-graphs
- span-metrics
usage_report:
reporting_enabled: false
My alloy.yml configuration file is provided below.
otelcol.receiver.otlp "test" {
http {
endpoint = "0.0.0.0:4318"
}
grpc {
endpoint = "0.0.0.0:4317"
}
output {
traces = [otelcol.processor.filter.test.input]
metrics = [otelcol.exporter.prometheus.test.input]
}
}
otelcol.processor.filter "test" {
traces {
span = [
`attributes["http.request.method"] != "POST"`,
]
}
output {
traces = [otelcol.processor.attributes.test.input]
}
}
otelcol.processor.attributes "test" {
action {
key = "CIty"
value = "aaaaa"
action = "insert"
}
action {
from_attribute = "network.peer.address"
key = "127.0.0.1"
action = "upsert"
}
action {
key = "http.status_code"
value = "ok"
action = "upsert"
}
output {
traces = [otelcol.processor.transform.test.input]
}
}
otelcol.processor.transform "test" {
trace_statements {
context = "span"
statements = [
`set(attributes["Kind"],"AAAAAAAAAAAAAAAAAAAAAAAA")`,
]
}
output {
traces = [otelcol.processor.batch.test.input]
}
}
otelcol.processor.batch "test" {
output {
traces = [otelcol.exporter.otlp.test.input]
}
}
otelcol.exporter.otlp "test" {
client {
endpoint = "http://tempo.d.vb.local:4317"
tls {
insecure = true
}
}
}
otelcol.exporter.prometheus "test" {
forward_to = [prometheus.remote_write.test.receiver]
}
prometheus.remote_write "test" {
endpoint { url = "http://<IP3>:9090/api/v1/write" }
}
Who can help me? Thanks!