How to make Tempo support handling 1000 traces per second?

I have set up a ​​2-node memberlist cluster​​ using VMs, each configured with ​​8 vCPUs and 8GB RAM​​.

Trace data is generated via the otel-ctlcommand-line tool, deployed across ​​15 VMs​​. ​​Two infinite loop scripts run on each VM​​, designed to produce approximately ​​1,000 traces simultaneously per second​​ at the same time.

Testing reveals the system ​​currently handles only ~300 traces per second​​. Traces exceeding this limit are lost (likely discarded), despite ​​low resource utilization (CPU: 5%, RAM: 25%)​​.

​​Objective:​​ Optimize the configuration to achieve a throughput of ​​1,000 traces per second​​.

Architecture: otel-ctl → Nginx → Alloy → Nginx → Tempo
My Tempo version is 2.8.
tempo.yml:

target: scalable-single-binary
server:
  http_listen_port: 3200
  grpc_listen_port: 9095
  grpc_server_max_recv_msg_size: 107374182400
  grpc_server_max_send_msg_size: 107374182400
  http_server_read_timeout: 777s
  http_server_write_timeout: 600s
  log_level: debug

distributor:
  ring:
    kvstore:
      store: memberlist
    instance_addr: ${HOSTNAME}
    instance_port: 9095
  receivers:
    otlp:
      protocols:
        grpc:
          endpoint: "0.0.0.0:4317"
        http:
          endpoint: "0.0.0.0:4318"
  retry_after_on_resource_exhausted: 5s

ingester:
  lifecycler:
    ring:
      kvstore:
        store: memberlist
      replication_factor: 2
      heartbeat_timeout: 1m
    address: ${HOSTNAME}
    port: 9095
    heartbeat_period: 2s
    heartbeat_timeout: 3s
    observe_period: 5s
    join_after: 5s
    final_sleep: 5s
    min_ready_duration: 3s
  flush_all_on_shutdown: true
  flush_check_period: 60s
  complete_block_timeout: 600s
  max_block_bytes: 10485760
  trace_idle_period: 60s
  max_block_duration: 600s


query_frontend:
  search:
    default_result_limit: 100
    max_duration:  720h
    duration_slo: 5s
    throughput_bytes_slo: 1073741824
    metadata_slo:
      duration_slo: 5s
      throughput_bytes_slo: 1073741824
  trace_by_id:
    duration_slo: 5s
  metrics:
    max_duration:  720h
    concurrent_jobs: 100
  multi_tenant_queries_enabled: false

metrics_generator:
  ring:
    kvstore:
      store: memberlist
    instance_addr: ${HOSTNAME}
    instance_port: 9095
  processor:
    service_graphs:
      enable_client_server_prefix: true
      enable_messaging_system_latency_histogram: true
      peer_attributes: []
      enable_virtual_node_label: true
      histogram_buckets: [0.5, 0.9, 1, 2, 3, 4, 5, 10,30, 60]
    span_metrics:
      intrinsic_dimensions:
        status_message: true
      enable_target_info: true
      histogram_buckets: [0.5, 0.9, 1, 2, 3, 4, 5, 10,30, 60]
    local_blocks:
      block:
        version: vParquet3
      filter_server_spans: true
      flush_to_storage: true
  storage:
    path: /export/tempo_data/metrics_storage
    remote_write:
      - url: http://<IP3>:9090/api/v1/write
        send_exemplars: true
  traces_storage:
    path: /export/tempo_data/metrics_traces_storage
    version: vParquet3

querier:
  frontend_worker:
    frontend_address: ${HOSTNAME}:9095
    grpc_client_config:
      max_recv_msg_size: 10737418240
      max_send_msg_size: 10737418240
  search:
    query_timeout: 600s
  max_concurrent_queries: 30

compactor:
  ring:
    kvstore:
      store: memberlist
    instance_addr: ${HOSTNAME}
    instance_port: 9095
  compaction:
    compaction_cycle: 120s
    compaction_window: 600s
    block_retention: 10m
    compacted_block_retention: 1h
    retention_concurrency: 10
    v2_out_buffer_bytes: 20971520

storage:
  trace:
    backend: azure
    azure:
      storage_account_name: xxxxx
      container_name: tempo
    blocklist_poll: 5m
    blocklist_poll_tolerate_consecutive_errors: 10
    blocklist_poll_tolerate_tenant_failures: 10
    wal:
      path: /export/tempo_data/storage_wal
      version: vParquet3
    block:
      version: vParquet3

memberlist:
  randomize_node_name: false
  retransmit_factor: 2
  gossip_nodes: 2
  message_history_buffer_bytes: 1000
  gossip_to_dead_nodes_time: 3s
  pull_push_interval: 3s
  leave_timeout: 3s
  bind_addr: ["0.0.0.0"]
  bind_port: 7946
  abort_if_cluster_join_fails: false
  join_members:
    - <IP1>:7946
    - <IP2>:7946

overrides:
  defaults:
    ingestion:
      rate_strategy: global
    metrics_generator:
      processors:
        - local-blocks
        - service-graphs
        - span-metrics

usage_report:
  reporting_enabled: false

My alloy.yml configuration file is provided below.


otelcol.receiver.otlp "test" {
    http {
        endpoint = "0.0.0.0:4318"
    }
    grpc {
        endpoint = "0.0.0.0:4317"
    }
    output {
        traces  = [otelcol.processor.filter.test.input]
        metrics = [otelcol.exporter.prometheus.test.input]
    }
}

otelcol.processor.filter "test" {
  traces {
    span = [
   `attributes["http.request.method"] != "POST"`,
    ]
  }
  output {
    traces = [otelcol.processor.attributes.test.input]
  }
}

otelcol.processor.attributes "test" {
    action {
        key = "CIty"
        value = "aaaaa"
        action = "insert"
    }
    action {
        from_attribute  = "network.peer.address"
        key = "127.0.0.1"
        action = "upsert"
    }
    action {
        key = "http.status_code"
        value = "ok"
        action = "upsert"
    }
    output {
         traces  = [otelcol.processor.transform.test.input]
    }
}

otelcol.processor.transform "test" {
    trace_statements {
        context = "span"
        statements = [
            `set(attributes["Kind"],"AAAAAAAAAAAAAAAAAAAAAAAA")`,
        ]
    }
    output {
        traces = [otelcol.processor.batch.test.input]
    }
}

otelcol.processor.batch "test" {
    output {
         traces  = [otelcol.exporter.otlp.test.input]
    }
}

otelcol.exporter.otlp "test" {
    client {
        endpoint = "http://tempo.d.vb.local:4317"
        tls {
            insecure = true
        }
    }
}

otelcol.exporter.prometheus "test" {
  forward_to = [prometheus.remote_write.test.receiver]
}

prometheus.remote_write "test" {
    endpoint { url = "http://<IP3>:9090/api/v1/write" }
}

​​Who can help me? Thanks!​​

This topic was automatically closed 365 days after the last reply. New replies are no longer allowed.