Skip to content

3. Global Section

Process security, performance tuning, debugging, and HTTP client settings

Parameters in the “global” section are process-wide and often OS-specific. They are generally set once for all and do not need being changed once correct. Some of them have command-line equivalents.

The following keywords are supported in the “global” section:

  • Process management and security

    • 51degrees-allow-unmatched
    • 51degrees-cache-size
    • 51degrees-data-file
    • 51degrees-difference
    • 51degrees-drift
    • 51degrees-property-name-list
    • 51degrees-property-separator
    • 51degrees-use-performance-graph
    • 51degrees-use-predictive-graph
    • ca-base
    • chroot
    • cluster-secret
    • cpu-affinity
    • cpu-map
    • cpu-policy
    • cpu-set
    • crt-base
    • daemon
    • default-path
    • description
    • deviceatlas-json-file
    • deviceatlas-log-level
    • deviceatlas-properties-cookie
    • deviceatlas-separator
    • dns-accept-family
    • expose-deprecated-directives
    • expose-experimental-directives
    • external-check
    • fd-hard-limit
    • gid
    • grace
    • group
    • h1-accept-payload-with-any-method
    • h1-case-adjust
    • h1-case-adjust-file
    • h1-do-not-close-on-insecure-transfer-encoding
    • h2-workaround-bogus-websocket-clients
    • hard-stop-after
    • harden.reject-privileged-ports.tcp
    • harden.reject-privileged-ports.quic
    • insecure-fork-wanted
    • insecure-setuid-wanted
    • issuers-chain-path
    • jwt.decrypt_alg_list
    • jwt.decrypt_enc_list
    • key-base
    • limited-quic
    • localpeer
    • log
    • log-send-hostname
    • log-tag
    • lua-load
    • lua-load-per-thread
    • lua-prepend-path
    • max-threads-per-group
    • mworker-max-reloads
    • nbthread
    • node
    • numa-cpu-mapping
    • ocsp-update.disable
    • ocsp-update.maxdelay
    • ocsp-update.mindelay
    • ocsp-update.httpproxy
    • ocsp-update.mode
    • pidfile
    • pp2-never-send-local
    • presetenv
    • prealloc-fd
    • resetenv
    • set-dumpable
    • set-var
    • setenv
    • ssl-default-bind-ciphers
    • ssl-default-bind-ciphersuites
    • ssl-default-bind-client-sigalgs
    • ssl-default-bind-curves
    • ssl-default-bind-options
    • ssl-default-bind-sigalgs
    • ssl-default-server-ciphers
    • ssl-default-server-ciphersuites
    • ssl-default-server-client-sigalgs
    • ssl-default-server-curves
    • ssl-default-server-options
    • ssl-default-server-sigalgs
    • ssl-dh-param-file
    • ssl-propquery
    • ssl-provider
    • ssl-provider-path
    • ssl-security-level
    • ssl-server-verify
    • ssl-skip-self-issued-ca
    • stats
    • stats-file
    • strict-limits
    • uid
    • ulimit-n
    • unix-bind
    • unsetenv
    • user
    • wurfl-cache-size
    • wurfl-data-file
    • wurfl-information-list
    • wurfl-information-list-separator
  • Performance tuning

    • busy-polling
    • max-spread-checks
    • maxcompcpuusage
    • maxcomprate
    • maxconn
    • maxconnrate
    • maxpipes
    • maxsessrate
    • maxsslconn
    • maxsslrate
    • maxzlibmem
    • no-memory-trimming
    • noepoll
    • noevports
    • nogetaddrinfo
    • nokqueue
    • noktls
    • nopoll
    • noreuseport
    • nosplice
    • profiling.memory
    • profiling.tasks
    • server-state-base
    • server-state-file
    • spread-checks
    • ssl-engine
    • ssl-mode-async
    • tune.applet.zero-copy-forwarding
    • tune.buffers.limit
    • tune.buffers.reserve
    • tune.bufsize
    • tune.bufsize.large
    • tune.bufsize.small
    • tune.cli.max-payload-size
    • tune.comp.maxlevel
    • tune.defaults.purge
    • tune.disable-fast-forward
    • tune.disable-zero-copy-forwarding
    • tune.epoll.mask-events
    • tune.events.max-events-at-once
    • tune.fail-alloc
    • tune.fd.edge-triggered
    • tune.h1.be.glitches-threshold
    • tune.h1.fe.glitches-threshold
    • tune.h1.zero-copy-fwd-recv
    • tune.h1.zero-copy-fwd-send
    • tune.h2.be.glitches-threshold
    • tune.h2.be.initial-window-size
    • tune.h2.be.max-concurrent-streams
    • tune.h2.be.max-frames-at-once
    • tune.h2.be.rxbuf
    • tune.h2.fe.glitches-threshold
    • tune.h2.fe.initial-window-size
    • tune.h2.fe.max-concurrent-streams
    • tune.h2.fe.max-frames-at-once
    • tune.h2.fe.max-rst-at-once
    • tune.h2.fe.max-total-streams
    • tune.h2.fe.rxbuf
    • tune.h2.header-table-size
    • tune.h2.initial-window-size
    • tune.h2.max-concurrent-streams
    • tune.h2.max-frame-size
    • tune.h2.zero-copy-fwd-send
    • tune.http.cookielen
    • tune.http.logurilen
    • tune.http.maxhdr
    • tune.idle-pool.shared
    • tune.idletimer
    • tune.lua.bool-sample-conversion
    • tune.lua.burst-timeout
    • tune.lua.forced-yield
    • tune.lua.log.loggers
    • tune.lua.log.stderr
    • tune.lua.maxmem
    • tune.lua.openlibs
    • tune.lua.service-timeout
    • tune.lua.session-timeout
    • tune.lua.task-timeout
    • tune.max-checks-per-thread
    • tune.maxaccept
    • tune.maxpollevents
    • tune.maxrewrite
    • tune.max-rules-at-once
    • tune.memory.hot-size
    • tune.pattern.cache-size
    • tune.peers.max-updates-at-once
    • tune.pipesize
    • tune.pool-high-fd-ratio
    • tune.pool-low-fd-ratio
    • tune.pt.zero-copy-forwarding
    • tune.quic.be.cc.cubic-min-losses
    • tune.quic.be.cc.hystart
    • tune.quic.be.cc.max-frame-loss
    • tune.quic.be.cc.max-win-size
    • tune.quic.be.cc.reorder-ratio
    • tune.quic.be.max-idle-timeout
    • tune.quic.be.sec.glitches-threshold
    • tune.quic.be.stream.data-ratio
    • tune.quic.be.stream.max-concurrent
    • tune.quic.be.stream.rxbuf
    • tune.quic.be.tx.pacing
    • tune.quic.be.tx.udp-gso
    • tune.quic.cc.cubic.min-losses (deprecated)
    • tune.quic.cc-hystart (deprecated)
    • tune.quic.disable-tx-pacing (deprecated)
    • tune.quic.disable-udp-gso (deprecated)
    • tune.quic.fe.cc.cubic-min-losses
    • tune.quic.fe.cc.hystart
    • tune.quic.fe.cc.max-frame-loss
    • tune.quic.fe.cc.max-win-size
    • tune.quic.fe.cc.reorder-ratio
    • tune.quic.fe.max-idle-timeout
    • tune.quic.fe.sec.glitches-threshold
    • tune.quic.fe.sec.retry-threshold
    • tune.quic.fe.sock-per-conn
    • tune.quic.fe.stream.data-ratio
    • tune.quic.fe.stream.max-concurrent
    • tune.quic.fe.stream.max-total
    • tune.quic.fe.stream.rxbuf
    • tune.quic.fe.tx.pacing
    • tune.quic.fe.tx.udp-gso
    • tune.quic.frontend.max-data-size (deprecated)
    • tune.quic.frontend.max-idle-timeout (deprecated)
    • tune.quic.frontend.max-streams-bidi (deprecated)
    • tune.quic.frontend.max-tx-mem (deprecated)
    • tune.quic.frontend.stream-data-ratio (deprecated)
    • tune.quic.frontend.default-max-window-size (deprecated)
    • tune.quic.listen
    • tune.quic.max-frame-loss (deprecated)
    • tune.quic.mem.tx-max
    • tune.quic.reorder-ratio (deprecated)
    • tune.quic.retry-threshold (deprecated)
    • tune.quic.socket-owner (deprecated)
    • tune.quic.zero-copy-fwd-send
    • tune.renice.runtime
    • tune.renice.startup
    • tune.rcvbuf.backend
    • tune.rcvbuf.client
    • tune.rcvbuf.frontend
    • tune.rcvbuf.server
    • tune.recv_enough
    • tune.ring.queues
    • tune.runqueue-depth
    • tune.sched.low-latency
    • tune.sndbuf.backend
    • tune.sndbuf.client
    • tune.sndbuf.frontend
    • tune.sndbuf.server
    • tune.streams-elasticity
    • tune.stick-counters
    • tune.ssl.cachesize
    • tune.ssl.capture-buffer-size
    • tune.ssl.capture-cipherlist-size (deprecated)
    • tune.ssl.certificate-compression
    • tune.ssl.default-dh-param
    • tune.ssl.force-private-cache
    • tune.ssl.hard-maxrecord
    • tune.ssl.keylog
    • tune.ssl.keyupdate-rate-limit
    • tune.ssl.lifetime
    • tune.ssl.maxrecord
    • tune.ssl.ssl-ctx-cache-size
    • tune.ssl.ocsp-update.maxdelay (deprecated)
    • tune.ssl.ocsp-update.mindelay (deprecated)
    • tune.takeover-other-tg-connections
    • tune.vars.global-max-size
    • tune.vars.proc-max-size
    • tune.vars.reqres-max-size
    • tune.vars.sess-max-size
    • tune.vars.txn-max-size
    • tune.zlib.memlevel
    • tune.zlib.windowsize
  • Debugging

    • anonkey
    • debug.counters
    • force-cfg-parser-pause
    • quiet
    • warn-blocked-traffic-after
    • zero-warning
  • HTTPClient

    • httpclient.resolvers.disabled
    • httpclient.resolvers.id
    • httpclient.resolvers.prefer
    • httpclient.retries
    • httpclient.ssl.ca-file
    • httpclient.ssl.verify
    • httpclient.timeout.connect

3.1. Process management and security

51degrees-data-file <file path>

51degrees-data-file <file path>

The path of the 51Degrees data file to provide device detection services. The file should be unzipped and accessible by HAProxy with relevant permissions.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES.

51degrees-property-name-list [<string> ...]

51degrees-property-name-list [<string> ...]

A list of 51Degrees property names to be load from the dataset. A full list of names is available on the 51Degrees website: https://51degrees.com/resources/property-dictionary

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES.

51degrees-property-separator <char>

51degrees-property-separator <char>

A char that will be appended to every property value in a response header containing 51Degrees results. If not set that will be set as ‘,’.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES.

51degrees-cache-size <number>

51degrees-cache-size <number>

Sets the size of the 51Degrees converter cache to <number> entries. This is an LRU cache which reminds previous device detections and their results. By default, this cache is disabled.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES.

51degrees-use-performance-graph { on | off }

51degrees-use-performance-graph { on | off }

Enables (‘on’) or disables (‘off’) the use of the performance graph in the detection process. The default value depends on 51Degrees library.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES and 51DEGREES_VER=4.

51degrees-use-predictive-graph { on | off }

51degrees-use-predictive-graph { on | off }

Enables (‘on’) or disables (‘off’) the use of the predictive graph in the detection process. The default value depends on 51Degrees library.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES and 51DEGREES_VER=4.

51degrees-drift <number>

51degrees-drift <number>

Sets the drift value that a detection can allow.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES and 51DEGREES_VER=4.

51degrees-difference <number>

51degrees-difference <number>

Sets the difference value that a detection can allow.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES and 51DEGREES_VER=4.

51degrees-allow-unmatched { on | off }

51degrees-allow-unmatched { on | off }

Enables (‘on’) or disables (‘off’) the use of unmatched nodes in the detection process. The default value depends on 51Degrees library.

Please note that this option is only available when HAProxy has been compiled with USE_51DEGREES and 51DEGREES_VER=4.

acme.scheduler { auto | off }

acme.scheduler { auto | off }

Enable or disable the ACME scheduler.

The ACME scheduler starts at HAProxy startup, it will loop over the certificates and start an ACME renewal task when the notAfter value is past curtime + (notAfter - notBefore) / 12, or 7 days if notBefore is not defined. The scheduler will then sleep and wakeup after 12 hours.

The default value is “auto”.

See also: acme

ca-base <dir>

ca-base <dir>

Assigns a default directory to fetch SSL CA certificates and CRLs from when a relative path is used with “ca-file”, “ca-verify-file” or “crl-file” directives. Absolute locations specified in “ca-file”, “ca-verify-file” and “crl-file” prevail and ignore “ca-base”.

chroot { <jail dir> | auto }

chroot { <jail dir> | auto }

Changes current directory to <jail dir> and performs a chroot() there before dropping privileges. This increases the security level in case an unknown vulnerability would be exploited, since it would make it very hard for the attacker to exploit the system. It is important to ensure that <jail dir> is both empty and non-writable to anyone. When the process is started with superuser privileges, the chroot() is performed directly. On Linux, when started unprivileged, haproxy attempts to perform it from inside a new user namespace created with unshare(CLONE_NEWUSER); if that mechanism is unavailable the chroot() will fail with the usual error.

As a special case, <jail dir> may be set to “auto”, in which case haproxy creates an anonymous temporary directory, unlinks it, and chroots into it. The resulting jail has no name in the filesystem and is empty and read-only, removing the need to prepare a dedicated jail directory.

When starting with superuser privileges, a warning will be displayed if no chroot is used, in order to encourage users to always use the mechanism. If for any reason there is a compelling reason not to use chroot (e.g. access to a server via a UNIX socket with an unconvenient path), it remains possible to silence the warning by adding an explicit “chroot /”, which has the benefit of being visible in a configuration.

close-spread-time <time>

close-spread-time <time>

Define a time window during which idle connections and active connections closing is spread in case of soft-stop. After a SIGUSR1 is received and the grace period is over (if any), the idle connections will all be closed at once if this option is not set, and active HTTP or HTTP2 connections will be ended after the next request is received, either by appending a “Connection: close” line to the HTTP response, or by sending a GOAWAY frame in case of HTTP2. When this option is set, connection closing will be spread over this set <time>. If the close-spread-time is set to “infinite”, active connection closing during a soft-stop will be disabled. The “Connection: close” header will not be added to HTTP responses (or GOAWAY for HTTP2) anymore and idle connections will only be closed once their timeout is reached (based on the various timeouts set in the configuration).

Arguments:

<time>  is a time window (by default in milliseconds) during which
        connection closing will be spread during a soft-stop operation, or
        "infinite" if active connection closing should be disabled.

It is recommended to set this setting to a value lower than the one used in the “hard-stop-after” option if this one is used, so that all connections have a chance to gracefully close before the process stops.

See also: grace, hard-stop-after, idle-close-on-response

cluster-secret <secret>

cluster-secret <secret>

Define an ASCII string secret shared between several nodes belonging to the same cluster. It could be used for different usages. It is at least used to derive stateless reset tokens for all the QUIC connections instantiated by this process. This is also the case to derive secrets used to encrypt Retry tokens.

If this parameter is not set, a random value will be selected on process startup. This allows to use features which rely on it, albeit with some limitations.

cpu-map [auto:]<thread-group>[/<thread-set>] <cpu-set>[,...] [...]

cpu-map [auto:]<thread-group>[/<thread-set>] <cpu-set>[,...] [...]

On some operating systems, it is possible to bind a thread group or a thread to a specific CPU set. This means that the designated threads will never run on other CPUs. The “cpu-map” directive specifies CPU sets for individual threads or thread groups. The first argument is a thread group range, optionally followed by a thread set. These ranges have the following format:

all | odd | even | number[-[number]]

<number> must be a number between 1 and 32 or 64, depending on the machine’s word size. Any group IDs above ’thread-groups’ and any thread IDs above the machine’s word size are ignored. All thread numbers are relative to the group they belong to. It is possible to specify a range with two such number delimited by a dash (’-’). It also is possible to specify all threads at once using “all”, only odd numbers using “odd” or even numbers using “even”, just like with the “thread” bind directive. The second and forthcoming arguments are CPU sets. Each CPU set is either a unique number starting at 0 for the first CPU or a range with two such numbers delimited by a dash (’-’). These CPU numbers and ranges may be repeated by delimiting them with commas or by passing more ranges as new arguments on the same line. Outside of Linux and BSD operating systems, there may be a limitation on the maximum CPU index to either 31 or 63. Multiple “cpu-map” directives may be specified, but each “cpu-map” directive will replace the previous ones when they overlap.

Ranges can be partially defined. The higher bound can be omitted. In such case, it is replaced by the corresponding maximum value, 32 or 64 depending on the machine’s word size.

The prefix “auto:” can be added before the thread set to let HAProxy automatically bind a set of threads to a CPU by incrementing threads and CPU sets. To be valid, both sets must have the same size. No matter the declaration order of the CPU sets, it will be bound from the lowest to the highest bound. Having both a group and a thread range with the “auto:” prefix is not supported. Only one range is supported, the other one must be a fixed number.

Note that group ranges are supported for historical reasons. Nowadays, a lone number designates a thread group and must be 1 if thread-groups are not used, and specifying a thread range or number requires to prepend “1/” in front of it if thread groups are not used. Finally, “1” is strictly equivalent to “1/all” and designates all threads in the group.

Examples:

cpu-map 1/all 0-3 # bind all threads of the first group on the
                  # first 4 CPUs

cpu-map 1/1- 0-   # will be replaced by "cpu-map 1/1-64 0-63"
                  # or "cpu-map 1/1-32 0-31" depending on the machine's
                  # word size.

# all these lines bind thread 1 to the cpu 0, the thread 2 to cpu 1
# and so on.
cpu-map auto:1/1-4   0-3
cpu-map auto:1/1-4   0-1 2-3
cpu-map auto:1/1-4   3 2 1 0
cpu-map auto:1/1-4   3,2,1,0

# bind each thread to exactly one CPU using all/odd/even keyword
cpu-map auto:1/all   0-63
cpu-map auto:1/even  0-31
cpu-map auto:1/odd   32-63

# invalid cpu-map because thread and CPU sets have different sizes.
cpu-map auto:1/1-4   0    # invalid
cpu-map auto:1/1     0-3  # invalid

# map 40 threads of those 4 groups to individual CPUs
cpu-map auto:1/1-10   0-9
cpu-map auto:2/1-10   10-19
cpu-map auto:3/1-10   20-29
cpu-map auto:4/1-10   30-39

# Map 80 threads to one physical socket and 80 others to another socket
# without forcing assignment. These are split into 4 groups since no
# group may have more than 64 threads.
cpu-map 1/1-40   0-39,80-119    # node0, siblings 0 & 1
cpu-map 2/1-40   0-39,80-119
cpu-map 3/1-40   40-79,120-159  # node1, siblings 0 & 1
cpu-map 4/1-40   40-79,120-159

cpu-affinity <affinity>

cpu-affinity <affinity>

Defines how you want threads to be bound to cpus. It currently accepts the following values:

  • per-core: each thread will be bound to all the hardware threads of one core.
  • per-group: each thread will be bound to all the hardware threads of the group. This is the default unless “threads-per-core 1” is used in “cpu-policy”. “per-group” accepts an optional argument, to specify how CPUs should be allocated. When a list of CPUs is larger than the maximum allowed number of CPUs per group and has to be split between multiple groups, an extra option allows to choose how the groups will be bound to those CPUs:
    • auto: each thread group will only be assigned a fair share of contiguous CPU cores that are dedicated to it and not shared with other groups. This is the default as it generally is more optimal.
    • loose: each group will still be allowed to use any CPU in the list. This generally causes more contention, but may sometimes help deal better with parasitic loads running on the same CPUs.
  • auto: “per-group” will be used, unless “threads-per-core 1” is used in “cpu-policy”, in which case “per-core” will be used. This is the default.
  • per-thread: that will bind one thread to one hardware thread only. If “threads-per-core 1” is used in “cpu-policy”, then each thread will be bound to one hardware thread of a different core.
  • per-ccx: each thread will be bound to all the hardware threads of a CCX.

cpu-policy <policy> [threads-per-core 1 | auto]

cpu-policy <policy> [threads-per-core 1 | auto]

Selects the CPU allocation policy to be used.

On multi-CPU systems, there can be plenty of reasons for not using all available CPU cores, and/or for grouping them into different thread groups, for performance, latency, cost, or system-wide resource management. The “cpu-set” directive already allows to evict a number of them, but once done, it is necessary to decide how to assign the remaining ones to threads and thread groups.

This mapping is normally performed using the “cpu-map” directive, though it can be particularly difficult to maintain on heterogeneous systems.

The “cpu-policy” directive chooses between a small number of allocation policies which one to use instead, when “cpu-map” is not used. The following policies are currently supported, with “performance” being the default one:

  • none no particular post-selection is performed. All enabled CPUs will be usable, and if the number of threads is not set, it will be set to the number of available CPUs but no more than 32 for 32-bit systems or 64 for 64-bit systems, per thread-group. The number of thread-groups, if not set, will be set to 1.

  • efficiency exactly like “group-by-ccx” below, except that CPU clusters composed of cores whose performance is more than 25% above that of the next less performant one are evicted. These are typically “big” or “performance” cores. This means that if more than one type of CPU cores are detected, only the efficient one will be used. This can make sense for use with moderate loads when the most powerful cores need to be available to the application or a security component. Some modern CPUs have a large number of such efficient CPU cores which can collectively deliver a decent level of performance while using less power.

  • first-usable-node if the CPUs were not previously restricted at boot (for example using the “taskset” utility), and if the “nbthread” directive was not set, then the first NUMA node with enabled CPUs will be used, and this number of CPUs will be used as the number of threads. A single thread group will be enabled with all of them, within the limit of 32 or 64 depending on the system.

  • group-by-2-ccx same as “group-by-ccx” below but create a group every two CCX. This can make sense on CPUs having many CCX of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between CCX is slow. This is generally recommended against.

  • group-by-2-clusters same as “group-by-cluster” but create a group every two clusters. This can make sense on CPUs having many clusters of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between clusters is slow. This is generally recommended against.

  • group-by-3-ccx same as “group-by-ccx” below but create a group every three CCX. This can make sense on CPUs having many CCX of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between CCX is slow. This is generally recommended against.

  • group-by-3-clusters same as “group-by-cluster” but create a group every three clusters. This can make sense on CPUs having many clusters of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between clusters is slow. This is generally recommended against.

  • group-by-4-ccx same as “group-by-ccx” below but create a group every four CCX. This can make sense on CPUs having many CCX of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between CCX is slow. This is generally recommended against.

  • group-by-4-clusters same as “group-by-cluster” but create a group every four clusters. This can make sense on CPUs having many clusters of few cores each, to avoid creating many groups, or to smooth the distribution a little bit when not all cores are in use. Please note that it can have very bad performance effects when the communication between clusters is slow. This is generally recommended against.

  • group-by-ccx if neither “nbthread” not “nbtgroups” were set, then one thread group is created for each CPU core complex (“CCX”) with available CPUs, each with as many threads as CPUs. A CCX groups CPUs having a similarly fast access to the last level cache (“LLC”), typically the L3 cache. On most modern machines, it is critical for performance not to mix CPUs from distant CCX in the same thread group. All threads of a group are then bound to all CPUs of the CCX so that intra-group communications remain local to the CCX without enforcing too strong a binding. The per-group thread limits and thread-group limits are respected. This is recommended on multi-socket and NUMA systems, as well as CPUs with bad inter-CCX latencies.

  • group-by-cluster if neither “nbthread” not “nbtgroups” were set, then one thread group is created for each CPU cluster with available CPUs, each with as many threads as CPUs. All threads of a group are bound to all CPUs of the cluster so that intra-group communications remain local to the cluster without enforcing too strong a binding. The per-group thread limits and thread-group limits are respected. This is recommended on multi-socket and NUMA systems, as well as CPUs with bad inter-CCX latencies. On most server machines, clusters and CCX are the same, but on heterogeneous machines (“performance” vs “efficiency” or “big” vs “little”), a cluster will generally be made of only a part of a CCX composed only of very similar CPUs (same type, +/-5% frequency difference max). The difference is visible on modern laptops and desktop machines used by developers and admins to validate setups.

  • performance exactly like “group-by-ccx” above, except that CPU clusters composed of cores whose performance is less than 80% of those of the next more performant one are evicted. These are typically “little” or “efficient” cores, whose addition generally doesn’t bring significant gains and can easily be counter-productive (e.g. TLS handshakes). Often, keeping such cores for other tasks such as network handling is much more effective. On development systems, these can also be used to run auxiliary tools such as load generators and monitoring tools. This is the default policy.

  • resource this is like “group-by-cluster” above, except that only the smallest and most efficient CPU cluster will be used, while all other ones will be ignored. This can be used to limit the resource usage to the strict minimum that still delivers decent performance, for example to try to further reduce power consumption or minimize the number of cores needed on some rented systems for a sidecar setup, in order to scale the system down more easily. Note that if a single cluster is present, it will still be fully used.

An optional keyword can be added, “threads-per-core”. It can accept two values, “1” and “auto”. If set to 1, then only one thread per core will be created, unrespective of how many hardware threads the core has. If set to auto, then one thread per hardware thread will be created. If no affinity is specified, and threads-per-core 1 is used, then by default the affinity will be per-core.

See also: “cpu-map”, “cpu-set”, “nbthread”

cpu-set <directive>...

cpu-set <directive>...

Allows to symbolically describe what sets of CPUs to run on. The directive supports the following keyword: - reset this undoes any previous limitation that could have been inherited by a service manager or a “taskset” command for example. - drop-cpu <set> do not bind to CPUs in this set - only-cpu <set> do not bind to CPUs not in this set - drop-node <set> do not bind to CPUs belonging to this NUMA node - only-node <set> do not bind to CPUs not belonging to this NUMA node - drop-cluster <set> do not bind to CPUs on this hardware cluster number - only-cluster <set> do not bind to CPUs on other hardware cluster number - drop-core <set> do not bind to CPUs on this hardware core number - only-core <set> do not bind to CPUs on other hardware core number - drop-thread <set> do not bind to CPUs on this hardware thread number - only-thread <set> do not bind to CPUs on other hardware thread number See also: “cpu-policy”

crt-base <dir>

crt-base <dir>

Assigns a default directory to fetch SSL certificates from when a relative path is used with “crtfile” or “crt” directives. Absolute locations specified prevail and ignore “crt-base”.

daemon

daemon

Makes the process fork into background. This is the recommended mode of operation. It is equivalent to the command line “-D” argument. It can be disabled by the command line “-db” argument. This option is ignored in systemd mode.

default-path { current | config | parent | origin <path> }

default-path { current | config | parent | origin <path> }

By default HAProxy loads all files designated by a relative path from the location the process is started in. In some circumstances it might be desirable to force all relative paths to start from a different location just as if the process was started from such locations. This is what this directive is made for. Technically it will perform a temporary chdir() to the designated location while processing each configuration file, and will return to the original directory after processing each file. It takes an argument indicating the policy to use when loading files whose path does not start with a slash (’/’): - “current” indicates that all relative files are to be loaded from the directory the process is started in; this is the default.

- "config" indicates that all relative files should be loaded from the
  directory containing the configuration file. More specifically, if the
  configuration file contains a slash ('/'), the longest part up to the
  last slash is used as the directory to change to, otherwise the current
  directory is used. This mode is convenient to bundle maps, errorfiles,
  certificates and Lua scripts together as relocatable packages. When
  multiple configuration files are loaded, the directory is updated for
  each of them.

- "parent" indicates that all relative files should be loaded from the
  parent of the directory containing the configuration file. More
  specifically, if the configuration file contains a slash ('/'), ".."
  is appended to the longest part up to the last slash is used as the
  directory to change to, otherwise the directory is "..". This mode is
  convenient to bundle maps, errorfiles,  certificates and Lua scripts
  together as relocatable packages, but where each part is located in a
  different subdirectory (e.g. "config/", "certs/", "maps/", ...).

- "origin" indicates that all relative files should be loaded from the
  designated (mandatory) path. This may be used to ease management of
  different HAProxy instances running in parallel on a system, where each
  instance uses a different prefix but where the rest of the sections are
  made easily relocatable.

Each “default-path” directive instantly replaces any previous one and will possibly result in switching to a different directory. While this should always result in the desired behavior, it is really not a good practice to use multiple default-path directives, and if used, the policy ought to remain consistent across all configuration files.

Warning: some configuration elements such as maps or certificates are uniquely identified by their configured path. By using a relocatable layout, it becomes possible for several of them to end up with the same unique name, making it difficult to update them at run time, especially when multiple configuration files are loaded from different directories. It is essential to observe a strict collision-free file naming scheme before adopting relative paths. A robust approach could consist in prefixing all files names with their respective site name, or in doing so at the directory level.

description <text>

description <text>

Add a text that describes the instance.

Please note that it is required to escape certain characters (# for example) and this text is inserted into a html page so you should avoid using “<” and “>” characters.

deviceatlas-json-file <path>

deviceatlas-json-file <path>

Sets the path of the DeviceAtlas JSON data file to be loaded by the API. The path must be a valid JSON data file and accessible by HAProxy process.

deviceatlas-log-level <value>

deviceatlas-log-level <value>

Sets the level of information returned by the API. This directive is optional and set to 0 by default if not set.

deviceatlas-properties-cookie <name>

deviceatlas-properties-cookie <name>

Sets the client cookie’s name used for the detection if the DeviceAtlas Client-side component was used during the request. This directive is optional and set to DAPROPS by default if not set.

deviceatlas-separator <char>

deviceatlas-separator <char>

Sets the character separator for the API properties results. This directive is optional and set to | by default if not set.

dns-accept-family <family>[,...]

dns-accept-family <family>[,...]

By default, DNS resolvers accept both IPv4 and IPv6 addresses. This can be influenced by the “resolve-prefer” keywords on server lines as well as the family argument to the “do-resolve” action, but that is only a preference, which does not block the other family from being used when it’s alone. In some environments where dual-stack is not usable, stumbling on an unreachable IPv6-only DNS record can cause significant trouble as it will replace a previous IPv4 one which would possibly have continued to work till next request. The “dns-accept-family” global option permits to enforce usage of only one (or both) address families. The argument is a comma-delimited list of the following words: - “ipv4”: query and accept IPv4 addresses (“A” records) - “ipv6”: query and accept IPv6 addresses (“AAAA” records) - “auto”: use IPv4, and IPv6 if the system has a default gateway for it. The result of the last check is cached for 30 seconds.

When a single family is used, no request will be sent to resolvers for the other family, and any response for the other family will be ignored. The default value since 3.3 is “auto”, which effectively enables both families only once IPv6 has been proven to be routable, otherwise sticks to IPv4. See also: “resolve-prefer”, “do-resolve”

expose-deprecated-directives

expose-deprecated-directives

This statement must appear before using some directives tagged as deprecated to silent warnings and make sure the config file will not be rejected. Not all deprecated directives are concerned, only those without any alternative solution.

expose-experimental-directives

expose-experimental-directives

This statement must appear before using directives tagged as experimental or the config file will be rejected. Please note that features covered by this option are not guaranteed to work well and may break during the maintenance cycle. Developers will maintain them in best effort mode while the next version is being worked on, and will deploy any reasonable effort to avoid breaking them but with no guarantee. For these reasons, these features are not expected to be supported beyond the release of the next LTS release. Users who want to try experimental features are expected to upgrade quickly to benefit from the improvements made to that feature. In order to know if this directive is still needed, it’s easy: if it is enabled without being used by any such feature, a warning will be emitted suggesting to turn it off. So without any warning, it means it’s still needed.

external-check [preserve-env]

external-check [preserve-env]

Allows the use of an external agent to perform health checks. This is disabled by default as a security precaution, and even when enabled, checks may still fail unless “insecure-fork-wanted” is enabled as well. If the program launched makes use of a setuid executable (it should really not), you may also need to set “insecure-setuid-wanted” in the global section. By default, the checks start with a clean environment which only contains variables defined in the “external-check” command in the backend section. It may sometimes be desirable to preserve the environment though, for example when complex scripts retrieve their extra paths or information there. This can be done by appending the “preserve-env” keyword. In this case however it is strongly advised not to run a setuid nor as a privileged user, as this exposes the check program to potential attacks. See “option external-check”, and “insecure-fork-wanted”, and “insecure-setuid-wanted” for extra details.

fd-hard-limit <number>

fd-hard-limit <number>

Sets an upper bound to the maximum number of file descriptors that the process will use, regardless of system limits. While “ulimit-n” and “maxconn” may be used to enforce a value, when they are not set, the process will be limited to the hard limit of the RLIMIT_NOFILE setting as reported by “ulimit -n -H”. But some modern operating systems are now allowing extremely large values here (in the order of 1 billion), which will consume way too much RAM for regular usage. The fd-hard-limit setting is provided to enforce a possibly lower bound to this limit. This means that it will always respect the system-imposed limits when they are below <number> but the specified value will be used if system-imposed limits are higher. By default fd-hard-limit is set to 1048576. This default could be changed via DEFAULT_MAXFD compile-time variable, that could serve as the maximum (kernel) system limit, if RLIMIT_NOFILE hard limit is extremely large. fd-hard-limit set in global section allows to temporarily override the value provided via DEFAULT_MAXFD at the build-time. In the example below, no other setting is specified and the maxconn value will automatically adapt to the lower of “fd-hard-limit” and the RLIMIT_NOFILE limit:

global
    # use as many FDs as possible but no more than 50000
    fd-hard-limit 50000

See also: ulimit-n, maxconn

gid <number>

gid <number>

Changes the process’s group ID to <number>. It is recommended that the group ID is dedicated to HAProxy or to a small set of similar daemons. HAProxy must be started with a user belonging to this group, or with superuser privileges. Note that if HAProxy is started from a user having supplementary groups, it will only be able to drop these groups if started with superuser privileges. See also “group” and “uid”.

grace <time>

grace <time>

Defines a delay between SIGUSR1 and real soft-stop.

Arguments:

<time>  is an extra delay (by default in milliseconds) after receipt of the
        SIGUSR1 signal that will be waited for before proceeding with the
        soft-stop operation.

This is used for compatibility with legacy environments where the haproxy process needs to be stopped but some external components need to detect the status before listeners are unbound. The principle is that the internal “stopping” variable (which is reported by the “stopping” sample fetch function) will be turned to true, but listeners will continue to accept connections undisturbed, until the delay expires, after what the regular soft-stop will proceed. This must not be used with processes that are reloaded, or this will prevent the old process from unbinding, and may prevent the new one from starting, or simply cause trouble.

Example:

global
  grace 10s

# Returns 200 OK until stopping is set via SIGUSR1
frontend ext-check
  bind:9999
  monitor-uri /ext-check
  monitor fail if { stopping }

Please note that a more flexible and durable approach would instead consist for an orchestration system in setting a global variable from the CLI, use that variable to respond to external checks, then after a delay send the SIGUSR1 signal.

Example:

# Returns 200 OK until proc.stopping is set to non-zero. May be done
# from HTTP using set-var(proc.stopping) or from the CLI using:
# > set var proc.stopping int(1)
frontend ext-check
  bind:9999
  monitor-uri /ext-check
  monitor fail if { var(proc.stopping) -m int gt 0 }

See also: hard-stop-after, monitor

group <group name>

group <group name>

Similar to “gid” but uses the GID of group name <group name> from /etc/group. See also “gid” and “user”.

h1-accept-payload-with-any-method

h1-accept-payload-with-any-method

Does not reject HTTP/1.0 GET/HEAD/DELETE requests with a payload with a 413 Payload Too Large HTTP response.

While It is explicitly allowed in HTTP/1.1, HTTP/1.0 is not clear on this point and some old servers don’t expect any payload and never look for body length (via Content-Length or Transfer-Encoding headers). It means that some intermediaries may properly handle the payload for HTTP/1.0 GET/HEAD/DELETE requests, while some others may totally ignore it. That may lead to security issues because a request smuggling attack is possible. Thus, by default, HAProxy rejects HTTP/1.0 GET/HEAD/DELETE requests with a payload.

However, it may be an issue with some old clients. In this case, this global option may be set.

h1-do-not-close-on-insecure-transfer-encoding

h1-do-not-close-on-insecure-transfer-encoding

As mandated by the HTTP/1.1 specification (RFC9112#6.1), the presence of both a Transfer-Encoding header field and a Content-Length header field in the same message represents a serious risk of conveying a content smuggling attack if there are any HTTP/1.0 agent anywhere in the upstream of downstream chain, and when facing this, an agent must absolutely close the connection after the response so as to prevent any exploitation. But this may have a performance impact on some very old clients, especially if they need to renegotiate a TLS connection for every request. This option is present to ask HAProxy not to enforce this rule, and to just sanitize the message but leave the connection alive after the response. This may only be done when absolutely certain that no HTTP/1.0 agents are present in the chain and that all implementations before HAProxy are fully HTTP/1.1 compliant regarding the rules that apply to these header fields. In any case, HAProxy will continue to ignore and drop the extraneous Content-Length header so as not to confuse the next hop.

When enabling this option to work around an old broken client or server, it is important to understand that regardless of the need or not for this option, such an agent violating this rule faces a risk to see its messages truncated by old agents that would consider Content-Length and ignore Transfer-Encoding, since the cumulated size of the encoded chunk sizes are not being accounted for. As such, the rule above is not just a matter of security but also of taking care of getting rid of agents that may face communication trouble due to incompatibilities with older ones.

h1-case-adjust <from> <to>

h1-case-adjust <from> <to>

Defines the case adjustment to apply, when enabled, to the header name <from>, to change it to <to> before sending it to HTTP/1 clients or servers. <from> must be in lower case, and <from> and <to> must not differ except for their case. It may be repeated if several header names need to be adjusted. Duplicate entries are not allowed. If a lot of header names have to be adjusted, it might be more convenient to use “h1-case-adjust-file”. Please note that no transformation will be applied unless “option h1-case-adjust-bogus-client” or “option h1-case-adjust-bogus-server” is specified in a proxy.

There is no standard case for header names because, as stated in RFC7230, they are case-insensitive. So applications must handle them in a case-insensitive manner. But some bogus applications violate the standards and erroneously rely on the cases most commonly used by browsers. This problem becomes critical with HTTP/2 because all header names must be exchanged in lower case, and HAProxy follows the same convention. All header names are sent in lower case to clients and servers, regardless of the HTTP version.

Applications which fail to properly process requests or responses may require to temporarily use such workarounds to adjust header names sent to them for the time it takes the application to be fixed. Please note that an application which requires such workarounds might be vulnerable to content smuggling attacks and must absolutely be fixed.

Example:

global
  h1-case-adjust content-length Content-Length

See “h1-case-adjust-file”, “option h1-case-adjust-bogus-client” and “option h1-case-adjust-bogus-server”.

h1-case-adjust-file <hdrs-file>

h1-case-adjust-file <hdrs-file>

Defines a file containing a list of key/value pairs used to adjust the case of some header names before sending them to HTTP/1 clients or servers. The file <hdrs-file> must contain 2 header names per line. The first one must be in lower case and both must not differ except for their case. Lines which start with ‘#’ are ignored, just like empty lines. Leading and trailing tabs and spaces are stripped. Duplicate entries are not allowed. Please note that no transformation will be applied unless “option h1-case-adjust-bogus-client” or “option h1-case-adjust-bogus-server” is specified in a proxy.

If this directive is repeated, only the last one will be processed. It is an alternative to the directive “h1-case-adjust” if a lot of header names need to be adjusted. Please read the risks associated with using this.

See “h1-case-adjust”, “option h1-case-adjust-bogus-client” and “option h1-case-adjust-bogus-server”.

h2-workaround-bogus-websocket-clients

h2-workaround-bogus-websocket-clients

This disables the announcement of the support for h2 websockets to clients. This can be use to overcome clients which have issues when implementing the relatively fresh RFC8441, such as Firefox 88. To allow clients to automatically downgrade to http/1.1 for the websocket tunnel, specify h2 support on the bind line using “alpn” without an explicit “proto” keyword. If this statement was previously activated, this can be disabled by prefixing the keyword with “no”.

hard-stop-after <time>

hard-stop-after <time>

Defines the maximum time allowed to perform a clean soft-stop.

Arguments:

<time>  is the maximum time (by default in milliseconds) for which the
        instance will remain alive when a soft-stop is received via the
        SIGUSR1 signal.

This may be used to ensure that the instance will quit even if connections remain opened during a soft-stop (for example with long timeouts for a proxy in tcp mode). It applies both in TCP and HTTP mode.

Example:

global
  hard-stop-after 30s

See also: grace

harden.reject-privileged-ports.tcp { on | off }

harden.reject-privileged-ports.tcp { on | off }
harden.reject-privileged-ports.quic { on | off }

Toggle per protocol protection which forbid communication with clients which use privileged ports as their source port. This range of ports is defined according to RFC 6335. By default, protection is active for QUIC protocol as this behavior is suspicious and may be used as a spoofing or DNS/NTP amplification attack.

http-err-codes [+-]<range>[,...] [...]

http-err-codes [+-]<range>[,...] [...]

Replace, reduce or extend the list of status codes that define an error as considered by the termination codes and the “http_err_cnt” counter in stick tables. The default range for errors is 400 to 499, but in certain contexts some users prefer to exclude specific codes, especially when tracking client errors (e.g. 404 on systems with dynamically generated contents). See also “http-fail-codes” and “http_err_cnt”.

A range specified without ‘+’ nor ‘-’ redefines the existing range to the new one. A range starting with ‘+’ extends the existing range to also include the specified one, which may or may not overlap with the existing one. A range starting with ‘-’ removes the specified range from the existing one. A range consists in a number from 100 to 599, optionally followed by “-” followed by another number greater than or equal to the first one to indicate the high boundary of the range. Multiple ranges may be delimited by commas for a same add/del/ replace operation.

Example:

http-err-codes 400,402-444,446-480,490   # sets exactly these codes
http-err-codes 400-499 -450 +500         # sets 400 to 500 except 450
http-err-codes -450-459                  # removes 450 to 459 from range
http-err-codes +501,505                  # adds 501 and 505 to range

http-fail-codes [+-]<range>[,...] [...]

http-fail-codes [+-]<range>[,...] [...]

Replace, reduce or extend the list of status codes that define a failure as considered by the termination codes and the “http_fail_cnt” counter in stick tables. The default range for failures is 500 to 599 except 501 and 505 which can be triggered by clients, and normally indicate a failure from the server to process the request. Some users prefer to exclude certain codes in certain contexts where it is known they’re not relevant, such as 500 in certain SOAP environments as it doesn’t translate a server fault there. The syntax is exactly the same as for http-err-codes above. See also “http-err-codes” and “http_fail_cnt”.

insecure-fork-wanted

insecure-fork-wanted

By default HAProxy tries hard to prevent any thread and process creation after it starts. Doing so is particularly important when using Lua files of uncertain origin, and when experimenting with development versions which may still contain bugs whose exploitability is uncertain. And generally speaking it’s good hygiene to make sure that no unexpected background activity can be triggered by traffic. But this prevents external checks from working, and may break some very specific Lua scripts which actively rely on the ability to fork. This option is there to disable this protection. Note that it is a bad idea to disable it, as a vulnerability in a library or within HAProxy itself will be easier to exploit once disabled. In addition, forking from Lua or anywhere else is not reliable as the forked process may randomly embed a lock set by another thread and never manage to finish an operation. As such it is highly recommended that this option is never used and that any workload requiring such a fork be reconsidered and moved to a safer solution (such as agents instead of external checks). This option supports the “no” prefix to disable it. This can also be activated with “-dI” on the haproxy command line.

insecure-setuid-wanted

insecure-setuid-wanted

HAProxy doesn’t need to call executables at run time (except when using external checks which are strongly recommended against), and is even expected to isolate itself into an empty chroot. As such, there basically is no valid reason to allow a setuid executable to be called without the user being fully aware of the risks. In a situation where HAProxy would need to call external checks and/or disable chroot, exploiting a vulnerability in a library or in HAProxy itself could lead to the execution of an external program. On Linux it is possible to lock the process so that any setuid bit present on such an executable is ignored. This significantly reduces the risk of privilege escalation in such a situation. This is what HAProxy does by default. In case this causes a problem to an external check (for example one which would need the “ping” command), then it is possible to disable this protection by explicitly adding this directive in the global section. If enabled, it is possible to turn it back off by prefixing it with the “no” keyword.

issuers-chain-path <dir>

issuers-chain-path <dir>

Assigns a directory to load certificate chain for issuer completion. All files must be in PEM format. For certificates loaded with “crt” or “crt-list”, if certificate chain is not included in PEM (also commonly known as intermediate certificate), HAProxy will complete chain if the issuer of the certificate corresponds to the first certificate of the chain loaded with “issuers-chain-path”. A “crt” file with PrivateKey+Certificate+IntermediateCA2+IntermediateCA1 could be replaced with PrivateKey+Certificate. HAProxy will complete the chain if a file with IntermediateCA2+IntermediateCA1 is present in “issuers-chain-path” directory. All other certificates with the same issuer will share the chain in memory.

The OCSP features are able to use the completed chain when no .issuer was used, or no chain was provided in the PEM.

jwt.decrypt_alg_list <list>

jwt.decrypt_alg_list <list>

Set the list of algorithms allowed in the jwt_decrypt_XXX converters. JWT tokens using an unsupported or disabled algorithms will never be decrypted. The specified algorithms must have the same format as in section 4.1 of RFC7518 and must be colon-separated. The special “ALL” name can be used to enable all the supported algorithms (see “jwt_decrypt_jwk” converter for a complete list) and a ‘!’ can be appended to an algorithm name to explicitly disable it. Please note that unless “ALL” is specified, using this option will disable any algorithm that is not explicitly mentioned in the provided list.

Examples:

# Enable all algorithms but the "ECDH-ES" one
jwt.decrypt_alg_list ALL:!ECDH-ES

# Only enable ECDH-ES algorithms
jwt.decrypt_alg_list ECDH-ES:ECDH-ES+A128KW:ECDH-ES+A192KW:ECDH-ES+A256KW

jwt.decrypt_enc_list <list>

jwt.decrypt_enc_list <list>

Set the list of encryption algorithms allowed in the jwt_decrypt_XXX converters. JWT tokens using an unsupported or disabled encryption algorithms will never be decrypted. The specified algorithms must have the same format as in section 5.1 of RFC7518 and must be colon-separated. The special “ALL” name can be used to enable all the supported algorithms (see “jwt_decrypt_jwk” converter for a complete list) and a ‘!’ can be appended to an algorithm name to explicitly disable it. Please note that unless “ALL” is specified, using this option will disable any algorithm that is not explicitly mentioned in the provided list.

Examples:

# Enable only AES GCM encrypting algorithms
jwt.decrypt_enc_list A128GCM:A192GCM:A256GCM

key-base <dir>

key-base <dir>

Assigns a default directory to fetch SSL private keys from when a relative path is used with “key” directives. Absolute locations specified prevail and ignore “key-base”. This option only works with a crt-store load line.

limited-quic

limited-quic

This setting must be used to explicitly enable the QUIC listener bindings when haproxy is compiled with a version of OpenSSL without QUIC support. It activates an haproxy internal compatibility layer which must have been selected at build time with USE_QUIC_OPENSSL_COMPAT=1. This compatibility layer supports most of the necessary TLS operations, albeit without QUIC 0-RTT capability.

This feature is primarily targeted for OpenSSL prior to version 3.5.2, where QUIC API was not implemented or only partially. The compatibility layer can still be activated for version 3.5.2 and above, but this is probably unnecessary.

If limited-quic is set but the compatibility layer was not selected at build time, the option is silently ignored and QUIC TLS operations rely on the TLS library.

localpeer <name>

localpeer <name>

Sets the local instance’s peer name. It will be ignored if the “-L” command line argument is specified or if used after “peers” section definitions. In such cases, a warning message will be emitted during the configuration parsing.

This option will also set the HAPROXY_LOCALPEER environment variable. See also “-L” in the management guide and “peers” section below.

log <target> [len <length>] [format <format>] [sample <ranges>:<sample_size>]

log <target> [len <length>] [format <format>] [sample <ranges>:<sample_size>]
    [profile <prof>] <facility> [max level [min level]]

Adds a global syslog server. Several global servers can be defined. They will receive logs for starts and exits, as well as all logs from proxies configured with “log global”. See “log” option for proxies for more details.

log-send-hostname [<string>]

log-send-hostname [<string>]

Sets the hostname field in the syslog header. If optional “string” parameter is set the header is set to the string contents, otherwise uses the hostname of the system. Generally used if one is not relaying logs through an intermediate syslog server or for simply customizing the hostname printed in the logs.

log-tag <string>

log-tag <string>

Sets the tag field in the syslog header to this string. It defaults to the program name as launched from the command line, which usually is “haproxy”. Sometimes it can be useful to differentiate between multiple processes running on the same host. See also the per-proxy “log-tag” directive.

lua-load <file> [ <arg1> [ <arg2> [ ... ] ] ]

lua-load <file> [ <arg1> [ <arg2> [ ... ] ] ]

This global directive loads and executes a Lua file in the shared context that is visible to all threads. Any variable set in such a context is visible from any thread. This is the easiest and recommended way to load Lua programs but it will not scale well if a lot of Lua calls are performed, as only one thread may be running on the global state at a time. A program loaded this way will always see 0 in the “core.thread” variable. This directive can be used multiple times.

args are available in the lua file using the code below in the body of the file. Do not forget that Lua arrays start at index 1. A “local” variable declared in a file is available in the entire file and not available on other files.

 local args = table.pack(...)

lua-load-per-thread <file> [ <arg1> [ <arg2> [ ... ] ] ]

lua-load-per-thread <file> [ <arg1> [ <arg2> [ ... ] ] ]

This global directive loads and executes a Lua file into each started thread. Any global variable has a thread-local visibility so that each thread could see a different value. As such it is strongly recommended not to use global variables in programs loaded this way. An independent copy is loaded and initialized for each thread, everything is done sequentially and in the thread’s numeric order from 1 to nbthread. If some operations need to be performed only once, the program should check the “core.thread” variable to figure what thread is being initialized. Programs loaded this way will run concurrently on all threads and will be highly scalable. This is the recommended way to load simple functions that register sample-fetches, converters, actions or services once it is certain the program doesn’t depend on global variables. For the sake of simplicity, the directive is available even if only one thread is used and even if threads are disabled (in which case it will be equivalent to lua-load). This directive can be used multiple times.

See lua-load for usage of args.

lua-prepend-path <string> [<type>]

lua-prepend-path <string> [<type>]

Prepends the given string followed by a semicolon to Lua’s package.<type> variable. <type> must either be “path” or “cpath”. If <type> is not given it defaults to “path”.

Lua’s paths are semicolon delimited lists of patterns that specify how the require function attempts to find the source file of a library. Question marks (?) within a pattern will be replaced by module name. The path is evaluated left to right. This implies that paths that are prepended later will be checked earlier.

As an example by specifying the following path:

lua-prepend-path /usr/share/haproxy-lua/?/init.lua
lua-prepend-path /usr/share/haproxy-lua/?.lua

When require "example" is being called Lua will first attempt to load the /usr/share/haproxy-lua/example.lua script, if that does not exist the /usr/share/haproxy-lua/example/init.lua will be attempted and the default paths if that does not exist either.

See https://www.lua.org/pil/8.1.html for the details within the Lua documentation.

master-worker (deprecated)

master-worker (deprecated)

Master-worker mode. It is equivalent to the command line “-W” argument.

This keyword is deprecated, please start in master-worker mode using “-W” or “-Ws”.

This mode will launch a “master” which will fork a “worker” after reading the configuration to process the traffic. The master is used as a process manager which will monitor the “workers”.

Using this mode, you can reload HAProxy directly by sending a SIGUSR2 signal to the master. Reloading will ask the master to read the configuration again and fork a new worker. The previous worker will be kept until the end of its jobs.

The master-worker mode is compatible either with the foreground or daemon mode.

By default, if a worker exits with a bad return code, in the case of a segfault for example, all workers will be killed, and the master will leave. It is convenient to combine this behavior with Restart=on-failure in a systemd unit file in order to relaunch the whole process. If you don’t want this behavior, you must use the keyword “no-exit-on-failure”.

See also “-W” in the management guide.

master-worker no-exit-on-failure

master-worker no-exit-on-failure

In master-worker mode, by default, if a worker exits with a bad return code, in the case of a segfault for example, all workers will be killed, and the master will leave. It is convenient to combine this behavior with Restart=on-failure in a systemd unit file in order to relaunch the whole process.

This keyword allows to keep the remaining processes alive when a worker crashed instead of killing everything. This need to be used with caution as it is only meant for debugging and could put the master process in an abnormal state.

max-threads-per-group <number>

max-threads-per-group <number>

Defines the maximum number of threads in a thread group. Unless the number of thread groups is fixed with the “thread-groups” directive, haproxy will create as many thread groups as needed to satisfy the requested number of threads. The minimum value is 1, and the maximum value is 64 (on 64-bit systems, or 32 on 32-bit systems). Lower values reduce contention caused by atomic operations on shared states, but can increase the number of sockets needed to create all listeners and to hold idle backend connections. Higher values will reduce these costs, at the expense of higher CPU usage under contented situations, and lower connection rates. The default value is 16, which provides the best tradeoff that was experimentally found on various tested systems, including x86_64 processors from multiple vendors, and large Arm64 systems, both on bare metal and hypervisors.

mworker-max-reloads <number>

mworker-max-reloads <number>

In master-worker mode, this option limits the number of time a worker can survive to a reload. If the worker did not leave after a reload, once its number of reloads is greater than this number, the worker will receive a SIGTERM. This option helps to keep under control the number of workers. See also “show proc” in the Management Guide.

By default this value is set to 50.

nbthread <number>

nbthread <number>

This setting is only available when support for threads was built in. It makes HAProxy run on <number> threads. “nbthread” also works when HAProxy is started in foreground. On some platforms supporting CPU affinity, the default “nbthread” value is automatically set to the number of CPUs the process is bound to upon startup. This means that the thread count can easily be adjusted from the calling process using commands like “taskset” or “cpuset”. Otherwise, this value defaults to 1. The default value is reported in the output of “haproxy -vv”. Note that values set here or automatically detected are subject to the limit set by “thread-hard-limit” (if set).

numa-cpu-mapping

numa-cpu-mapping

When running on a NUMA-aware platform, this enables the “cpu-policy” directive to inspect the topology and figure the best set of CPUs to use and the corresponding number of threads. However, if the applied binding is non optimal on a particular architecture, it can be disabled with the statement ’no numa-cpu-mapping’. This automatic binding is also not applied if a ’nbthread’ statement is present in the configuration, if the affinity of the process is already specified, for example via the ‘cpu-map’ directive or the taskset utility, or if the cpu-policy is set to any other value. See also “cpu-map”, “cpu-policy”, “cpu-set”.

ocsp-update.disable [ on | off ]

ocsp-update.disable [ on | off ]

Disable completely the ocsp-update in HAProxy. Any ocsp-update configuration will be ignored. Default is “off”. See option “ocsp-update” for more information about the auto update mechanism.

ocsp-update.httpproxy <address>[:port]

ocsp-update.httpproxy <address>[:port]

Allow to use an HTTP proxy for the OCSP updates. This only works with HTTP, HTTPS is not supported. This option will allow the OCSP updater to send absolute URI in the request to the proxy.

ocsp-update.maxdelay <number>

ocsp-update.maxdelay <number>
tune.ssl.ocsp-update.maxdelay <number> (deprecated)

Sets the maximum interval between two automatic updates of the same OCSP response. This time is expressed in seconds and defaults to 3600 (1 hour). It must be set to a higher value than “ocsp-update.mindelay”. See option “ocsp-update” for more information about the auto update mechanism.

ocsp-update.mindelay <number>

ocsp-update.mindelay <number>
tune.ssl.ocsp-update.mindelay <number> (deprecated)

Sets the minimum interval between two automatic updates of the same OCSP response. This time is expressed in seconds and defaults to 300 (5 minutes). It is particularly useful for OCSP response that do not have explicit expiration times. It must be set to a lower value than “ocsp-update.maxdelay”. See option “ocsp-update” for more information about the auto update mechanism.

ocsp-update.mode [ on | off ]

ocsp-update.mode [ on | off ]

Sets the default ocsp-update mode for all certificates used in the configuration. This global option can be superseded by the crt-list “ocsp-update” option. This option is set to “off” by default. See option “ocsp-update” for more information about the auto update mechanism.

pidfile <pidfile>

pidfile <pidfile>

Writes PIDs of all daemons into file <pidfile> when daemon mode or writes PID of master process into file <pidfile> when master-worker mode. This option is equivalent to the “-p” command line argument. The file must be accessible to the user starting the process. See also “daemon” and “master-worker”.

pp2-never-send-local

pp2-never-send-local

A bug in the PROXY protocol v2 implementation was present in HAProxy up to version 2.1, causing it to emit a PROXY command instead of a LOCAL command for health checks. This is particularly minor but confuses some servers’ logs. Sadly, the bug was discovered very late and revealed that some servers which possibly only tested their PROXY protocol implementation against HAProxy fail to properly handle the LOCAL command, and permanently remain in the “down” state when HAProxy checks them. When this happens, it is possible to enable this global option to revert to the older (bogus) behavior for the time it takes to contact the affected components’ vendors and get them fixed. This option is disabled by default and acts on all servers having the “send-proxy-v2” statement.

presetenv <name> <value>

presetenv <name> <value>

Sets environment variable <name> to value <value>. If the variable exists, it is NOT overwritten. The changes immediately take effect so that the next line in the configuration file sees the new value. See also “setenv”, “resetenv”, and “unsetenv”.

prealloc-fd

prealloc-fd

Performs a one-time open of the maximum file descriptor which results in a pre-allocation of the kernel’s data structures. This prevents short pauses when nbthread>1 and HAProxy opens a file descriptor which requires the kernel to expand its data structures.

resetenv [<name> ...]

resetenv [<name> ...]

Removes all environment variables except the ones specified in argument. It allows to use a clean controlled environment before setting new values with setenv or unsetenv. Please note that some internal functions may make use of some environment variables, such as time manipulation functions, but also OpenSSL or even external checks. This must be used with extreme care and only after complete validation. The changes immediately take effect so that the next line in the configuration file sees the new environment. See also “setenv”, “presetenv”, and “unsetenv”.

server-state-base <directory>

server-state-base <directory>

Specifies the directory prefix to be prepended in front of all servers state file names which do not start with a ‘/’. See also “server-state-file”, “load-server-state-from-file” and “server-state-file-name”.

server-state-file <file>

server-state-file <file>

Specifies the path to the file containing state of servers. If the path starts with a slash (’/’), it is considered absolute, otherwise it is considered relative to the directory specified using “server-state-base” (if set) or to the current directory. Before reloading HAProxy, it is possible to save the servers’ current state using the stats command “show servers state”. The output of this command must be written in the file pointed by <file>. When starting up, before handling traffic, HAProxy will read, load and apply state for each server found in the file and available in its current running configuration. See also “server-state-base” and “show servers state”, “load-server-state-from-file” and “server-state-file-name”

set-dumpable [ on | off | libs ]

set-dumpable [ on | off | libs ]

This option helps choose the core dump behavior in case of process crash. Available options are:

  • on this enables core dumping at the process level if it was previously disabled.

  • off this disables a previously enabled core dumping.

  • libs this enables core dumping with an embedded copy of the binaries and libraries that are required for debugging. This may be requested by developers. In this case haproxy will try to load the libraries it depends on into memory and keep them preciously. If the process crashes, they will be dumped into the core so there is no need for retrieving them from the file system anymore and no risk that they do not match the core. This takes a few megabytes to a few tens of megabytes of additional RAM, so it is better not to use it on small systems.

This option is better left disabled by default and enabled only upon a developer’s request. By default it is disabled. Without argument, it defaults to “on”. If it has been enabled, it may still be forcibly disabled by prefixing it with the “no” keyword or by setting it to “off”. It has no impact on performance nor stability but will try hard to re-enable core dumps that were possibly disabled by file size limitations (ulimit -f), core size limitations (ulimit -c), or “dumpability” of a process after changing its UID/GID (such as /proc/sys/fs/suid_dumpable on Linux). Core dumps might still be limited by the current directory’s permissions (check what directory the file is started from), the chroot directory’s permission (it may be needed to temporarily disable the chroot directive or to move it to a dedicated writable location), or any other system-specific constraint. For example, some Linux flavours are notorious for replacing the default core file with a path to an executable not even installed on the system (check /proc/sys/kernel/core_pattern). Often, simply writing “core”, “core.%p” or “/var/log/core/core.%p” addresses the issue. When trying to enable this option waiting for a rare issue to re-appear, it’s often a good idea to first try to obtain such a dump by issuing, for example, “kill -11” to the “haproxy” process and verify that it leaves a core where expected when dying.

set-var <var-name> <expr>

set-var <var-name> <expr>

Sets the process-wide variable ‘<var-name>’ to the result of the evaluation of the sample expression <expr>. The variable ‘<var-name>’ may only be a process-wide variable (using the ‘proc.’ prefix). It works exactly like the ‘set-var’ action in TCP or HTTP rules except that the expression is evaluated at configuration parsing time and that the variable is instantly set. The sample fetch functions and converters permitted in the expression are only those using internal data, typically ‘int(value)’ or ‘str(value)’. It is possible to reference previously allocated variables as well. These variables will then be readable (and modifiable) from the regular rule sets.

Example:

global
    set-var proc.current_state str(primary)
    set-var proc.prio int(100)
    set-var proc.threshold int(200),sub(proc.prio)

set-var-fmt <var-name> <fmt>

set-var-fmt <var-name> <fmt>

Sets the process-wide variable ‘<var-name>’ to the string resulting from the evaluation of the log-format <fmt>. The variable ‘<var-name>’ may only be a process-wide variable (using the ‘proc.’ prefix). It works exactly like the ‘set-var-fmt’ action in TCP or HTTP rules except that the expression is evaluated at configuration parsing time and that the variable is instantly set. The sample fetch functions and converters permitted in the expression are only those using internal data, typically ‘int(value)’ or ‘str(value)’. It is possible to reference previously allocated variables as well. These variables will then be readable (and modifiable) from the regular rule sets. Please see section 8.2.6 for details on the Custom log format syntax.

Example:

global
    set-var-fmt proc.current_state "primary"
    set-var-fmt proc.bootid        "%pid|%t"

setcap <name>[,<name>...]

setcap <name>[,<name>...]

Sets a list of capabilities that must be preserved when starting and running either as a non-root user (uid > 0), or when starting with uid 0 (root) and switching then to a non-root. By default all permissions are lost by the uid switch, but some are often needed when trying to connect to a server from a foreign address during transparent proxying, or when binding to a port below 1024, e.g. when using “tune.quic.fe.sock-per-conn default-on”, resulting in setups running entirely under uid 0. Setting capabilities generally is a safer alternative, as only the required capabilities will be preserved. The feature is OS-specific and only enabled on Linux when USE_LINUX_CAP=1 is set at build time. The list of supported capabilities also depends on the OS and is enumerated by the error message displayed when an invalid capability name or an empty one is passed. Multiple capabilities may be passed, delimited by commas. Among those commonly used, “cap_net_raw” allows to transparently bind to a foreign address, and “cap_net_bind_service” allows to bind to a privileged port and may be used by QUIC. If the process is started and run under the same non-root user, needed capabilities should be set on haproxy binary file with setcap along with this keyword. For more details about setting capabilities on haproxy binary, please see chapter 13.1 Linux capabilities support in the Management guide.

Example:

global
    setcap cap_net_bind_service,cap_net_admin

setenv <name> <value>

setenv <name> <value>

Sets environment variable <name> to value <value>. If the variable exists, it is overwritten. The changes immediately take effect so that the next line in the configuration file sees the new value. See also “presetenv”, “resetenv”, and “unsetenv”.

shm-stats-file <name>

shm-stats-file <name>

When this directive is set, it enables the use of shared memory for storing stats counters. <name> is used as argument to shm_open() to open the shared memory at a unique location. It also means that the directive is only available on systems which support shm_open(). When SHM is used for stats, all shareable counters for frontends, backends, listeners and servers will be stored in the SHM, provided that they have a GUID set. When reloading haproxy, new process will try to scan the SHM for objects that could be associated to objects defined in the configuration based on GUID and type, the goal is to be able to preserve some counters’ values upon reload. On the other hand, when haproxy is properly stopped, the SHM objects are released, which means counters are effectively reset. It is also possible to manually remove the file before starting a fresh process to force a reset.

See also “guid”, “guid-prefix” and “shm-stats-file-max-objects”

shm-stats-file-max-objects <number>

shm-stats-file-max-objects <number>

This setting defines the maximum number of objects the shared memory used for shared counters will be able to store per thread group. It is directly related to the maximum memory size of the shm and is used to “premap” the shm to a given size in order to avoid runtime re-mapping. It defaults to 2k, which should suit for most setups without risking unsuitable memory usage, but can be easily changed if needed. haproxy will complain during startup if this value is to low to register objects that are expected to be stored in the shared memory. It is only relevant when “shm-stats-file” was defined.

See also “thread-groups”

ssl-default-bind-ciphers <ciphers>

ssl-default-bind-ciphers <ciphers>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of cipher algorithms (“cipher suite”) that are negotiated during the SSL/TLS handshake up to TLSv1.2 for all “bind” lines which do not explicitly define theirs. The format of the string is defined in “man 1 ciphers” from OpenSSL man pages. For background information and recommendations see e.g. (https://wiki.mozilla.org/Security/Server_Side_TLS ) and (https://mozilla.github.io/server-side-tls/ssl-config-generator/ ). For TLSv1.3 cipher configuration, please check the “ssl-default-bind-ciphersuites” keyword. Please check the “bind” keyword for more information.

ssl-default-bind-ciphersuites <ciphersuites>

ssl-default-bind-ciphersuites <ciphersuites>

This setting is only available when support for OpenSSL was built in and OpenSSL 1.1.1 or later was used to build HAProxy. It sets the default string describing the list of cipher algorithms (“cipher suite”) that are negotiated during the TLSv1.3 handshake for all “bind” lines which do not explicitly define theirs. The format of the string is defined in “man 1 ciphers” from OpenSSL man pages under the section “ciphersuites”. For cipher configuration for TLSv1.2 and earlier, please check the “ssl-default-bind-ciphers” keyword. This setting might accept TLSv1.2 ciphersuites however this is an undocumented behavior and not recommended as it could be inconsistent or buggy. The default TLSv1.3 ciphersuites of OpenSSL are: “TLS_AES_256_GCM_SHA384:TLS_CHACHA20_POLY1305_SHA256:TLS_AES_128_GCM_SHA256”

TLSv1.3 only supports 5 ciphersuites:

  • TLS_AES_128_GCM_SHA256
  • TLS_AES_256_GCM_SHA384
  • TLS_CHACHA20_POLY1305_SHA256
  • TLS_AES_128_CCM_SHA256
  • TLS_AES_128_CCM_8_SHA256

Please check the “bind” keyword for more information.

Example:

global
    ssl-default-bind-ciphers ECDHE-RSA-AES256-GCM-SHA384:ECDHE-RSA-CHACHA20-POLY1305:ECDHE-RSA-AES128-GCM-SHA256
    ssl-default-bind-ciphersuites TLS_AES_256_GCM_SHA384:TLS_CHACHA20_POLY1305_SHA256:TLS_AES_128_GCM_SHA256

ssl-default-bind-client-sigalgs <sigalgs>

ssl-default-bind-client-sigalgs <sigalgs>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of signature algorithms related to client authentication for all “bind” lines which do not explicitly define theirs. The format of the string is a colon-delimited list of signature algorithms. Each signature algorithm can use one of two forms: TLS1.3 signature scheme names (“rsa_pss_rsae_sha256”) or the public key algorithm + digest form (“ECDSA+SHA256”). A list can contain both forms. For more information on the format, see SSL_CTX_set1_client_sigalgs(3). A list of signature algorithms is also available in RFC8446 section 4.2.3 and in OpenSSL in the ssl/t1_lib.c file. This setting is not applicable to TLSv1.1 and earlier versions of the protocol as the signature algorithms aren’t separately negotiated in these versions. It is not recommended to change this setting unless compatibility with a middlebox is required.

ssl-default-bind-curves <curves>

ssl-default-bind-curves <curves>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of elliptic curves algorithms (“curve suite”) that are negotiated during the SSL/TLS handshake with ECDHE. The format of the string is a colon-delimited list of curve name. Please check the “bind” keyword for more information.

ssl-default-bind-options [<option>]...

ssl-default-bind-options [<option>]...

This setting is only available when support for OpenSSL was built in. It sets default ssl-options to force on all “bind” lines. Please check the “bind” keyword to see available options.

Example:

global
   ssl-default-bind-options ssl-min-ver TLSv1.0 no-tls-tickets

ssl-default-bind-sigalgs <sigalgs>

ssl-default-bind-sigalgs <sigalgs>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of signature algorithms that are negotiated during the TLSv1.2 and TLSv1.3 handshake for all “bind” lines which do not explicitly define theirs. The format of the string is a colon-delimited list of signature algorithms. Each signature algorithm can use one of two forms: TLS1.3 signature scheme names (“rsa_pss_rsae_sha256”) or the public key algorithm + digest form (“ECDSA+SHA256”). A list can contain both forms. For more information on the format, see SSL_CTX_set1_sigalgs(3). A list of signature algorithms is also available in RFC8446 section 4.2.3 and in OpenSSL in the ssl/t1_lib.c file. This setting is not applicable to TLSv1.1 and earlier versions of the protocol as the signature algorithms aren’t separately negotiated in these versions. It is not recommended to change this setting unless compatibility with a middlebox is required.

ssl-default-server-ciphers <ciphers>

ssl-default-server-ciphers <ciphers>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of cipher algorithms that are negotiated during the SSL/TLS handshake up to TLSv1.2 with the server, for all “server” lines which do not explicitly define theirs. The format of the string is defined in “man 1 ciphers” from OpenSSL man pages. For background information and recommendations see e.g. (https://wiki.mozilla.org/Security/Server_Side_TLS ) and (https://mozilla.github.io/server-side-tls/ssl-config-generator/ ). For TLSv1.3 cipher configuration, please check the “ssl-default-server-ciphersuites” keyword. Please check the “server” keyword for more information.

ssl-default-server-ciphersuites <ciphersuites>

ssl-default-server-ciphersuites <ciphersuites>

This setting is only available when support for OpenSSL was built in and OpenSSL 1.1.1 or later was used to build HAProxy. It sets the default string describing the list of cipher algorithms that are negotiated during the TLSv1.3 handshake with the server, for all “server” lines which do not explicitly define theirs. The format of the string is defined in “man 1 ciphers” from OpenSSL man pages under the section “ciphersuites”. For cipher configuration for TLSv1.2 and earlier, please check the “ssl-default-server-ciphers” keyword. Please check the “server” keyword for more information.

ssl-default-server-client-sigalgs <sigalgs>

ssl-default-server-client-sigalgs <sigalgs>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of signature algorithms related to client authentication for all “server” lines which do not explicitly define theirs. The format of the string is a colon-delimited list of signature algorithms. Each signature algorithm can use one of two forms: TLS1.3 signature scheme names (“rsa_pss_rsae_sha256”) or the public key algorithm + digest form (“ECDSA+SHA256”). A list can contain both forms. For more information on the format, see SSL_CTX_set1_client_sigalgs(3). A list of signature algorithms is also available in RFC8446 section 4.2.3 and in OpenSSL in the ssl/t1_lib.c file. This setting is not applicable to TLSv1.1 and earlier versions of the protocol as the signature algorithms aren’t separately negotiated in these versions. It is not recommended to change this setting unless compatibility with a middlebox is required.

ssl-default-server-curves <curves>

ssl-default-server-curves <curves>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of elliptic curves algorithms (“curve suite”) that are negotiated during the SSL/TLS handshake with ECDHE. The format of the string is a colon-delimited list of curve name. Please check the “server” keyword for more information.

ssl-default-server-options [<option>]...

ssl-default-server-options [<option>]...

This setting is only available when support for OpenSSL was built in. It sets default ssl-options to force on all “server” lines. Please check the “server” keyword to see available options.

ssl-default-server-sigalgs <sigalgs>

ssl-default-server-sigalgs <sigalgs>

This setting is only available when support for OpenSSL was built in. It sets the default string describing the list of signature algorithms that are negotiated during the TLSv1.2 and TLSv1.3 handshake for all “server” lines which do not explicitly define theirs. The format of the string is a colon-delimited list of signature algorithms. Each signature algorithm can use one of two forms: TLS1.3 signature scheme names (“rsa_pss_rsae_sha256”) or the public key algorithm + digest form (“ECDSA+SHA256”). A list can contain both forms. For more information on the format, see SSL_CTX_set1_sigalgs(3). A list of signature algorithms is also available in RFC8446 section 4.2.3 and in OpenSSL in the ssl/t1_lib.c file. This setting is not applicable to TLSv1.1 and earlier versions of the protocol as the signature algorithms aren’t separately negotiated in these versions. It is not recommended to change this setting unless compatibility with a middlebox is required.

ssl-dh-param-file <file>

ssl-dh-param-file <file>

This setting is only available when support for OpenSSL was built in. It sets the default DH parameters that are used during the SSL/TLS handshake when ephemeral Diffie-Hellman (DHE) key exchange is used, for all “bind” lines which do not explicitly define theirs. It will be overridden by custom DH parameters found in a bind certificate file if any. If custom DH parameters are not specified either by using ssl-dh-param-file or by setting them directly in the certificate file, DHE ciphers will not be used, unless tune.ssl.default-dh-param is set. In this latter case, pre-defined DH parameters of the specified size will be used. Custom parameters are known to be more secure and therefore their use is recommended. Custom DH parameters may be generated by using the OpenSSL command “openssl dhparam <size>”, where size should be at least 2048, as 1024-bit DH parameters should not be considered secure anymore.

ssl-passphrase-cmd <cmd> <args> ...

ssl-passphrase-cmd <cmd> <args> ...

This settings is only available when support for OpenSSL was built in. It allows to define a full command line that will be called when an encrypted certificate is loaded during init. The command could be a script or any other program. It will be provided with the encrypted private key path as first parameter and the user-defined “args” parameters then and should dump the passphrase that allows to decode the encrypted private key on the standard output. For every new encrypted private key loaded during init, HAProxy will first try every other already known passphrase to decode the private key and will ultimately call the passphrase command again if none works.

ssl-propquery <query>

ssl-propquery <query>

This setting is only available when support for OpenSSL was built in and when OpenSSL’s version is at least 3.0. It allows to define a default property string used when fetching algorithms in providers. It behave the same way as the openssl propquery option and it follows the same syntax (described in https://www.openssl.org/docs/man3.0/man7/property.html ). For instance, if you have two providers loaded, the foo one and the default one, the propquery “?provider=foo” allows to pick the algorithm implementations provided by the foo provider by default, and to fallback on the default provider’s one if it was not found.

ssl-provider <name>

ssl-provider <name>

This setting is only available when support for OpenSSL was built in and when OpenSSL’s version is at least 3.0. It allows to load a provider during init. If loading is successful, any capabilities provided by the loaded provider might be used by HAProxy. Multiple ‘ssl-provider’ options can be specified in a configuration file. The providers will be loaded in their order of appearance.

Please note that loading a provider explicitly prevents OpenSSL from loading the ‘default’ provider automatically. OpenSSL also allows to define the providers that should be loaded directly in its configuration file (openssl.cnf for instance) so it is not necessary to use this ‘ssl-provider’ option to load providers. The “show ssl providers” CLI command can be used to show all the providers that were successfully loaded.

The default search path of OpenSSL provider can be found in the output of the “openssl version -a” command. If the provider is in another directory, you can set the OPENSSL_MODULES environment variable, which takes the directory where your provider can be found.

See also “ssl-propquery” and “ssl-provider-path”.

ssl-provider-path <path>

ssl-provider-path <path>

This setting is only available when support for OpenSSL was built in and when OpenSSL’s version is at least 3.0. It allows to specify the search path that is to be used by OpenSSL for looking for providers. It behaves the same way as the OPENSSL_MODULES environment variable. It will be used for any following ‘ssl-provider’ option or until a new ‘ssl-provider-path’ is defined. See also “ssl-provider”.

ssl-load-extra-del-ext

ssl-load-extra-del-ext

This setting allows to configure the way HAProxy does the lookup for the extra SSL files. By default HAProxy adds a new extension to the filename. (ex: with “foobar.crt” load “foobar.crt.key”). With this option enabled, HAProxy removes the extension before adding the new one (ex: with “foobar.crt” load “foobar.key”).

Your crt file must have a “.crt” extension for this option to work.

This option is not compatible with bundle extensions (.ecdsa, .rsa. .dsa) and won’t try to remove them.

This option is disabled by default. See also “ssl-load-extra-files”.

ssl-load-extra-files <none|all|bundle|sctl|ocsp|issuer|key>*

ssl-load-extra-files <none|all|bundle|sctl|ocsp|issuer|key>*

This setting alters the way HAProxy will look for unspecified files during the loading of the SSL certificates. This option applies to certificates associated to “bind” lines as well as “server” lines but some of the extra files will not have any functional impact for “server” line certificates.

By default, HAProxy discovers automatically a lot of files not specified in the configuration, and you may want to disable this behavior if you want to optimize the startup time.

“none”: Only load the files specified in the configuration. Don’t try to load a certificate bundle if the file does not exist. In the case of a directory, it won’t try to bundle the certificates if they have the same basename.

“all”: This is the default behavior, it will try to load everything, bundles, sctl, ocsp, issuer, key.

“bundle”: When a file specified in the configuration does not exist, HAProxy will try to load a “cert bundle”. Certificate bundles are only managed on the frontend side and will not work for backend certificates.

Starting from HAProxy 2.3, the bundles are not loaded in the same OpenSSL certificate store, instead it will loads each certificate in a separate store which is equivalent to declaring multiple “crt”. OpenSSL 1.1.1 is required to achieve this. Which means that bundles are now used only for backward compatibility and are not mandatory anymore to do an hybrid RSA/ECC bind configuration.

To associate these PEM files into a “cert bundle” that is recognized by HAProxy, they must be named in the following way: All PEM files that are to be bundled must have the same base name, with a suffix indicating the key type. Currently, three suffixes are supported: rsa, dsa and ecdsa. For example, if www.example.com has two PEM files, an RSA file and an ECDSA file, they must be named: “example.pem.rsa” and “example.pem.ecdsa”. The first part of the filename is arbitrary; only the suffix matters. To load this bundle into HAProxy, specify the base name only:

Example: bind:8443 ssl crt example.pem

Note that the suffix is not given to HAProxy; this tells HAProxy to look for a cert bundle.

HAProxy will load all PEM files in the bundle as if they were configured separately in several “crt”.

The bundle loading does not have an impact anymore on the directory loading since files are loading separately.

On the CLI, bundles are seen as separate files, and the bundle extension is required to commit them.

OCSP files (.ocsp), issuer files (.issuer), Certificate Transparency (.sctl) as well as private keys (.key) are supported with multi-cert bundling.

“sctl”: Try to load “<basename>.sctl” for each crt keyword. If provided for a backend certificate, it will be loaded but will not have any functional impact.

“ocsp”: Try to load “<basename>.ocsp” for each crt keyword. If provided for a backend certificate, it will be loaded but will not have any functional impact.

“issuer”: Try to load “<basename>.issuer” if the issuer of the OCSP file is not provided in the PEM file. If provided for a backend certificate, it will be loaded but will not have any functional impact.

“key”: If the private key was not provided by the PEM file, try to load a file “<basename>.key” containing a private key.

The default behavior is “all”.

Example:

ssl-load-extra-files bundle sctl
ssl-load-extra-files sctl ocsp issuer
ssl-load-extra-files none

See also: “crt”, section 5.1 about bind options and section 5.2 about server options.

ssl-security-level <number>

ssl-security-level <number>

This directive allows to chose the OpenSSL security level as described in https://www.openssl.org/docs/man1.1.1/man3/SSL_CTX_set_security_level.html The security level will be applied to every SSL contextes in HAProxy. Only a value between 0 and 5 is supported.

The default value depends on your OpenSSL version, distribution and how was compiled the library.

This directive requires at least OpenSSL 1.1.1.

ssl-server-verify [none|required]

ssl-server-verify [none|required]

The default behavior for SSL verify on servers side. If specified to ’none’, servers certificates are not verified. The default is ‘required’ except if forced using cmdline option ‘-dV’.

ssl-skip-self-issued-ca

ssl-skip-self-issued-ca

Self issued CA, aka x509 root CA, is the anchor for chain validation: as a server is useless to send it, client must have it. Standard configuration need to not include such CA in PEM file. This option allows you to keep such CA in PEM file without sending it to the client. Use case is to provide issuer for ocsp without the need for ‘.issuer’ file and be able to share it with ‘issuers-chain-path’. This concerns all certificates without intermediate certificates. It’s useless for BoringSSL, .issuer is ignored because ocsp bits does not need it. Requires at least OpenSSL 1.0.2.

stats calculate-max-counters [on|off]

stats calculate-max-counters [on|off]

Activates or deactivates the calculation of stats max counters. If you don’t need them, deactivating them may increase performances a bit. The default is on.

stats maxconn <connections>

stats maxconn <connections>

By default, the stats socket is limited to 10 concurrent connections. It is possible to change this value with “stats maxconn”.

stats socket [<address:port>|<path>] [param*]

stats socket [<address:port>|<path>] [param*]

Binds a UNIX socket to <path> or a TCPv4/v6 address to <address:port>. Connections to this socket will return various statistics outputs and even allow some commands to be issued to change some runtime settings. Please consult section 9.3 “Unix Socket commands” of Management Guide for more details.

All parameters supported by “bind” lines are supported, for instance to restrict access to some users or their access rights. Please consult section 5.1 for more information.

stats timeout <timeout, in milliseconds>

stats timeout <timeout, in milliseconds>

The default timeout on the stats socket is set to 10 seconds. It is possible to change this value with “stats timeout”. The value must be passed in milliseconds, or be suffixed by a time unit among { us, ms, s, m, h, d }.

stats-file <path>

stats-file <path>

Path to a generated haproxy stats-file. On startup haproxy will preload the values to its internal counters. Use the CLI command “dump stats-file” to produce such stats-file. See the management manual for more details.

stress-level <level>

stress-level <level>

Activate alternative code to stress haproxy binary. Level is an integer from 0 to 9. The default value 0 disable any stressing execution. Levels from 1 to 9 will increase the stress pressure on the haproxy binary. Note that using any positive level can significantly hurt performance. As such it should never be activated unless for debugging purpose and on a developer request.

strict-limits

strict-limits

Makes process fail at startup when a setrlimit fails. HAProxy tries to set the best setrlimit according to what has been calculated. If it fails, it will emit a warning. This option is here to guarantee an explicit failure of HAProxy when those limits fail. It is enabled by default. It may still be forcibly disabled by prefixing it with the “no” keyword.

thread-group <group> [<thread-range>...]

thread-group <group> [<thread-range>...]

This setting is only available when support for threads was built in. It enumerates the list of threads that will compose thread group <group>. Thread numbers and group numbers start at 1. Thread ranges are defined either using a single thread number at once, or by specifying the lower and upper bounds delimited by a dash ‘-’ (e.g. “1-16”). Unassigned threads will be automatically assigned to unassigned thread groups, and thread groups defined with this directive will never receive more threads than those defined. Defining the same group multiple times overrides previous definitions with the new one. See also “nbthread” and “thread-groups”.

thread-groups <number>

thread-groups <number>

This setting is only available when support for threads was built in. It makes HAProxy split its threads into <number> independent groups. At the moment, the default value is 1. Thread groups make it possible to reduce sharing between threads to limit contention, at the expense of some extra configuration efforts. It is also the only way to use more than 64 threads since up to 64 threads per group may be configured. The maximum number of groups is configured at compile time and defaults to 16. See also “nbthread”.

thread-hard-limit <number>

thread-hard-limit <number>

This setting is used to enforce a limit to the number of threads, either detected, or configured. This is particularly useful on operating systems where the number of threads is automatically detected, where a number of threads lower than the number of CPUs is desired in generic and portable configurations. Indeed, while “nbthread” enforces a number of threads that will result in a warning and bad performance if higher than CPUs available, thread-hard-limit will only cap the maximum value and automatically limit the number of threads to no higher than this value, but will not raise lower values. If “nbthread” is forced to a higher value, thread-hard-limit wins, and a warning is emitted in so that the configuration anomaly can be fixed. By default there is no limit. See also “nbthread”.

uid <number>

uid <number>

Changes the process’s user ID to <number>. It is recommended that the user ID is dedicated to HAProxy or to a small set of similar daemons. HAProxy must be started with superuser privileges in order to be able to switch to another one. See also “gid” and “user”.

ulimit-n <number>

ulimit-n <number>

Sets the maximum number of per-process file-descriptors to <number>. By default, it is automatically computed, so it is recommended not to use this option. If the intent is only to limit the number of file descriptors, better use “fd-hard-limit” instead.

Note that the dynamic servers are not taken into account in this automatic resource calculation. If using a large number of them, it may be needed to manually specify this value.

See also: fd-hard-limit, maxconn

unix-bind [ prefix <prefix> ] [ mode <mode> ] [ user <user> ] [ uid <uid> ] [ group <group> ] [ gid <gid> ]

Fixes common settings to UNIX listening sockets declared in “bind” statements. This is mainly used to simplify declaration of those UNIX sockets and reduce the risk of errors, since those settings are most commonly required but are also process-specific. The <prefix> setting can be used to force all socket path to be relative to that directory. This might be needed to access another component’s chroot. Note that those paths are resolved before HAProxy chroots itself, so they are absolute. The <mode>, <user>, <uid>, <group> and <gid> all have the same meaning as their homonyms used by the “bind” statement. If both are specified, the “bind” statement has priority, meaning that the “unix-bind” settings may be seen as process-wide default settings.

unsetenv [<name> ...]

unsetenv [<name> ...]

Removes environment variables specified in arguments. This can be useful to hide some sensitive information that are occasionally inherited from the user’s environment during some operations. Variables which did not exist are silently ignored so that after the operation, it is certain that none of these variables remain. The changes immediately take effect so that the next line in the configuration file will not see these variables. See also “setenv”, “presetenv”, and “resetenv”.

user <user name>

user <user name>

Similar to “uid” but uses the UID of user name <user name> from /etc/passwd. See also “uid” and “group”.

node <name>

node <name>

Only letters, digits, hyphen and underscore are allowed, like in DNS names.

This statement is useful in HA configurations where two or more processes or servers share the same IP address. By setting a different node-name on all nodes, it becomes easy to immediately spot what server is handling the traffic.

wurfl-cache-size <size>

wurfl-cache-size <size>

Sets the WURFL Useragent cache size. For faster lookups, already processed user agents are kept in a LRU cache:

  • “0” : no cache is used.
  • <size> : size of lru cache in elements.

Please note that this option is only available when HAProxy has been compiled with USE_WURFL=1.

wurfl-data-file <file path>

wurfl-data-file <file path>

The path of the WURFL data file to provide device detection services. The file should be accessible by HAProxy with relevant permissions.

Please note that this option is only available when HAProxy has been compiled with USE_WURFL=1.

wurfl-information-list [<capability>]*

wurfl-information-list [<capability>]*

A space-delimited list of WURFL capabilities, virtual capabilities, property names we plan to use in injected headers. A full list of capability and virtual capability names is available on the Scientiamobile website:

https://www.scientiamobile.com/wurflCapability

Valid WURFL properties are:

  • wurfl_id Contains the device ID of the matched device.

  • wurfl_root_id Contains the device root ID of the matched device.

  • wurfl_isdevroot Tells if the matched device is a root device. Possible values are “TRUE” or “FALSE”.

  • wurfl_useragent The original useragent coming with this particular web request.

  • wurfl_api_version Contains a string representing the currently used Libwurfl API version.

  • wurfl_info A string containing information on the parsed wurfl.xml and its full path.

  • wurfl_last_load_time Contains the UNIX timestamp of the last time WURFL has been loaded successfully.

  • wurfl_normalized_useragent The normalized useragent.

Please note that this option is only available when HAProxy has been compiled with USE_WURFL=1.

wurfl-information-list-separator <char>

wurfl-information-list-separator <char>

A char that will be used to separate values in a response header containing WURFL results. If not set that a comma (’,’) will be used by default.

Please note that this option is only available when HAProxy has been compiled with USE_WURFL=1.

wurfl-patch-file [<file path>]

wurfl-patch-file [<file path>]

A list of WURFL patch file paths. Note that patches are loaded during startup thus before the chroot.

Please note that this option is only available when HAProxy has been compiled with USE_WURFL=1.

3.2. Performance tuning

busy-polling

busy-polling

In some situations, especially when dealing with low latency on processors supporting a variable frequency or when running inside virtual machines, each time the process waits for an I/O using the poller, the processor goes back to sleep or is offered to another VM for a long time, and it causes excessively high latencies. This option provides a solution preventing the processor from sleeping by always using a null timeout on the pollers. This results in a significant latency reduction (30 to 100 microseconds observed) at the expense of a risk to overheat the processor. It may even be used with threads, in which case improperly bound threads may heavily conflict, resulting in a worse performance and high values for the CPU stolen fields in “show info” output, indicating which threads are misconfigured. It is important not to let the process run on the same processor as the network interrupts when this option is used. It is also better to avoid using it on multiple CPU threads sharing the same core. This option is disabled by default. If it has been enabled, it may still be forcibly disabled by prefixing it with the “no” keyword. It is ignored by the “select” and “poll” pollers.

This option is automatically disabled on old processes in the context of seamless reload; it avoids too much cpu conflicts when multiple processes stay around for some time waiting for the end of their current connections.

max-spread-checks <delay in milliseconds>

max-spread-checks <delay in milliseconds>

By default, HAProxy tries to spread the start of health checks across the smallest health check interval of all the servers in a farm. The principle is to avoid hammering services running on the same server. But when using large check intervals (10 seconds or more), the last servers in the farm take some time before starting to be tested, which can be a problem. This parameter is used to enforce an upper bound on delay between the first and the last check, even if the servers’ check intervals are larger. When servers run with shorter intervals, their intervals will be respected though.

maxcompcpuusage <number>

maxcompcpuusage <number>

Sets the maximum CPU usage HAProxy can reach before stopping the compression for new requests or decreasing the compression level of current requests. It works like ‘maxcomprate’ but measures CPU usage instead of incoming data bandwidth. The value is expressed in percent of the CPU used by HAProxy. A value of 100 disable the limit. The default value is 100. Setting a lower value will prevent the compression work from slowing the whole process down and from introducing high latencies.

maxcomprate <number>

maxcomprate <number>

Sets the maximum per-process input compression rate to <number> kilobytes per second. For each stream, if the maximum is reached, the compression level will be decreased during the stream. If the maximum is reached at the beginning of a stream, the stream will not compress at all. If the maximum is not reached, the compression level will be increased up to tune.comp.maxlevel. A value of zero means there is no limit, this is the default value.

maxconn <number>

maxconn <number>

Sets the maximum per-process number of concurrent connections to <number>. It is equivalent to the command-line argument “-n”. The value provided in command-line argument via “-n” takes the precedence over the maxconn value set in the global section. Haproxy process could be also compiled with SYSTEM_MAXCONN compile-time variable, which is served in this case as the system maxconn maximum. Again, the command-line “-n” argument allows at runtime to bypass SYSTEM_MAXCONN limit, if set. Proxies will stop accepting connections when maxconn is reached. The process soft file descriptor limit (could be obtained with “ulimit -n” command) is automatically adjusted according to provided maxconn. See also “ulimit-n”. Note: the “select” poller cannot reliably use more than 1024 file descriptors on some platforms. If your platform only supports select and reports “select FAILED” on startup, you need to reduce the maxconn until it works (slightly below 500 in general). If maxconn value is not set, it will be automatically calculated based on the current file descriptors limits, reported by the “ulimit -nH” command (we take the maximum between the hard and soft values), then automatic value will be possibly reduced by “fd-hard-limit” and by memory limit, if the latter was enforced via “-m” command line option. Automatic value is also dependent from the buffer size, memory allocated to compression, SSL cache size, and the use or not of SSL and the associated maxsslconn (which can also be automatic).

See also: fd-hard-limit, ulimit-n

maxconnrate <number>

maxconnrate <number>

Sets the maximum per-process number of connections per second to <number>. Proxies will stop accepting connections when this limit is reached. It can be used to limit the global capacity regardless of each frontend capacity. It is important to note that this can only be used as a service protection measure, as there will not necessarily be a fair share between frontends when the limit is reached, so it’s a good idea to also limit each frontend to some value close to its expected share. Also, lowering tune.maxaccept can improve fairness.

maxpipes <number>

maxpipes <number>

Sets the maximum per-process number of pipes to <number>. Currently, pipes are only used by kernel-based tcp splicing. Since a pipe contains two file descriptors, the “ulimit-n” value will be increased accordingly. The default value is maxconn/4, which seems to be more than enough for most heavy usages. The splice code dynamically allocates and releases pipes, and can fall back to standard copy, so setting this value too low may only impact performance.

maxsessrate <number>

maxsessrate <number>

Sets the maximum per-process number of sessions per second to <number>. Proxies will stop accepting connections when this limit is reached. It can be used to limit the global capacity regardless of each frontend capacity. It is important to note that this can only be used as a service protection measure, as there will not necessarily be a fair share between frontends when the limit is reached, so it’s a good idea to also limit each frontend to some value close to its expected share. Also, lowering tune.maxaccept can improve fairness.

maxsslconn <number>

maxsslconn <number>

Sets the maximum per-process number of concurrent SSL connections to <number>. By default there is no SSL-specific limit, which means that the global maxconn setting will apply to all connections. Setting this limit avoids having openssl use too much memory and crash when malloc returns NULL (since it unfortunately does not reliably check for such conditions). Note that the limit applies both to incoming and outgoing connections, so one connection which is deciphered then ciphered accounts for 2 SSL connections. If this value is not set, but a memory limit is enforced, this value will be automatically computed based on the memory limit, maxconn, the buffer size, memory allocated to compression, SSL cache size, and use of SSL in either frontends, backends or both. If neither maxconn nor maxsslconn are specified when there is a memory limit, HAProxy will automatically adjust these values so that 100% of the connections can be made over SSL with no risk, and will consider the sides where it is enabled (frontend, backend, both).

maxsslrate <number>

maxsslrate <number>

Sets the maximum per-process number of SSL sessions per second to <number>. SSL listeners will stop accepting connections when this limit is reached. It can be used to limit the global SSL CPU usage regardless of each frontend capacity. It is important to note that this can only be used as a service protection measure, as there will not necessarily be a fair share between frontends when the limit is reached, so it’s a good idea to also limit each frontend to some value close to its expected share. It is also important to note that the sessions are accounted before they enter the SSL stack and not after, which also protects the stack against bad handshakes. Also, lowering tune.maxaccept can improve fairness.

maxzlibmem <number>

maxzlibmem <number>

Sets the maximum amount of RAM in megabytes per process usable by the zlib. When the maximum amount is reached, future streams will not compress as long as RAM is unavailable. When sets to 0, there is no limit. The default value is 0. The value is available in bytes on the UNIX socket with “show info” on the line “MaxZlibMemUsage”, the memory used by zlib is “ZlibMemUsage” in bytes.

no-memory-trimming

no-memory-trimming

Disables memory trimming (“malloc_trim”) at a few moments where attempts are made to reclaim lots of memory (on memory shortage or on reload). Trimming memory forces the system’s allocator to scan all unused areas and to release them. This is generally seen as nice action to leave more available memory to a new process while the old one is unlikely to make significant use of it. But some systems dealing with tens to hundreds of thousands of concurrent connections may experience a lot of memory fragmentation, that may render this release operation extremely long. During this time, no more traffic passes through the process, new connections are not accepted anymore, some health checks may even fail, and the watchdog may even trigger and kill the unresponsive process, leaving a huge core dump. If this ever happens, then it is suggested to use this option to disable trimming and stop trying to be nice with the new process. Note that advanced memory allocators usually do not suffer from such a problem.

noepoll

noepoll

Disables the use of the “epoll” event polling system on Linux. It is equivalent to the command-line argument “-de”. The next polling system used will generally be “poll”. See also “nopoll”.

noevports

noevports

Disables the use of the event ports event polling system on SunOS systems derived from Solaris 10 and later. It is equivalent to the command-line argument “-dv”. The next polling system used will generally be “poll”. See also “nopoll”.

nogetaddrinfo

nogetaddrinfo

Disables the use of getaddrinfo(3) for name resolving. It is equivalent to the command line argument “-dG”. Deprecated gethostbyname(3) will be used.

nokqueue

nokqueue

Disables the use of the “kqueue” event polling system on BSD. It is equivalent to the command-line argument “-dk”. The next polling system used will generally be “poll”. See also “nopoll”.

noktls

noktls

Disables the use of ktls. It is equivalent to the command line argument “-dT”.

nopoll

nopoll

Disables the use of the “poll” event polling system. It is equivalent to the command-line argument “-dp”. The next polling system used will be “select”. It should never be needed to disable “poll” since it’s available on all platforms supported by HAProxy. See also “nokqueue”, “noepoll” and “noevports”.

noreuseport

noreuseport

Disables the use of SO_REUSEPORT - see socket(7). It is equivalent to the command line argument “-dR”.

nosplice

nosplice

Disables the use of kernel tcp splicing between sockets on Linux. It is equivalent to the command line argument “-dS”. Data will then be copied using conventional and more portable recv/send calls. Kernel tcp splicing is limited to some very recent instances of kernel 2.6. Most versions between 2.6.25 and 2.6.28 are buggy and will forward corrupted data, so they must not be used. This option makes it easier to globally disable kernel splicing in case of doubt. See also “option splice-auto”, “option splice-request” and “option splice-response”.

profiling.memory { on | off }

profiling.memory { on | off }

Enables (‘on’) or disables (‘off’) per-function memory profiling. This will keep usage statistics of malloc/calloc/realloc/free calls anywhere in the process (including libraries) which will be reported on the CLI using the “show profiling” command. This is essentially meant to be used when an abnormal memory usage is observed that cannot be explained by the pools and other info are required. The performance hit will typically be around 1%, maybe a bit more on highly threaded machines, so it is normally suitable for use in production. The same may be achieved at run time on the CLI using the “set profiling memory” command, please consult the management manual.

profiling.tasks { auto | on | off | lock | no-lock | memory | no-memory }*

profiling.tasks { auto | on | off | lock | no-lock | memory | no-memory }*

Enables (‘on’) or disables (‘off’) per-task CPU profiling. When set to ‘auto’ the profiling automatically turns on a thread when it starts to suffer from an average latency of 1000 microseconds or higher as reported in the “avg_loop_us” activity field, and automatically turns off when the latency returns below 990 microseconds (this value is an average over the last 1024 loops so it does not vary quickly and tends to significantly smooth short spikes). It may also spontaneously trigger from time to time on overloaded systems, containers, or virtual machines, or when the system swaps (which must absolutely never happen on a load balancer).

When task profiling is enabled, HAProxy can also collect the time each task spends with a lock held or waiting for a lock, as well as the time spent waiting for a memory allocation to succeed in case of a pool cache miss. This can sometimes help understand certain causes of latency. For this, the extra keywords “lock” (to enable lock time collection), “no-lock” (to disable it), “memory” (to enable memory allocation time collection) or “no-memory” (to disable it) may additionally be passed. By default they are not enabled since they can have a non-negligible CPU impact on highly loaded systems (3-10%). Note that the overhead is only taken when profiling is effectively running, so that when running in “auto” mode, it will only appear when HAProxy decides to turn it on.

CPU profiling per task can be very convenient to report where the time is spent and which requests have what effect on which other request. Enabling it will typically affect the overall’s performance by less than 1%, thus it is recommended to leave it to the default ‘auto’ value so that it only operates when a problem is identified. This feature requires a system supporting the clock_gettime(2) syscall with clock identifiers CLOCK_MONOTONIC and CLOCK_THREAD_CPUTIME_ID, otherwise the reported time will be zero. This option may be changed at run time using “set profiling” on the CLI.

spread-checks <0..50, in percent>

spread-checks <0..50, in percent>

Sometimes it is desirable to avoid sending agent and health checks to servers at exact intervals, for instance when many logical servers are located on the same physical server. With the help of this parameter, it becomes possible to add some randomness in the check interval between 0 and +/- 50%. A value between 2 and 5 seems to show good results. The default value remains at 0.

ssl-engine <name> [algo <comma-separated list of algorithms>]

ssl-engine <name> [algo <comma-separated list of algorithms>]

Sets the OpenSSL engine to <name>. List of valid values for <name> may be obtained using the command “openssl engine”. This statement may be used multiple times, it will simply enable multiple crypto engines. Referencing an unsupported engine will prevent HAProxy from starting. Note that many engines will lead to lower HTTPS performance than pure software with recent processors. The optional command “algo” sets the default algorithms an ENGINE will supply using the OPENSSL function ENGINE_set_default_string(). A value of “ALL” uses the engine for all cryptographic operations. If no list of algo is specified then the value of “ALL” is used. A comma-separated list of different algorithms may be specified, including: RSA, DSA, DH, EC, RAND, CIPHERS, DIGESTS, PKEY, PKEY_CRYPTO, PKEY_ASN1. This is the same format that openssl configuration file uses: https://www.openssl.org/docs/man1.0.2/apps/config.html

HAProxy Version 2.6 disabled the support for engines in the default build. This option is only available when HAProxy has been built with support for it. In case the ssl-engine is required HAProxy can be rebuild with the USE_ENGINE=1 flag.

ssl-mode-async

ssl-mode-async

Adds SSL_MODE_ASYNC mode to the SSL context. This enables asynchronous TLS I/O operations if asynchronous capable SSL engines are used. The current implementation supports a maximum of 32 engines. The Openssl ASYNC API doesn’t support moving read/write buffers and is not compliant with HAProxy’s buffer management. So the asynchronous mode is disabled on read/write operations (it is only enabled during initial and renegotiation handshakes).

tune.applet.zero-copy-forwarding { on | off }

tune.applet.zero-copy-forwarding { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy forwarding of data for the applets. It is enabled by default.

See also: tune.disable-zero-copy-forwarding.

tune.buffers.limit <number>

tune.buffers.limit <number>

Sets a hard limit on the number of buffers which may be allocated per process. The default value is zero which means unlimited. The limit will automatically be re-adjusted to satisfy the reserved buffers for emergency situations so that the user doesn’t have to perform complicated calculations. Forcing this value can be particularly useful to limit the amount of memory a process may take, while retaining a sane behavior. When this limit is reached, a task that requests a buffer waits for another one to be released first. Most of the time the waiting time is very short and not perceptible provided that limits remain reasonable. However, some historical limitations have weakened this mechanism over versions and it is known that in certain situations of sustained shortage, some tasks may freeze until their timeout expires, so it is safer to avoid using this when not strictly necessary.

tune.buffers.reserve <number>

tune.buffers.reserve <number>

Sets the number of per-thread buffers which are pre-allocated and reserved for use only during memory shortage conditions resulting in failed memory allocations. The minimum value is 0 and the default is 4. There is no reason a user would want to change this value, unless a core developer suggests to change it for a very specific reason.

tune.bufsize <size>

tune.bufsize <size>

Sets the buffer size to this size (in bytes). Lower values allow more streams to coexist in the same amount of RAM, and higher values allow some applications with very large cookies to work. The default value is 16384 and can be changed at build time. It is strongly recommended not to change this from the default value, as very low values will break some services such as statistics, and values larger than default size will increase memory usage, possibly causing the system to run out of memory. At least the global maxconn parameter should be decreased by the same factor as this one is increased. In addition, use of HTTP/2 mandates that this value must be 16384 or more. If an HTTP request is larger than (tune.bufsize - tune.maxrewrite), HAProxy will return HTTP 400 (Bad Request) error. Similarly if an HTTP response is larger than this size, HAProxy will return HTTP 502 (Bad Gateway). Note that the value set using this parameter will automatically be rounded up to the next multiple of 8 on 32-bit machines and 16 on 64-bit machines.

tune.bufsize.large <size>

tune.bufsize.large <size>

Sets the size in bytes for large buffers. By defaults, support for large buffers is not enabled, it must explicitly be enable by setting this value.

These buffers are designed to be used in some specific contexts where more data must be bufferized without changing the size of regular buffers. The large buffers are not implicitly used.

Note that when large buffers are configured, three special large buffers will be allocated for each threads during startup for internal usage.

tune.bufsize.small <size>

tune.bufsize.small <size>

Sets the size in bytes for small buffers. The defaults value is 1024.

These buffers are designed to be used in some specific contexts where memory consumption is restrained but it seems unnecessary to allocate a full buffer. If however a small buffer is not sufficient, a reallocation is automatically done to switch to a standard size buffer.

For the moment, it is automatically used only by HTTP/3 protocol to emit the response headers. Otherwise, small buffers support can be enabled for specific proxies via the “use-small-buffers” option.

See also: option use-small-buffers

tune.cli.max-payload-size <size>

tune.cli.max-payload-size <size>

Sets the maximum size allowed for the payload passed to a command on the CLI.

On the CLI, a command line is limited by the buffer size. It means all commands and their arguments must fit in a buffer to be processed, excluding the payload that can be passed to the last command of the command line. This payload can be allocated into a dedicated area if necessary. Its size is limited by this parameter. The default value is 128KB.

While it should be high enough for most usage, if this value is changed, it must be carefully chosen. A huge value can have impact on the HAProxy performance. Depending on the command, a huge payload can be quite long to process and can possibly trigger the watchdog.

Please consult the management manual for details about the CLI.

tune.comp.maxlevel <number>

tune.comp.maxlevel <number>

Sets the maximum compression level. The compression level affects CPU usage during compression. This value affects CPU usage during compression. Each stream using compression initializes the compression algorithm with this value. The default value is 1.

tune.defaults.purge

tune.defaults.purge

For dynamic backends support, all named defaults sections are now kept in memory after parsing. This is necessary as backend added at runtime must be based on a named defaults for its configuration.

This may consume significant memory if the number of defaults instances is important. In this case and if dynamic backend feature is unnecessary, it’s possible to use this option to force deletion of defaults section after parsing. It is still mandatory though to keep referenced defaults section which contain settings whose cannot be copied by their referencing proxies. For example, this is the case if the defaults section defines TCP/HTTP rules or a tcpcheck ruleset.

tune.disable-fast-forward

tune.disable-fast-forward

Disables the data fast-forwarding. It is a mechanism to optimize the data forwarding by passing data directly from a side to the other one without waking the stream up. Thanks to this directive, it is possible to disable this optimization. Note it also disable any kernel tcp splicing but also the zero-copy forwarding. This command is not meant for regular use, it will generally only be suggested by developers along complex debugging sessions.

tune.disable-zero-copy-forwarding

tune.disable-zero-copy-forwarding

Globally disables the zero-copy forwarding of data. It is a mechanism to optimize the data fast-forwarding by avoiding to use the channel’s buffer. Thanks to this directive, it is possible to disable this optimization. Note it also disable any kernel tcp splicing.

See also: tune.pt.zero-copy-forwarding, tune.applet.zero-copy-forwarding, tune.h1.zero-copy-fwd-recv, tune.h1.zero-copy-fwd-send, tune.h2.zero-copy-fwd-send, tune.quic.zero-copy-fwd-send

tune.epoll.mask-events <event[,...]>

tune.epoll.mask-events <event[,...]>

Along HAProxy’s history, a few complex issues were met that were caused by bugs in the epoll mechanism in the Linux kernel. These ones usually are very rare and unreproducible outside the reporter’s environment, and may only be worked around by disabling epoll and switching to poll instead, which is not very satisfying for high performance environments. Each time, issues affect only very specific (and rare) event types, and offering the ability to mask them can constitute a more acceptable work-around. This options offers this possibility by permitting to silently ignore events a few uncommon events and replace them with an input (which reports an unspecified incoming event). The effect is to avoid the fast error processing paths in certain places and only use the common paths. This should never be used unless being invited to do so by an expert in order to diagnose or work around a kernel bug.

The option takes a single argument which is a comma-delimited list of words each designating an event to be masked. The currently supported list of events is: - “err”: mask the EPOLLERR event - “hup”: mask the EPOLLHUP events - “rdhup”: mask the EPOLLRDHUP events

Example:

# mask all non-traffic epoll events:
tune.epoll.mask-events err,hup,rdhup

tune.events.max-events-at-once <number>

tune.events.max-events-at-once <number>

Sets the number of events that may be processed at once by an asynchronous task handler (from event_hdl API). <number> should be included between 1 and 10000. Large number could cause thread contention as a result of the task doing heavy work without interruption, and on the other hand, small number could result in the task being constantly rescheduled because it cannot consume enough events per run and is not able to catch up with the event producer. The default value may be forced at build time, otherwise defaults to 100.

tune.fail-alloc

tune.fail-alloc

If compiled with DEBUG_FAIL_ALLOC or started with “-dMfail”, gives the percentage of chances an allocation attempt fails. Must be between 0 (no failure) and 100 (no success). This is useful to debug and make sure memory failures are handled gracefully. When not set, the ratio is 0. However the command-line “-dMfail” option automatically sets it to 1% failure rate so that it is not necessary to change the configuration for testing.

tune.fd.edge-triggered { on | off } [ EXPERIMENTAL ]

tune.fd.edge-triggered { on | off }  [ EXPERIMENTAL ]

Enables (‘on’) or disables (‘off’) the edge-triggered polling mode for FDs that support it. This is currently only support with epoll. It may noticeably reduce the number of epoll_ctl() calls and slightly improve performance in certain scenarios. This is still experimental, it may result in frozen connections if bugs are still present, and is disabled by default.

tune.glitches.kill.cpu-usage <number>

tune.glitches.kill.cpu-usage <number>

Sets the minimum CPU usage between 0 and 100, at which connections showing too many glitches will be killed. This applies to connections that have reached their glitches-threshold limit. In environments where very long connections often behave badly without causing any performance impact, it might be desirable to keep them regardless of their misbehavior as long as they do not hurt, and to only start to kill such connections when the CPU is getting busy. This parameters allows to specify that a connection reaching its glitches threshold will be actively killed when the CPU usage is at this level or above, but never when it’s below. Note that the CPU usage is measured per thread, so a single misbehaving connection might be killed. The default is zero, meaning that a connection reaching its glitches-threshold will automatically get killed. A rule of thumb would be to set this value to twice the usually observed CPU usage, or the commonly observed CPU usage plus half the idle one (i.e. if CPU commonly reaches 60%, setting 80 here can make sense). This parameter has no effect without tune.h2.fe.glitches-threshold, tune.quic.fe.sec.glitches-threshold or tune.h1.fe.glitches-threshold. See also the global parameters “tune.h2.fe.glitches-threshold”, “tune.h1.fe.glitches-threshold” and “tune.quic.fe.sec.glitches-threshold”.

tune.h1.be.glitches-threshold <number>

tune.h1.be.glitches-threshold <number>

Sets the threshold for the number of glitches on a HTTP/1 backend connection, after which that connection will automatically be killed. This allows to automatically kill misbehaving connections without having to write explicit rules for them. The default value is zero, indicating that no threshold is set so that no event will cause a connection to be closed. Typical events include improperly formatted headers that had been nevertheless accepted by “accept-unsafe-violations-in-http-response”. Any non-zero value here should probably be in the hundreds or thousands to be effective without affecting slightly bogus servers. It is also possible to only kill connections when the CPU usage crosses a certain level, by using “tune.glitches.kill.cpu-usage”. Note that a graceful close is attempted at 75% of the configured threshold by advertising a GOAWAY for a future stream. This ensures that a slightly faulty connection will stop being used after some time without risking to interrupt ongoing transfers.

See also: tune.h1.fe.glitches-threshold, bc_glitches, and tune.glitches.kill.cpu-usage

tune.h1.fe.glitches-threshold <number>

tune.h1.fe.glitches-threshold <number>

Sets the threshold for the number of glitches on a HTTP/1 frontend connection after which that connection will automatically be killed. This allows to automatically kill misbehaving connections without having to write explicit rules for them. The default value is zero, indicating that no threshold is set so that no event will cause a connection to be closed. Typical events include improperly formatted headers that had been nevertheless accepted by “accept-unsafe-violations-in-http-request”. Any non-zero value here should probably be in the hundreds or thousands to be effective without affecting slightly bogus clients. It is also possible to only kill connections when the CPU usage crosses a certain level, by using “tune.glitches.kill.cpu-usage”. Note that a graceful close is attempted at 75% of the configured threshold by advertising a GOAWAY for a future stream. This ensures that a slightly non-compliant client will have the opportunity to create a new connection and continue to work unaffected without ever triggering the hard close thus risking to interrupt ongoing transfers.

See also: tune.h1.be.glitches-threshold, fc_glitches, and tune.glitches.kill.cpu-usage

tune.h1.zero-copy-fwd-recv { on | off }

tune.h1.zero-copy-fwd-recv { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy receives of data for the H1 multiplexer. It is enabled by default.

See also: tune.disable-zero-copy-forwarding, tune.h1.zero-copy-fwd-send

tune.h1.zero-copy-fwd-send { on | off }

tune.h1.zero-copy-fwd-send { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy sends of data for the H1 multiplexer. It is enabled by default.

See also: tune.disable-zero-copy-forwarding, tune.h1.zero-copy-fwd-recv

tune.h2.be.glitches-threshold <number>

tune.h2.be.glitches-threshold <number>

Sets the threshold for the number of glitches on a backend connection, where that connection will automatically be killed. This allows to automatically kill misbehaving connections without having to write explicit rules for them. The default value is zero, indicating that no threshold is set so that no event will cause a connection to be closed. Beware that some H2 servers may occasionally cause a few glitches over long lasting connection, so any non-zero value here should probably be in the hundreds or thousands to be effective without affecting slightly bogus servers. It is also possible to only kill connections when the CPU usage crosses a certain level, by using “tune.glitches.kill.cpu-usage”. Note that a graceful close is attempted at 75% of the configured threshold by advertising a GOAWAY for a future stream. This ensures that a slightly faulty connection will stop being used after some time without risking to interrupt ongoing transfers.

See also: tune.h2.fe.glitches-threshold, bc_glitches, and tune.glitches.kill.cpu-usage

tune.h2.be.initial-window-size <number>

tune.h2.be.initial-window-size <number>

Sets the HTTP/2 initial window size for outgoing connections, which is the number of bytes the server can respond before waiting for an acknowledgment from HAProxy. This setting only affects payload contents, not headers. When not set, the common default value set by tune.h2.initial-window-size applies. It can make sense to slightly increase this value to allow faster downloads or to reduce CPU usage on the servers, at the expense of creating unfairness between clients. It is better to use tune.h2.be.rxbuf instead, which does not cause any unfairness. It doesn’t affect resource usage.

See also: tune.h2.initial-window-size.

tune.h2.be.max-concurrent-streams <number>

tune.h2.be.max-concurrent-streams <number>

Sets the HTTP/2 maximum number of concurrent streams per outgoing connection (i.e. the number of outstanding requests on a single connection to a server). When not set, the default set by tune.h2.max-concurrent-streams applies. A smaller value than the default 100 may improve a site’s responsiveness at the expense of maintaining more established connections to the servers. When the “http-reuse” setting is set to “always”, it is recommended to reduce this value so as not to mix too many different clients over the same connection, because if a client is slower than others, a mechanism known as “head of line blocking” tends to cause cascade effect on download speed for all clients sharing a connection (keep tune.h2.be.initial-window-size low in this case). It is highly recommended not to increase this value; some might find it optimal to run at low values (1..5 typically).

tune.h2.be.max-frames-at-once <number>

tune.h2.be.max-frames-at-once <number>

Sets the maximum number of HTTP/2 incoming frames that will be processed at once on a backend connection. It can be useful to set this to a low value (a few tens to a few hundreds) when dealing with very large buffers in order to maintain a low latency and a better fairness between multiple connections. The default value is zero, which means that no limitation is enforced.

tune.h2.be.rxbuf <size>

tune.h2.be.rxbuf <size>

Sets the HTTP/2 receive buffer size for outgoing connections, in bytes. This size will be rounded up to the next multiple of tune.bufsize and will be shared between all streams uploading data (both HEADERS and DATA frames). In any case, one buffer will always be granted to each stream, and 7/8 of the unused buffers will be shared between streams downloading payload, allowing to significantly improve upload performance and avoid head-of-line blocking (HoL) on backend connections shared between multiple clients when http-reuse is set to “always”. The advertised per-stream window is automatically adjusted to reflect the available space so that in practice it should not be required to touch tune.h2.be.initial-window-size. If less than the size required to deal with all streams is set, this minimum will be used. The default value is about 1600k (100 streams with 16kB buffers each).

See also: tune.h2.be.initial-window-size, tune.h2.fe.rxbuf, http-reuse.

tune.h2.fe.glitches-threshold <number>

tune.h2.fe.glitches-threshold <number>

Sets the threshold for the number of glitches on a frontend connection, where that connection will automatically be killed. This allows to automatically kill misbehaving connections without having to write explicit rules for them. The default value is zero, indicating that no threshold is set so that no event will cause a connection to be closed. Beware that some H2 clientss may occasionally cause a few glitches over long lasting connection, so any non-zero value here should probably be in the hundreds or thousands to be effective without affecting slightly bogus clients. It is also possible to only kill connections when the CPU usage crosses a certain level, by using “tune.glitches.kill.cpu-usage”. Note that a graceful close is attempted at 75% of the configured threshold by advertising a GOAWAY for a future stream. This ensures that a slightly non-compliant client will have the opportunity to create a new connection and continue to work unaffected without ever triggering the hard close thus risking to interrupt ongoing transfers.

See also: tune.h2.be.glitches-threshold, fc_glitches, and tune.glitches.kill.cpu-usage

tune.h2.fe.initial-window-size <number>

tune.h2.fe.initial-window-size <number>

Sets the HTTP/2 initial window size for incoming connections, which is the number of bytes the client can upload before waiting for an acknowledgment from HAProxy. This setting only affects payload contents (i.e. the body of POST requests), not headers. When not set, the common default value set by tune.h2.initial-window-size applies. It can make sense to increase this value to allow faster uploads. The default value equals tune.bufsize (16384) and allows at least 1.25 Mbps of bandwidth per stream over a 100 ms ping time, and 125 Mbps for 1 ms ping time. It doesn’t affect resource usage. Using too large values may cause clients to experience a lack of responsiveness if pages are accessed in parallel to large uploads. It is better to use tune.h2.fe.rxbuf instead, which does not cause any unfairness.

See also: tune.h2.initial-window-size.

tune.h2.fe.max-concurrent-streams <number> [args...]

tune.h2.fe.max-concurrent-streams <number> [args...]

Sets the HTTP/2 maximum number of concurrent streams per incoming connection (i.e. the number of outstanding requests on a single connection from a client). When not set, the default set by tune.h2.max-concurrent-streams applies. A larger value than the default 100 may sometimes slightly improve the page load time for complex sites with lots of small objects over high latency networks but can also result in using more memory by allowing a client to allocate more resources at once. The default value of 100 is generally good and it is recommended not to change this value. A larger concurrency also has an impact on the processing load and latency when dealing with large numbers of connections which are themselves using many streams, and it may lower the barrier to denial of service attacks. The command supports the following optional arguments after the number:

  • rq-load { <number> | auto | ignore }:
The optional argument "rq-load" permits to dynamically adjust the
advertised concurrency based on the executing thread's run-queue load:
as long as the thread's load remains below the indicated threshold, the
configured streams limit will be advertised. When the thread's load
increases beyond the configured limit, the advertised streams limit will be
decreased proportionally to the square of the excess ratio. Target load
levels between 50 and 100 generally show very good moderation under heavy
loads. Alternately, instead of specifying an explicit number, the keyword
accepts "ignore", which is the default and means that the thread's
run-queue load will not be considered to moderate the advertised streams
limit, and "auto", which sets the limit to the "tune.runqueue-depth"
value, which generally provides good results without having to tweak
the configuration any further.
  • min <number>:
This sets the minimum advertised concurrency level when rq-load is used,
even if this results in a higher load than the configured target. This
allows to maintain a good level of interactivity on a site under very
heavy load. The minimum and default value is 1, but values between 5
and 15 can improve user experience.

Example:

tune.h2.fe.max-concurrent-streams 100 rq-load auto min 15

tune.h2.fe.max-frames-at-once <number>

tune.h2.fe.max-frames-at-once <number>

Sets the maximum number of HTTP/2 incoming frames that will be processed at once on a frontend connection. It can be useful to set this to a low value (a few tens to a few hundreds) when dealing with very large buffers in order to maintain a low latency and a better fairness between multiple connections. The default value is zero, which means that no limitation is enforced.

tune.h2.fe.max-rst-at-once <number>

tune.h2.fe.max-rst-at-once <number>

Sets the maximum number of HTTP/2 incoming RST_STREAM that will be processed at once on a frontend connection. Once the specified number of RST_STREAM frames are received, the connection handler will be placed in a low priority queue and be processed after all other tasks. It can be useful to set this to a very low value (1 or a few units) to significantly reduce the impacts of RST_STREAM floods. RST_STREAM do happen when a user clicks on the Stop button in their browser, but the few extra milliseconds caused by this requeuing are generally unnoticeable, however they are generally effective at significantly lowering the load caused from such floods. The default value is zero, which means that no limitation is enforced.

tune.h2.fe.max-total-streams <number>

tune.h2.fe.max-total-streams <number>

Sets the HTTP/2 maximum number of total streams processed per incoming connection. Once this limit is reached, HAProxy will send a graceful GOAWAY frame informing the client that it will close the connection after all pending streams have been closed. In practice, clients tend to close as fast as possible when receiving this, and to establish a new connection for next requests. Doing this is sometimes useful and desired in situations where clients stay connected for a very long time and cause some imbalance inside a farm. For example, in some highly dynamic environments, it is possible that new load balancers are instantiated on the fly to adapt to a load increase, and that once the load goes down they should be stopped without breaking established connections. By setting a limit here, the connections will have a limited lifetime and will be frequently renewed, with some possibly being established to other nodes, so that existing resources are quickly released.

It’s important to understand that there is an implicit relation between this limit and “tune.h2.fe.max-concurrent-streams” above. Indeed, HAProxy will always accept to process any possibly pending streams that might be in flight between the client and the frontend, so the advertised limit will always automatically be raised by the value configured in max-concurrent-streams, and this value will serve as a hard limit above which a violation by a non-compliant client will result in the connection being closed. Thus when counting the number of requests per connection from the logs, any number between max-total-streams and (max-total-streams + max-concurrent-streams) may be observed depending on how fast streams are created by the client.

The default value is zero, which enforces no limit beyond those implied by the protocol (2^30 ~= 1.07 billion). Values around 1000 may already cause frequent connection renewal without causing any perceptible latency to most clients. Setting it too low may result in an increase of CPU usage due to frequent TLS reconnections, in addition to increased page load time. Please note that some load testing tools do not support reconnections and may report errors with this setting; as such it may be needed to disable it when running performance benchmarks. See also “tune.h2.fe.max-concurrent-streams”.

tune.h2.fe.rxbuf <size>

tune.h2.fe.rxbuf <size>

Sets the HTTP/2 receive buffer size for incoming connections, in bytes. This size will be rounded up to the next multiple of tune.bufsize and will be shared between all streams uploading data (both HEADERS and DATA frames). In any case, one buffer will always be granted to each stream, and 7/8 of the unused buffers will be shared between streams uploading payload, allowing to significantly improve upload performance. The advertised per-stream window is automatically adjusted to reflect the available space so that in practice it should not be required to touch tune.h2.fe.initial-window-size. If less than the size required to deal with all streams is set, this minimum will be used. The default value of 1600k (100 streams with 16kB buffers each) permits roughly 130 Mbps of upload speed for a client with a 100ms RTT.

See also: tune.h2.fe.initial-window-size and tune.h2.be.rxbuf.

tune.h2.header-table-size <number>

tune.h2.header-table-size <number>

Sets the HTTP/2 dynamic header table size. It defaults to 4096 bytes and cannot be larger than 65536 bytes. A larger value may help certain clients send more compact requests, depending on their capabilities. This amount of memory is consumed for each HTTP/2 connection. It is recommended not to change it.

tune.h2.initial-window-size <number>

tune.h2.initial-window-size <number>

Sets the default value for the HTTP/2 initial window size, on both incoming and outgoing connections. This value is used for incoming connections when tune.h2.fe.initial-window-size is not set, and by outgoing connections when tune.h2.be.initial-window-size is not set. This setting is used both as the initial value and as a minimum per stream. The default value equals 16384 (tune.bufsize), which for uploads roughly allows at least 1.25 Mbps of bandwidth per stream over a network showing a 100 ms ping time, or 125 Mbps over a 1-ms local network. When less receive buffers than the maximum are in use, within the limits defined by tune.h2.be.rxbuf and tune.h2.fe.rxbuf, unused buffers will be shared between receiving streams. As such there is normally no point in changing this default setting. Given that changing this default value will both increase upload speeds and cause more unfairness between clients on downloads, it is recommended to instead use the side-specific settings tune.h2.fe.initial-window-size and tune.h2.be.initial-window-size.

tune.h2.log-errors { none | connection | stream }

tune.h2.log-errors { none | connection | stream }

Sets the level of errors in the H2 demultiplexer that will generate a log. The default is “stream”, which means that any decoding error encountered in the demultiplexer will lead to the emission of a log. The “connection” value indicates that only logs that result in invalidating the connection will produce a log. Finally, “none” indicates that no decoding error will produce any log. It is recommended to set at least “connection” in order to detect protocol anomalies, even if this means temporarily switching to “none” during difficult periods.

tune.h2.max-concurrent-streams <number>

tune.h2.max-concurrent-streams <number>

Sets the default HTTP/2 maximum number of concurrent streams per connection (i.e. the number of outstanding requests on a single connection). This value is used for incoming connections when tune.h2.fe.max-concurrent-streams is not set, and for outgoing connections when tune.h2.be.max-concurrent-streams is not set. The default value is 100. The impact varies depending on the side so please see the two settings above for more details. It is recommended not to use this setting and to switch to the per-side ones instead. A value of zero disables the limit so a single client may create as many streams as allocatable by HAProxy. It is highly recommended not to change this value.

tune.h2.max-frame-size <number>

tune.h2.max-frame-size <number>

Sets the HTTP/2 maximum frame size that HAProxy announces it is willing to receive to its peers. The default value is the largest between 16384 and the buffer size (tune.bufsize). In any case, HAProxy will not announce support for frame sizes larger than buffers. The main purpose of this setting is to allow to limit the maximum frame size setting when using large buffers. Too large frame sizes might have performance impact or cause some peers to misbehave. It is highly recommended not to change this value.

tune.h2.zero-copy-fwd-send { on | off }

tune.h2.zero-copy-fwd-send { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy sends of data for the H2 multiplexer. It is enabled by default.

See also: tune.disable-zero-copy-forwarding

tune.http.cookielen <number>

tune.http.cookielen <number>

Sets the maximum length of captured cookies. This is the maximum value that the “capture cookie xxx len yyy” will be allowed to take, and any upper value will automatically be truncated to this one. It is important not to set too high a value because all cookie captures still allocate this size whatever their configured value (they share a same pool). This value is per request per response, so the memory allocated is twice this value per connection. When not specified, the limit is set to 63 characters. It is recommended not to change this value.

tune.http.logurilen <number>

tune.http.logurilen <number>

Sets the maximum length of request URI in logs. This prevents truncating long request URIs with valuable query strings in log lines. This is not related to syslog limits. If you increase this limit, you may also increase the ’log … len yyy’ parameter. Your syslog daemon may also need specific configuration directives too. The default value is 1024.

tune.http.maxhdr <number>

tune.http.maxhdr <number>

Sets the maximum number of headers allowed in received HTTP messages. When a message comes with a number of headers greater than this value (including the first line), it is rejected with a “400 Bad Request” status code for a request, or “502 Bad Gateway” for a response. The default value is 101, which is enough for all usages, considering that the widely deployed Apache server uses the same limit. It can be useful to push this limit further to temporarily allow a buggy application to work by the time it gets fixed. The accepted range is 1..32767. Keep in mind that each new header consumes 32bits of memory for each stream, so don’t push this limit too high.

Note that HTTP/1.1 is a text protocol, so there is no special limit when the message is sent. The limit during the message parsing is sufficient. HTTP/2 and HTTP/3 are binary protocols and require an encoding step. A limit is set too when headers are encoded to comply to limitation imposed by the protocols. This limit is large enough but not documented on purpose. The same limit is applied on the first steps of the decoding for the same reason.

tune.idle-pool.shared { full | on | off }

tune.idle-pool.shared { full | on | off }

Controls sharing idle connection pools between threads for a same server. It can be enabled for all threads in a same thread group (‘on’), enabled for all threads (‘full’) or disabled (‘off’). The default is to share them between threads in the same thread group (‘on’), in order to minimize the number of persistent connections to a server, and to optimize the connection reuse rate. Sharing with threads from other thread groups can have a performance impact, and is not enabled by default, but can be useful if maximizing connection reuse is a priority. To help with debugging or when suspecting a bug in HAProxy around connection reuse, it can be convenient to forcefully disable this idle pool sharing between multiple threads, and force this option to “off”. It is strongly recommended against disabling this option without setting a conservative value on “pool-low-conn” for all servers relying on connection reuse to achieve a high performance level, otherwise connections might be closed very often as the thread count increases.

tune.idletimer <timeout>

tune.idletimer <timeout>

Sets the duration after which HAProxy will consider that an empty buffer is probably associated with an idle stream. This is used to optimally adjust some packet sizes while forwarding large and small data alternatively. The decision to use splice() or to send large buffers in SSL is modulated by this parameter. The value is in milliseconds between 0 and 65535. A value of zero means that HAProxy will not try to detect idle streams. The default is 1000, which seems to correctly detect end user pauses (e.g. read a page before clicking). There should be no reason for changing this value. Please check tune.ssl.maxrecord below.

tune.listener.default-shards { by-process | by-thread | by-group }

tune.listener.default-shards { by-process | by-thread | by-group }

Normally, all “bind” lines will create a single shard, that is, a single socket that all threads of the process will listen to. With many threads, this is not very efficient, and may even induce some important overhead in the kernel for updating the polling state or even distributing events to the various threads. Modern operating systems support balancing of incoming connections, a mechanism that will consist in permitting multiple sockets to be bound to the same address and port, and to evenly distribute all incoming connections to these sockets so that each thread only sees the connections that are waiting in the socket it is bound to. This significantly reduces kernel-side overhead and increases performance in the incoming connection path. This is usually enabled in HAProxy using the “shards” setting on “bind” lines, which defaults to 1, meaning that each listener will be unique in the process. On systems with many processors, it may be more convenient to change the default setting to “by-thread” in order to always create one listening socket per thread, or “by-group” in order to always create one listening socket per thread group. Be careful about the file descriptor usage with “by-thread” as each listener will need as many sockets as there are threads. Also some operating systems (e.g. FreeBSD) are limited to no more than 256 sockets on a same address. Note that “by-group” will remain equivalent to “by-process” for default configurations involving a single thread group, and will fall back to sharing the same socket on systems that do not support this mechanism. The default is “by-group” with a fallback to “by-process” for systems or socket families that do not support multiple bindings.

tune.listener.multi-queue { on | fair | off }

tune.listener.multi-queue { on | fair | off }

Enables (‘on’ / ‘fair’) or disables (‘off’) the listener’s multi-queue accept which spreads the incoming traffic to all threads a “bind” line is allowed to run on instead of taking them for itself. This provides a smoother traffic distribution and scales much better, especially in environments where threads may be unevenly loaded due to external activity (network interrupts colliding with one thread for example). The default mode, “on”, optimizes the choice of a thread by picking in a sample the one with the less connections. It is often the best choice when connections are long-lived as it manages to keep all threads busy. A second mode, “fair”, instead cycles through all threads regardless of their instant load level. It can be better suited for short-lived connections, or on machines with very large numbers of threads where the probability to find the least loaded thread with the first mode is low. Finally it is possible to forcefully disable the redistribution mechanism using “off” for troubleshooting, or for situations where connections are short-lived and it is estimated that the operating system already provides a good enough distribution. The default is “on”.

tune.lua.bool-sample-conversion { normal | pre-3.1-bug }

tune.lua.bool-sample-conversion { normal | pre-3.1-bug }

Explicitly tell haproxy how haproxy sample objects should be handled when pushed to Lua. Indeed, when leveraging native converters, sample fetches or variables from Lua script (to name a few), haproxy converts the internal smp type to equivalent Lua type. Because of historical implementation, there is an ambiguity around boolean handling: when doing Lua -> haproxy smp conversion, booleans are properly preserved, but when doing haproxy smp -> Lua conversion, booleans were converted to integers by mistake. This means that a sample fetch or converter returning a boolean would return an integer 0 or 1 when leveraged from Lua. Unfortunately, in Lua, booleans and integers are not interchangeable. Thus, to avoid ambiguities, “tune.lua.bool-sample-conversion” must explicitly be set to either “normal” (which means dropping the historical behavior for better consistency) or “pre-3.1-bug” (enforce historical behavior to prevent existing script logic from misbehaving). If the option is not set explicitly and a Lua script is loaded from the configuration, haproxy will emit a warning, and the option will implicitly default to “pre-3.1-bug” to match with the historical behavior. It is recommended to set this option to “normal” after ensuring that in-use Lua scripts are properly handling bool haproxy samples as booleans.

This setting must be set before any “lua-load” or “lua-load-per-thread” directive for it to be considered, else it is ignored.

tune.lua.burst-timeout <timeout>

tune.lua.burst-timeout <timeout>

The “burst” execution timeout applies to any Lua handler. If the handler fails to finish or yield before timeout is reached, it will be aborted to prevent thread contention, to prevent traffic from not being served for too long, and ultimately to prevent the process from crashing because of the watchdog kicking in. Unlike other lua timeouts which are yield-cumulative, burst-timeout will ensure that the time spent in a single lua execution window does not exceed the configured timeout.

Yielding here means that the lua execution is effectively interrupted either through an explicit call to lua-yielding function such as core.(m)sleep() or core.yield(), or following an automatic forced-yield (see tune.lua.forced-yield) and that it will be resumed later when the related task is set for rescheduling. Not all lua handlers may yield: we have to make a distinction between yieldable handlers and unyieldable handlers.

For yieldable handlers (tasks, actions..), reaching the timeout means “tune.lua.forced-yield” might be too high for the system, reducing it could improve the situation, but it could also be a good idea to check if adding manual yields at some key points within the lua function helps or not. It may also indicate that the handler is spending too much time in a specific lua library function that cannot be interrupted.

For unyieldable handlers (lua converters, sample fetches), it could simply indicate that the handler is doing too much computation, which could result from an improper design given that such handlers, which often block the request execution flow, are expected to terminate quickly to allow the request processing to go through. A common resolution approach here would be to try to better optimize the lua function for speed since decreasing “tune.lua.forced-yield” won’t help.

This timeout only counts the pure Lua runtime. If the Lua does a core.sleep, the sleeping time is not taken in account. The default timeout is 1000ms.

Note: if a lua GC cycle is initiated from the handler (either explicitly requested or automatically triggered by lua after some time), the GC cycle time will also be accounted for.

Indeed, there is no way to deduce the GC cycle time, so this could lead to some false positives on saturated systems (where GC is having hard time to catch up and consumes most of the available execution runtime). If it were to be the case, here are some resolution leads:

- checking if the script could be optimized to reduce lua memory footprint
- fine-tuning lua GC parameters and / or requesting manual GC cycles
  (see: https://www.lua.org/manual/5.4/manual.html#pdf-collectgarbage)
- increasing tune.lua.burst-timeout

Setting value to 0 completely disables this protection.

tune.lua.forced-yield <number>

tune.lua.forced-yield <number>

This directive forces the Lua engine to execute a yield each <number> of instructions executed. This permits interrupting a long script and allows the HAProxy scheduler to process other tasks like accepting connections or forwarding traffic. The default value is 10000 instructions for scripts loaded using “lua-load-per-thread” and MAX(500, 10000 / nbthread) instructions for scripts loaded using “lua-load” (it was found to be an optimal value for performance while taking care of not creating thread contention with multiple threads competing for the global lua lock).

If HAProxy often executes some Lua code but more responsiveness is required, this value can be lowered. If the Lua code is quite long and its result is absolutely required to process the data, the <number> can be increased, but the value should be set wisely as in multithreading context it could increase contention.

tune.lua.log.loggers { on | off }

tune.lua.log.loggers { on | off }

Enables (‘on’) or disables (‘off’) logging the output of LUA scripts via the loggers applicable to the current proxy, if any.

Defaults to ‘on’.

tune.lua.log.stderr { on | auto | off }

tune.lua.log.stderr { on | auto | off }

Enables (‘on’) or disables (‘off’) logging the output of LUA scripts via stderr. When set to ‘auto’, logging via stderr is conditionally ‘on’ if any of:

- tune.lua.log.loggers is set to 'off'
- the script is executed in a non-proxy context with no global logger
- the script is executed in a proxy context with no logger attached

Please note that, when enabled, this logging is in addition to the logging configured via tune.lua.log.loggers.

Defaults to ‘auto’.

tune.lua.maxmem <number>

tune.lua.maxmem <number>

Sets the maximum amount of RAM in megabytes per process usable by Lua. By default it is zero which means unlimited. It is important to set a limit to ensure that a bug in a script will not result in the system running out of memory.

tune.lua.openlibs [all | none | <lib>[,<lib>...]]

tune.lua.openlibs [all | none | <lib>[,<lib>...]]

Selects which Lua standard libraries are loaded when initialising the Lua state. The argument is a comma-separated list of library names taken from the following set: table, io, os, string, math, utf8, package, debug. The special values “all” and “none” may be used instead of a list. “none” cannot be combined with library names. The default value is “all”.

The base and coroutine libraries are always loaded regardless of this setting: base provides core Lua functions that HAProxy relies on, and coroutine is required because HAProxy overrides coroutine.create() with its own safe implementation.

Note that fork() and new thread creation are already blocked by default in HAProxy regardless of this setting, and can only be re-enabled via the “insecure-fork-wanted” global directive. Restricting the set of loaded libraries further reduces the attack surface exposed to Lua scripts. In particular: - omitting “os” prevents os.execute() and os.exit() - omitting “io” prevents io.open() and io.popen() - omitting “package” prevents loading native C modules via require() - omitting “debug” prevents introspection of HAProxy internals via debug.getupvalue(), debug.getmetatable(), or debug.sethook()

Examples:

tune.lua.openlibs none                    # only base + coroutine
tune.lua.openlibs string,math,table,utf8  # safe subset, no I/O or OS
tune.lua.openlibs all                     # default, load everything

This setting must be set before any “lua-load”, “lua-load-per-thread” or “lua-prepend-path” directive, otherwise a parse error is returned.

tune.lua.service-timeout <timeout>

tune.lua.service-timeout <timeout>

This is the execution timeout for the Lua services. This is useful for preventing infinite loops or spending too much time in Lua. This timeout counts only the pure Lua runtime. If the Lua does a sleep, the sleep is not taken in account. The default timeout is 4s.

tune.lua.session-timeout <timeout>

tune.lua.session-timeout <timeout>

This is the execution timeout for the Lua sessions. This is useful for preventing infinite loops or spending too much time in Lua. This timeout counts only the pure Lua runtime. If the Lua does a sleep, the sleep is not taken in account. The default timeout is 4s.

tune.lua.task-timeout <timeout>

tune.lua.task-timeout <timeout>

Purpose is the same as “tune.lua.session-timeout”, but this timeout is dedicated to the tasks. By default, this timeout isn’t set because a task may remain alive during of the lifetime of HAProxy. For example, a task used to check servers.

tune.max-checks-per-thread <number>

tune.max-checks-per-thread <number>

Sets the number of active checks per thread above which a thread will actively try to search a less loaded thread to run the health check, or queue it until the number of active checks running on it diminishes. The default value is zero, meaning no such limit is set. It may be needed in certain environments running an extremely large number of expensive checks with many threads when the load appears unequal and may make health checks to randomly time out on startup, typically when using OpenSSL 3.0 which is about 20 times more CPU-intensive on health checks than older ones. This will have for result to try to level the health check work across all threads. The vast majority of configurations do not need to touch this parameter. Please note that too low values may significantly slow down the health checking if checks are slow to execute.

tune.maxaccept <number>

tune.maxaccept <number>

Sets the maximum number of consecutive connections a process may accept in a row before switching to other work. In single process mode, higher numbers used to give better performance at high connection rates, though this is not the case anymore with the multi-queue. This value applies individually to each listener, so that the number of processes a listener is bound to is taken into account. This value defaults to 4 which showed best results. If a significantly higher value was inherited from an ancient config, it might be worth removing it as it will both increase performance and lower response time. In multi-process mode, it is divided by twice the number of processes the listener is bound to. Setting this value to -1 completely disables the limitation. It should normally not be needed to tweak this value.

tune.maxpollevents <number>

tune.maxpollevents <number>

Sets the maximum amount of events that can be processed at once in a call to the polling system. The default value is adapted to the operating system. It has been noticed that reducing it below 200 tends to slightly decrease latency at the expense of network bandwidth, and increasing it above 200 tends to trade latency for slightly increased bandwidth. The configured value must be lower than or equal to 1000000.

tune.maxrewrite <number>

tune.maxrewrite <number>

Sets the reserved buffer space to this size in bytes. The reserved space is used for header rewriting or appending. The first reads on sockets will never fill more than bufsize-maxrewrite. Historically it has defaulted to half of bufsize, though that does not make much sense since there are rarely large numbers of headers to add. Setting it too high prevents processing of large requests or responses. Setting it too low prevents addition of new headers to already large requests or to POST requests. It is generally wise to set it to about 1024. It is automatically readjusted to half of bufsize if it is larger than that. This means you don’t have to worry about it when changing bufsize.

tune.max-rules-at-once <number>

tune.max-rules-at-once <number>

Sets the maximum number of rules that can be evaluated at once in ruleset evaluating functions, provided that they support yielding. Indeed, it is not rare to see configurations with a large number of “tcp-request content” or “http-request” rules for instance. A large number of rules combined with cpu-demanding actions (e.g.: actions that work on content) may create thread contention as all the rules from a given ruleset are evaluated under the same polling loop if the evaluation is not interrupted. This option ensures that no more than <number> number of rules may be executed under the same polling loop for content-oriented rulesets (those that already support yielding due to content inspection). What it does is that it forces the evaluating function to yield, so that it comes back on the next polling loop to continues the evaluation.

Affected rulesets are:

  • “tcp-request content”
  • “tcp-response content”
  • “http-request”
  • “http-response”

The default value is 50.

tune.memory.hot-size <number>

tune.memory.hot-size <number>

Sets the per-thread amount of memory that will be kept hot in the local cache and will never be recoverable by other threads. Access to this memory is very fast (lockless), and having enough is critical to maintain a good performance level under extreme thread contention. The value is expressed in bytes, and the default value is configured at build time via CONFIG_HAP_POOL_CACHE_SIZE which defaults to 524288 (512 kB). A larger value may increase performance in some usage scenarios, especially when performance profiles show that memory allocation is stressed a lot. Experience shows that a good value sits between once to twice the per CPU core L2 cache size. Too large values will have a negative impact on performance by making inefficient use of the L3 caches in the CPUs, and will consume larger amounts of memory. It is recommended not to change this value, or to proceed in small increments. In order to completely disable the per-thread CPU caches, using a very small value could work, but it is better to use “-dMno-cache” on the command-line.

tune.notsent-lowat.client <size>

tune.notsent-lowat.client <size>
tune.notsent-lowat.server <size>

Adjusts the kernel’s per-socket buffering so as to report that the sending side of a socket is full once the amount of buffered data equals this value plus the measured window size. The principle is to let the strict minimum needed amount of bytes in socket buffers, plus a small margin corresponding to what would be sent by the time haproxy tries to send again. Setting this to a low value (typically around tune.bufsize) allows to significantly reduce the memory consumption in system buffers, and reduce the application level latency incurred by flushing buffered data. This generally represents a more effective and more accurate setting than tune.sndbuf.client and tune.sndbuf.client for systems supporting it. This applies per connection (connection from a client or connection to a server depending on the setting) and is only used by TCP connections. The default is zero, which means unlimited. This is only available on Linux.

tune.pattern.cache-size <number>

tune.pattern.cache-size <number>

Sets the size of the pattern lookup cache to <number> entries. This is an LRU cache which reminds previous lookups and their results. It is used by ACLs and maps on slow pattern lookups, namely the ones using the “sub”, “reg”, “dir”, “dom”, “end”, “bin” match methods as well as the case-insensitive strings. It applies to pattern expressions which means that it will be able to memorize the result of a lookup among all the patterns specified on a configuration line (including all those loaded from files). It automatically invalidates entries which are updated using HTTP actions or on the CLI. The default cache size is set to 10000 entries, which limits its footprint to about 5 MB per process/thread on 32-bit systems and 8 MB per process/thread on 64-bit systems, as caches are thread/process local. There is a very low risk of collision in this cache, which is in the order of the size of the cache divided by 2^64. Typically, at 10000 requests per second with the default cache size of 10000 entries, there’s 1% chance that a brute force attack could cause a single collision after 60 years, or 0.1% after 6 years. This is considered much lower than the risk of a memory corruption caused by aging components. If this is not acceptable, the cache can be disabled by setting this parameter to 0.

tune.peers.max-updates-at-once <number>

tune.peers.max-updates-at-once <number>

Sets the maximum number of stick-table updates that haproxy will try to process at once when sending messages. Retrieving the data for these updates requires some locking operations which can be CPU intensive on highly threaded machines if unbound, and may also increase the traffic latency during the initial batched transfer between an older and a newer process. Conversely low values may also incur higher CPU overhead, and take longer to complete. The default value is 200 and it is suggested not to change it.

tune.pipesize <size>

tune.pipesize <size>

Sets the kernel pipe buffer size to this size (in bytes). By default, pipes are the default size for the system. But sometimes when using TCP splicing, it can improve performance to increase pipe sizes, especially if it is suspected that pipes are not filled and that many calls to splice() are performed. This has an impact on the kernel’s memory footprint, so this must not be changed if impacts are not understood.

tune.pool-high-fd-ratio <number>

tune.pool-high-fd-ratio <number>

This setting sets the max number of file descriptors (in percentage) used by HAProxy globally against the maximum number of file descriptors HAProxy can use before we start killing idle connections when we can’t reuse a connection and we have to create a new one. The default is 25 (one quarter of the file descriptor will mean that roughly half of the maximum front connections can keep an idle connection behind, anything beyond this probably doesn’t make much sense in the general case when targeting connection reuse).

tune.pool-low-fd-ratio <number>

tune.pool-low-fd-ratio <number>

This setting sets the max number of file descriptors (in percentage) used by HAProxy globally against the maximum number of file descriptors HAProxy can use before we stop putting connection into the idle pool for reuse. The default is 20.

tune.pt.zero-copy-forwarding { on | off }

tune.pt.zero-copy-forwarding { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy forwarding of data for the pass-through multiplexer. To be used, the kernel splicing must also be configured. It is enabled by default.

See also: tune.disable-zero-copy-forwarding, option splice-auto, option splice-request and option splice-response

tune.quic.be.cc.cubic-min-losses <number>

tune.quic.be.cc.cubic-min-losses <number>
tune.quic.fe.cc.cubic-min-losses <number>

Defines how many lost packets are needed for the Cubic congestion control algorithm to really consider a loss event. Normally, any loss event is considered as the result of a congestion and is sufficient for Cubic to restart from a smaller window. But experiments show that there can be a variety of causes for losses that are not at all caused by congestion and that can simply be qualified of spurious losses, and for which adjusting the window will have no effect, except slowing communication down. Poor radio signal, out-of-order delivery, high CPU usage on a client causing random delays, as well as system timer imprecision can be among the common causes for this. This setting allows to make Cubic a bit more tolerant to spurious losses, by changing the minimum number of cumulated losses between two ACKs to be considered as a loss event, which defaults to 1. Some significant gains have been observed experimentally, but always accompanied with an aggravation of the bandwidth wasted on retransmits, and an increased risk of saturation of congested links. The value 2 may be used for short periods of time to compare some metrics. Never go beyond 2 without an expert’s prior analysis of the situation. The default and minimum value is 1. Always use 1.

tune.quic.cc.cubic.min-losses <number> (deprecated)

tune.quic.cc.cubic.min-losses <number> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.cc.hystart { on | off }

tune.quic.be.cc.hystart { on | off }
tune.quic.fe.cc.hystart { on | off }

Enables (‘on’) or disabled (‘off’) the HyStart++ (RFC 9406) algorithm for QUIC connections used as a replacement for the slow start phase of congestion control algorithms which may cause high packet loss. It is disabled by default.

tune.quic.cc-hystart { on | off } (deprecated)

tune.quic.cc-hystart { on | off } (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.cc.max-frame-loss <number>

tune.quic.be.cc.max-frame-loss <number>
tune.quic.fe.cc.max-frame-loss <number>

Sets the limit for which a single QUIC frame can be marked as lost. If exceeded, the connection is considered as failing and is closed immediately.

The default value is 10.

tune.quic.max-frame-loss <number> (deprecated)

tune.quic.max-frame-loss <number> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.cc.max-win-size <size>

tune.quic.be.cc.max-win-size <size>
tune.quic.fe.cc.max-win-size <size>

Sets the default maximum window size for the congestion controller of a single QUIC connection either on frontend or backend side. The value must be written as an integer with an optional suffix ‘k’, ’m’ or ‘g’. It must be between 10k and 4g.

QUIC multiplexer also uses the current congestion window size to determine if it can allocate new stream buffers on data emission. As such, the maximum congestion window size also serves as a limit on this allocator.

The default value is 480k.

See also the “quic-cc-algo” bind and server options.

tune.quic.frontend.default-max-window-size <size> (deprecated)

tune.quic.frontend.default-max-window-size <size> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.cc.reorder-ratio <0..100, in percent>

tune.quic.be.cc.reorder-ratio <0..100, in percent>
tune.quic.fe.cc.reorder-ratio <0..100, in percent>

The ratio applied to the packet reordering threshold calculated. It may trigger a high packet loss detection when too small.

The default value is 50.

tune.quic.reorder-ratio <0..100, in percent> (deprecated)

tune.quic.reorder-ratio <0..100, in percent> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.max-idle-timeout <timeout>

tune.quic.be.max-idle-timeout <timeout>
tune.quic.fe.max-idle-timeout <timeout>

Sets the QUIC max_idle_timeout transport parameters on either frontend or backend side. It follows the HAProxy time format and is expressed in milliseconds. This determines the period of time after which a connection is silently closed if it has remained inactive during an effective period of time. Both endpoints relies on the same negotiated value: - the minimum of the two parameters if both are not null, - the maximum if only one of them is not null, - if both parameters are null, this feature is disabled.

The default value is 30s.

tune.quic.frontend.max-idle-timeout <timeout> (deprecated)

tune.quic.frontend.max-idle-timeout <timeout> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.sec.glitches-threshold <number>

tune.quic.be.sec.glitches-threshold <number>
tune.quic.fe.sec.glitches-threshold <number>

Sets the threshold for the number of glitches per connection either on frontend or backend side, where that connection will automatically be killed. This allows to automatically kill misbehaving connections without having to write explicit rules for them. The default value is zero, indicating that no threshold is set so that no event will cause a connection to be closed. Beware that some QUIC clients may occasionally cause a few glitches over long lasting connection, so any non- zero value here should probably be in the hundreds or thousands to be effective without affecting slightly bogus clients. It is also possible to only kill connections when the CPU usage crosses a certain level, by using “tune.glitches.kill.cpu-usage”.

See also: fc_glitches, tune.glitches.kill.cpu-usage

tune.quic.frontend.glitches-threshold <number> (deprecated)

tune.quic.frontend.glitches-threshold <number> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.fe.sec.retry-threshold <number>

tune.quic.fe.sec.retry-threshold <number>

Dynamically enables the Retry feature for all the configured QUIC listeners as soon as this number of half open connections is reached. A half open connection is a connection whose handshake has not already successfully completed or failed. To be functional this setting needs a cluster secret to be set, if not it will be silently ignored (see “cluster-secret” setting). This setting will be also silently ignored if the use of QUIC Retry was forced (see “quic-force-retry”).

The default value is 100.

See https://www.rfc-editor.org/rfc/rfc9000.html#section-8.1.2 for more information about QUIC retry.

tune.quic.retry-threshold <number> (deprecated)

tune.quic.retry-threshold <number> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.fe.sock-per-conn { default-on | force-off }

tune.quic.fe.sock-per-conn { default-on | force-off }

Specifies globally how QUIC frontend connections will use socket for receive/send operations. Connections can share listener socket or each connection can allocate its own socket.

The default value is “default-on”. This is used to allocate a dedicated socket for every QUIC connections. This option is the preferred one to achieve the best performance with a large QUIC traffic. This is also the only way to ensure soft-stop is conducted properly without data loss for QUIC connections and cases of transient errors during sendto() operation are handled efficiently. However, this relies on some advanced features from the UDP network stack. If your platform is deemed not compatible, haproxy will automatically switch to “force-off” mode on startup. Please note that QUIC listeners running on privileged ports may require to run as uid 0, or some OS-specific tuning to permit the target uid to bind such ports, such as system capabilities. See also the “setcap” global directive.

The “force-off” value indicates that QUIC transfers will occur on the shared listener socket. This option can be a good compromise for small traffic as it allows to reduce FD consumption. However, performance won’t be optimal due to a higher CPU usage if listeners are shared across a lot of threads or a large number of QUIC connections can be used simultaneously.

This setting is applied in conjunction with each “quic-socket” bind options. If “default-on” mode is used on global tuning, it will be activated for each listener, except for the ones with “quic-socket listener”. However, if “force-off” is used globally, it will be applied on every listener instance, regardless of their individual configuration.

tune.quic.socket-owner { connection | listener } (deprecated)

tune.quic.socket-owner { connection | listener } (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. The newer option is named “tune.quic.fe.sock-per-conn”, with legacy value “connection” corresponding to “default-on” and “listener” to “force-off”.

tune.quic.be.stream.data-ratio <0..100, in percent>

tune.quic.be.stream.data-ratio <0..100, in percent>
tune.quic.fe.stream.data-ratio <0..100, in percent>

This setting allows to configure the hard limit of the number of data bytes in flight over each stream. It is expressed as a percentage relative to the QUIC stream rxbuf connection setting, with the result rounded up to bufsize.

The default value is 90. This is suitable with the most frequent web scenario, where uploads is performed only for one or a few streams, whereas the rest are used for download only. If the stream rxbuf connection limit remains at a reasonable level, it ensures that only a portion of opened streams can allocate to their maximum capacity.

In the case of an application using many uploading streams in parallel and suffering from unfairness between these streams, it can make sense to reduce this ratio, to increase fairness and reduce the per-stream bandwidth.

See also: “tune.quic.be.stream.rxbuf”, “tune.quic.fe.stream.rxbuf”, “tune.quic.be.stream.max-concurrent”, “tune.quic.fe.stream.max-concurrent”

tune.quic.frontend.stream-data-ratio <0..100, in percent> (deprecated)

tune.quic.frontend.stream-data-ratio  <0..100, in percent> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.stream.max-concurrent <number>

tune.quic.be.stream.max-concurrent <number>
tune.quic.fe.stream.max-concurrent <number>

On frontend side, this is used as the value for the advertised initial_max_streams_bidi transport parameter. This is enforced as the maximum number of bidirectional streams that the remote peer will be authorized to open concurrently during the connection lifetime. This effectively limits the number of concurrent HTTP/3 client requests.

The default value is 100. Note that if you reduces it, it can restrict the buffering capabilities of streams on receive, which would result in poor upload throughput. It can be corrected by increasing the QUIC stream rxbuf connection setting.

On backend side, this is enforced locally by haproxy to limit the number of concurrent requests multiplexed over a single connection. This may be further restricted by the peer flow control. It may be necessary to reduce the default value of 100 to improve a site’s responsiveness at the expense of a higher number of opened backend connections. Similarly to the frontend side, this setting also directly impacts the Rx buffering capability, this time though limiting the HTTP download capacity. QUIC stream rxbuf setting can be increased when dealing mostly with HTTP responses larger than “tune.bufsize”.

See also: “tune.quic.be.stream.rxbuf”, “tune.quic.fe.stream.rxbuf”, “tune.quic.be.stream.data-ratio”, “tune.quic.fe.stream.data-ratio”

tune.quic.fe.stream.max-total <number>

tune.quic.fe.stream.max-total <number>

Sets the maximum number of requests that can be handled by a single QUIC connection. Once this total is reached, the connection will be gracefully shutdown. In HTTP/3, this translates to a GOAWAY frame. The connection is finally closed when all remaining transfers are completed.

This setting is applied as a hard limit on the connection via the QUIC flow control mechanism. If a peer violates it, the connection will be immediately closed.

This setting can be used to force clients to open new connections once in a while to continue the emission of requests and avoid maintaining connections for too many times. However, low values will increase latency on the client side, as well as CPU consumption on both sides due to TLS handshakes.

The default value is 0 which implies no specific limit outside of the QUIC protocol encoding limitation (2^60, more than a billion billion).

tune.quic.frontend.max-streams-bidi <number> (deprecated)

tune.quic.frontend.max-streams-bidi <number> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.stream.rxbuf <size>

tune.quic.be.stream.rxbuf <size>
tune.quic.fe.stream.rxbuf <size>

This setting is the hard limit for the number of data bytes in flight over a QUIC frontend connection. It is reused as the value for the initial_max_data transport parameter. It directly impacts the upload bandwidth for the peer depending on the latency and the per-connection memory consumption in haproxy.

By default, the value is set to 0, which indicates that it must be automatically generated as the product between max-concurrent and bufsize. This can be increased for example if a backend application relies on massive uploads over high latency networks.

See also: “tune.quic.be.stream.max-concurrent”, “tune.quic.fe.stream.max-concurrent”, “tune.quic.be.stream.data-ratio”, “tune.quic.fe.stream.data-ratio”

tune.quic.frontend.max-data-size <size> (deprecated)

tune.quic.frontend.max-data-size <size> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.tx.pacing { on | off }

tune.quic.be.tx.pacing { on | off }
tune.quic.fe.tx.pacing { on | off }

Enables (‘on’) or disables (‘off’) pacing support for QUIC emission. By default, it is active. The purpose of pacing is to smooth emission of data to reduce network losses. In most scenario, it will significantly improve network throughput by avoiding retransmissions. However, it can be useful to deactivate it for networks with very high bandwidth/low latency characteristics to prevent unwanted delay and reduce CPU consumption.

See also the “quic-cc-algo” bind and server options.

tune.quic.disable-tx-pacing (deprecated)

tune.quic.disable-tx-pacing (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.be.tx.udp-gso { on | off }

tune.quic.be.tx.udp-gso { on | off }
tune.quic.fe.tx.udp-gso { on | off }

Enables (‘on’) or disables (‘off’) UDP GSO support for QUIC emission. By default, it is active. This kernel feature allows to emit multiple datagrams via a single system call which is more efficient for large transfer. It may be useful to disable it on developers suggestion when suspecting an issue on emission.

tune.quic.disable-udp-gso (deprecated)

tune.quic.disable-udp-gso (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.listen { on | off }

tune.quic.listen { on | off }

Disable QUIC transport protocol on the frontend side. All the QUIC listeners will still be created, but they won’t listen for incoming datagrams. Hence, no QUIC traffic will be processed by haproxy on the frontend side.

The default value is “on”. If an issue is suspected with QUIC traffic, this option can be used to easily toggle QUIC listeners without messing with each individual config lines.

See also “quic_enabled” sample fetch.

tune.quic.mem.tx-max <size>

tune.quic.mem.tx-max <size>

Sets the maximum amount of memory usable by QUIC stack at the transport layer for emission. This serves both as a limit of in flight bytes and multiplexer output buffers. Note that to prevent threads contention this limit is not strictly enforced so it can be exceeded on some occasions. Also, each connection will always be able to use a window of at least 2 datagrams, so a proper maxconn should be used in conjunction.

tune.quic.frontend.max-tx-mem <size> (deprecated)

tune.quic.frontend.max-tx-mem <size> (deprecated)

This keyword has been deprecated in 3.3 and will be removed in 3.5. It is part of the streamlining process apply on QUIC configuration. If used, this setting will only be applied on frontend connections.

tune.quic.zero-copy-fwd-send { on | off }

tune.quic.zero-copy-fwd-send { on | off }

Enables (‘on’) of disabled (‘off’) the zero-copy sends of data for the QUIC multiplexer. It is enabled by default.

See also: tune.disable-zero-copy-forwarding

tune.renice.runtime <number>

tune.renice.runtime <number>

This configuration option takes a value between -20 and 19. It applies a scheduling priority as documented in man 2 setpriority. This priority is applied after the configuration parsing, which means only the worker or the standalone process will apply it. It is usually configured to set a higher priority than a process doing configuration parsing (tune.renice.startup).

See also: tune.renice.startup

tune.renice.startup <number>

tune.renice.startup <number>

This configuration option takes a value between -20 and 19. It applies a scheduling priority as documented in man 2 setpriority. This priority is applied before applying the rest of the configuration which can be useful if you want to lower the priority for configuration parsing. This is applied on the standalone process or the worker before configuration parsing. Once the configuration is parsed, the previous priority is restored unless tune.renice.runtime is used.

See also: tune.renice.runtime

tune.rcvbuf.backend <size>

tune.rcvbuf.backend  <size>
tune.rcvbuf.frontend <size>

For the kernel socket receive buffer size on non-connected sockets to this size. This can be used QUIC in listener mode and log-forward on the frontend. The default system buffers might sometimes be too small for sockets receiving lots of aggregated traffic, causing some losses and possibly retransmits (in case of QUIC), possibly slowing down connection establishment under heavy traffic. The value is expressed in bytes, applied to each socket. In listener mode, sockets are shared between all connections, and the total number of sockets depends on the “shards” value of the “bind” line. There’s no good value, a good one corresponds to an expected size per connection multiplied by the expected number of connections. The kernel may trim large values. See also “tune.rcvbuf.client” and “tune.rcvbuf.server” for their connected socket counter parts, as well as “tune.sndbuf.backend” and “tune.sndbuf.frontend” for the send setting.

tune.rcvbuf.client <size>

tune.rcvbuf.client <size>
tune.rcvbuf.server <size>

Forces the kernel socket receive buffer size on the client or the server side to the specified value in bytes. This value applies to all TCP/HTTP frontends and backends. It should normally never be set, and the default size (0) lets the kernel auto-tune this value depending on the amount of available memory. However it can sometimes help to set it to very low values (e.g. 4096) in order to save kernel memory by preventing it from buffering too large amounts of received data. Lower values will significantly increase CPU usage though.

tune.recv_enough <size>

tune.recv_enough <size>

HAProxy uses some hints to detect that a short read indicates the end of the socket buffers. One of them is that a read returns more than <recv_enough> bytes, which defaults to 10136 (7 segments of 1448 each). This default value may be changed by this setting to better deal with workloads involving lots of short messages such as telnet or SSH sessions.

tune.ring.queues <number>

tune.ring.queues <number>

Sets the number of write queues in front of ring buffers. This can have an effect on the CPU usage of traces during debugging sessions, and both too low or too large a value can have an important effect. The good value was determined experimentally by developers and there should be no reason to try to change it unless instructed to do so in order to try to address specific issues. Such a setting should not be left in the configuration across version upgrades because its optimal value may evolve over time.

tune.runqueue-depth <number>

tune.runqueue-depth <number>

Sets the maximum amount of task that can be processed at once when running tasks. The default value depends on the number of threads but sits between 35 and 280, which tend to show the highest request rates and lowest latencies. Increasing it may incur latency when dealing with I/Os, making it too small can incur extra overhead. Higher thread counts benefit from lower values. When experimenting with much larger values, it may be useful to also enable tune.sched.low-latency and possibly tune.fd.edge-triggered to limit the maximum latency to the lowest possible.

tune.sched.low-latency { on | off }

tune.sched.low-latency { on | off }

Enables (‘on’) or disables (‘off’) the low-latency task scheduler. By default HAProxy processes tasks from several classes one class at a time as this is the most efficient. But when running with large values of tune.runqueue-depth this can have a measurable effect on request or connection latency. When this low-latency setting is enabled, tasks of lower priority classes will always be executed before other ones if they exist. This will permit to lower the maximum latency experienced by new requests or connections in the middle of massive traffic, at the expense of a higher impact on this large traffic. For regular usage it is better to leave this off. The default value is off.

tune.sndbuf.backend <size>

tune.sndbuf.backend  <size>
tune.sndbuf.frontend <size>

For the kernel socket send buffer size on non-connected sockets to this size. This can be used for UNIX socket and UDP logging on the backend side, and for QUIC in listener mode on the frontend. The default system buffers might sometimes be too small for sockets shared between many connections (or log senders), causing some losses and possibly retransmits, slowing down new connection establishment under high traffic. The value is expressed in bytes, applied to each socket. In listener mode, sockets are shared between all connections, and the total number of sockets depends on the “shards” value of the “bind” line. There’s no good value, a good one corresponds to an expected size per connection multiplied by the expected number of connections. The kernel may trim large values. See also “tune.sndbuf.client” and “tune.sndbuf.server” for their connected socket counter parts, as well as “tune.rcvbuf.backend” and “tune.rcvbuf.frontend” for the receive setting.

tune.sndbuf.client <size>

tune.sndbuf.client <size>
tune.sndbuf.server <size>

Forces the kernel socket send buffer size on the client or the server side to the specified value in bytes. This value applies to all TCP/HTTP frontends and backends. It should normally never be set, and the default size (0) lets the kernel auto-tune this value depending on the amount of available memory. However it can sometimes help to set it to very low values (e.g. 4096) in order to save kernel memory by preventing it from buffering too large amounts of received data. Lower values will significantly increase CPU usage though. Another use case is to prevent write timeouts with extremely slow clients due to the kernel waiting for a large part of the buffer to be read before notifying HAProxy again. See also tune.notsent-lowat.client and tune.notsent-lowat.server for more effective settings to more finely control memory usage and responsiveness on Linux without hurting performance.

tune.ssl.cachesize <number>

tune.ssl.cachesize <number>

Sets the size of the global SSL session cache, in a number of blocks. A block is large enough to contain an encoded session without peer certificate. An encoded session with peer certificate is stored in multiple blocks depending on the size of the peer certificate. A block uses approximately 200 bytes of memory (based on sizeof(struct sh_ssl_sess_hdr) + SHSESS_BLOCK_MIN_SIZE calculation used for shctx_init function). The default value may be forced at build time, otherwise defaults to 20000. When the cache is full, the most idle entries are purged and reassigned. Higher values reduce the occurrence of such a purge, hence the number of CPU-intensive SSL handshakes by ensuring that all users keep their session as long as possible. All entries are pre-allocated upon startup. Setting this value to 0 disables the SSL session cache.

tune.ssl.capture-buffer-size <number>

tune.ssl.capture-buffer-size <number>
tune.ssl.capture-cipherlist-size <number> (deprecated)

Sets the maximum size of the buffer used for capturing client hello cipher list, extensions list, elliptic curves list and elliptic curve point formats. If the value is 0 (default value) the capture is disabled, otherwise a buffer is allocated for each SSL/TLS connection.

tune.ssl.certificate-compression { auto | off }

tune.ssl.certificate-compression { auto | off }

This setting allows to configure the certificate compression support which is an extension (RFC 8879) to TLS 1.3.

When set to “auto” it uses the default value of the TLS library.

With “off” it tries to explicitly disable the support of the feature. HAProxy won’t try to send compressed certificates anymore nor accept compressed certificates.

Configures both backend and frontend sides.

This keyword is supported by OpenSSL >= 3.2.0.

The default value is auto.

tune.ssl.default-dh-param <number>

tune.ssl.default-dh-param <number>

Sets the maximum size of the Diffie-Hellman parameters used for generating the ephemeral/temporary Diffie-Hellman key in case of DHE key exchange. The final size will try to match the size of the server’s RSA (or DSA) key (e.g, a 2048 bits temporary DH key for a 2048 bits RSA key), but will not exceed this maximum value. Only 1024 or higher values are allowed. Higher values will increase the CPU load, and values greater than 1024 bits are not supported by Java 7 and earlier clients. This value is not used if static Diffie-Hellman parameters are supplied either directly in the certificate file or by using the ssl-dh-param-file parameter. If there is neither a default-dh-param nor a ssl-dh-param-file defined, and if the server’s PEM file of a given frontend does not specify its own DH parameters, then DHE ciphers will be unavailable for this frontend.

tune.ssl.force-private-cache

tune.ssl.force-private-cache

This option disables SSL session cache sharing between all processes. It should normally not be used since it will force many renegotiations due to clients hitting a random process. But it may be required on some operating systems where none of the SSL cache synchronization method may be used. In this case, adding a first layer of hash-based load balancing before the SSL layer might limit the impact of the lack of session sharing.

tune.ssl.hard-maxrecord <number>

tune.ssl.hard-maxrecord <number>

Sets the maximum amount of bytes passed to SSL_write() at any time. Default value 0 means there is no limit. In contrast to tune.ssl.maxrecord this settings will not be adjusted dynamically. Smaller records may decrease throughput, but may be required when dealing with low-footprint clients.

tune.ssl.keylog { on | off }

tune.ssl.keylog { on | off }

This option activates the logging of the TLS keys. It should be used with care as it will consume more memory per SSL session and could decrease performances. This is disabled by default.

These sample fetches should be used to generate the SSLKEYLOGFILE that is required to decipher traffic with wireshark.

https://tlswg.org/sslkeylogfile/draft-ietf-tls-keylogfile.html

The SSLKEYLOG is a series of lines which are formatted this way:

<Label> <space> <ClientRandom> <space> <Secret>

The ClientRandom is provided by the %[ssl_fc_client_random,hex] sample fetch, the secret and the Label could be find in the array below. You need to generate a SSLKEYLOGFILE with all the labels in this array.

The following sample fetches are hexadecimal strings and does not need to be converted.

  SSLKEYLOGFILE Label             |  Sample fetches for the Secrets
  --------------------------------|-----------------------------------------
  CLIENT_EARLY_TRAFFIC_SECRET     |  %[ssl_xx_client_early_traffic_secret]
  CLIENT_HANDSHAKE_TRAFFIC_SECRET |  %[ssl_xx_client_handshake_traffic_secret]
  SERVER_HANDSHAKE_TRAFFIC_SECRET |  %[ssl_xx_server_handshake_traffic_secret]
  CLIENT_TRAFFIC_SECRET_0         |  %[ssl_xx_client_traffic_secret_0]
  SERVER_TRAFFIC_SECRET_0         |  %[ssl_xx_server_traffic_secret_0]
  EXPORTER_SECRET                 |  %[ssl_xx_exporter_secret]
  EARLY_EXPORTER_SECRET           |  %[ssl_xx_early_exporter_secret]

These fetches exists for frontend (fc) or backend (bc) sides, replace “xx” by “fc” or “bc” to use the right side.

This is only available with OpenSSL 1.1.1, and useful with TLS1.3 session.

If you want to generate the content of a SSLKEYLOGFILE with TLS < 1.3, you only need this line:

“CLIENT_RANDOM %[ssl_fc_client_random,hex] %[ssl_fc_session_key,hex]”

A complete keylog could be generate with a log-format these way, even though this is not ideal for syslog:

log-format "CLIENT_EARLY_TRAFFIC_SECRET %[ssl_bc_client_random,hex] %[ssl_bc_client_early_traffic_secret]\n
            CLIENT_HANDSHAKE_TRAFFIC_SECRET %[ssl_bc_client_random,hex] %[ssl_bc_client_handshake_traffic_secret]\n
            SERVER_HANDSHAKE_TRAFFIC_SECRET %[ssl_bc_client_random,hex] %[ssl_bc_server_handshake_traffic_secret]\n
            CLIENT_TRAFFIC_SECRET_0 %[ssl_bc_client_random,hex] %[ssl_bc_client_traffic_secret_0]\n
            SERVER_TRAFFIC_SECRET_0 %[ssl_bc_client_random,hex] %[ssl_bc_server_traffic_secret_0]\n
            EXPORTER_SECRET %[ssl_bc_client_random,hex] %[ssl_bc_exporter_secret]\n
            EARLY_EXPORTER_SECRET %[ssl_bc_client_random,hex] %[ssl_bc_early_exporter_secret]"

HAProxy also provides the above formats as predefined environment variables that can be used directly in a “log-format” directive:

$HAPROXY_KEYLOG_FC_LOG_FMT   frontend (client-facing) connection keys
$HAPROXY_KEYLOG_BC_LOG_FMT   backend (server-facing) connection keys

tune.ssl.keyupdate-rate-limit <limit>

tune.ssl.keyupdate-rate-limit <limit>

Limit the amount of KeyUpdate per second we’re willing to accept to <limit> before considering it flood, and killing the connection. Dealing with KeyUpdate is cpu-expensive, and there is little reason to receive a lot of them. Using a value of “0” disables the rate limiting. The default value is 100.

tune.ssl.lifetime <timeout>

tune.ssl.lifetime <timeout>

Sets how long a cached SSL session may remain valid. This time is expressed in seconds and defaults to 300 (5 min). It is important to understand that it does not guarantee that sessions will last that long, because if the cache is full, the longest idle sessions will be purged despite their configured lifetime. The real usefulness of this setting is to prevent sessions from being used for too long.

tune.ssl.maxrecord <number>

tune.ssl.maxrecord <number>

Sets the maximum amount of bytes passed to SSL_write() at the beginning of the data transfer. Default value 0 means there is no limit. Over SSL/TLS, the client can decipher the data only once it has received a full record. With large records, it means that clients might have to download up to 16kB of data before starting to process them. Limiting the value can improve page load times on browsers located over high latency or low bandwidth networks. It is suggested to find optimal values which fit into 1 or 2 TCP segments (generally 1448 bytes over Ethernet with TCP timestamps enabled, or 1460 when timestamps are disabled), keeping in mind that SSL/TLS add some overhead. Typical values of 1419 and 2859 gave good results during tests. Use “strace -e trace=write” to find the best value. HAProxy will automatically switch to this setting after an idle stream has been detected (see tune.idletimer above). See also tune.ssl.hard-maxrecord.

tune.ssl.ssl-ctx-cache-size <number>

tune.ssl.ssl-ctx-cache-size <number>

Sets the size of the cache used to store generated certificates to <number> entries. This is a LRU cache. Because generating a SSL certificate dynamically is expensive, they are cached. The default cache size is set to 1000 entries.

tune.streams-elasticity <number>

tune.streams-elasticity <number>

Defines a target percentage of streams per frontend connection relative to the maximum number of concurrent connections (maxconn) when all connections are established. This metric applies to multiplexed protocols like HTTP/2 or QUIC, where each connection may receive multiple streams. At least one is always guaranteed, so the percentage must be at least 100%. During connection setup, HAProxy dynamically advertises additional streams up to the configured limit, maintaining the target ratio. At connection establishment, every frontend connection receives at least one stream; extra streams are assigned based on the target percentage and configured stream limits. This ensures efficient stream allocation under varying load conditions (more streams at low loads, fewer at high loads).

Highly dynamic sites with many objects per page benefit from high ratios, enabling many streams per connection. Sites using fewer streams on average (WebSocket, application code) may prefer small ratios closer to 120 or 150 (20 to 50% more streams than connections) preventing excessive stream counts under sustained loads.

The default value is 0, meaning no enforcement at this level, so only H2 and QUIC configurations apply (with the default setting of 100 streams per connection, this corresponds to 10000%). This remains the recommended setting for small deployments (maxconn around a thousand). Moderately sized setups (few thousands to tens of thousands connections) typically set the ratio between 1000 and 5000, allowing 10 to 50 streams per connection at full load. Large-scale deployments (hundreds of thousands to millions connections) might use lower values (120 to 200) to support 1.2 to 2 streams per connection on average at full load.

Contrary to HTTP/2, QUIC is capable to dynamically adjust the number of concurrent streams during the connection lifetime. However, QUIC flow control is stricter than HTTP/2, thus it is preferable when using it to specify values big enough to prevent extra latency on the connection. There is also a limitation for QUIC listeners with enabled 0-RTT. In this case, the initial value advertised to the peer will ignore stream elasticity and instead rely solely on the “tune.quic.fe.stream.max-concurrent” setting. However, the stream elasticity principle will still be effective past this initial annoucement during the connection lifetime.

Monitoring the total number of active streams on backends, including queues, provides a practical indicator of a sustainable target load and helps avoid over-provisioning.

tune.stick-counters <number>

tune.stick-counters <number>

Sets the number of stick-counters that may be tracked at the same time by a connection or a request via “track-sc*” actions in “tcp-request” or “http-request” rules. The default value is set at build time by the macro MAX_SESS_STK_CTR, and defaults to 3. With this setting it is possible to change the value and ignore the one passed at build time, but it cannot be set to a value greater than 100. Increasing this value may be needed when porting complex configurations to haproxy, but users are warned against the costs: each entry takes 16 bytes per connection and 16 bytes per request, all of which need to be allocated and zeroed for all requests even when not used. As such a value of 10 will inflate the memory consumption per request by 320 bytes and will cause this memory to be erased for each request, which does have measurable CPU impacts. Conversely, when no “track-sc” rules are used, the value may be lowered (0 being valid to entirely disable stick-counters).

tune.takeover-other-tg-connections <value>

tune.takeover-other-tg-connections <value>

By default, we won’t attempt to use idle connections from other thread groups. This can however be changed. Valid values for <value> are: “none”, the default, if used, no attempt will be made to use idle connections from other thread groups, “restricted” where we will only attempt to get an idle connection from another thread if we’re using protocols that can’t create new connections, such as reverse http, as well as when using strict-maxconn, and “full” where we will always look in other thread groups for idle connections. Note that using connections from other thread groups can occur performance penalties, so it should not be used unless really needed. Note that this behavior is now controlled by tune.idle-pool.shared, and this keyword is just there for compatibility with older configurations, and will be deprecated.

tune.vars.global-max-size <size>

tune.vars.global-max-size <size>
tune.vars.proc-max-size <size>
tune.vars.reqres-max-size <size>
tune.vars.sess-max-size <size>
tune.vars.txn-max-size <size>

These five tunes help to manage the maximum amount of memory used by the variables system. “global” limits the overall amount of memory available for all scopes. “proc” limits the memory for the process scope, “sess” limits the memory for the session scope, “txn” for the transaction scope, and “reqres” limits the memory for each request or response processing. Memory accounting is hierarchical, meaning more coarse grained limits include the finer grained ones: “proc” includes “sess”, “sess” includes “txn”, and “txn” includes “reqres”.

For example, when “tune.vars.sess-max-size” is limited to 100, “tune.vars.txn-max-size” and “tune.vars.reqres-max-size” cannot exceed 100 either. If we create a variable “txn.var” that contains 100 bytes, all available space is consumed. Notice that exceeding the limits at runtime will not result in an error message, but values might be cut off or corrupted. So make sure to accurately plan for the amount of space needed to store all your variables.

tune.zlib.memlevel <number>

tune.zlib.memlevel <number>

Sets the memLevel parameter in zlib initialization for each stream. It defines how much memory should be allocated for the internal compression state. A value of 1 uses minimum memory but is slow and reduces compression ratio, a value of 9 uses maximum memory for optimal speed. Can be a value between 1 and 9. The default value is 8.

tune.zlib.windowsize <number>

tune.zlib.windowsize <number>

Sets the window size (the size of the history buffer) as a parameter of the zlib initialization for each stream. Larger values of this parameter result in better compression at the expense of memory usage. Can be a value between 8 and 15. The default value is 15.

3.3. Debugging

anonkey <key>

anonkey <key>

This sets the global anonymizing key to <key>, which must be a 32-bit number between 0 and 4294967295. This is the key that will be used by default by CLI commands when anonymized mode is enabled. This key may also be set at runtime from the CLI command “set anon global-key”. See also command line argument “-dC” in the management manual.

debug.counters { on | off }

debug.counters { on | off }

Enables (‘on’) or disables (‘off’) the updating of event counters in the code. These are the counters reported under the type “CNT” in the CLI command “debug counters”. These counters are only available when the code was build with DEBUG_COUNTERS set to a value 1 or above. With the value 1, the counters are not updated by default (“debug.counters off”), and with value 2, they are updated by default (“debug.counters on”). There is normally no reason to change this setting unless a developer requests it, or unless it is suspected to consume abnormal amounts of CPU (in which case a report to developers is necessary with a dump of the counters). It is also possible to change this status at run time using the “debug counters” CLI command. Please consult the management manual.

force-cfg-parser-pause <timeout>

force-cfg-parser-pause <timeout>

This command is pausing the configuration parser for <timeout> milliseconds. This is useful for development or for testing timeouts of init scripts, particularly to simulate a very long reload. It requires the expose-experimental-directives to be set.

<timeout> is the timeout value specified in milliseconds by default, but can be in any other unit if the number is suffixed by the unit, as explained at the top of this document.

Example:

global
    expose-experimental-directives
    force-cfg-parser-pause 10s

quick-exit

quick-exit

This speeds up the old process exit upon reload by skipping the releasing of memory objects and listeners, since all of these are reclaimed by the operating system at the process’ death. The gains are only marginal (in the order of a few hundred milliseconds for huge configurations at most). The main target usage in fact is when a bug is spotted in the deinit() code, as this allows to bypass it. It is better not to use this unless instructed to do so by developers.

quiet

quiet

Do not display any message during startup. It is equivalent to the command-line argument “-q”.

warn-blocked-traffic-after <time>

warn-blocked-traffic-after <time>

This allows to adjust the delay after which a stuck task blocking the traffic will trigger the emission of a warning on the standard error output. The delay is expressed in milliseconds and defaults to 100 ms. Permitted values must be comprised between 1 ms and 1000 ms included. Lower values will trigger warnings frequently and higher ones will rarely. The watchdog will kill a runaway task that fails to respond twice for one second anyway, so a 1000 ms warning delay will normally not trigger any warning. It is recommended to stay with values between 10 and 100ms to detect configuration anomalies that may degrade the user’s experience, causing long response times or jerkiness on interactive sessions. For example, a poorly designed Lua sample-fetch function doing heavy computations, or a very large map_reg or map_regm map file with a very high evaluation cost may cause such trouble. For comparison a TLS handshake can eat between one and two milliseconds, and compressing a 16kB HTTP response buffer is around one millisecond. The output contains a thread dump of the offending task with a backtrace and some context that helps figure where the time is being spent.

zero-warning

zero-warning

When this option is set, HAProxy will refuse to start if any warning was emitted while processing the configuration and applying it. It means that warnings about bad combinations of parameters, warnings about very high limits that couldn’t be set, and so on, make the process exit with an error during startup. A few late startup warnings cannot be caught by this option, such as the failure to drop supplementary groups when changing the group ID in “daemon” or “master-worker” modes, or the failure to mark the process dumpable after the fork(). This option does not catch warnings emitted at runtime. It is highly recommended to set this option on configurations that are not changed often, as it helps to detect subtle mistakes and keep the configuration clean and forward-compatible. Note that “haproxy -c” will also report errors in such a case. This option is equivalent to command line argument “-dW”.

3.4. HTTPClient tuning

HTTPClient is an internal HTTP library, it can be used by various subsystems, for example in LUA scripts. HTTPClient is not used in the data path, in other words it has nothing with HTTP traffic passing through HAProxy.

httpclient.resolvers.disabled <on|off>

httpclient.resolvers.disabled <on|off>

Disable the DNS resolution of the httpclient. Prevent the creation of the “default” resolvers section.

Default value is off.

httpclient.resolvers.id <resolvers id>

httpclient.resolvers.id <resolvers id>

This option defines the resolvers section with which the httpclient will try to resolve.

Default option is the “default” resolvers ID. By default, if this option is not used, it will simply disable the resolving if the section is not found.

However, when this option is explicitly enabled it will trigger a configuration error if it fails to load.

httpclient.resolvers.prefer <ipv4|ipv6>

httpclient.resolvers.prefer <ipv4|ipv6>

This option allows to chose which family of IP you want when resolving, which is convenient when IPv6 is not available on your network. Default option is “ipv6”.

httpclient.retries <number>

httpclient.retries <number>

This option allows to configure the number of retries attempt of the httpclient when a request failed. This does the same as the “retries” keyword in a backend.

Default value is 3.

httpclient.ssl.ca-file <cafile>

httpclient.ssl.ca-file <cafile>

This option defines the ca-file which should be used to verify the server certificate. It takes the same parameters as the “ca-file” option on the server line.

By default and when this option is not used, the value is “@system-ca” which tries to load the CA of the system. If it fails the SSL will be disabled for the httpclient.

However, when this option is explicitly enabled it will trigger a configuration error if it fails.

httpclient.ssl.verify [none|required]

httpclient.ssl.verify [none|required]

Works the same way as the verify option on server lines. If specified to ’none’, servers certificates are not verified. Default option is “required”.

By default and when this option is not used, the value is “required”. If it fails the SSL will be disabled for the httpclient.

However, when this option is explicitly enabled it will trigger a configuration error if it fails.

httpclient.timeout.connect <timeout>

httpclient.timeout.connect <timeout>

Set the maximum time to wait for a connection attempt by default for the httpclient.

Arguments:

<timeout> is the timeout value specified in milliseconds by default, but
          can be in any other unit if the number is suffixed by the unit,
          as explained at the top of this document.

The default value is 5000ms.