Networking · Concept · 12 min read

What a Load Balancer Does, and the Two Decisions Behind It

Every glossary leads with the algorithms, which is the part that changes least. What decides whether the thing helps at all is the health check, and the first two depths of check will call a dead server healthy.

Written by Marko Ristic, Editor Updated Sep 12, 2026
2Check depths that will call a dead server healthy
3Error codes that come from the middle, not the application
2Load balancers you need, because one is a single point of failure
0Availability that sticky sessions leave you for existing users
Short answer

A load balancer sits in front of a group of servers and decides which one handles each incoming request. That is the easy half.

The half that decides whether it improves anything is the health check: it only adds availability if it can tell a broken server from a working one, and a check that confirms a port is open will keep sending users to a machine whose application has died.

  • It distributes traffic across a pool and removes members that fail a check
  • Layer 4 balances connections; layer 7 reads the request and routes on it
  • The health check is what turns distribution into actual availability
  • Sticky sessions undo much of the benefit, and the fix is in the application
  • One load balancer is a single point of failure, so they run in pairs
On this page

The three formsHardware, software, or a cloud service

The same job ships in three forms, and the choice is mostly about who operates it.

Hardware load balancers are dedicated appliances, historically from F5 or Citrix, sold as application delivery controllers. They are fast, they terminate a great deal of TLS in silicon, and they are bought in pairs with a support contract. This is still the normal answer in a data center with its own racks.

Software load balancers run on ordinary servers or virtual machines. HAProxy and NGINX are the common ones, and Envoy underneath most service meshes. They handle the same traffic as an appliance across multiple hosts. Same features, no appliance, and the operational burden moves to your team.

Cloud load balancing services are the managed form: an AWS Application or Network Load Balancer, an Azure Load Balancer or Application Gateway, a Google Cloud load balancer. The provider runs the redundancy and the scaling, and the trade is that the configuration surface is theirs rather than yours.

What follows applies to all three, because the decisions below are about load balancing itself rather than about which product does it.

F5 LTMF5 BIG-IP LTM, and how its names map onto this page

F5 sells the hardware form above as BIG-IP, and the part of it that load balances is Local Traffic Manager, LTM.

An LTM load balancer is the same machine this page describes, in F5's vocabulary, and the vocabulary is most of what makes F5 documentation hard going at first. Five of its terms map straight onto ideas explained further down.

F5 termWhat F5 means by itThe idea on this page
Virtual serverAn IP address and a service, such as 192.168.20.10:80, that clients send traffic toThe front end clients connect to
PoolA set of devices, such as web servers, grouped to receive and process trafficThe backend servers
Pool memberOne server in the pool, standing for a node on the networkA single backend
MonitorA check on pool members and nodes, repeated at a set intervalThe health check
Load balancing methodThe algorithm that picks a pool member for each requestThe algorithm

The flow, in F5's own terms: once a pool is assigned to a virtual server, LTM directs traffic coming into the virtual server to a member of that pool.

Which member is decided by the pool's load balancing method, and the default is Round Robin, each incoming request going to the next available member. When a member does not answer its monitor within the timeout, the system can send the traffic to another member instead.

Nothing later on this page stops applying on an F5 LTM. The health check, the algorithm and the failure modes are the same decisions under different names, which is why the rest of this page is written about load balancing rather than about any one product.

GTM and LTMGTM vs LTM: one answers DNS, the other carries the connection

The other half of the F5 pair is the Global Traffic Manager, GTM, which F5 now sells as BIG-IP DNS. Older material calls it F5 GTM, and the product and the job are the same.

F5 describes it as a system that monitors the availability and performance of global resources and uses that information to manage network traffic patterns, and the mechanism is in the new name: it distributes DNS name resolution requests, first to the best available pool in what F5 calls a wide IP, then to the best available virtual server within that pool.

So the two make different parts of the same decision. GTM chooses which site's address a client is given when it looks the name up. LTM, at that address, chooses which server handles the connection.

A GTM load balancer never carries the application traffic itself: by the time the client connects, its part is over. That follows from how it works, through DNS answers, rather than from any F5 sentence.

How many sites serve the application decides which you need. One data center needs LTM and no GTM. Two data centers that should both serve users, or one standing by for the other, is the case GTM exists for, and the virtual servers it chooses between are the ones at each site.

The DNS cache is the catch. A client that already has an answer keeps using it until the cached record expires, the same caching every DNS lookup goes through.

A GTM failover is therefore only as fast as the time to live on the name allows, and a long one means some users keep going to the failed site until it runs out.

The layerLayer 4 or layer 7, and what it costs

This is the first decision in any load balancing design and it constrains every other one.

Layer 4 load balancing works at the transport layer. It sees source and destination addresses and port numbers and nothing above that. It picks a server for a connection and forwards the packets.

It is fast, it barely touches the traffic, and it balances anything that runs over TCP or UDP across multiple servers, including databases and protocols nobody wrote an HTTP parser for.

What it cannot do is anything that requires reading the request. It cannot send /api to one pool and /images to another, cannot route by hostname, cannot add a header, cannot retry a failed request on another server, because by the time the failure happens the connection belongs to a server it already chose.

Layer 7 load balancing works at the application layer. It terminates the connection, reads the HTTP request, then opens its own connection to a backend. That single change makes almost everything possible: host and path based routing, header inspection and rewriting, compression, caching, request retries, rate limiting, and web application firewalling.

Content based load balancing only exists at this layer, and it is why the reverse proxy vs load balancer question has so little in it: at layer 7 the two are largely the same machine, as the comparison further down shows.

The cost is real and worth naming. It has to decrypt, which means the TLS certificate and private key live on the load balancer rather than only on the servers. It does more work per request. And it becomes a place where application behavior is configured, which is a place where application behavior gets misconfigured.

The practical rule for layer 4 vs layer 7 load balancing: layer 7 for anything HTTP, layer 4 for everything else, and layer 4 when the traffic must stay opaque to the middle. Most cloud providers sell both, under different names, and load balancers in software support both in one process.

Health checksThe health check is the whole point

The honest answer to what is a load balancer is not the distribution, it is the removal. Load balancing traffic evenly across a pool containing one dead server has made things worse, not better, because now every user has a one in three chance of failing instead of a specific third of them failing consistently.

Availability comes from removing the dead member from the load balancing pool, and that means the check has to detect death accurately.

Checks come in three depths, and the difference between them is the difference between an outage that is handled and an outage that is amplified.

A TCP check opens a connection to the port and closes it. It proves something is listening. It does not prove the application works, and a web server whose application pool has crashed will still accept connections all day.

An HTTP check requests a URL and decides based on the status code. Better, and still fooled by a server returning 200 from a static page while the part that matters is broken.

An application check requests a path the application owns, one that touches its real dependencies, usually a health endpoint that queries the database and returns a status. This is the one that actually works, and building it is the application team's job rather than the network team's.

Two settings decide how the check behaves over time under stress. The interval is how often the check runs, and the threshold is how many consecutive failures remove a member.

Aggressive settings eject a healthy server that had one slow moment; relaxed settings leave users on a dead server for a minute. Two failures at five second intervals is a common starting point and the right value depends on how expensive a wrong ejection is.

The failure mode worth designing for is all members failing at once, usually because they share a dependency the check touches. A health check that queries the database will eject the entire pool the moment the database hiccups, turning a slow database into a total outage.

AlgorithmsThe load balancing algorithms, and how little they matter

AlgorithmHow it picksBest when
Round robinThe next server in the listServers are identical, requests are similar
Weighted round robinProportionally, by assigned weightServers have different capacity
Least connectionsFewest open connectionsRequest duration varies a lot
Weighted least connectionsBoth of the above combinedMixed hardware and mixed workloads
Source IP hashSame client to the same serverPersistence is needed with no cookies
URL or header hashBy a value in the requestCache locality on a layer 7 pool

The honest summary is that the choice of algorithm matters far less than vendor documentation implies. Round robin load balancing is correct for most HTTP pools because requests are short and servers are usually identical.

Static algorithms like round robin and hashing decide without looking at the servers; dynamic ones like least connections look first, and that is the only distinction between them that changes anything. Least connections earns its keep when request duration varies widely, because round robin will happily hand a new request to the server already grinding through a slow one.

Hashing is different in kind. Hash based algorithms are not there for balance, they are there for stickiness, and stickiness has its own costs.

PersistenceSticky sessions, and why they undercut the point

Session persistence pins a user to one server for the duration of their session. It exists because an application kept the session in the memory of whichever server handled the first request, so any other server would not recognize the user.

It works, and it quietly gives back most of what load balancing was bought for.

A pool of four servers with sticky sessions is not four interchangeable servers. It is four servers each holding a quarter of the logged-in users, and losing one logs those users out.

Traffic is still spread across multiple machines; the availability is not. Scaling out helps only new sessions. Draining a server for maintenance means either dropping its sessions or waiting for them to expire.

The fix is not on the load balancer. It is to make the application stateless: keep sessions in a shared store, in a cache both servers read, or in a signed cookie the client carries. Then any server can answer any request, ejecting a member costs nothing, and the load balancer does what it was meant to do.

Where persistence really is needed, cookie based persistence is more precise than source IP, because many users share one public address behind a NAT and a source hash sends all of them to the same server.

Error codesWhat the error codes are telling you

A load balancer generates its own errors, and they say precisely where the failure was.

502 Bad Gateway. The load balancer was acting as a gateway and got an invalid response from the upstream server. Something answered and what it said was not usable. Look at the backend application, not the network.

504 Gateway Timeout. No response arrived from the upstream at all within the timeout. The distinction from 502 is exact: 502 means a bad answer, 504 means no answer. Look at whether the backend is overloaded or hung, and at whether the load balancer's timeout is shorter than the request legitimately takes.

503 Service Unavailable. Commonly what a load balancer returns when no pool member is healthy. This one is a health check story: either everything really is down, or the check is wrong and has ejected servers that work.

The useful habit is to notice that all three come from the middle. If the origin sends a valid HTTP error, that error should be passed through to the client rather than replaced, so a 500 reaching the browser came from the application and a 502 came from the thing in front of it.

PitfallsWhere deployments go wrong

One load balancer. Putting a single device in front of a redundant pool moves the single point of failure rather than removing it. Load balancers are deployed as a pair, active and standby or both active, sharing a virtual address. Cloud load balancing services do this for you, which is a large part of what you are paying for.

A health check that checks the wrong thing. The most common real defect, and the one that makes load balancing worse than no load balancing during a partial failure.

A health check that shares a dependency with the application. It ejects the whole pool the moment the shared thing is slow.

Timeouts that disagree. When the load balancer times out at 30 seconds and the application takes 45, users get a 504 on a request that would have succeeded, and the backend keeps working on a response nobody will read. Load balancers and applications need their time limits set together.

Sticky sessions as a permanent answer. Reasonable as a stopgap while the application is made stateless. Expensive as an architecture, because every availability property degrades.

Forgetting the client address is now hidden. Behind a layer 7 load balancer every request appears to come from the load balancer. Application logs, rate limits and geolocation all break until X-Forwarded-For or the PROXY protocol is configured and, importantly, trusted only from the balancer.

ONE BROKEN SERVER, THREE DEPTHS OF HEALTH CHECKThe application has crashed. The web server is still listening on 443.CHECKWHAT IT SENDSVERDICTWHYTCP checkconnect to :443healthySomething is listening. That is all it proves.HTTP checkGET /healthyA static page still returns 200.App checkGET /healthejectedTouches the database. It fails, so it ejects.With either of the first two, the load balancer keeps sending users to the dead server.A pool of three then fails one request in three, which is worse than having no pool.
The application is dead in all three rows. Only the third check notices, and only the third one ejects the member.

ComparisonLayer 4, layer 7, a reverse proxy and DNS round robin

CriterionLayer 4 LBLayer 7 LBReverse proxyDNS round robin
Sees the requestNoYesYesNo
Removes dead membersYes, by checkYes, by checkSometimesNo
Terminates TLSNoYesYesNo
Routes by path or hostNoYesYesNo
Cost per requestVery lowHigherHigherNone
Works for non-HTTPYesNoNoYes
Failover speedSecondsSecondsVariesCache TTL, minutes

The last row is why DNS round robin is not a substitute for load balancing. It spreads traffic across multiple addresses and it cannot remove a dead server in any useful time, because clients and resolvers hold the answer for the length of the record's TTL, which is a property of the DNS zone rather than of the failure.

A layer 7 load balancer and a reverse proxy are largely the same machine described from two directions. The proxy framing emphasizes what it does to the request; the balancer framing emphasizes how it picks a destination. Most products do both.

FAQFrequently asked questions

What is a load balancer?

Hardware, software or a cloud service that sits in front of a pool of servers and decides which one handles each request, while removing members that fail a health check so traffic stops reaching a broken server.

Is a load balancer hardware or software?

Either, and increasingly neither. Appliances from vendors like F5 still run in data centers, software load balancers such as HAProxy and NGINX run on ordinary servers, and cloud load balancing services are managed by the provider.

What is the difference between layer 4 and layer 7 load balancing?

Layer 4 balances TCP and UDP connections and cannot see inside the request. Layer 7 terminates the connection, reads the HTTP request, and can route on host, path or header, at the cost of decrypting the traffic and doing more work per request.

Do I need a load balancer for two servers?

If both must be usable and a failure of one should be invisible, yes, because something has to detect the failure and stop sending traffic there. If the second server is a manual standby, a documented switch may be enough.

What load balancing algorithm should I use?

Round robin for most HTTP pools with identical servers. Least connections when request duration varies a lot. Static algorithms decide without looking at the servers and dynamic ones look first, and the choice matters much less than the health check does.

What is a health check?

A request the load balancer sends to each pool member to decide whether it is usable. Depth matters: a TCP check proves a port is open, an HTTP check proves a page answers, and an application check proves the parts the application depends on are working.

Why did the load balancer return 502?

Because it received an invalid response from the upstream server while acting as a gateway. Something answered and the answer was not usable, so the problem is in the backend rather than in the network path.

What is the difference between 502 and 504?

502 means the upstream sent a response the gateway could not use. 504 means the upstream sent nothing within the timeout. Bad answer against no answer.

Why am I getting 503 from the load balancer?

Usually because no pool member is passing its health check. Either the backends are genuinely down, or the check is wrong and has ejected servers that work.

What are sticky sessions?

Session persistence, which pins a user to one server for their session. It exists because the application stored session state in one server's memory, and it costs you most of the availability the pool was meant to provide.

How do I get rid of sticky sessions?

Move session state out of the individual server: a shared cache, a database, or a signed cookie the client carries. Once any server can answer any request, persistence is no longer needed.

Is a load balancer the same as a reverse proxy?

Largely the same machine described from two directions. Reverse proxy emphasizes what it does to the request, load balancer emphasizes how it chooses a destination, and most products do both.

Can DNS do load balancing instead?

It can spread traffic across multiple addresses and it cannot remove a failed server quickly, because clients hold the answer for the record's TTL. DNS based load balancing is a distribution mechanism, not an availability one.

Why do my application logs show the load balancer's address?

Because a layer 7 balancer opens its own connection to the backend. The original client address has to be carried in X-Forwarded-For or the PROXY protocol, and the application has to be configured to trust that header only from the balancer.

Should the load balancer handle TLS?

Terminating there is common and it is what makes layer 7 features possible. It also puts the private key on the balancer and sends traffic onward in the clear unless it is re-encrypted, so the decision is about what the network between the balancer and the servers is.

What is an LTM load balancer?

F5's BIG-IP Local Traffic Manager, the part of an F5 BIG-IP that load balances within a site. Clients connect to a virtual server, an address and service such as 192.168.20.10:80, and LTM sends each connection to a member of the pool behind it, chosen by the pool's load balancing method, Round Robin by default, and checked by monitors.

What is the difference between GTM and LTM?

GTM, now sold as BIG-IP DNS, balances between sites by choosing which address a DNS lookup returns. LTM balances within a site by choosing which server handles each connection to that address. One works before the connection exists and the other carries it, and an application served from several sites can use both.

Read next · DNS The SOA Record, Field by Field Why DNS round robin cannot remove a dead server quickly: the answer is cached for the length of a TTL you do not control. Open this next10 min
Also worth reading
One packet a weekA short, illustrated explainer every Tuesday. No vendor pitches, unsubscribe in one click.