How one hanging API takes down a whole Spring Boot app

We pointed a Spring Boot app at a partner API that never answers. With the defaults, the app went dark in 10 seconds and stayed dark. Two lines of config fixed it. Virtual threads only hid it.

Picture a partner API, say the one that prices your shipping, that starts accepting connections and never answering. Not down. Just silent. In our test, that silence took a Spring Boot app with default settings completely offline in 10 seconds, health check included, and the app stayed offline after the traffic stopped.

Our founder has had to fix thread exhaustion in a Java HTTP client before, and adding missing timeouts was part of that fix. For this post we built a test app from scratch and measured five setups side by side, including two that surprised us, plus a few follow-up tests.

In short

  • With the JDK’s HTTP client, which Spring Boot 4.1 uses unless another client is on the classpath, RestClient gets no timeouts at all. A partner that accepts connections but never answers held every Tomcat worker thread: at 20 requests a second, the whole app stopped answering after 10 seconds and was still down 45 seconds after the traffic stopped.
  • Two properties fixed it: spring.http.clients.connect-timeout and spring.http.clients.read-timeout. Every call failed fast, in 2 seconds (a 503, with a small exception handler), and the app stayed up. Upgrading from Spring Boot 3.4 or 3.5? The old spring.http.client.* names are silently ignored in 4.1.
  • Virtual threads kept the health check green, but all 600 partner calls were still hanging when the test ended: green health checks, stuck customers.
  • With Apache HttpClient on the classpath, timeouts weren’t enough: its pool allows 5 connections per host and waits up to 3 minutes for one. Raise the pool and cap the wait.

The experiment: one app, one silent partner, five setups

A Spring Boot 4.1.1 app on Java 21 with two endpoints. /quote calls a partner API through a RestClient built from Spring Boot’s auto-configured RestClient.Builder. With no other HTTP client library on the classpath, Spring Boot builds it on the JDK’s own HttpClient and sets neither a connect nor a read timeout. /ping returns “ok” and touches nothing, so if it stops answering, the whole app is down. The partner was a stub that accepts the connection, reads the request and never replies: one of the nastiest ways an API can fail. A server that refuses the connection fails fast. A silent one drags you down with it.

We sent 20 requests a second to /quote for 30 seconds, 600 in all, and checked /ping every 250 ms for 75 seconds, counting it as down if it took longer than 2 seconds. Each test request opened its own connection to the app.

SetupHealth checkPartner call
Spring Boot defaultsDown No answer from 10 s to the end of the runHung 600 of 600 unanswered at the end: 200 stuck on the partner, 400 queued for a thread
Timeouts: connect 1 s, read 2 sUp 300 of 300 checks answeredFast 503 600 of 600 in 2.0 s
Defaults + virtual threadsUp 300 of 300 checks answeredHung 600 of 600 still waiting at the end
Apache HttpClient + the same timeoutsDown No answer from 11.5 s to the end of the runQueued 190 got a 503 after a median wait of about 35 s; 410 still waiting at the end
Apache HttpClient, bigger pool, 1 s pool waitUp 300 of 300 checks answeredFast 503 600 of 600 in 2.0 s
Did the app answer its health check? Five setups, same traffic The app's /ping endpoint was checked every 250 milliseconds for 75 seconds, while /quote received 20 requests a second for the first 30 seconds, each calling a partner API that never answers. Defaults: answered until 10 seconds, then nothing for the rest of the run, including the 45 seconds after the traffic stopped. Timeouts: answered all 300 checks. Virtual threads: answered all 300 checks. Apache HttpClient with timeouts: answered until 11.5 seconds, then nothing. Apache HttpClient with a larger pool and a 1 second pool wait: answered all 300 checks. Did /ping answer within 2 s? answered no answer 20 requests/s Defaults Defaults: answered from 0 to 10 s (40 of 300 checks) Defaults: no answer within 2 s from 10 to 75 s (260 of 300 checks) down from 10 s, no recovery Timeouts Timeouts: answered from 0 to 75 s (300 of 300 checks) answered all 300 checks Virtual threads Virtual threads: answered from 0 to 75 s (300 of 300 checks) answered all 300 checks Apache client Apache client: answered from 0 to 11.5 s (46 of 300 checks) Apache client: no answer within 2 s from 11.5 to 75 s (254 of 300 checks) down from 11.5 s, no recovery Apache, pool fixed Apache, pool fixed: answered from 0 to 75 s (300 of 300 checks) answered all 300 checks 0 15 30 45 60 75 s
Same traffic, five setups. /ping never touches the partner API, so when it stops answering, the whole app is down. Traffic ran for the first 30 seconds.

Why the whole app goes down

Spring Boot’s embedded Tomcat serves requests with a pool of 200 worker threads. Each /quote request holds its thread while it waits on the partner, and with no read timeout, it waits for as long as the partner keeps quiet. At 20 requests a second, all 200 threads are taken after 10 seconds. From then on, every new request queues for a thread, /ping included.

Ten seconds in: every Tomcat worker thread is waiting on the partner A grid of 200 squares stands for Tomcat's 200 worker threads, and every square is marked as waiting on the partner API. Next to it, new requests wait in line for a free thread, including /ping, the health check. None comes free, so none of them is answered: /ping times out and the queued /quote requests hang. Ten seconds in, at 20 requests a second 200 of 200 worker threads waiting on a partner that never answers waiting for a thread /quote /quote /ping /quote /quote your health check, stuck in line too
Each request holds a thread while it waits on the partner. At 20 requests a second, all 200 are taken after 10 seconds, and everything after that, including /ping, queues for a thread that never comes free.

The part that hurts most: stopping the traffic doesn’t help. A thread comes back only when the partner answers or hangs up, and ours never did. A thread dump at the end of the run showed all 200 worker threads parked inside the HTTP call. If /ping were your liveness probe, an orchestrator would restart the pod, and at this traffic the fresh one would be stuck again about 10 seconds later, for as long as the partner stays silent.

The fix: two lines

# Applies to every client built from Spring Boot's RestClient.Builder
spring.http.clients.connect-timeout=1s
spring.http.clients.read-timeout=2s

Upgrade trap

Spring Boot 4 renamed these from spring.http.client.*, the names Spring Boot 3.4 and 3.5 used. The old names are still listed as deprecated, but 4.1 no longer reads them: with spring.http.client.read-timeout=2s, our call still hung, and nothing in the log said why.

Search your configuration for the old names, or add spring-boot-properties-migrator for one release. With it on the classpath, our app logged a warning naming each renamed key and applied it, so the call timed out after 2 seconds.

Inject Spring Boot’s RestClient.Builder rather than calling RestClient.create(), which builds its own client and ignores these properties. Then turn the timeout into an answer:

@RestController
class QuoteController {

    private final RestClient partner;

    // Inject Spring Boot's builder: RestClient.create() would ignore the timeout properties.
    QuoteController(RestClient.Builder builder, @Value("${partner.url}") String partnerUrl) {
        this.partner = builder.baseUrl(partnerUrl).build();
    }

    @GetMapping("/quote")
    String quote() {
        return partner.get().uri("/rates").retrieve().body(String.class);
    }

    // A timed-out call becomes a quick, honest 503 instead of a hang.
    @ExceptionHandler(ResourceAccessException.class)
    ResponseEntity<String> partnerUnavailable() {
        return ResponseEntity.status(503).body("Rates are unavailable right now. Please try again shortly.");
    }
}

With those in place, every one of the 600 calls got its 503 in 2.0 seconds, /ping answered all 300 checks, and about 40 calls were waiting on the partner at any moment: 20 requests a second times 2 seconds.

The read timeout did the work in that run, because our silent partner accepted every connection. The connect timeout is for a partner that never answers the connection attempt itself. Unlike a refusal, which fails at once, the attempt simply gets no reply. We simulated it with a server whose connection queue was full, so new attempts went unanswered. With no timeouts, the call was still hanging after 20 seconds. With the JDK client, whose read timeout covers the whole request, connecting included, the read timeout alone already returned a 503 in 2 seconds, and adding the connect timeout cut that to 1. With Apache HttpClient (more on that below), the read timeout alone was still hanging after 20 seconds; only the connect timeout stopped it. Set both.

Plot twist

Virtual threads kept the app answering every health check, with no timeouts at all. Every single partner call still hung: 600 of 600.

Virtual threads hide the problem

With spring.threads.virtual.enabled=true, Tomcat runs each request on a virtual thread, so there’s no pool of 200 to run out of. /ping answered every time. But nothing ever finished: when the test ended, 600 requests were still waiting and 600 connections to the partner were still open.

Your health checks would stay green while every customer who asked for a quote stared at a spinner. Virtual threads make waiting cheap. They don’t make it end. You still need the timeouts.

Gotcha

Add Apache HttpClient to your project, directly or through another library, and Spring Boot quietly switches to it. Your 2-second read timeout stops protecting the app.

The Apache HttpClient trap

Spring Boot picks the HTTP client by what’s on the classpath: Apache HttpClient first, then Jetty, Reactor Netty and finally the JDK’s own client. Apache does have a default read timeout of its own, but it’s 3 minutes: at 20 requests a second, all 200 threads are gone long before that. Its connection pool allows 5 connections per host and 25 in total, and a request waits up to 3 minutes for a free connection. The read timeout doesn’t cover that wait.

With the same 2-second timeouts, only 5 calls reached the partner at a time. A thread dump showed 195 of the 200 worker threads waiting for a pooled connection, and the app went down at 11.5 seconds. The calls that did fail took a median of about 35 seconds to do it. A bigger pool and a short wait for a connection fixed it:

import java.util.concurrent.TimeUnit;

import org.springframework.boot.http.client.HttpComponentsClientHttpRequestFactoryBuilder;
import org.springframework.boot.http.client.autoconfigure.ClientHttpRequestFactoryBuilderCustomizer;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;

@Configuration
class PartnerClientConfig {

    @Bean
    ClientHttpRequestFactoryBuilderCustomizer<HttpComponentsClientHttpRequestFactoryBuilder> partnerPool() {
        return builder -> builder
            .withConnectionManagerCustomizer(pool -> pool.setMaxConnPerRoute(50).setMaxConnTotal(200))
            .withDefaultRequestConfigCustomizer(request -> request.setConnectionRequestTimeout(1, TimeUnit.SECONDS));
    }
}

Same traffic, same timeouts: 600 fast 503s, every health check answered, and about 40 connections to the partner at a time. The 1-second cap never came into play there: about 40 connections fit in a pool of 50. To test the cap, we shrank the pool to 10. Of the 600 calls, 440 got their 503 after waiting 1 second for a connection, the other 160 got theirs in 2 to 3 seconds, and every health check was still answered.

Picking the numbers

Beyond timeouts

Timeouts make a silent partner cost you two seconds a request instead of everything. Two patterns go further. A circuit breaker stops calling a partner that keeps failing and answers straight away for a while, and a bulkhead caps how many threads one partner may hold at once. Resilience4j provides both. Spring Framework 7, which Spring Boot 4 is built on, also has a @ConcurrencyLimit annotation, enabled with @EnableResilientMethods. Set policy = REJECT: by default it makes extra callers wait, which is the problem you’re trying to avoid. Where you can, serve the last good answer from a cache, and alert on the partner’s response time so you hear about it before your customers do.

Checklist

  1. Set spring.http.clients.connect-timeout and spring.http.clients.read-timeout. Coming from Spring Boot 3.4 or 3.5, rename the old spring.http.client.* keys.
  2. Build clients from Spring Boot’s RestClient.Builder, so the settings apply.
  3. Turn a timeout into a fast, honest error, or a cached answer.
  4. Using Apache HttpClient? Size the pool and cap the wait for a connection.
  5. Don’t count on virtual threads: they keep the app up, not your users’ requests.
  6. Point a test at a partner that never answers, before a real one does it for you.