Skip to content
Cumps.
Diagram of Traefik routing client traffic to backend servers, with its routing table stored in Valkey

Traefik as the Front Door

  • Sep 22, 2026

Everything web-facing on this platform goes through Traefik: TLS termination, Let’s Encrypt certificates, HTTP-to-HTTPS redirects, and routing to backends. The mail protocols themselves, ports 25, 465, and 993, deliberately do not; Stalwart terminates those directly, and I removed an early attempt to proxy them once it became clear it added a hop without adding value. Traefik owns HTTP, Stalwart owns mail, and the boundary has stayed clean since January.

Two design choices from that first deployment turned out to matter far more than I understood at the time.

Certificates without a reachable server

Let’s Encrypt normally proves you control a domain by fetching a file from it over HTTP. That works when the thing requesting the certificate is a public web server. Half of my backends aren’t: they’re mesh-only services with .cumps.internal addresses, as covered last week, reachable through the proxy or not at all.

So issuance runs on the DNS-01 challenge instead: Traefik proves control by planting a TXT record through the registrar’s API. No inbound HTTP required, no port 80 ceremony, and it works identically for a public site and for a service that will never be directly reachable. The certificate machinery stopped caring about network topology, which is exactly what certificate machinery should do.

The mail server complicated this pleasantly. Stalwart needs certificates too, for SMTP and IMAP, and it has its own ACME client. Running two ACME clients for one hostname is how you meet Let’s Encrypt’s rate limiter. Instead, Traefik is the single source of certificate truth, and traefik-certs-dumper watches Traefik’s certificate store and extracts PEM files that Stalwart consumes. One issuance path, every protocol served. The dumper is pinned to a specific version because it sits in the critical path of mail TLS. It earns a return appearance in a much later post, not an entirely happy one.

The routing table is a database

Traefik usually reads routes from container labels or config files. Mine reads them from Valkey, the same managed Redis-compatible service the mail cluster coordinates through. The traefik_config Ansible role renders every router, service, and middleware into KV entries and writes them in one pass; Traefik watches the keyspace and applies changes live.

The reason was the two-node mail cluster. Webmail sat behind Traefik on both mx1 and mx2 for the same DNS round-robin availability as mail itself, and I wanted those two proxies to be incapable of disagreeing about routes. A config file per host is two files that drift; one KV store is one routing table that happens to be read twice. Adding a route means pushing to Valkey once, and every Traefik instance in the fleet picks it up within seconds, no restarts, no rolling deploys, no “did I update both?”.

Combined with the sidecar mesh, this produced a property I didn’t fully appreciate until the fleet grew: the routing config contains no hostnames of physical machines. Backends are mesh names like stalwart.cumps.internal, so the same KV entries are valid from any node running Traefik. When apps1 and monitoring1 joined months later, they ran the identical proxy setup on day one.

Was Valkey-as-config-store overkill for two nodes? Probably. It also bit me twice much later, once with a push timeout and once with a database-index collision, both stories for the fleet-era posts. I’d still choose it again: every alternative I’ve sketched reintroduces config drift between proxies, and drift is the failure mode I fear most in a one-person operation.

Hardening the edge

The edge config carries the security posture you’d expect, most of it added in one hardening pass: explicit TLS cipher suites with 1.0 and 1.1 gone entirely, HSTS on everything, and a redirectregex middleware canonicalizing stray domains onto the main one. The Traefik dashboard started life behind HTTP basic auth, which was fine for January and wrong for the long term; it eventually moved behind real single sign-on with the rest of the admin surfaces.

Which is the cue for next week: lldap, Authelia, and the identity layer, including the OIDC decision I got to make twice.

You May Also Like