ProtonVPN returns 16-byte success responses (last field is uint32 lifetime,
not uint16). The previous !BBHIHHH format expected 14 bytes causing unpack
failure. Fixed to !BBHIHHI (16 bytes).
ProtonVPN grants 60s leases; request 60s and renew every 45s instead of
sleeping 240s which let leases expire between renewals.
The protonvpn provider mode uses gluetun's embedded server database which
doesn't contain US-TX#253 or US-TX#34. Our WireGuard keys are registered
for specific ProtonVPN servers, so connecting to other servers (US-TX#179,
US-TX#220) results in successful WireGuard handshake but ProtonVPN drops all
internet-bound traffic.
Fix: use VPN_SERVICE_PROVIDER=custom to directly configure the correct server
peer public key and endpoint IP for each worker. ExternalSecrets updated to
also pull public_key and endpoint_ip from Vault.
- tx253: endpoint=95.173.217.29, peer=mngiSxBpH7GU24nnWdBEcnhDnCPn2jq5+ZP3zwPwISA=
- tx34: endpoint=146.70.58.130, peer=wqJcz4akzVFxx35aJ5B7G/IJ9qsRvpcGNub3rLHcqXo=
Will refresh gluetun's embedded server list from ProtonVPN API.
Need to extract public key for US-TX#253 (95.173.217.29) which
is not in gluetun's embedded database but should be in the live list.
US-TX#253 and US-TX#34 not in gluetun embedded server list — gluetun
crashes on startup. Reverted to valid names while investigating correct
server mapping. Need to find server public keys for Vault endpoint IPs.
Keys in Vault are registered for US-TX#253 (tx253 worker) and US-TX#34
(tx34 worker), but deployment was connecting to US-TX#179 and US-TX#220.
WireGuard handshake succeeded globally but ProtonVPN only routes internet
traffic through the server the key was registered with.
With VPN_PORT_FORWARDING=off, gluetun startup check fails with TLS EOF
for generic IP-check services. Previous session showed startup check
PASSED with VPN_PORT_FORWARDING=on (only port-forwarding API failed).
Testing to confirm VPN routing works when startup check uses ProtonVPN
API instead of generic HTTPS services.
eth0 MTU in pods is 1450 (Flannel VXLAN). gluetun's MTU discovery
was setting tun0 to 1440, making WireGuard outer packets ~1500 bytes
which exceeds the 1450 limit. TLS ClientHello gets dropped, causing
startup check to fail and loop. Fix: set MTU=1320 so WireGuard UDP
packets stay well under 1450.
Multiline NATPMP variable in tx34 had column-0 lines that broke
Go's YAML parser (same bug that was fixed in tx253). Replace with
identical single-line python3 -c approach.
Custom provider hardcodes 10.2.0.2/32 but ProtonVPN assigns a different
IP per session via their API. WireGuard handshake succeeds but ProtonVPN
doesn't route traffic for the wrong IP. Native protonvpn provider fetches
the correct IP assignment automatically and enables port forwarding.
gluetun adds 'ip rule priority 100 from <pod-IP> lookup 200' which routes
all pod traffic via eth0. The iptables OUTPUT DROP policy then blocks it.
Rule 101 (not fwmark 0xca6c -> tun0) never fires because rule 100 matches
first. Result: qBittorrent has no internet access through the VPN.
Fix: portforward-helper waits for gluetun to be ready (polls :9999), then
deletes rule 100. Application traffic then falls through to rule 101 and
routes correctly via tun0.
The fornax-worker-tx253 and fornax-worker-tx34 Argo CD apps were fighting
the media app (which also manages the same Deployments, Services, PVCs,
and ExternalSecrets via fornax-workers.yaml). This caused 73+ rolling
restarts per 2 hours. Since neither app had a cascade-delete finalizer,
removing these Application CRDs leaves existing resources intact and
transfers ownership to the media app.
ProtonVPN does not NAT ICMP to external IPs (1.1.1.1, 8.8.8.8), causing
the periodic small health check to always fail and restart the VPN loop.
10.2.0.1 (WireGuard gateway) responds to ICMP at ~28ms and remains
reachable as long as the WireGuard tunnel itself is up.
DOT=off alone does not override DNS_UPSTREAM_RESOLVER_TYPE in this version
of gluetun — it defaults to DoT regardless, causing DNS failures through
ProtonVPN (port 853 connection reset). Setting the resolver type directly
fixes plain DNS routing to k8s CoreDNS.
Change HEALTH_TARGET_ADDRESS to ProtonVPN API endpoints so TLS startup
check doesn't fail against cloudflare.com/github.com (rejected from
ProtonVPN exit IPs). Rename deprecated DNS_ADDRESS to
DNS_UPSTREAM_PLAIN_ADDRESSES in fornax-workers.yaml.
DoT (port 853) to 1.1.1.1 through ProtonVPN VPN tunnel gets TCP RST,
causing gluetun healthcheck to fail (can't resolve github.com /
cloudflare.com). k8s coredns at 10.43.0.10 is reachable via eth0
within FIREWALL_OUTBOUND_SUBNETS, bypassing the VPN for DNS while
letting all other traffic tunnel correctly.
Also rename VPN_ENDPOINT_IP/PORT to WIREGUARD_ENDPOINT_IP/PORT to
suppress gluetun deprecation warnings.
gluetun's bundled protonvpn server list has stale IPs for US-TX#179
(37.19.200.26) and US-TX#220 (95.173.217.2). The WIREGUARD_ENDPOINT_IP
env var is a filter not an override, so there's no way to redirect
protonvpn provider to a new IP.
Switch to custom WireGuard provider with hardcoded endpoint IPs from
fresh ProtonVPN configs (95.173.217.29 / 146.70.58.130). Add
DNS_KEEP_NAMESERVER=on so gluetun leaves k8s DNS intact instead of
routing DNS through its own proxy (which breaks in-cluster).
Port forwarding is not available with custom provider; will restore
once gluetun releases an updated server list image.
gluetun rejects VPN_ENDPOINT_PORT when SERVER_NAMES is used (server
selection mode), and warns that VPN_ENDPOINT_IP is deprecated in
favour of WIREGUARD_ENDPOINT_IP. Use only WIREGUARD_ENDPOINT_IP;
port 51820 is ProtonVPN's default and doesn't need to be set.
ProtonVPN migrated US-TX#179 (37.19.200.26 → 95.173.217.29) and
US-TX#220 (95.173.217.2 → 146.70.58.130). gluetun's bundled server
list still has the old dead IPs. Override via VPN_ENDPOINT_IP and
VPN_ENDPOINT_PORT so gluetun uses the correct endpoints while keeping
the protonvpn provider (and port forwarding) intact.
Avoids gluetun's ProtonVPN server name lookup entirely. All WireGuard
parameters (endpoint, peer key, client address, keepalive) come directly
from the downloaded ProtonVPN config.
US-TX#179 and US-TX#220 were failing — WireGuard handshake completing
but no traffic passing (silent drop). Switching both workers to US-IL#267
and pinning the endpoint IP and peer public key directly to avoid relying
on gluetun's built-in server list lookup.
Python 3.14 raises PatternError for backslash sequences like \P in
re.sub replacement strings. Switch to lambda replacements which bypass
that interpretation. Also pin to python:3.13-alpine to avoid future
surprises from pulling :3-alpine.
The /api/ui path is now routed to the fornax-ui nginx pod (which proxies
internally to the coordinator) rather than directly to the coordinator.
This fixes the browser fetch being blocked by Cloudflare Access.
Both coordinator and UI deployments get a restartedAt bump to pull the
new images after CI builds.
- fornax service: targetPort 8080→3000 (Express listens on 3000, not 8080)
This was silently dropping all Radarr/Sonarr requests
- coordinator deployment: add QBT_WORKER_ADDRS, QBT_USER, QBT_PASS
so /api/ui/state can query real worker qBittorrent instances
- remove Python coordinator ConfigMap and Deployment (superseded)