Envoy Gateway has released version 1.9.1, a maintenance update that reverses a critical configuration change from v1.9.0 after reports emerged of production outages where long-running proxies lost their TLS certificates and stopped serving HTTPS traffic. Released on August 28, the update addresses security vulnerabilities, restores a 15-second timeout that was removed in the previous version, and introduces more detailed observability into gateway translation processes. The release highlights the operational complexity of running Kubernetes-native gateways at scale, where configuration adjustments can disrupt live proxy state and interrupt active connections.

The most significant change concerns the initial fetch timeout used for Secret Discovery Service and Route Discovery Service. Envoy Gateway 1.9.0 set this timeout to zero, but the project discovered that a cluster waiting indefinitely for a missing secret or endpoint could remain in a warming state, which prevented updates from progressing and delayed health checks. Version 1.9.1 restores Envoy's default 15-second timeout, returning behavior to the v1.8.x model. For organizations already running v1.9.0, the upgrade is more complicated: changing the SDS configuration during a controller upgrade can trigger an Envoy issue affecting proxies that remain running while the new controller is deployed, causing TLS listeners to become active without their certificates and new TLS handshakes to fail. Backend TLS configurations and the global rate-limit service can experience similar certificate or CA problems. A GitHub issue from a user running Envoy Gateway in production described an outage in which long-running Envoy proxies lost their downstream TLS certificates and stopped serving HTTPS traffic, with recovery requiring manual proxy restarts.

Security is another major focus in the release. Envoy Gateway now enables AES-256-GCM for OAuth2/OIDC session-cookie encryption and removes support for the legacy AES-256-CBC decryption path, addressing a padding-oracle vulnerability identified as CVE-2026-47775. Existing sessions using the older encryption mechanism will require users to authenticate again after upgrading. The release also strengthens several other security boundaries: OIDC issuer URLs now receive additional validation, OCI Wasm image pulls no longer silently fall back from HTTPS to HTTP, and a security-context issue involving tenant-supplied EnvoyProxy configuration has been fixed. The Wasm changes are particularly relevant to supply-chain security—previously, an OCI registry that rejected an HTTPS request could cause Envoy Gateway to fall back to plain HTTP, potentially allowing an on-path attacker to provide arbitrary Wasm code. Version 1.9.1 removes that implicit fallback, with HTTP now permitted only where a registry has explicitly been configured as insecure.

For teams evaluating or already running Envoy Gateway, the upgrade path matters. Users on v1.8.x can move directly to v1.9.1 and skip v1.9.0 entirely. Existing v1.9.0 deployments require more care: Envoy Gateway recommends a rolling-update configuration that replaces proxy pods quickly, although this requires sufficient cluster capacity and can terminate existing connections as pods are replaced. The release notes specifically warn that long-lived WebSocket and gRPC connections can be interrupted during replacement. One of the quieter changes may prove useful for platform engineering teams: Envoy Gateway now adds per-phase tracing spans to Gateway API and xDS translation, allowing operators to identify where processing time is being spent across listener processing, HTTP and gRPC routes, policy processing, EnvoyPatchPolicy handling, extension hooks, and xDS validation. The release illustrates that configuration changes can affect live proxy state, security boundaries extend into Wasm and tenant-controlled resources, and control-plane performance increasingly needs its own telemetry. Gateway operators who prioritize deployment velocity may find themselves weighing that speed against the risk of service interruption when core infrastructure changes ripple through live traffic patterns. Organizations with strict uptime requirements will need to factor rollback complexity into their evaluation of when and how to adopt new gateway versions.