NEWS
Containers as Security Boundary?
“Containers are not a security boundary” is a sentence many security professionals repeat like a mantra. In a way, it is true. The Linux kernel has a huge attack surface and container escape vulnerabilities are more common than one would hope for. However, Linux container escape zero-days are not something your run-of-the-mill ransomware attacker will waste on you. And if you ever gained code execution in a container during a pentest, you are often pretty sad that there is little to do in there. So effectively, containers are a security boundary. Maybe, just not a particularly strong one.
This same container technology underpins Kubernetes, one of the most widely used platforms for deploying and managing scalable services. Kubernetes adds a whole abstraction layer on top of containers to build distributed systems. Its fundamental execution unit is a Pod, which bundles multiple containers. A Pod is a functional unit, and the containers inside it typically depend on each other. A very common setup runs your service container next to several sidecar containers. These sidecars handle things like authentication, TLS termination, and telemetry.
If you take away only one thing from this post it should be: a container in a Kubernetes Pod is not a security boundary. The containers inside a Pod share more than people think. “It’s a different container” is not the isolation guarantee it sounds like. And that is true even aside from the whole Linux kernel zero-day narrative.
In this blog post, we describe how we confirmed this lesson once again. During an audit of a customer’s Kubernetes deployment, we found a particularly interesting memory corruption bug in the Envoy proxy. Whether this bug is a security vulnerability depends entirely on your setup. Specifically, it depends on whether all containers in your Pod really do have the same privileges.
How we got here
During an engagement we got root access to a container in a Kubernetes cluster.
We started poking at shared files and found that our container could see a file
created by the Envoy proxy.
Envoy is an extremely popular edge/service proxy. It shows up as a sidecar container
in basically every other Kubernetes cluster, often as the data plane for the
Istio service mesh. Envoy has a neat feature
called hot restart that lets it swap binaries and reload config without
dropping a single connection. To coordinate between the old and new process, it
uses file-backed shared memory using a file in /dev/shm.
Nothing wrong with that so far. Sharing memory between two processes is literally the point of file-based shared memory. However, since this file was visible from the container we had access to, we started poking at this file and how Envoy uses it. The interesting part is what lives inside that file.
Envoy’s Shared Memory File
The shared region is described by this struct:
// source/server/hot_restart_impl.h
struct SharedMemory {
uint64_t size_;
uint64_t version_;
pthread_mutex_t log_lock_;
pthread_mutex_t access_log_lock_;
std::atomic<uint64_t> flags_;
};
Two pthread_mutex_t fields, which are initialized
as PTHREAD_PROCESS_SHARED and PTHREAD_MUTEX_ROBUST.
The “robust” part is where it gets interesting. A robust mutex is one where, if a
thread dies while holding the lock, glibc can recover instead of leaving the
lock set forever. To implement this, glibc keeps a per-thread robust list:
a doubly-linked list of the mutexes a thread currently holds. Which means the
mutex struct contains pointers:
struct __pthread_mutex_s {
int __lock;
/* [...] */
__pthread_list_t __list = {
// raw pointers into Envoy's address space
void *__prev;
void *__next;
};
};
Read that again: there are raw pointers into Envoy’s address space sitting in a file. The moment another process can read that file, ASLR is basically decorative. The moment another process can write that file, things get much worse. We were smelling a fun exercise in binary exploitation.
Turning it into a real exploit
Once you can see and touch those mutex internals, a few things fall out naturally:
1. Bypass ASLR. When Envoy is holding one of the locks, __list.__prev and
__list.__next point into its live address space. Read them out of the file and
you’ve defeated address-space layout randomization. Free info leak.
2. Denial of service. Poison the __lock futex word and the next time Envoy
tries to take the lock, it hangs. Forever.
3. An arbitrary write. This is the fun one. Corrupt the __list pointers
while Envoy holds the lock, and when Envoy calls pthread_mutex_unlock(), glibc
dutifully unlinks the mutex from the robust list:
// glibc nptl/pthread_mutex_unlock.c - unlink from robust list
list->__prev->__next = list->__next; // write #1
list->__next->__prev = list->__prev; // write #2
Both __prev and __next came from the file we control. So those two lines
are, from our point of view, two attacker-controlled pointer writes into
Envoy’s memory. That’s a classic linked-list unlink write primitive.
Now the question is: How often is this unlocking triggered by Envoy during normal operation?
log_lock_is taken on every log write to stderr (when--log-pathisn’t set).access_log_lock_fires on the periodic access-log flush (every ~10s by default).- Both are also used during hot-restart coordination.
We haven’t built this out into a full code execution exploit for Envoy, but all the building blocks are there. We have a pointer leak, so we can infer the address space layout. ASLR is defeated. Then we have an arbitrary write. It is kind of “racey”, because we have to corrupt the pointer values in the small window after Envoy locks the mutex but before it unlocks the mutex again. However, our testing showed that this is not hard to hit this race window. From there, we could point the write primitive at a C++ vtable pointer, some heap metadata, functions pointers, or another thread’s stack. Anything the pointer leak reveals is fair game.
PoC||GTFO
We wrote a PoC that demonstrates the pointer leak and the arbitrary write.
The debugger output below shows Envoy crashed inside glibc’s unlink code. It is
about to write a value we chose to an address we chose. rax is the value
written and rdx is the target address. Both come straight out of the __list
pointers we planted in shared memory.
pwndbg> x/i $rip
=> 0x7ffff7d01c16 <__pthread_mutex_unlock_full+758>: mov QWORD PTR [rdx-0x8],rax
pwndbg> i r rax rdx
rax 0x4142434445464748 0x4142434445464748
rdx 0xfefefefefefeff06 0xfefefefefefeff06
pwndbg> bt
#0 0x00007ffff7d01c16 in __pthread_mutex_unlock_full (mutex=0x7ffff7fa9038, decr=0x1) at pthread_mutex_unlock.c:152
#1 0x000055555a1b6c80 in ?? ()
#2 0x000055555a1b7062 in ?? ()
#3 0x000055555af2e2c6 in ?? ()
#4 0x00007ffff7cfcc19 in start_thread (arg=<optimized out>) at pthread_create.c:454
#5 0x00007ffff7d805cc in __GI___clone3 () at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:78
We confirmed all of this on Envoy 1.38.3 and the latest 1.39.1. The shared-memory design has been essentially unchanged since at least 2016. So it is fair to assume that any version of Envoy currently in use is also affected.
Reporting this Bug Upstream
We reported this, and the Envoy maintainers closed it as not a security issue. From their vantage point, they’re not entirely wrong.
Envoy already creates the shared memory file with tight permissions. By default only the user running Envoy gets read-write access. So within a s ingle container, an attacker who isn’t already Envoy can’t touch the file, and there is no security issue.
But that argument assumes the file permissions mean something across the boundary you care about. Often they don’t. Container services frequently run as root, and Kubernetes is very liberal about sharing data within a Pod. This brings us back to the whole point of this post.
Lateral Movement within a Kubernetes Pod
By default, /dev/shm is shared across all containers in a Kubernetes Pod.
Recall the sidecar pattern: your app runs in one container and Envoy runs right next to it in
another. Both belong to the same Pod, so they share /dev/shm.
Envoy’s carefully-permissioned shared memory file is sitting there, writable by all the other containers in the Pod.
Container user IDs are just numbers, and those numbers are shared across the Pod’s view of /dev/shm. So the permission check evaporates in two common cases: the attacker is root in the compromised neighbor container, or the attacker runs under the same UID as Envoy. Same file, same numeric owner, so write access is granted.
So the realistic attack path looks like this:
- An attacker compromises some other container in the Pod. For example, the application container, an init/debug sidecar, etc.
- That container shares
/dev/shmwith Envoy (the default). - The attacker is root there (or shares Envoy’s UID).
- They read Envoy’s mutex pointers (ASLR bypassed), corrupt the robust list, and turn it into arbitrary code execution within the Envoy container.
Why is that a potential privilege gain? Because the Envoy sidecar can often reach things the compromised container can’t: mTLS client certs, service-mesh identity, upstream services behind the mesh, credentials, and network segments. Moving into the Envoy container means gaining access to all of that. That is textbook lateral movement. Whether any of this is useful, depends entirely on how the overall service architecture is designed.
That’s the lesson, stated plainly: if two containers share a Pod, treat them as sharing
a trust domain. They share a kernel. They share /dev/shm. They might share other
resources enabling attacks.
Workaround
We don’t know if this Envoy issue ever gets assigned a CVE or will be patched upstream. Luckily, for this concrete vulnerability, you have two options to make this bug unreachable in your deployment.
Option 1: Turn off hot restart
The whole attack hinges on that shared memory file existing, and it exists because hot restart is on by default. If you don’t need it, turn it off:
envoy --disable-hot-restart ...
No file, no mutexes in shared memory, no attack path.
Option 2: Don’t share /dev/shm across the Pod
/dev/shm being shared is a default, not a hard requirement. Give each container its own
tmpfs and the cross-container path disappears:
# Kubernetes Pod spec - separate tmpfs per container
apiVersion: v1
kind: Pod
spec:
containers:
- name: service
volumeMounts:
- name: service-shm
mountPath: /dev/shm
- name: envoy
volumeMounts:
- name: envoy-shm
mountPath: /dev/shm
volumes:
- name: service-shm
emptyDir:
medium: Memory
- name: envoy-shm
emptyDir:
medium: Memory
More broadly: you should go look at what your Pods actually share and whether this potentially enables any attacks. Shared /dev/shm is just the example that enabled this attack.
A Note on MicroVMs
Now you might think: “No problem, let’s use virtualization instead of containers
for better isolation.” Historically, virtual machines are a stronger
security boundary. You can even wire up micro VMs into Kubernetes
using kata containers. But that does
not help at all here. A Pod is a functional unit, so kata containers deploy all
of a Pod’s containers to the same virtual machine. They still share their /dev/shm,
and the attack still works.
Timeline
- 2026-09-08 - Reported to Envoy via GitHub Security Advisory (GHSA-vp84-m2pc-c5mp).
- 2026-09-16 - Closed by the maintainers: Envoy already restricts the file to the Envoy user, so within a single container it’s not writable by an attacker.
- 2026-09-17 - We clarified that the problem is specifically about
Kubernetes deployments where
/dev/shmis shared across the Pod, and let them know we’d publish this information on our homepage. - 2026-10-09 - This post is released.
Conclusion
The bug is a genuinely fun one: a shared file full of live pointers that turns into a pointer leak and an arbitrary write. But the reason it’s worth writing up isn’t the exploit. It’s the assumption the bug breaks. Many threat models draw their trust boundary at the container edge. They quietly assume a sidecar in the same Pod is walled off from the service container next to it. It isn’t. A Kubernetes Pod is a unit of co-location and shared resources, not a security boundary. Make sure your security architecture reflects this.