I've used Linux containers directly and indirectly for years, but I wanted to become more familiar with them. So I wrote some code. This used to be 500 lines of code, I swear, but I've revised it some since publishing; I've ended up with about 70 lines more.
I wanted specifically to find a minimal set of restrictions to run untrusted code. This isn't how you should approach containers on anything with any exposure: you should restrict everything you can. But I think it's important to know which permissions are categorically unsafe! I've tried to back up things I'm saying with links to code or people I trust, but I'd love to know if I missed anything.
This is a noweb-style piece of literate code. References named
<<x>> will be expanded to the code block named x. You can find the
tangled source here. This document is an orgmode document, you can
find its source here. This document and this code are licensed under
the GPLv3; you can find its source here.
Container setup
There are several complementary and overlapping mechanisms that make up modern Linux containers. Roughly,
namespacesare used to group kernel objects into different sets that can be accessed by specific process trees. For example, pid namespaces limit the view of the process list to the processes within the namespace. There are a couple of different kind of namespaces. I'll go into this more later.capabilitiesare used here to set some coarse limits on what uid 0 can do.cgroupsis a mechanism to limit usage of resources like memory, disk io, and cpu-time.setrlimitis another mechanism for limiting resource usage. It's older than cgroups, but can do some things cgroups can't.
These are all Linux kernel mechanisms. Seccomp, capabilities, and
setrlimit are all done with system calls. cgroups is accessed
through a filesystem.
There's a lot here, and the scope of each mechanism is pretty unclear. They overlap a lot and it's tricky to find the best way to limit things. User namespaces are somewhat new, and promise to unify a lot of this behavior. But unfortunately compiling the kernel with user namespaces enabled complicates things. Compiling with user namespaces changes the semantics of capabilities system-wide, which could cause more problems or at least confusion1. There have been a large number of privilege-escalation bugs exposed by user namespaces. "Understanding and Hardening Linux Containers" explains
Despite the large upsides the user namespace provides in terms of security, due to the sensitive nature of the user namespace, somewhat conflicting security models and large amount of new code, several serious vulnerabilities have been discovered and new vulnerabilities have unfortunately continued to be discovered. These deal with both the implementation of user namespaces itself or allow the illegitimate or unintended use of the user namespace to perform a privilege escalation. Often these issues present themselves on systems where containers are not being used, and where the kernel version is recent enough to support user namespaces.
It's turned off by default in Linux at the time of this writing2, but many distributions apply patches to turn it on in a limited way3.
But all of these issues apply to hosts with user namespaces compiled in; it doesn't really matter whether we use user namespaces or not, especially since I'll be preventing nested user namespaces. So I'll only use a user namespace if they're available.
(The user-namespace handling in this code was originally pretty broken. Jann Horn in particular gave great feedback. Thanks!)
contained.c
This program can be used like this, to run /misc/img/bin/sh in
/misc/img as root:
[lizzie@empress l-c-i-500-l]$ sudo ./contained -m ~/misc/busybox-img/ -u 0 -c /bin/sh => validating Linux version...4.7.10.201610222037-1-grsec on x86_64. => setting cgroups...memory...cpu...pids...blkio...done. => setting rlimit...done. => remounting everything with MS_PRIVATE...remounted. => making a temp directory and a bind mount there...done. => pivoting root...done. => unmounting /oldroot.oQ5jOY...done. => trying a user namespace...writing /proc/32627/uid_map...writing /proc/32627/gid_map...done. => switching to uid 0 / gid 0...done. => dropping capabilities...bounding...inheritable...done. => filtering syscalls...done. / # whoami root / # hostname 05fe5c-three-of-pentacles / # exit => cleaning cgroups...done.
So, a skeleton for it:
/* -*- compile-command: "gcc -Wall -Werror -lcap -lseccomp contained.c -o contained" -*- */
/* This code is licensed under the GPLv3. You can find its text here:
https://www.gnu.org/licenses/gpl-3.0.en.html */
#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <grp.h>
#include <pwd.h>
#include <sched.h>
#include <seccomp.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
#include <unistd.h>
#include <sys/capability.h>
#include <sys/mount.h>
#include <sys/prctl.h>
#include <sys/resource.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <sys/utsname.h>
#include <sys/wait.h>
#include <linux/capability.h>
#include <linux/limits.h>
struct child_config {
int argc;
uid_t uid;
int fd;
char *hostname;
char **argv;
char *mount_dir;
};
<<capabilities>>
<<mounts>>
<<syscalls>>
<<resources>>
<<child>>
<<choose-hostname>>
int main (int argc, char **argv)
{
struct child_config config = {0};
int err = 0;
int option = 0;
int sockets[2] = {0};
pid_t child_pid = 0;
int last_optind = 0;
while ((option = getopt(argc, argv, "c:m:u:"))) {
switch (option) {
case 'c':
config.argc = argc - last_optind - 1;
config.argv = &argv[argc - config.argc];
goto finish_options;
case 'm':
config.mount_dir = optarg;
break;
case 'u':
if (sscanf(optarg, "%d", &config.uid) != 1) {
fprintf(stderr, "badly-formatted uid: %s\n", optarg);
goto usage;
}
break;
default:
goto usage;
}
last_optind = optind;
}
finish_options:
if (!config.argc) goto usage;
if (!config.mount_dir) goto usage;
<<check-linux-version>>
char hostname[256] = {0};
if (choose_hostname(hostname, sizeof(hostname)))
goto error;
config.hostname = hostname;
<<namespaces>>
goto cleanup;
usage:
fprintf(stderr, "Usage: %s -u -1 -m . -c /bin/sh ~\n", argv[0]);
error:
err = 1;
cleanup:
if (sockets[0]) close(sockets[0]);
if (sockets[1]) close(sockets[1]);
return err;
}
Since I'll be blacklisting system calls and capabilities, it's important to make sure there aren't any new ones.
fprintf(stderr, "=> validating Linux version...");
struct utsname host = {0};
if (uname(&host)) {
fprintf(stderr, "failed: %m\n");
goto cleanup;
}
int major = -1;
int minor = -1;
if (sscanf(host.release, "%u.%u.", &major, &minor) != 2) {
fprintf(stderr, "weird release format: %s\n", host.release);
goto cleanup;
}
if (major != 4 || (minor != 7 && minor != 8)) {
fprintf(stderr, "expected 4.7.x or 4.8.x: %s\n", host.release);
goto cleanup;
}
if (strcmp("x86_64", host.machine)) {
fprintf(stderr, "expected x86_64: %s\n", host.machine);
goto cleanup;
}
fprintf(stderr, "%s on %s.\n", host.release, host.machine);
(This had a bug. captainjey on reddit let me know. Thanks!)
And I wasn't quite at 500 lines of code, so I thought I had some space to build nice hostnames.
int choose_hostname(char *buff, size_t len)
{
static const char *suits[] = { "swords", "wands", "pentacles", "cups" };
static const char *minor[] = {
"ace", "two", "three", "four", "five", "six", "seven", "eight",
"nine", "ten", "page", "knight", "queen", "king"
};
static const char *major[] = {
"fool", "magician", "high-priestess", "empress", "emperor",
"hierophant", "lovers", "chariot", "strength", "hermit",
"wheel", "justice", "hanged-man", "death", "temperance",
"devil", "tower", "star", "moon", "sun", "judgment", "world"
};
struct timespec now = {0};
clock_gettime(CLOCK_MONOTONIC, &now);
size_t ix = now.tv_nsec % 78;
if (ix < sizeof(major) / sizeof(*major)) {
snprintf(buff, len, "%05lx-%s", now.tv_sec, major[ix]);
} else {
ix -= sizeof(major) / sizeof(*major);
snprintf(buff, len,
"%05lxc-%s-of-%s",
now.tv_sec,
minor[ix % (sizeof(minor) / sizeof(*minor))],
suits[ix / (sizeof(minor) / sizeof(*minor))]);
}
return 0;
}
Namespaces
clone is the system call behind fork() et al. It's also the key to
all of this. Conceptually we want to create a process with different
properties than its parent: it should be able to mount a different
/, set its own hostname, and do other things. We'll specify all of
this by passing flags to clone 4.
The child needs to send some messages to the parent, so we'll initialize a socketpair, and then make sure the child only receives access to one.
if (socketpair(AF_LOCAL, SOCK_SEQPACKET, 0, sockets)) {
fprintf(stderr, "socketpair failed: %m\n");
goto error;
}
if (fcntl(sockets[0], F_SETFD, FD_CLOEXEC)) {
fprintf(stderr, "fcntl failed: %m\n");
goto error;
}
config.fd = sockets[1];
But first we need to set up room for a stack. We'll execve later,
which will actually set up the stack again, so this is only
temporary.5
#define STACK_SIZE (1024 * 1024)
char *stack = 0;
if (!(stack = malloc(STACK_SIZE))) {
fprintf(stderr, "=> malloc failed, out of memory?\n");
goto error;
}
We'll also prepare the cgroup for this process tree. More on this later.
if (resources(&config)) {
err = 1;
goto clear_resources;
}
We'll namespace the mounts, pids, IPC data structures, network devices, and hostname / domain name. I'll go into these more in the code for capabilities, cgroups, and syscalls.
int flags = CLONE_NEWNS | CLONE_NEWCGROUP | CLONE_NEWPID | CLONE_NEWIPC | CLONE_NEWNET | CLONE_NEWUTS;
Stacks on x86, and almost everything else Linux runs on, grow
downwards, so we'll add STACK_SIZE to get a pointer just below the
end.6 We also | the flags with SIGCHLD so
that we can wait on it.
if ((child_pid = clone(child, stack + STACK_SIZE, flags | SIGCHLD, &config)) == -1) {
fprintf(stderr, "=> clone failed! %m\n");
err = 1;
goto clear_resources;
}
Close and zero the child's socket, so that if something breaks then we don't leave an open fd, possibly causing the child to or the parent to hang.
close(sockets[1]); sockets[1] = 0;
The parent process will configure the child's user namespace and then pause until the child process tree exits7.
#define USERNS_OFFSET 10000
#define USERNS_COUNT 2000
int handle_child_uid_map (pid_t child_pid, int fd)
{
int uid_map = 0;
int has_userns = -1;
if (read(fd, &has_userns, sizeof(has_userns)) != sizeof(has_userns)) {
fprintf(stderr, "couldn't read from child!\n");
return -1;
}
if (has_userns) {
char path[PATH_MAX] = {0};
for (char **file = (char *[]) { "uid_map", "gid_map", 0 }; *file; file++) {
if (snprintf(path, sizeof(path), "/proc/%d/%s", child_pid, *file)
> sizeof(path)) {
fprintf(stderr, "snprintf too big? %m\n");
return -1;
}
fprintf(stderr, "writing %s...", path);
if ((uid_map = open(path, O_WRONLY)) == -1) {
fprintf(stderr, "open failed: %m\n");
return -1;
}
if (dprintf(uid_map, "0 %d %d\n", USERNS_OFFSET, USERNS_COUNT) == -1) {
fprintf(stderr, "dprintf failed: %m\n");
close(uid_map);
return -1;
}
close(uid_map);
}
}
if (write(fd, & (int) { 0 }, sizeof(int)) != sizeof(int)) {
fprintf(stderr, "couldn't write: %m\n");
return -1;
}
return 0;
}
The child process will send a message to the parent process about
whether it should set uid and gid mappings. If that works, it will
setgroups, setresgid, and setresuid. Both setgroups and
setresgid are necessary here since there are two separate group
mechanisms on Linux9. I'm also assuming here
that every uid has a corresponding gid, which is common but not
necessarily universal.
int userns(struct child_config *config)
{
fprintf(stderr, "=> trying a user namespace...");
int has_userns = !unshare(CLONE_NEWUSER);
if (write(config->fd, &has_userns, sizeof(has_userns)) != sizeof(has_userns)) {
fprintf(stderr, "couldn't write: %m\n");
return -1;
}
int result = 0;
if (read(config->fd, &result, sizeof(result)) != sizeof(result)) {
fprintf(stderr, "couldn't read: %m\n");
return -1;
}
if (result) return -1;
if (has_userns) {
fprintf(stderr, "done.\n");
} else {
fprintf(stderr, "unsupported? continuing.\n");
}
fprintf(stderr, "=> switching to uid %d / gid %d...", config->uid, config->uid);
if (setgroups(1, & (gid_t) { config->uid }) ||
setresgid(config->uid, config->uid, config->uid) ||
setresuid(config->uid, config->uid, config->uid)) {
fprintf(stderr, "%m\n");
return -1;
}
fprintf(stderr, "done.\n");
return 0;
}
And this is where the child process from clone will end up. We'll
perform all of our setup, switch users and groups, and then load the
executable. The order is important here: we can't change mounts
without certain capabilities, we can't unshare after we limit the
syscalls, etc.
int child(void *arg)
{
struct child_config *config = arg;
if (sethostname(config->hostname, strlen(config->hostname))
|| mounts(config)
|| userns(config)
|| capabilities()
|| syscalls()) {
close(config->fd);
return -1;
}
if (close(config->fd)) {
fprintf(stderr, "close failed: %m\n");
return -1;
}
if (execve(config->argv[0], config->argv, NULL)) {
fprintf(stderr, "execve failed! %m.\n");
return -1;
}
return 0;
}
Capabilties
capabilities subdivide the property of "being root" on Linux. It's
useful to compartmentalize privileges so that, for example a process
can allocate network devices (CAP_NET_ADMIN) but not read all files
(CAP_DAC_OVERRIDE). I'll use them here to drop the ones we don't
want.
But not all of "being root" is subvidivided into capabilities. For example, writing to parts of procfs is allowed by root even after having dropped capabilities10. There are a lot of things like this: this is part of why need other restrictions beside capabilities.
It's also important to think about how we're dropping capabilities. man 7
capabilities has an algorithm for us:
During an execve(2), the kernel calculates the new capabilities of the process using the following algorithm: P'(ambient) = (file is privileged) ? 0 : P(ambient) P'(permitted) = (P(inheritable) & F(inheritable)) | (F(permitted) & cap_bset) | P'(ambient) P'(effective) = F(effective) ? P'(permitted) : P'(ambient) P'(inheritable) = P(inheritable) [i.e., unchanged] where: P denotes the value of a thread capability set before the execve(2) P' denotes the value of a thread capability set after the execve(2) F denotes a file capability set cap_bset is the value of the capability bounding set (described below).
We'd like P'(ambient) and P(inheritable) to be empty, and
P'(permitted) and P(effective) to only include the capabilities
above. This is achievable by doing the following
- Clearing our own inheritable set. This clears the ambient set;
man 7 capabilitiessays "The ambient capability set obeys the invariant that no capability can ever be ambient if it is not both permitted and inheritable." This also clears the child's inheritable set. - Clearing the bounding set. This limits the file capabilities we'll
gain when we
execve, and the rest are limited by clearing the inheritable and ambient sets.
If we were to only drop our own effective, permitted and inheritable
sets, we'd regain the permissions in the child file's capabilities.
This is how bash can call ping, for example.11
Dropped capabilities
int capabilities()
{
fprintf(stderr, "=> dropping capabilities...");
CAP_AUDIT_CONTROL, _READ, and _WRITE allow access to the audit
system of the kernel (i.e. functions like audit_set_enabled, usually
used with auditctl). The kernel prevents messages that normally
require CAP_AUDIT_CONTROL outside of the first pid namespace, but it
does allow messages that would require CAP_AUDIT_READ and
CAP_AUDIT_WRITE from any namespace.12 So
let's drop them all. We especially want to drop CAP_AUDIT_READ,
since it isn't namespaced13 and may contain important
information, but CAP_AUDIT_WRITE may also allow the contained
process to falsify logs or DOS the audit system.
int drop_caps[] = {
CAP_AUDIT_CONTROL,
CAP_AUDIT_READ,
CAP_AUDIT_WRITE,
CAP_BLOCK_SUSPEND lets programs prevent the system from suspending,
either with EPOLLWAKEUP or
/proc/sys/wake_lock.14 Supend isn't namespaced, so
we'd like to prevent this.
CAP_BLOCK_SUSPEND,
CAP_DAC_READ_SEARCH lets programs call open_by_handle_at with an
arbitrary struct file_handle *. struct file_handle is in theory an
opaque type, but in practice it corresponds to inode numbers. So it's
easy to brute-force them, and read arbitrary files. This was used by
Sebastian Krahmer to write a program to read arbitrary system files
from within Docker in 2014.15
CAP_DAC_READ_SEARCH,
CAP_FSETID, without user namespacing, allows the process to modify a
setuid executable without removing the setuid bit. This is pretty
dangerous! It means that if we include a setuid binary in a container,
it's easy for us to accidentally leave a dangerous setuid root binary
on our disk, which any user can use to escalate
privileges.16
CAP_FSETID,
CAP_IPC_LOCK can be used to lock more of a process' own memory than
would normally be allowed17, which could be a way to deny service.
CAP_IPC_LOCK,
CAP_MAC_ADMIN and CAP_MAC_OVERRIDE are used by the mandatory acess
control systems Apparmor, SELinux, and SMACK to restrict access to
their settings. These aren't namespaced, so they could be used by the
contained programs to circumvent system-wide access control.
CAP_MAC_ADMIN, CAP_MAC_OVERRIDE,
CAP_MKNOD, without user namespacing, allows programs to create
device files corresponding to real-world devices. This includes
creating new device files for existing hardware. If this capability
were not dropped, a contained process could re-create the hard disk
device, remount it, and read or write to it.18
CAP_MKNOD,
I was worried that CAP_SETFCAP could be used to add a capability to
an executable and execve it, but it's not actually possible for a
process to set capabilities it doesn't have19. But!
An executable altered this way could be executed by any unsandboxed
user, so I think it unacceptably undermines the security of the
system.
CAP_SETFCAP,
CAP_SYSLOG lets users perform destructive actions against the
syslog. Importantly, it doesn't prevent contained processes from
reading the syslog, which could be risky. It also exposes kernel
addresses, which could be used to circumvent kernel address layout
randomization20.
CAP_SYSLOG,
CAP_SYS_ADMIN allows many behaviors! We don't want most of them
(mount, vm86, etc). Some would be nice to have (sethostname,
mount for bind mounts…) but the extra complexity doesn't seem
worth it.
CAP_SYS_ADMIN,
CAP_SYS_BOOT allows programs to restart the system (the reboot
syscall) and load new kernels (the kexec_load and kexec_file
syscalls)21. We absolutely don't want
this. reboot is user-namespaced, and the kexec* functions only work
in the root user namespace, but neither of those help us.
CAP_SYS_BOOT,
CAP_SYS_MODULE is used by the syscalls delete_module,
init_module, finit_module 22, by the code for kmod 23,
and by the code for loading device modules with ioctl24.
CAP_SYS_MODULE,
CAP_SYS_NICE allows processes to set higher priority on given pids
than the default25. The default kernel scheduler
doesn't know anything about pid namespaces, so it's possible for a
contained process to deny service to the rest of the system26.
CAP_SYS_NICE,
CAP_SYS_RAWIO allows full access to the host systems memory with
/proc/kcore, /dev/mem, and /dev/kmem 27, but a
contained process would need mknod to access these within the
namespace.28. But it also allows things like iopl
and ioperm, which give raw access to the IO ports29.
CAP_SYS_RAWIO,
CAP_SYS_RESOURCE specifically allows circumventing kernel-wide
limits, so we probably should drop it30. But I
don't think this can do more than DOS the
kernel, in general31.
CAP_SYS_RESOURCE,
CAP_SYS_TIME: setting the time isn't namespaced, so we should prevent
contained processes from altering the system-wide
time32.
CAP_SYS_TIME,
CAP_WAKE_ALARM, like CAP_BLOCK_SUSPEND, lets the contained process
interfere with suspend33, and we'd like to prevent that.
CAP_WAKE_ALARM };
size_t num_caps = sizeof(drop_caps) / sizeof(*drop_caps);
fprintf(stderr, "bounding...");
for (size_t i = 0; i < num_caps; i++) {
if (prctl(PR_CAPBSET_DROP, drop_caps[i], 0, 0, 0)) {
fprintf(stderr, "prctl failed: %m\n");
return 1;
}
}
fprintf(stderr, "inheritable...");
cap_t caps = NULL;
if (!(caps = cap_get_proc())
|| cap_set_flag(caps, CAP_INHERITABLE, num_caps, drop_caps, CAP_CLEAR)
|| cap_set_proc(caps)) {
fprintf(stderr, "failed: %m\n");
if (caps) cap_free(caps);
return 1;
}
cap_free(caps);
fprintf(stderr, "done.\n");
return 0;
}
Retained Capabilities
It's important to keep track of the capabilities I'm not dropping, too.
I've heard multiple places34 that CAP_DAC_OVERRIDE
might expose the same functionality as CAP_DAC_READ_SEARCH
(i.e. open_by_handle_at), but as far as I can tell that isn't
true. shocker.c doesn't get anywhere with only
CAP_DAC_OVERRIDE 35, and the
only usage in the kernel is in the Unix permission-checking
code36. So my understanding is that
CAP_DAC_OVERRIDE on its own doesn't allow processes to read outside
of their mount namespaces ("DAC" or "Discretionary Access Control"
refers here to ordinary unix permissions).
CAP_FOWNER, CAP_LEASE, and CAP_LINUX_IMMUTABLE all operate on
files inside of the mount namespace.
Likewise, CAP_SYS_PACCT allows processes to switch accounting on and
off for itself. The acct system call takes a path to log to (which
must be within the mount namespace), and only operates on the calling
process. We're not using process accounting in our containerization,
so turning it off should be harmless as well.37
CAP_IPC_OWNER is only used by functions that respect IPC
namespaces38; since we're in a separate IPC namespace
from the host, we can allow this.
CAP_NET_ADMIN lets processes create network devices;
CAP_NET_BIND_SERVICE lets processes bind to low ports on those
devices; CAP_NET_RAW lets processes send raw packets on those
devices. Since we're going to isolate the networking with a virtual
bridge, and the contained process is inside of a network namespace,
these shouldn't be an issue39. I was wondering
whether we could recreate an existing device like mknod does, but I
don't think it's possible 40.
CAP_SYS_PTRACE doesn't allow ptrace across pid
namespaces41. CAP_KILL doesn't allow signals across
pid namespaces42.
CAP_SETUID and CAPSETGID have similar behaviors43:
Make arbitrary manipulations of process UIDS and GIDs and supplementary GID list, which will only apply to pids in the namespace.forge UID (GID) when passing socket credentials via UNIX domain socketsthe mount namespace should prevent us from reading the host system's unix domain sockets.write a user(group ID) mapping in a user namespace (see user_namespaces(7)): this is/proc/self/uid_map, which will be hidden inside the container.
CAP_SETPCAP only lets processes add or drop capabilities they
already effectively have; man 7 capabilities says
If file capabilities are supported: add any capability from the calling thread's bounding set to its inheritable set; drop capabilities from the bounding set (via prctl(2) PR_CAPBSET_DROP); make changes to the securebits flags.
We've dropped everything relevant from the bounding set, and dropping further capabilities should be harmless.
CAP_SYS_CHROOT is traditionally abused by changing root to a
directory with a setuid root binary and tampered-with dynamic
libraries44. Additionally, it can be used
to escape a chroot "jail"45. Neither of those
should be relevant in our setup so this should be harmless.
Brad Spengler, in "False Boundaries and Arbitrary Code Execution" says
that CAP_SYS_TTYCONFIG can "temporarily change the keyboard
mapping of an administrator's tty via the KDSETKEYCODE ioctl to cause
a different command to be executed than intended", but again this is
an ioctl against a device that should be impossible to access within
the mount namespace.
Mounts
The child process is in its own mount namespace, so we can unmount things that it specifically shouldn't have access to. Here's how:
- Create a temporary directory, and one inside of it.
- Bind mount of the user argument onto the temporary directory
pivot_root, making the bind mount our root and mounting the old root onto the inner temporary directory.umountthe old root, and remove the inner temporary directory.
But first we'll remount everything with MS_PRIVATE. This is mostly a
convenience, so that the bind mount is invisible outside of our
namespace.
<<pivot-root>>
int mounts(struct child_config *config)
{
fprintf(stderr, "=> remounting everything with MS_PRIVATE...");
if (mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL)) {
fprintf(stderr, "failed! %m\n");
return -1;
}
fprintf(stderr, "remounted.\n");
fprintf(stderr, "=> making a temp directory and a bind mount there...");
char mount_dir[] = "/tmp/tmp.XXXXXX";
if (!mkdtemp(mount_dir)) {
fprintf(stderr, "failed making a directory!\n");
return -1;
}
if (mount(config->mount_dir, mount_dir, NULL, MS_BIND | MS_PRIVATE, NULL)) {
fprintf(stderr, "bind mount failed!\n");
return -1;
}
char inner_mount_dir[] = "/tmp/tmp.XXXXXX/oldroot.XXXXXX";
memcpy(inner_mount_dir, mount_dir, sizeof(mount_dir) - 1);
if (!mkdtemp(inner_mount_dir)) {
fprintf(stderr, "failed making the inner directory!\n");
return -1;
}
fprintf(stderr, "done.\n");
fprintf(stderr, "=> pivoting root...");
if (pivot_root(mount_dir, inner_mount_dir)) {
fprintf(stderr, "failed!\n");
return -1;
}
fprintf(stderr, "done.\n");
char *old_root_dir = basename(inner_mount_dir);
char old_root[sizeof(inner_mount_dir) + 1] = { "/" };
strcpy(&old_root[1], old_root_dir);
fprintf(stderr, "=> unmounting %s...", old_root);
if (chdir("/")) {
fprintf(stderr, "chdir failed! %m\n");
return -1;
}
if (umount2(old_root, MNT_DETACH)) {
fprintf(stderr, "umount failed! %m\n");
return -1;
}
if (rmdir(old_root)) {
fprintf(stderr, "rmdir failed! %m\n");
return -1;
}
fprintf(stderr, "done.\n");
return 0;
}
pivot_root is a system call lets us swap the mount at / with
another. Glibc doesn't provide a wrapper for it, but includes a
prototype in the man page. I don't really understand, but OK, we'll
include our own.
int pivot_root(const char *new_root, const char *put_old)
{
return syscall(SYS_pivot_root, new_root, put_old);
}
It's worth noting that I'm avoiding packing and unpackaging containers. This is fertile ground for vulnerabilities46; I'll count on the user to ensure that the mounted directory doesn't contain trusted or sensitive files or hard links.
System Calls
I'll be blacklisting system calls that I can demonstrate causing harm or sandbox escapes. Again this isn't the best way to do this, but it seems like the most illustrative.
Docker's documentation and default seccomp profile are reasonable sources for dangerous system calls47. They also include obsolete sytem calls and calls that overlap with restricted capabilities; I'll ignore those.
Disallowed System Calls
#define SCMP_FAIL SCMP_ACT_ERRNO(EPERM)
int syscalls()
{
scmp_filter_ctx ctx = NULL;
fprintf(stderr, "=> filtering syscalls...");
if (!(ctx = seccomp_init(SCMP_ACT_ALLOW))
We want to prevent new setuid / setgid executables from being created, since in the absence of user namespaces the contained process could create a setuid binary that could be used by any user to get root.48
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(chmod), 1, SCMP_A1(SCMP_CMP_MASKED_EQ, S_ISUID, S_ISUID)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(chmod), 1, SCMP_A1(SCMP_CMP_MASKED_EQ, S_ISGID, S_ISGID)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(fchmod), 1, SCMP_A1(SCMP_CMP_MASKED_EQ, S_ISUID, S_ISUID)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(fchmod), 1, SCMP_A1(SCMP_CMP_MASKED_EQ, S_ISGID, S_ISGID)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(fchmodat), 1, SCMP_A2(SCMP_CMP_MASKED_EQ, S_ISUID, S_ISUID)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(fchmodat), 1, SCMP_A2(SCMP_CMP_MASKED_EQ, S_ISGID, S_ISGID))
Allowing contained processes to start new user namespaces can allow processes to gain new (albeit limited) capabilities, so we prevent it.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(unshare), 1, SCMP_A0(SCMP_CMP_MASKED_EQ, CLONE_NEWUSER, CLONE_NEWUSER)) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(clone), 1, SCMP_A0(SCMP_CMP_MASKED_EQ, CLONE_NEWUSER, CLONE_NEWUSER))
TIOCSTI allows contained processes to write to the controlling
terminal49.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(ioctl), 1, SCMP_A1(SCMP_CMP_MASKED_EQ, TIOCSTI, TIOCSTI))
The kernel keyring system isn't namespaced.50
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(keyctl), 0) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(add_key), 0) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(request_key), 0)
Before Linux 4.8, ptrace totally breaks seccomp51.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(ptrace), 0)
These system calls let processes assign NUMA nodes. I don't have anything specific in mind, but I could see these being used to deny service to some other NUMA-aware application on the host.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(mbind), 0) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(migrate_pages), 0) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(move_pages), 0) || seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(set_mempolicy), 0)
userfaultd allows userspace to handle page
faults52. It doesn't require any privileges, so in
theory it should be safe to be called by an unprivileged user. But it
can be used to pause execution in the kernel by triggering page faults
in system calls. This is an important part in some kernel
exploits53. It's only rarely used legitimately, so
I'll disable it.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(userfaultfd), 0)
I was initially worried about perf_event_open because the Docker
documentation says it "could leak a lot of information on the host",
but it can't be used in our system to see information for
out-of-namespace processes54. But, if
/proc/sys/kernel/perf_event_paranoid is less than 2, it can be used
to discover kernel addresses and possibly uninitialized memory. 2 is
the default since is the default since 4.6, but it can be changed, and
relying on it seems like a bad idea55.
|| seccomp_rule_add(ctx, SCMP_FAIL, SCMP_SYS(perf_event_open), 0)
We'll set PR_SET_NO_NEW_PRIVS to 0. The name is a little vague: it
specifically prevents setuid and setcap'd binaries from being
executed with their additional privileges. This has some security
benefits (it makes it harder for an unprivileged user in-container to
exploit a vulnerability in a setuid or setcap executable to become
in-container root, for example). But it's a little weird, and means
that, for example, ping won't work in a container for an
unprivileged user56.
|| seccomp_attr_set(ctx, SCMP_FLTATR_CTL_NNP, 0)
And we'll actually apply it to the process, and release the context.
|| seccomp_load(ctx)) {
if (ctx) seccomp_release(ctx);
fprintf(stderr, "failed: %m\n");
return 1;
}
seccomp_release(ctx);
fprintf(stderr, "done.\n");
return 0;
}
Allowed System Calls
Here are the system calls that are disallowed by the default Docker policy but permitted by this code:
_sysctl is obsolete and disabled by
default57. alloc_hugepages and
free_hugepages 58, bdflush 59,
create_module 60, nfsservctl 61,
perfctr 62, get_kernel_syms 63, and
setup 64 are not present on modern Linux.
clock_adjtime, clock_settime 65, and
adjtime 66 depend on CAP_SYS_TIME.
pciconfig_read and pciconfig_write 67 and all of the
side-effecting operations of quotactl 68 are prevented by
CAP_SYS_ADMIN.
get_mempolicy and getpagesize reveal information about the memory
layout of the system, but they can be made by unprivileged processes,
and are probably harmless. pciconfig_iobase can be made by
unprivileged processes, and reveals information about PCI decvices.
ustat 69 and sysfs 70 leak some information about
the filesystems, but are nothing that I see as critical. uselib is
more-or-less obsolete, but is just used for loading a shared library
in userspace 71
sync_file_range2 is sync_file_range with swapped argument
order72.
readdir is mostly obsolete, but probably harmless73.
kexec_file_load and kexec_load are prevented by
CAP_SYS_BOOT 74.
nice can only be used to lower priority without
CAP_SYS_NICE 75.
oldfstat, oldlstat, oldolduname, oldstat, and olduname are
just older versions of their respective functions. I expect them to
have the same security properties as the modern ones.
perfmonctl 76 is only available on
IA-64. ppc_rtas 77, spu_create 78 and
spu_run 79, and subpage_prot 80 are only
avaiable on PowerPC. utrap_install is only available on
Sparc81. kern_features is only available on
Sparc64, and should be harmless anyway82.
I don't believe pivot_root is a problem in our setup (but it could
probably be used to circumvent path-based MAC).
preadv2 and pwritev2 are just extensions to preadv and pwritev
/ readv and writev, which are "scatter input" / "gather output"
extensions to read and write 83.
Resources
We'd like to prevent badly-behaved child processes from denying service to the rest of the system84. Cgroups let us limit memory and cpu time in particular; limiting the pid count and IO usage is also useful. There's a very useful document in the kernel tree about it.
The cgroup and cgroup2 filesystems are the canonical interfaces to
the cgroup system. cgroup2 is a little different, and unitialized
on my system, so I'll use the first version here.
Cgroup namespaces are a little different from, for example, mount
namespaces. We need to create the cgroup before we enter a cgroup
namespace; once we do, that cgroup will behave like the root cgroup
inside of the namespace85. This isn't the most
relevant, since a contained process can't mount the cgroup filesystem
or /proc for introspection, but it's nice to be thorough.
I'll set up a struct so I don't have to repeat myself too much, with the following instructions:
- Set
memory/$hostname/memory.limit_in_bytes, so the contained process and its child processes can't total more than 1GB memory in userspace86. - Set
memory/$hostname/memory.kmem.limit_in_bytes, so that the contained process and its child processes can't total more than 1GB memory in userspace87. - Set
cpu/$hostname/cpu.sharesto 256. CPU shares are chunks of 1024; 256 * 4 = 1024, so this lets the contained process take a quarter of cpu-time on a busy system at most88. - Set the
pids/$hostname/pid.max, allowing the contained process and its children to have 64 pids at most. This is useful because there are per-user pid limits that we could hit on the host if the contained process occupies too many89. - Set
blkio/$hostname/weightto 50, so that it's lower than the rest of the system and prioritized accordingly90.
I'll also add the calling process for each of
{memory,cpu,blkio,pids}/$hostname/tasks by writing '0' to it.
#define MEMORY "1073741824"
#define SHARES "256"
#define PIDS "64"
#define WEIGHT "10"
#define FD_COUNT 64
struct cgrp_control {
char control[256];
struct cgrp_setting {
char name[256];
char value[256];
} **settings;
};
struct cgrp_setting add_to_tasks = {
.name = "tasks",
.value = "0"
};
struct cgrp_control *cgrps[] = {
& (struct cgrp_control) {
.control = "memory",
.settings = (struct cgrp_setting *[]) {
& (struct cgrp_setting) {
.name = "memory.limit_in_bytes",
.value = MEMORY
},
& (struct cgrp_setting) {
.name = "memory.kmem.limit_in_bytes",
.value = MEMORY
},
&add_to_tasks,
NULL
}
},
& (struct cgrp_control) {
.control = "cpu",
.settings = (struct cgrp_setting *[]) {
& (struct cgrp_setting) {
.name = "cpu.shares",
.value = SHARES
},
&add_to_tasks,
NULL
}
},
& (struct cgrp_control) {
.control = "pids",
.settings = (struct cgrp_setting *[]) {
& (struct cgrp_setting) {
.name = "pids.max",
.value = PIDS
},
&add_to_tasks,
NULL
}
},
& (struct cgrp_control) {
.control = "blkio",
.settings = (struct cgrp_setting *[]) {
& (struct cgrp_setting) {
.name = "blkio.weight",
.value = PIDS
},
&add_to_tasks,
NULL
}
},
NULL
};
Writing to the cgroups version 1 filesystem works like this91:
- In each controller, you can create a cgroup with a name with
mkdir. For memory,mkdir /sys/fs/cgroup/memory/$hostname. - Inside of that you can write to the individual files to set
values. For example,
echo $MEMORY > /sys/fs/cgroup/memory/$hostname/memory.limit_in_bytes. - You can a pid to
tasksto add the process tree to the cgroup. "0" is a special value that means "the writing process".
so I'll iterate over that structure and fill in the values.
int resources(struct child_config *config)
{
fprintf(stderr, "=> setting cgroups...");
for (struct cgrp_control **cgrp = cgrps; *cgrp; cgrp++) {
char dir[PATH_MAX] = {0};
fprintf(stderr, "%s...", (*cgrp)->control);
if (snprintf(dir, sizeof(dir), "/sys/fs/cgroup/%s/%s",
(*cgrp)->control, config->hostname) == -1) {
return -1;
}
if (mkdir(dir, S_IRUSR | S_IWUSR | S_IXUSR)) {
fprintf(stderr, "mkdir %s failed: %m\n", dir);
return -1;
}
for (struct cgrp_setting **setting = (*cgrp)->settings; *setting; setting++) {
char path[PATH_MAX] = {0};
int fd = 0;
if (snprintf(path, sizeof(path), "%s/%s", dir,
(*setting)->name) == -1) {
fprintf(stderr, "snprintf failed: %m\n");
return -1;
}
if ((fd = open(path, O_WRONLY)) == -1) {
fprintf(stderr, "opening %s failed: %m\n", path);
return -1;
}
if (write(fd, (*setting)->value, strlen((*setting)->value)) == -1) {
fprintf(stderr, "writing to %s failed: %m\n", path);
close(fd);
return -1;
}
close(fd);
}
}
fprintf(stderr, "done.\n");
I'll also lower the hard limit on the number of file descriptors. The
file descriptor number, like the number of pids, is per-user, and so
we want to prevent in-container process from occupying all of
them. Setting the hard limit sets a permanent upper bound for this
process tree, since I've dropped
CAP_SYS_RESOURCE 92.
fprintf(stderr, "=> setting rlimit...");
if (setrlimit(RLIMIT_NOFILE,
& (struct rlimit) {
.rlim_max = FD_COUNT,
.rlim_cur = FD_COUNT,
})) {
fprintf(stderr, "failed: %m\n");
return 1;
}
fprintf(stderr, "done.\n");
return 0;
}
We'd also like to clean up the cgroup for this hostname. There's
built-in functionality for this, but we would need to change
system-wide values to do it cleanly93. Since we
have the contained process waiting on the contained process, it's
simple to do it this way. First we move the contained process back
into the root tasks; then, since the child process is finished, and
leaving the pid namespace SIGKILLS its children, the tasks is
empty. We can safely rmdir at this point.
int free_resources(struct child_config *config)
{
fprintf(stderr, "=> cleaning cgroups...");
for (struct cgrp_control **cgrp = cgrps; *cgrp; cgrp++) {
char dir[PATH_MAX] = {0};
char task[PATH_MAX] = {0};
int task_fd = 0;
if (snprintf(dir, sizeof(dir), "/sys/fs/cgroup/%s/%s",
(*cgrp)->control, config->hostname) == -1
|| snprintf(task, sizeof(task), "/sys/fs/cgroup/%s/tasks",
(*cgrp)->control) == -1) {
fprintf(stderr, "snprintf failed: %m\n");
return -1;
}
if ((task_fd = open(task, O_WRONLY)) == -1) {
fprintf(stderr, "opening %s failed: %m\n", task);
return -1;
}
if (write(task_fd, "0", 2) == -1) {
fprintf(stderr, "writing to %s failed: %m\n", task);
close(task_fd);
return -1;
}
close(task_fd);
if (rmdir(dir)) {
fprintf(stderr, "rmdir %s failed: %m", dir);
return -1;
}
}
fprintf(stderr, "done.\n");
return 0;
}
Networking
Container networking takes a little too much explanation for this space. It usually works like this:
- Create a bridge device.
- Create a virtual ethernet pair and attach one end to the bridge.
- Put the other end in the network namespace.
- For outside networking access, the host needs to be set to forward (and possibly NAT) packets.
Having multiple contained processes sharing a bridge device would mean they're both on the same LAN from the host's perspective. So ARP spoofing is a recurring issue with containers that work this way94.
The canonical way to do this from C is the rtnetlink interface; it
would probably be easier to use ip link ....
We could also limit the network usage with the net_prio cgroup
controller95.