用 500 行代码实现 Linux 容器
Linux containers in 500 lines of code (2016)

原始链接: https://blog.lizzie.io/linux-containers-in-500-loc.html

`capabilities()` 函数通过从进程的能力边界集(capability bounding set)中移除危险的 Linux capabilities,并使用 `prctl(PR_CAPBSET_DROP)` 和 libcap 清除其可继承标志,来增强容器的安全性。 被移除的 capabilities 包括: - 审计控制、读取和写入能力,防止审计被操纵或泄露。 - 挂起和唤醒锁控制能力,保护主机可用性。 - `CAP_DAC_READ_SEARCH`,该能力可通过 `open_by_handle_at()` 实现任意文件访问。 - `CAP_FSETID` 和 `CAP_SETFCAP`,可用于保留或为可执行文件添加危险权限。 - MAC 管理/覆盖、设备创建、内存锁定和内核管理能力。 - 系统重启、内核加载、模块操作、调度器优先级更改、原始 I/O/内存访问、资源限制绕过以及系统时钟修改能力。 总体而言,该代码可防止容器内进程削弱主机范围内的安全性、访问主机资源、篡改日志或可执行文件、干扰服务,或利用未完全隔离的命名空间内核功能获取权限。

# 黑客新闻 最新 | 往期 | 评论 | 提问 | 展示 | 工作机会 | 提交 登录 **用 500 行代码实现 Linux 容器** (lizzie.io) 12 分 由 **mkornaukhov** 提交 3 小时前 隐藏 | 往期 | 收藏 | 1 条评论 **help** **js2** 18 分钟前 [–] (2016)。此前带有评论的投稿: https://news.ycombinator.com/item?id=30623372 (250 分 | 2022 年 3 月 10 日 | 27 条评论) https://news.ycombinator.com/item?id=22232705 (267 分 | 2020 年 2 月 4 日 | 29 条评论) https://news.ycombinator.com/item?id=15608435 (440 分 | 2017 年 11 月 2 日 | 53 条评论) 回复 考虑申请 YC 的 2027 年冬季批次! 申请截止日期为 11 月 2 日。 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请加入 YC | 联系方式 搜索:
相关文章

原文
int capabilities()
{
	fprintf(stderr, "=> dropping capabilities...");

CAP_AUDIT_CONTROL, _READ, and _WRITE allow access to the audit system of the kernel (i.e. functions like audit_set_enabled, usually used with auditctl). The kernel prevents messages that normally require CAP_AUDIT_CONTROL outside of the first pid namespace, but it does allow messages that would require CAP_AUDIT_READ and CAP_AUDIT_WRITE from any namespace. So let's drop them all. We especially want to drop CAP_AUDIT_READ, since it isn't namespaced and may contain important information, but CAP_AUDIT_WRITE may also allow the contained process to falsify logs or DOS the audit system.

	int drop_caps[] = {
		CAP_AUDIT_CONTROL,
		CAP_AUDIT_READ,
		CAP_AUDIT_WRITE,

CAP_BLOCK_SUSPEND lets programs prevent the system from suspending, either with EPOLLWAKEUP or /proc/sys/wake_lock. Supend isn't namespaced, so we'd like to prevent this.

		CAP_BLOCK_SUSPEND,

CAP_DAC_READ_SEARCH lets programs call open_by_handle_at with an arbitrary struct file_handle *. struct file_handle is in theory an opaque type, but in practice it corresponds to inode numbers. So it's easy to brute-force them, and read arbitrary files. This was used by Sebastian Krahmer to write a program to read arbitrary system files from within Docker in 2014.

		CAP_DAC_READ_SEARCH,

CAP_FSETID, without user namespacing, allows the process to modify a setuid executable without removing the setuid bit. This is pretty dangerous! It means that if we include a setuid binary in a container, it's easy for us to accidentally leave a dangerous setuid root binary on our disk, which any user can use to escalate privileges.

		CAP_FSETID,

CAP_IPC_LOCK can be used to lock more of a process' own memory than would normally be allowed, which could be a way to deny service.

		CAP_IPC_LOCK,

CAP_MAC_ADMIN and CAP_MAC_OVERRIDE are used by the mandatory acess control systems Apparmor, SELinux, and SMACK to restrict access to their settings. These aren't namespaced, so they could be used by the contained programs to circumvent system-wide access control.

		CAP_MAC_ADMIN,
		CAP_MAC_OVERRIDE,

CAP_MKNOD, without user namespacing, allows programs to create device files corresponding to real-world devices. This includes creating new device files for existing hardware. If this capability were not dropped, a contained process could re-create the hard disk device, remount it, and read or write to it.

		CAP_MKNOD,

I was worried that CAP_SETFCAP could be used to add a capability to an executable and execve it, but it's not actually possible for a process to set capabilities it doesn't have. But! An executable altered this way could be executed by any unsandboxed user, so I think it unacceptably undermines the security of the system.

		CAP_SETFCAP,

CAP_SYSLOG lets users perform destructive actions against the syslog. Importantly, it doesn't prevent contained processes from reading the syslog, which could be risky. It also exposes kernel addresses, which could be used to circumvent kernel address layout randomization.

		CAP_SYSLOG,

CAP_SYS_ADMIN allows many behaviors! We don't want most of them (mount, vm86, etc). Some would be nice to have (sethostname, mount for bind mounts…) but the extra complexity doesn't seem worth it.

		CAP_SYS_ADMIN,

CAP_SYS_BOOT allows programs to restart the system (the reboot syscall) and load new kernels (the kexec_load and kexec_file syscalls). We absolutely don't want this. reboot is user-namespaced, and the kexec* functions only work in the root user namespace, but neither of those help us.

		CAP_SYS_BOOT,

CAP_SYS_MODULE is used by the syscalls delete_module, init_module, finit_module , by the code for kmod , and by the code for loading device modules with ioctl.

		CAP_SYS_MODULE,

CAP_SYS_NICE allows processes to set higher priority on given pids than the default. The default kernel scheduler doesn't know anything about pid namespaces, so it's possible for a contained process to deny service to the rest of the system.

		CAP_SYS_NICE,

CAP_SYS_RAWIO allows full access to the host systems memory with /proc/kcore, /dev/mem, and /dev/kmem , but a contained process would need mknod to access these within the namespace.. But it also allows things like iopl and ioperm, which give raw access to the IO ports.

		CAP_SYS_RAWIO,

CAP_SYS_RESOURCE specifically allows circumventing kernel-wide limits, so we probably should drop it. But I don't think this can do more than DOS the kernel, in general.

		CAP_SYS_RESOURCE,

CAP_SYS_TIME: setting the time isn't namespaced, so we should prevent contained processes from altering the system-wide time.

		CAP_SYS_TIME,

CAP_WAKE_ALARM, like CAP_BLOCK_SUSPEND, lets the contained process interfere with suspend, and we'd like to prevent that.

		CAP_WAKE_ALARM
	};
	size_t num_caps = sizeof(drop_caps) / sizeof(*drop_caps);
	fprintf(stderr, "bounding...");
	for (size_t i = 0; i < num_caps; i++) {
		if (prctl(PR_CAPBSET_DROP, drop_caps[i], 0, 0, 0)) {
			fprintf(stderr, "prctl failed: %m\n");
			return 1;
		}
	}
	fprintf(stderr, "inheritable...");
	cap_t caps = NULL;
	if (!(caps = cap_get_proc())
	    || cap_set_flag(caps, CAP_INHERITABLE, num_caps, drop_caps, CAP_CLEAR)
	    || cap_set_proc(caps)) {
		fprintf(stderr, "failed: %m\n");
		if (caps) cap_free(caps);
		return 1;
	}
	cap_free(caps);
	fprintf(stderr, "done.\n");
	return 0;
}
联系我们 contact @ memedata.com