How and Why fork() Uses Copy-on-Write
In the last video, we learned about copy-on-write (CoW) and understood its mechanics by digging inside the mmap pathway in the kernel. This is a continuation of that where we look at the fork() system call and how it uses copy-on-write behind the scenes.
Following is a timestamped list of chapters in the video to help you navigate quickly. Also, I recommend watching the video at a higher speed to get a better experience.
- (00:00) Introduction: Copy-on-write with
fork()and the Instagram example. - (02:54) Copy-on-write recap: Sharing physical frames and making private copies when a process writes.
- (10:08) What fork() does: Creating a child process and inheriting the parent’s state.
- (13:15) A code example: Return values, inherited memory, and independent writes in parent and child.
- (18:26) Why use copy-on-write?: The cost of copying the parent’s memory.
- (21:05) Page tables after fork(): How parent and child initially share physical frames.
- (23:45) The shell and exec(): Why copying memory upfront would often be wasted work.
- (26:04) Handling a write: Read-only page-table entries, page faults, VMA permissions, and copying a frame.
- (30:01) Benefits and tradeoffs: Memory savings and the performance cost of deferred copying.
- (33:41) Instagram’s memory-sharing problem: Shared memory gradually becoming private in Python worker processes.
- (41:14) Python objects and reference counting: Why an expression like
if x is Nonecan cause memory writes. - (45:16) Inside the interpreter: How evaluating an expression changes reference counts.
- (48:12) Immortal objects: Avoiding reference-count updates for objects that live throughout the runtime.
- (50:19) Wrap-up: Plans for a hands-on demonstration using Linux memory debugging tools.
If you are new to this series, it is based on my ebook called “Virtual Memory from First Principles”. It is available to read for free online and also available to purchase from Gumroad (PDF/Epub) and Amazon (Kindle edition).
And, if you want to watch the previous videos in this series, the following is what has been published so far:
- What is virtual memory and why do we need it
- Size of the virtual address space of a process
- Address space of a process
- Structure of page tables and how page table walks work
- What is inside a page table entry
- Size of a physical address
- Demand paging
- Demand Paging in Action: mmap, Page Faults, and RSS
- How Copy-on-Write Works with Memory-Mapped Files