Amid multiple ongoing innovations being pursued in the Linux memory management space amid sky high RAM prices, a new patch series posted on Monday aim to enhance the Zswap performance for this compressed RAM cache for swap pages. Linux developer Usama Arif posted the set of two patches for Zswap in working to reduce request contention on loads for Zswap. Following similar work done to zRAM, the Zswap code is being adapted to address the current shortcoming of stores and loads sharing a per-CPU acomp request and mutex. This current handling is ultimately proving to be a bottleneck while the two patches help in reducing contention on loads. Usama Arif explained further in the patch cover letter: "Stores and loads share a per-CPU acomp request and mutex. A low-priority store can be preempted right after the compressor drops its stream lock, while it still holds the zswap mutex, and a higher-priority load on that CPU then waits for the store to run again. This follows the work from Sergey Senozhatsky's zram series which splits it for the same reason. Patch 1 gives compression and decompression separate requests, waits and mutexes, so loads no longer wait for stores, though they can still wait for each other. Patch 2 decompresses with an on-stack request when the algorithm is synchronous and needs no request context, which covers all in-tree software compressors, so those loads take no zswap lock. Asynchronous algorithms keep the per-CPU request and mutex. For software compressors the series allocates the same number of requests as before; each per-CPU context grows by 72 bytes, and the load path is about 270 bytes deeper on x86-64." What's impressive are the end results: Usama explained further with the patch series: "The numbers below are the slowest read per run, as a median (min-max) of 5 runs. Each run is 12 seconds in a zstd VM with lazy preemption, vm.page-cluster=0 and swap on /dev/ram0. With 1 vCPU, four nice +10 workers page memory out and read it back while a nice 0 task spins. A nice -19 reader pages out its own buffer and measures how long each read of it takes. With 8 vCPUs there are 16 workers, 8 spinning tasks and 8 readers. Before series (ms) With series (ms) 1 vCPU 22.3 (21.6-22.6) 0.97 (0.72-1.4) 8 vCPUs 314 (97-2542) 7.0 (5.0-98) Reads over 10 ms fell from 26-35 per run to none with 1 vCPU, and from 3-18 per run to at most one with 8 vCPUs. The benchmark and test programs were written with the help of an LLM." Overall this looks to be a very nice improvement to Zswap that is now under review on the Linux kernel mailing list.
New Patches For Linux's Zswap Show Major Improvements
Full Article
Original Source
Read the full article at Phoronix →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.