Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OMGzip: Optimized Multithreaded Gzip decompressor

Up to 2x faster than rapidgzip. A standalone C program for Linux x86-64, providing parallel gzip/BGZF decompression with strictly bounded memory requirements.

Why another gzip decompressor?

Gzip is everywhere. Spare CPU cores usually are, too. Getting them to cooperate on decoding is the difficult part: gzip's underlying DEFLATE format allows references to earlier output, so workers cannot simply start decoding independently at arbitrary positions.

OMGzip was inspired by rapidgzip, an innovative, state-of-the-art parallel decompressor, described by Maximilian Knespel and Holger Brunst in Rapidgzip: Parallel Decompression and Seeking in Gzip Files Using Cache Prefetching (HPDC 2023). OMGzip explores a different way to decode speculatively: let Intel's exceptionally fast ISA-L do the difficult work, even if that means running its inflater twice.

The goal is to be boringly better at one job: turn compressed bytes into the right uncompressed bytes, quickly. That means a plain executable, sensible pipeline behavior, an explicit memory budget, and treating truncated or corrupt input as an error rather than a successful shortcut. No new format, index, or recompression is needed.

Features

  • Fast native decoding. Intel ISA-L supplies inflate and CRC routines, many written in highly optimized assembly; OMGzip adds its own block finding, symbolic decoding and SIMD dependency resolution. C keeps the interface to ISA-L direct and the runtime small. Scalar, AVX2 and AVX-512 implementations are selected separately for OMGzip's kernels.
  • Ordinary gzip and BGZF. Text, FASTQ and binary contents are all supported, including concatenated gzip members. BGZF is detected automatically and uses a dedicated parallel engine. Read from a file or stdin; write to a file or stdout.
  • Bounded, user-controlled memory. Allocation limits depend on your settings, not the input length or an unexpectedly high compression ratio. Inspect the budget before decoding with --show-config.
  • Validation stays on. CRC32 and ISIZE checks are always enabled. Malformed DEFLATE, truncation, read errors and write errors produce a nonzero exit status.
  • No pinning required. Worker-local inflate buffers and locality-aware reuse were designed and tested for unpinned execution, including dual-socket NUMA servers. Explicit affinity remains an optional tuning tool.

Performance

Here is OMGzip against the official PyPI distribution of rapidgzip 0.16.0 on a six-core, twelve-thread Xeon W-2133, decompressing 21.04 GB of concatenated FASTQ data:

OMGzip and rapidgzip throughput on Intel Xeon W-2133, with mean throughput and 95% confidence intervals.

At P=4 through P=12, OMGzip delivers approximately 1.87–1.99x the throughput, using the ratio of median wall times. The twofold headline does not require AVX-512: on the AVX2-only Xeon E3-1230 v5, P=4 took 16.28 seconds with OMGzip versus 33.88 seconds with rapidgzip, a 2.08x ratio.

These are warmed regular-file-to-/dev/null measurements, with defaults apart from P and no CPU pinning. They measure decompression throughput, not disk speed or the speed of a downstream pipeline. The advantage depends on the data and hardware; it is not a promise that every gzip file will decompress twice as fast.

Benchmark details, dual-socket scaling, and reproducing the input

Systems and method

The same production OMGzip executable was used on all five systems, with its source code identical to public commit c6dd86a. It was built with GCC 14.2.0, conservative baseline flags and separately compiled SIMD kernels, against the pinned ISA-L revision 11f2004. There were no benchmark-specific code changes. All hosts used Debian 13 and Linux 6.12; rapidgzip was version 0.16.0 from PyPI.

System Physical cores / logical CPUs, total Decoded input Runs per program and P
Intel Xeon E3-1230 v5 4 / 8 21.04 GB 50
Intel Xeon W-2133 6 / 12 21.04 GB 50
Intel Xeon E-2386G 6 / 12 21.04 GB 50
2x Intel Xeon 6517P 32 / 64 168.3 GB 25
2x AMD EPYC 9474F 96 / 192 168.3 GB 25

Before timing, compressed-input hashes were checked, and both programs' decompressed output was checked against the expected SHA-256 and byte count at the highest requested P. A cache read and at least five minutes of OMGzip warmup followed. All program × P × repetition combinations were then shuffled together, rather than running one program's measurements first. Every observation was retained: no outlier removal, retries or selection of lucky minimum times.

The timed commands were simply:

omgzip -d -c -P "$P" "$FILE" > /dev/null
rapidgzip -d -c -P "$P" "$FILE" > /dev/null

GNU time recorded elapsed time and resource use. Both programs retained normal gzip validation. Linux managed placement across all available CPUs; neither internal pinning nor taskset/numactl was used. The plots show the arithmetic mean of each run's decoded GB/s, with Student-t 95% confidence intervals; GB means 1,000,000,000 bytes. Numerical speedup statements use median wall times.

P is the argument passed to each program, not a guarantee of identical total thread counts. For ordinary file input, OMGzip creates P workers plus its main and output threads. Their extra CPU use matters particularly at P=2, so the headline does not rely on the larger ratios seen there.

Across two sockets

OMGzip and rapidgzip throughput on two Intel Xeon 6517P processors.

On the dual Xeon 6517P, both programs benefit substantially from using the SMT range. At P=64, median times are 10.35 seconds for OMGzip and 17.60 seconds for rapidgzip: approximately 16.26 versus 9.56 GB/s, or 1.70x.

OMGzip and rapidgzip throughput on two AMD EPYC 9474F processors.

On the dual EPYC, OMGzip's advantage is largest at moderate P; the two programs approach parity around 21 GB/s at P=192. That server had only four populated memory channels per socket. A shared memory-system ceiling is plausible, but this benchmark did not measure DRAM or interconnect traffic, so it does not establish the cause. The complete curve is more useful than hiding the endpoint.

The other single-socket results are available as plots for the Xeon E3-1230 v5 and Xeon E-2386G.

Recreating the fixtures

The input comes from the publicly available GIAB NA12878 Garvan HiSeq exome dataset. Download these four original compressed files from that directory, then concatenate them without decompressing or recompressing:

cat \
    NIST7035_TAAGGCGA_L001_R2_001_trimmed.fastq.gz \
    NIST7035_TAAGGCGA_L002_R2_001_trimmed.fastq.gz \
    NIST7086_CGTACTAG_L001_R2_001_trimmed.fastq.gz \
    NIST7086_CGTACTAG_L002_R2_001_trimmed.fastq.gz \
    > fastq4.gz

for i in 1 2 3 4 5 6 7 8; do
    cat fastq4.gz
done > fastq32.gz

fastq4.gz contains four gzip members and decodes to 21,037,400,609 bytes. fastq32.gz contains eight copies of that compressed stream, hence 32 members and 168,299,204,872 decoded bytes. The longer input keeps the server runs long enough to measure usefully.

Expected compressed SHA-256 values:

eb8bcfd8612b0ccadb2252f42a5645f5923e2f043403ebf1e5c91aa82cfb817d  fastq4.gz
63b56bae57dd0d546a249c00ca852ec9cd5bac3bf5c92018f11ff6f4fb9369b9  fastq32.gz

Expected decompressed SHA-256 values, obtainable with gzip -dc FILE | sha256sum:

d519fc54a4634cae5896641e0884a46a382ceb54ce78502e1513558bcaa65914  fastq4
7ef9486a305c91364c27341d5522ed9731fc31def4518bb06a738be7597a91f5  fastq32

The benchmark inputs are not bundled with the repository; download them from the linked dataset.

Installation

The supported platform is Linux x86-64. Building requires a recent GCC or Clang, GNU Make, CMake 3.12 or newer, and NASM 2.14.01 or newer for ISA-L's assembly. GCC 14 and Clang 19 have been used for development and validation.

On Debian-based systems:

sudo apt-get update
sudo apt-get install --no-install-recommends build-essential git cmake nasm

git clone --recursive https://github.com/clinbiolab/OMGzip.git
cd OMGzip
make -j4

The executable is build/omgzip. Run it there or copy it to a directory on your PATH. ISA-L is built from the pinned submodule and linked statically; it does not need to be installed separately. The production executable needs no Python, Perl, C++ runtime or separately installed compression library. The benchmark binary's only shared-library dependency is glibc, requiring version 2.34 or newer; a local build uses the host toolchain and libc.

For Clang, use separate build directories:

make -j4 CC=clang-19 BUILD_DIR=build/clang ISAL_BUILD_DIR=build/isal-clang

The compiler name can be changed to the Clang installed on your machine. Builds do not download dependencies or update submodules. If the initial clone omitted them, initialize them explicitly with git submodule update --init --recursive before building.

Usage

Examples below assume omgzip is on your PATH; otherwise use ./build/omgzip.

# Decompress to a named file; retain the compressed input.
omgzip -P 4 -o reads.fastq reads.fastq.gz

# Feed a downstream program.
omgzip -P 8 reads.fastq.gz | wc -l

# Receive compressed input through a pipe.
cat lane1.fastq.gz lane2.fastq.gz | omgzip -P 8 > combined.fastq

# BGZF is detected automatically; the filename suffix does not matter.
omgzip -P 8 -o variants.vcf variants.vcf.gz

# Use a genuinely single-threaded ISA-L path.
omgzip -P 0 data.gz > data

# Check the settings and allocation budget without opening the input.
omgzip -P 8 --show-config

# Show all options.
omgzip --help

The default is four workers and output to stdout. -d and -c, including -dc, are accepted for compatibility with familiar decompression commands. A named output selected with -o must not already exist unless -f is given. Shell redirection follows the shell's overwrite rules. OMGzip takes one input operand; with no filename, or with -, it reads stdin.

In Bash pipelines, set -o pipefail ensures that an error from the decompressor is not hidden by a successful downstream command. OMGzip exits with 0 on success, 1 on a decoding or runtime error, and 2 on an invocation error.

More threads are not always the right answer. A slow disk or consumer sets its own limit, and very small files may not repay parallel startup. Ordinary files with no useful dynamic-block boundaries, notably typical igzip -0 output, decode correctly but offer little or no speculative parallelism. igzip -dc is an excellent lightweight single-threaded alternative; omgzip -P 0 uses the same backend with comparable performance. By contrast, -P 1 really creates one parallel-engine worker.

See the user manual for every option, memory tuning, BGZF details, CPU placement with --affinity, taskset or numactl, and pipeline examples.

Memory: bounded does not mean tiny

At the ordinary-gzip defaults, the main storage arena reserves 256 MiB per worker, plus smaller buffers and metadata. Four workers require about 1.02 GiB of allocation payload for file input or 1.06 GiB for a stream; BGZF uses a different, smaller budget at that worker count. The single-threaded mode needs only a few MiB.

Actual RSS is often considerably lower, but do not size a job on that assumption: a system with restrictive overcommit can require backing for the full allocation. --show-config reports the exact maximum live allocation payload for the chosen settings and each input mode. Thread stacks, allocator/runtime overhead and kernel buffers are additional. The manual explains how to reduce the budget and what that costs in throughput.

Highly compressible input does not silently enlarge a speculative chunk. If it exceeds the configured capacity, OMGzip continues with bounded serial decoding. Users can explicitly trade memory for more parallel work, for example with --capacity-multiplier 16 on highly compressible binary data.

How it works

Speculative decoding without a custom DEFLATE decoder

A worker starting in the middle of a gzip member has a problem: DEFLATE back-references may refer to the preceding 32 KiB of history, which that worker does not yet know. Literal bytes are immediately known; some copied bytes are not.

Rapidgzip addresses this with a custom decoder that can retain unresolved references alongside literals, then resolve them later. OMGzip takes a different route: it turns ISA-L's existing inflater into a symbolic decoder by running the same segment twice, with two specially constructed synthetic histories, A and B.

Every position in the unknown history is assigned a unique pair of byte values, one in each dictionary. DEFLATE only copies bytes; it does not transform or combine them. Those pairs therefore propagate through the decoded output, even through overlapping matches and copies of copies.

For corresponding output bytes:

  • If A equals B, the byte is already independent of the unknown history.
  • Otherwise, A together with A XOR B identifies the exact incoming-history position from which that byte originated.

This is an exact encoding, not a hash or a probabilistic guess. More concretely, for history index i = 256*h + l, the dictionaries contain A[i] = l and B[i] = l XOR (h + 1). A nonzero delta = A XOR B gives the original index as 256*(delta - 1) + A. Once the real preceding history is known, a lookup replaces each symbolic byte with its final value.

Running inflate twice sounds extravagant. The payoff is that both passes use ISA-L's highly optimized decoder, while pairing and resolution are much simpler operations than DEFLATE decoding. On the workloads shown above, the complete implementation beats rapidgzip despite that extra pass. The scheduler and memory layout matter too; the benchmark is not a claim that two ISA-L calls alone explain the whole speedup.

Nor does every byte have to be decoded twice. Once a full 32-KiB slab of A and B agrees, all subsequent history is known: the rest of the chunk needs only A decoding, without further symbolic pairing or resolution. CRC checking still covers every output byte.

Keeping the cores supplied

Workers search for plausible dynamic-block starts and prepare chunks ahead of the authoritative decoder. A successful speculative decode is not enough to trust a chunk: its start must meet the real decoding frontier. The main thread then propagates the next chunk's history using only the necessary tail, allowing whole-body resolution and CRC calculation to proceed in parallel. Ordered output combines those checksums and validates every gzip member.

Decoding and resolution share a worker pool, with resolution taking priority and a preference for reusing the original producer. Small private inflate workspaces separate the decoder's frequent memory accesses from the larger retained bodies. Queues, slots and input rings are bounded; file and streaming input have separate compiled hot paths. BGZF bypasses symbolic work entirely and groups independent members into batches for its own simpler parallel engine.

The full explanation, including correctness invariants, is in ALGORITHMS.md and ARCHITECTURE.md.

Testing and diagnostics

The test suite generates its own synthetic inputs from readable recipes. No prebuilt gzip/BGZF fixtures or external datasets are included or needed.

Tests require Bash, GNU gzip, coreutils and Perl, including Compress::Raw::Zlib and Digest::SHA. The full tier also needs zlib development headers and util-linux's script. On Debian-based systems, install the test prerequisites with:

sudo apt-get install --no-install-recommends \
    bash coreutils gzip perl libcompress-raw-zlib-perl libdigest-sha-perl \
    zlib1g-dev util-linux

These are test-only dependencies, not requirements for building or running the production executable. No CPAN setup is needed.

make -j4 test          # CLI smoke tests, clean and stats builds
make -j4 test-full     # Components, malformed input, arguments, live streams
make -j4 test-stress   # Full tier plus extended scheduling and boundary checks
make -j4 test-asan     # Full tier with AddressSanitizer and UBSan
make -j4 test-tsan     # Full tier with ThreadSanitizer

Sanitizer targets need the selected compiler's sanitizer runtimes and a host that permits them to run. They instrument OMGzip's code, not ISA-L. Generated fixtures and logs remain under build/. See tests/README.md for coverage, prerequisites and reproducible stress settings; make help lists all supported targets.

For performance investigation, make stats builds build/omgzip-stats, which accepts the same options and emits detailed accounting to stderr. Normal omgzip compiles that instrumentation out. The output is documented in STATISTICS.md.

AI use disclosure

OMGzip was written in C, but the programming was done in English.

Agentic AI performed the implementation and played a substantial role in brainstorming, debugging, testing and documentation. Human direction supplied the requirements, algorithmic ideas, experimental questions and final design decisions. The development process involved repeated measurements, failed hypotheses, differential tests and rather a lot of discussion about memory traffic.

No, this was not written in one prompt.

Acknowledgements and third-party notices

The largest thank-you goes to the authors and contributors of Intel ISA-L. It is the decoding engine that makes this approach practical. OMGzip includes a focused fork of ISA-L 2.32.1 with resumable DEFLATE-boundary stopping, rather than maintaining its own inflate implementation.

Rapidgzip provided the inspiration and a state-of-the-art comparison point. OMGzip is not a rewrite, translation or fork of rapidgzip: its core A/B symbolic machinery and schedulers were developed independently. Two specific source-informed connections deserve explicit credit: the deep dynamic-block validator is a C reimplementation of validation logic studied in rapidgzip 0.16.0, and the small ISA-L stop/resume extension was informed by the interface and deferred-output handling in rapidgzip's ISA-L fork. The rapidgzip library is not linked or required at runtime.

omgzip --oss-attributions prints the third-party notices and license texts, including ISA-L's BSD-3-Clause license and rapidgzip's MIT/Apache-2.0 licensing. Those components and notices retain their respective terms.

Copyright and license

Copyright (c) 2026 Fedor Konovalov.

OMGzip is licensed under the MIT License.

About

OMGzip: Optimized Multithreaded Gzip decompressor. Up to 2x faster than rapidgzip.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages