mothergod — chevrons compressing the golden byte

mothergod

A general-purpose lossless compressor in Rust, built for compression ratio rather than speed, and meant to be judged on bits per byte against zstd -19 and xz -9e.

Wins Canterbury, vs zstd -19 and xz -9e
Loses Silesia, vs zstd -19 and xz -9e
Not yet pre-alpha, no release
Status

Pre-alpha: no release, no packaged binary, no version tag. The container format (version 10) carries Stored and Lz. Lz finds repeated stretches of the input, works out the cheapest way to describe each one rather than taking the first match, then encodes what is left with an arithmetic coder whose byte predictions blend several models and keep adjusting as they read. That format is specified (format spec) and versioned, not frozen: until 1.0 a version can be retired, and a frame from a build before the first release has no promise at all. A version a release has written is retired only after a later release that still reads it and writes its successor, named in the changelog, so you can re-compress first. From 1.0 on, no version is ever retired. Everything around it still moves: the API, the CLI, and the ratio below. Do not use this for data you care about yet.

What this is

Compression runs in three stages. First, reversible transforms rearrange bytes to expose patterns a plain byte-by-byte view hides. Second, an LZ stage finds repeated stretches of the input and works out the cheapest way to describe them, including cheap references back to a handful of recently used distances. Third, an arithmetic coder encodes what remains, blending predictions from several models and adjusting that blend as it reads. Ratio came first, so judge it on bits per byte; the speed section below says what that order cost. Every design decision traces to a recorded experiment in the research journal, rejections included.

Measured

Aggregate bits per byte on the two named corpora, lower is better. Canterbury is 11 files and 2.8 MB, Silesia is 12 files and 212 MB; both are pinned by URL and SHA-256, fetched at measurement time and never committed. Measured 2026-10-03 against gzip 1.12, Zstandard 1.5.7 and XZ Utils 5.4.5, at the flags in the header row.

Aggregate bits per byte by corpus and compressor
corpus mothergod gzip -9 zstd -19 xz -9e
Canterbury 1.366 2.081 1.470 1.403
Silesia 2.052 2.553 1.997 1.829

Read it honestly: mothergod beats both zstd -19 and xz -9e in aggregate on Canterbury, and loses to both in aggregate on Silesia. Per file, against whichever of the two is stronger on that file, it wins 6 of 11 on Canterbury and 1 of 12 on Silesia. Closing Silesia is the current milestone. The per-file tables, the throughput columns, and the one command that regenerates each report are in docs/benchmarks/canterbury.md and silesia.md.

Speed

Slow, and that is the honest headline. mothergod's codec has no internal parallelism; the reports measure it one thread per file, several files at once, on one CI machine (AMD EPYC 7763, 4 logical cores), more threads than cores, so the rate is a lower bound on what an isolated single-file run would show.

Aggregate encode and decode throughput by corpus, same run as the Measured table
corpus encode MB/s decode MB/s
Canterbury 0.126 3.016
Silesia 0.055 1.215

Silesia's 212 MB took about an hour to compress and about three minutes to read back. Encoding carries the optimal parse, so it runs roughly twenty times slower than decoding.

One caveat the reports state themselves: these are single-run figures from one machine, not a cross-machine claim. One more worth adding: an aggregate is total bytes over total time, so it is not a typical file. Per-file decode rates in those reports run from 0.2 to 24.1 MB/s. Speed is measured on every benchmark run and becomes a target of its own at the speed-tiers milestone on the roadmap; nothing before then is tuned for it.

Try it

There is no download: no release, no tag, no package. Building it yourself is the only way to run it, and the core crate has zero runtime dependencies, so this compiles exactly one crate.

git clone https://github.com/bugabinga/mothergod
cd mothergod
cargo build --release --bin mothergod

./target/release/mothergod compress < FILE > FILE.mgdc
./target/release/mothergod decompress < FILE.mgdc > FILE.out
cmp FILE FILE.out

The CLI has the shape gzip -c and zstd -c already taught you: two subcommands, stdin to stdout when you name no file. That last line is the point of running it at all, and cmp printing nothing is it passing. Given a file argument instead, compress writes <file>.mgdc and decompress reads one back to the name with the suffix stripped; neither deletes its input, and neither overwrites an existing file. mothergod --help is the entire interface: no flags, no compression levels, no tuning knobs. Expect the encode to take a while, at the rates above.

Who builds it

Day-to-day development is done by Claude agents running on GitHub Actions: triage, implementation, adversarial code review, research, releases. Slowly, in public, like a real team would. A human operator holds the veto and the keys. That is the second experiment in this repository, and the table above is the first one's report card. How it works: agents/GOVERNANCE.md, and the seats are listed on the agents page.

Principles

Follow along