crossbind
GitHub

Zstandard

v1.5.7Compression

Zstandard 1.5.7, zstandard compression, packaged by crossbind as @crossbind/port-zstd and one package per target. Only a variant that is actually on npm beta is listed as published.

npm install @crossbind/port-zstd-wasm@beta
LIVE · 3 APPS · RUNS IN THIS TAB

zstd in your browser, including what no browser API does

Browsers decode zstd on the network, but none gives JavaScript a zstd compressor by default. These apps run the upstream C library, compiled by crossbind: they train dictionaries, send an update as its difference from the last version, and stream .tar.zst archives. The first run downloads 0.9 MB of WebAssembly once; every app on this page shares it, and nothing is uploaded.

APP 01

Teach zstd your data, then watch small messages shrink

A single JSON record barely compresses: there is too little of it to learn from. Train a dictionary on a thousand records and every new record compresses about four times further. The training runs here, in this tab; none of the JavaScript zstd packages we checked can train one.

The records look like a public API's user objects and are generated in the module, so the numbers are the same on every machine.

Train a dictionary to compare zstd with and without it, record by record.
SHOW THE CODE
src/native/dictionary_lab.h
// src/native/dictionary_lab.h (excerpt)
int train(int capacity) {
std::string joined;
std::vector<size_t> sizes;
for (const std::string& record : training) {
joined += record;
sizes.push_back(record.size());
}
std::string trained(static_cast<size_t>(capacity), '\0');
const size_t size = ZDICT_trainFromBuffer(&trained[0], trained.size(), joined.data(),
sizes.data(), static_cast<unsigned>(sizes.size()));
if (ZDICT_isError(size)) throw std::runtime_error(ZDICT_getErrorName(size));
trained.resize(size);
dictionary = trained;
return static_cast<int>(size);
}
main.js
const m = await initNative();
const lab = await new m.DictionaryLab(1000); // learn from 1,000 records
await lab.train(4096); // a 4 KB dictionary
 
const result = JSON.parse(await lab.evaluate(3, false));
// 1,000 unseen records, one frame each:
// result.plain 294,389 B (1.82x)
// result.withDictionary 70,823 B (7.55x)
APP 02

Send a 1 MB update in about 1 KB

Version two of a 1 MB dataset changes a few numbers and adds twenty records. zstd compresses it with version one as its dictionary, and the page wraps the result in RFC 9842 dcz framing: the Compression Dictionary Transport format Chrome decodes natively over HTTP.

Both versions are generated in the module: v1 is 2,000 records, and v2 raises every fiftieth follower count and appends 20 records.

Build the update to compare gzip, zstd, and zstd with the previous version as its dictionary.
SHOW THE CODE
src/native/delta_lab.h
// src/native/delta_lab.h (excerpt)
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_compressionLevel, level));
if (windowLog) check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_windowLog, windowLog));
// the previous version is the dictionary
if (withPrefix) check(ZSTD_CCtx_refPrefix(cctx.get(), previousText.data(), previousText.size()));
out.resize(check(ZSTD_compress2(cctx.get(), &out[0], out.size(), nextText.data(), nextText.size())));
 
// RFC 9842 dcz: a 40-byte skippable frame naming v1 by its SHA-256, then the frame
header.resize(check(ZSTD_writeSkippableFrame(&header[0], header.size(),
digest.data(), digest.size(), 0xE)));
main.js
const m = await initNative();
const lab = await new m.DeltaLab();
const v1 = Uint8Array.from(await lab.v1(), (c) => c.charCodeAt(0));
const sha = new Uint8Array(await crypto.subtle.digest('SHA-256', v1));
 
const dcz = await lab.dcz(String.fromCharCode(...sha), 19, false);
dcz.length; // 1022: 40 bytes of header, 982 of zstd
await lab.verify(dcz); // true: decodes back to v2 byte for byte
APP 03

Look inside a .tar.zst without unpacking it

Open a Linux package, a backup or a dataset archive and list what is inside. zstd streams the file in 128 KB steps while the tar headers are read as they pass, so nothing is unpacked, and windows up to 1 GiB decode, far past the 32 MB that fzstd, the popular pure-JavaScript decoder, supports.

The sample is a three-file .tar.zst the module writes itself. Your own file stays in this tab: it is mounted into the module's in-memory filesystem, never uploaded.

Open the sample, or a .zst or .tar.zst of your own, to see its frame header and what is inside.
SHOW THE CODE
src/support/tar.h
// src/support/tar.h (excerpt): stream the file through zstd
// `zstd --long=30` frames need a 1 GiB window, the most wasm32 can address.
check(ZSTD_DCtx_setParameter(dctx.get(), ZSTD_d_windowLogMax, 30));
std::vector<char> in(ZSTD_DStreamInSize());
std::vector<char> out(ZSTD_DStreamOutSize());
while ((count = std::fread(in.data(), 1, in.size(), file.get())) > 0) {
ZSTD_inBuffer input = {in.data(), count, 0};
while (input.pos < input.size) {
ZSTD_outBuffer output = {out.data(), out.size(), 0};
pending = check(ZSTD_decompressStream(dctx.get(), &output, &input));
if (output.pos && !sink(out.data(), output.pos)) return;
}
}
main.js
const m = await initNative();
const [path] = await m.autoMountFiles([file], await m.getRandomPath('/memfs')); // the dropped file
 
const header = JSON.parse(await m.ZstOpener.frameInfo(path));
const listing = JSON.parse(await m.ZstOpener.listTar(path));
// listing.entries: [{ name, size, mtime, type }, ...]
const text = await m.ZstOpener.extractText(path, listing.entries[0].name, 2048);

Usage

The calls most Zstandard code makes, each a small C++ header crossbind binds and the JavaScript that uses it. Every example runs here in WebAssembly and prints what the site build checked; the same headers and calls work on Android and iOS.

Each example also has a JavaScript only tab: the same task with no C++ file, calling Zstandard's own headers from @crossbind/port-zstd directly. 3 of 4 work that way; the other says what stops it.

Compress and decompress a buffer

The two calls most zstd code makes: ZSTD_compress, and ZSTD_decompress with the original size read back from the frame.

src/native/zstd_codec.h
#pragma once
 
#include <zstd.h>
 
#include <stdexcept>
#include <string>
 
// One-shot Zstandard. Bytes cross the binding as a byte string: one UTF-16 code unit (0-255) per byte.
class Zstd {
public:
static std::string version() { return ZSTD_versionString(); }
 
static std::u16string compress(const std::string& text, int level) {
std::string out(ZSTD_compressBound(text.size()), '\0');
const size_t size = ZSTD_compress(&out[0], out.size(), text.data(), text.size(), level);
if (ZSTD_isError(size)) throw std::runtime_error(ZSTD_getErrorName(size));
std::u16string bytes(size, u'\0');
for (size_t i = 0; i < size; ++i) bytes[i] = static_cast<unsigned char>(out[i]);
return bytes;
}
 
static std::string decompress(const std::u16string& bytes) {
std::string in(bytes.size(), '\0');
for (size_t i = 0; i < bytes.size(); ++i) {
if (bytes[i] > 0xFF) throw std::invalid_argument("not a byte string");
in[i] = static_cast<char>(bytes[i]);
}
const unsigned long long size = ZSTD_getFrameContentSize(in.data(), in.size());
if (size == ZSTD_CONTENTSIZE_ERROR) throw std::runtime_error("not a zstd frame");
if (size == ZSTD_CONTENTSIZE_UNKNOWN) throw std::runtime_error("size not stored in the frame; use streaming");
if (size > (256u << 20)) throw std::runtime_error("refusing to allocate more than 256 MiB");
std::string out(static_cast<size_t>(size), '\0');
const size_t got = ZSTD_decompress(&out[0], out.size(), in.data(), in.size());
if (ZSTD_isError(got)) throw std::runtime_error(ZSTD_getErrorName(got));
out.resize(got);
return out;
}
};
main.js
import { initNative, Zstd } from './native/zstd_codec.h';
 
await initNative();
const text = 'crossbind '.repeat(100);
const frame = await Zstd.compress(text, 19);
const bytes = Uint8Array.from(frame, (c) => c.charCodeAt(0));
console.log(await Zstd.version(), bytes.length, [...bytes.slice(0, 4)].map((b) => b.toString(16)).join(' '));
console.log((await Zstd.decompress(frame)) === text);
PRINTSfirst run downloads 0.9 MB
1.5.7 27 28 b5 2f fd
true

Stream a file through zstd

For data you should not hold in one buffer: ZSTD_compressStream2 and ZSTD_decompressStream work file to file in 128 KB steps.

src/native/zstd_stream.h
#pragma once
 
#include <zstd.h>
 
#include <cstdio>
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>
 
// Streaming Zstandard, file to file. Memory stays at two small buffers whatever the file size,
// the pattern of zstd's own examples/streaming_compression.c.
class ZstdStream {
public:
// Compresses `input` into `output` with a content checksum; returns the compressed size.
static double compressFile(const std::string& input, const std::string& output, int level) {
File in = open(input, "rb");
File out = open(output, "wb");
std::unique_ptr<ZSTD_CCtx, size_t (*)(ZSTD_CCtx*)> cctx(ZSTD_createCCtx(), ZSTD_freeCCtx);
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_compressionLevel, level));
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_checksumFlag, 1));
std::vector<char> inBuffer(ZSTD_CStreamInSize());
std::vector<char> outBuffer(ZSTD_CStreamOutSize());
double written = 0;
for (;;) {
const size_t read = std::fread(inBuffer.data(), 1, inBuffer.size(), in.get());
const bool last = read < inBuffer.size();
ZSTD_inBuffer source = {inBuffer.data(), read, 0};
bool finished = false;
while (!finished) {
ZSTD_outBuffer target = {outBuffer.data(), outBuffer.size(), 0};
const size_t remaining = check(ZSTD_compressStream2(cctx.get(), &target, &source, last ? ZSTD_e_end : ZSTD_e_continue));
write(out.get(), outBuffer.data(), target.pos);
written += static_cast<double>(target.pos);
finished = last ? remaining == 0 : source.pos == source.size;
}
if (last) return written;
}
}
 
// Decompresses `input` into `output`; returns the decompressed size.
static double decompressFile(const std::string& input, const std::string& output) {
File in = open(input, "rb");
File out = open(output, "wb");
std::unique_ptr<ZSTD_DCtx, size_t (*)(ZSTD_DCtx*)> dctx(ZSTD_createDCtx(), ZSTD_freeDCtx);
std::vector<char> inBuffer(ZSTD_DStreamInSize());
std::vector<char> outBuffer(ZSTD_DStreamOutSize());
double written = 0;
size_t pending = 0;
size_t read = 0;
while ((read = std::fread(inBuffer.data(), 1, inBuffer.size(), in.get())) > 0) {
ZSTD_inBuffer source = {inBuffer.data(), read, 0};
while (source.pos < source.size) {
ZSTD_outBuffer target = {outBuffer.data(), outBuffer.size(), 0};
pending = check(ZSTD_decompressStream(dctx.get(), &target, &source));
write(out.get(), outBuffer.data(), target.pos);
written += static_cast<double>(target.pos);
}
}
if (pending != 0) throw std::runtime_error("the input ends in the middle of a zstd frame");
return written;
}
 
private:
using File = std::unique_ptr<FILE, int (*)(FILE*)>;
 
static File open(const std::string& path, const char* mode) {
File file(std::fopen(path.c_str(), mode), std::fclose);
if (!file) throw std::runtime_error("cannot open " + path);
return file;
}
 
static void write(FILE* file, const char* data, size_t size) {
if (size && std::fwrite(data, 1, size, file) != size) throw std::runtime_error("write failed");
}
 
static size_t check(size_t code) {
if (ZSTD_isError(code)) throw std::runtime_error(ZSTD_getErrorName(code));
return code;
}
};
main.js
import { initNative } from './native/zstd_stream.h';
 
const m = await initNative();
const { ZstdStream } = m;
let seed = 42;
const random = (n) => (seed = (seed * 48271) % 2147483647) % n;
const lines = Array.from({ length: 50000 }, (_, i) => `2026-09-24T12:00:${String(i % 60).padStart(2, '0')}Z GET /api/items/${random(9000)} ${random(10) ? 200 : 404} ${random(900)}ms`);
// m.FS.writeFile adds to a file that already exists, so every run gets a fresh directory.
const dir = await m.getRandomPath('/memfs');
await m.FS.writeFile(`${dir}/access.log`, lines.join('\n'));
 
const packed = await ZstdStream.compressFile(`${dir}/access.log`, `${dir}/access.log.zst`, 3);
const unpacked = await ZstdStream.decompressFile(`${dir}/access.log.zst`, `${dir}/access.copy.log`);
const original = await m.getFileBytes(`${dir}/access.log`);
const copy = await m.getFileBytes(`${dir}/access.copy.log`);
console.log(`${original.length} B -> ${packed} B -> ${unpacked} B`);
console.log(copy.length === original.length && copy.every((byte, i) => byte === original[i]));
PRINTSfirst run downloads 0.9 MB
2537578 B -> 378699 B -> 2537578 B
true

Choose a level, a checksum and a window

A context set up with ZSTD_CCtx_setParameter decides what every frame it writes looks like; ZSTD_getFrameHeader reads the choices back.

src/native/zstd_frame.h
#pragma once
 
// ZSTD_getFrameHeader is in zstd's static API; safe with this statically linked, pinned library.
#define ZSTD_STATIC_LINKING_ONLY
#include <zstd.h>
 
#include <memory>
#include <stdexcept>
#include <string>
 
// A compression context with explicit parameters, and a reader for what they put in the frame header.
class ZstdFrame {
public:
// A windowLog of 0 keeps the level's default window.
static std::u16string compress(const std::string& text, int level, bool checksum, int windowLog) {
std::unique_ptr<ZSTD_CCtx, size_t (*)(ZSTD_CCtx*)> cctx(ZSTD_createCCtx(), ZSTD_freeCCtx);
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_compressionLevel, level));
check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_checksumFlag, checksum ? 1 : 0));
if (windowLog) check(ZSTD_CCtx_setParameter(cctx.get(), ZSTD_c_windowLog, windowLog));
std::string out(ZSTD_compressBound(text.size()), '\0');
out.resize(check(ZSTD_compress2(cctx.get(), &out[0], out.size(), text.data(), text.size())));
std::u16string bytes(out.size(), u'\0');
for (size_t i = 0; i < out.size(); ++i) bytes[i] = static_cast<unsigned char>(out[i]);
return bytes;
}
 
static std::string header(const std::u16string& frame) {
std::string head;
for (size_t i = 0; i < frame.size() && i < ZSTD_FRAMEHEADERSIZE_MAX; ++i) head += static_cast<char>(frame[i]);
ZSTD_FrameHeader info;
const size_t status = ZSTD_getFrameHeader(&info, head.data(), head.size());
if (ZSTD_isError(status) || status > 0) throw std::runtime_error("not a zstd frame header");
const std::string content = info.frameContentSize == ZSTD_CONTENTSIZE_UNKNOWN ? "not stored" : std::to_string(info.frameContentSize) + " B";
return "content " + content + ", window " + std::to_string(info.windowSize) + " B, checksum " + (info.checksumFlag ? "yes" : "no");
}
 
private:
static size_t check(size_t code) {
if (ZSTD_isError(code)) throw std::runtime_error(ZSTD_getErrorName(code));
return code;
}
};
main.js
import { initNative, ZstdFrame } from './native/zstd_frame.h';
 
await initNative();
let seed = 7;
const random = (n) => (seed = (seed * 48271) % 2147483647) % n;
const text = Array.from({ length: 400 }, (_, i) => `{"id":${i},"user":"user${random(5000)}","score":${random(1000)}}`).join('\n');
for (const level of [3, 9, 19]) {
const frame = await ZstdFrame.compress(text, level, false, 0);
console.log(`level ${level}: ${text.length} B -> ${frame.length} B`);
}
const small = await ZstdFrame.compress(text, 19, true, 10);
console.log(`level 19, 1 KiB window, checksum: ${small.length} B`);
console.log(await ZstdFrame.header(small));
PRINTSfirst run downloads 0.9 MB
level 3: 16160 B -> 3065 B
level 9: 16160 B -> 2749 B
level 19: 16160 B -> 2219 B
level 19, 1 KiB window, checksum: 2469 B
content 16160 B, window 1024 B, checksum yes

Compress small messages with a dictionary

Train once with ZDICT_trainFromBuffer, digest it once with ZSTD_createCDict and ZSTD_createDDict, then compress every message against it.

src/native/zstd_dictionary.h
#pragma once
 
#include <zdict.h>
#include <zstd.h>
 
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>
 
// Dictionary compression for small messages: train once on samples, prepare the dictionary once,
// then compress and decompress each message against it.
class ZstdDictionary {
public:
// Trains on newline-separated samples; returns at most `capacity` bytes of dictionary.
static std::u16string train(const std::string& samples, int capacity) {
std::string joined;
std::vector<size_t> sizes;
for (size_t start = 0; start < samples.size();) {
size_t end = samples.find('\n', start);
if (end == std::string::npos) end = samples.size();
joined.append(samples, start, end - start);
sizes.push_back(end - start);
start = end + 1;
}
std::string dictionary(static_cast<size_t>(capacity), '\0');
const size_t size = ZDICT_trainFromBuffer(&dictionary[0], dictionary.size(), joined.data(), sizes.data(), static_cast<unsigned>(sizes.size()));
if (ZDICT_isError(size)) throw std::runtime_error(ZDICT_getErrorName(size));
dictionary.resize(size);
return toUnits(dictionary);
}
 
// Digests the dictionary once for compression and once for decompression.
ZstdDictionary(const std::u16string& dictionary, int level)
: compressDictionary(nullptr, ZSTD_freeCDict), decompressDictionary(nullptr, ZSTD_freeDDict) {
const std::string bytes = fromUnits(dictionary);
compressDictionary.reset(ZSTD_createCDict(bytes.data(), bytes.size(), level));
decompressDictionary.reset(ZSTD_createDDict(bytes.data(), bytes.size()));
if (!compressDictionary || !decompressDictionary) throw std::runtime_error("not a usable dictionary");
}
 
std::u16string compress(const std::string& message) const {
std::unique_ptr<ZSTD_CCtx, size_t (*)(ZSTD_CCtx*)> cctx(ZSTD_createCCtx(), ZSTD_freeCCtx);
std::string out(ZSTD_compressBound(message.size()), '\0');
out.resize(check(ZSTD_compress_usingCDict(cctx.get(), &out[0], out.size(), message.data(), message.size(), compressDictionary.get())));
return toUnits(out);
}
 
std::string decompress(const std::u16string& frame) const {
const std::string in = fromUnits(frame);
const unsigned long long size = ZSTD_getFrameContentSize(in.data(), in.size());
if (size == ZSTD_CONTENTSIZE_ERROR || size == ZSTD_CONTENTSIZE_UNKNOWN) throw std::runtime_error("not a zstd frame with its size");
if (size > (1u << 20)) throw std::runtime_error("a message above 1 MiB is not a small message");
std::unique_ptr<ZSTD_DCtx, size_t (*)(ZSTD_DCtx*)> dctx(ZSTD_createDCtx(), ZSTD_freeDCtx);
std::string out(static_cast<size_t>(size), '\0');
out.resize(check(ZSTD_decompress_usingDDict(dctx.get(), &out[0], out.size(), in.data(), in.size(), decompressDictionary.get())));
return out;
}
 
private:
static size_t check(size_t code) {
if (ZSTD_isError(code)) throw std::runtime_error(ZSTD_getErrorName(code));
return code;
}
 
static std::u16string toUnits(const std::string& data) {
std::u16string units(data.size(), u'\0');
for (size_t i = 0; i < data.size(); ++i) units[i] = static_cast<unsigned char>(data[i]);
return units;
}
 
static std::string fromUnits(const std::u16string& units) {
std::string data(units.size(), '\0');
for (size_t i = 0; i < units.size(); ++i) {
if (units[i] > 0xFF) throw std::invalid_argument("not a byte string");
data[i] = static_cast<char>(units[i]);
}
return data;
}
 
std::unique_ptr<ZSTD_CDict, size_t (*)(ZSTD_CDict*)> compressDictionary;
std::unique_ptr<ZSTD_DDict, size_t (*)(ZSTD_DDict*)> decompressDictionary;
};
main.js
import { initNative, ZstdDictionary } from './native/zstd_dictionary.h';
import { Zstd } from './native/zstd_codec.h';
 
await initNative();
const event = (i) => `{"event":"click","user":${1000 + ((i * 37) % 900)},"page":"/products/${i % 12}","ms":${(i * 7919) % 400}}`;
const samples = Array.from({ length: 4000 }, (_, i) => event(i)).join('\n');
const dictionary = await ZstdDictionary.train(samples, 2048);
const codec = await new ZstdDictionary(dictionary, 3);
 
const message = event(4321);
const alone = await Zstd.compress(message, 3); // the one-shot wrapper from the first example
const frame = await codec.compress(message);
console.log(`dictionary: ${dictionary.length} B`);
console.log(`${message.length} B message: ${alone.length} B alone, ${frame.length} B with the dictionary`);
console.log((await codec.decompress(frame)) === message);
PRINTSfirst run downloads 0.9 MB
dictionary: 2048 B
59 B message: 68 B alone, 29 B with the dictionary
true

Add it to your project

One package per platform: install the ones you build for and list each in crossbind.config.js; crossbind compiles only the one that matches the build target. Your C++ goes in src/native, next to the headers it binds. Libraries explains the whole flow.

shell
npm install @crossbind/port-zstd-wasm@beta
crossbind.config.js
import zstdWasm from '@crossbind/port-zstd-wasm/crossbind.config.js';
 
export default {
dependencies: [zstdWasm],
paths: { config: import.meta.url },
};

Platforms

PlatformRuns inBuildsPage
WebAssemblybrowsers, Node.js and edge runtimeswasm32, single-threaded and multi-threadedZstandard for WebAssembly
AndroidReact Native apps on Androidarm64-v8a devices and the x86_64 emulatorZstandard for Android
iOSReact Native apps on iOSarm64 devices and simulatorsZstandard for iOS
macOSnative Node.js addons and Electron on macOSarm64 and x64, macOS 11 or laterZstandard for macOS
Linuxnative Node.js addons on Linuxx64 and arm64, glibc 2.28 or laterZstandard for Linux
Windowsnative Node.js addons on Windowsx64 and arm64, Windows 10 or laterZstandard for Windows
WASIcommand-line programs under wasmtimewasm32-wasip3, single-threadedZstandard for WASI

Packages

TargetPackagenpm `beta`
Meta package@crossbind/port-zstd2.0.0-beta.62
Web and Node.js@crossbind/port-zstd-wasm2.0.0-beta.62
WASI library@crossbind/port-zstd-wasi2.0.0-beta.62
WASI commands@crossbind/port-zstd-standalone-wasinot published
Android@crossbind/port-zstd-android2.0.0-beta.62
iOS@crossbind/port-zstd-ios2.0.0-beta.62
macOS@crossbind/port-zstd-darwin2.0.0-beta.62
Linux@crossbind/port-zstd-linux2.0.0-beta.62
Linux (musl)@crossbind/port-zstd-linuxmuslnot published
Windows@crossbind/port-zstd-win322.0.0-beta.62

Licence

  • npm license field of @crossbind/port-zstd: BSD-3-Clause.
  • Upstream declares BSD-3-Clause OR GPL-2.0-only; the npm field normalises it to SPDX.
  • The licence files that ship with the package, and the port recipe, are in the port directory.

Facts on this page come from the port manifests in the repository and from what npm served on beta when the site was built. See the Libraries guide for the full consumer flow.

MORE LIBRARIES
cURLExpatGDALGEOSGeoTIFFiconvLERClibjpeg-turbolibTIFFOpenSSLPROJSpatiaLiteSQLiteWebPzlib
Type to search every guide page and section.
↑↓ navigate↵ openesc close