MethScope

cli latest ↗

Ultra-fast analysis of sparse DNA methylomes via MRMP (Most Recurrent Methylation Pattern) encoding — cell-type annotation, deconvolution, and genome-wide CpG upscaling.

Install — the binary, and what it can fetch

Every model is one self-contained bundle carrying its own MRMP feature definition (and labels), so a query .cg runs directly — no separate annotation files.

# conda — installs the `methscope` binary, and `yame` with it: methscope
# declares yame as a runtime dependency, so one install covers both. It is a
# separate package because methscope bundles YAME as a LIBRARY and ships no
# `yame` binary of its own, and the Upscale and MRMP examples below call that
# binary to index, subset and view .cg stores. (Models and the CpG coordinate
# reference come from `methscope fetch`; yame is for the stores themselves.)
conda install -c zhou-lab -c conda-forge methscope

# optional, linux-64 only: adds a `methscope-cuda` binary with the CUDA
# backend for `upscale-train`. Everything else — upscale, classify,
# deconv — is pure C and needs no GPU, so most users can skip this.
conda install -c zhou-lab -c conda-forge methscope-cuda

# or build from source (bundles YAME). libxgboost is the one external
# dependency: the Makefile reads $CONDA_PREFIX, so activate the env first
# or pass XGB_PREFIX=/path/to/env.
conda create -n methscope -c conda-forge libxgboost
conda activate methscope
git clone --recurse-submodules https://github.com/zhou-lab/methscope-cli
cd methscope-cli && make          # add CUDA=1 for the training backend
# the source build links libyame.a and emits no `yame` binary either; for the
# examples that call it, `make -C YAME` builds one at YAME/yame.

Fetching data & models — digest-checked, container-safe

Data and models are fetched with methscope fetch; each workflow below opens by fetching just the file it needs into the working directory (methscope fetch -c, digest-checked). It never prompts, so it is safe inside a container build. The catalogue and the model tag are compiled into this binary (methscope --version prints both), so what a release documents is exactly what it fetches; no other tool needs to be installed.

A name is a store path — what the browser shows and where the file lands, the same string in both places. A directory takes that directory's own files, so hg38 is the genome annotation and not the models beneath it; name those explicitly — or reach them all at once with -R.

formwhat it does
methscope fetch browse the catalogue; piped or redirected it dumps TSV instead, so a script never blocks on a keystroke
methscope fetch -l the registry as TSV, with what the store already has
methscope fetch -n hg38/models what that would fetch, and stop — exits 0, so a script can check first
methscope fetch -y hg38/models the whole directory, unattended. A directory is confirmed first because a short name can reach a lot: on a terminal, fetch hg38/models opens the browser on that folder with the planned files checked (f fetches, q leaves); off one, -y is that confirmation
methscope fetch hg38/models/hg38_sex.clfx one file, no -y needed — naming the file IS the confirmation, so a documented fetch line runs in a script as it stands
methscope fetch -y -g celltype hg38/models only the files in that directory matching every term
methscope fetch -y -R hg38 that directory and everything beneath it — hg38, hg38/data and hg38/models together, 16 files. Without -R a bare hg38 is the one coordinate file

Atomic downloads. Downloads land on a .part sibling and are renamed only once the bytes verify, so an interrupted methscope fetch can never leave a short file that later reads as a valid model. Without -c files go to a shared store ($METHSCOPE_DATA_HOME, else $YAME_DATA_HOME, else ~/.local/share/yame) that other zhou-lab tools read too.

On disk — MRMPIDX1

hg38 · 35 samples · 29,401,795 CpGs · 2,345,190 patterns · 155 MB

header       136 B   reference, sample names, binstring params, checksum
patterns  16 B × 2,345,190   base-3 key + count, ranked
             11111111111111111111111111111111111  10,736,116  P1
             11111111111011111111111111111111111   2,859,193  P2
             00000000000000000000000000000000000   2,214,235  P3
membership 4 B × 29,401,795  rank per CpG
             CpG 0 → 0  P1    CpG 7 → 4294967295  PNA
MRMPIDX1 — every candidate ranked, so the top-K cut belongs to the consumer: mrmp-build --top N for a single set (default 10,000, the N patterns with the most CpGs, the rest folded into PNA), mrmp-pool --pooled-top N when sets from a chain must compete for one budget. Ranking is deterministic, so the build is byte-reproducible from the reference alone.