Ultra-fast analysis of sparse DNA methylomes via MRMP (Most Recurrent Methylation Pattern) encoding — cell-type annotation, deconvolution, and genome-wide CpG upscaling.
GitHub· models· For coding agents· hg38 · mm10 · WGBS · single-cell · spatial
Every model is one self-contained bundle carrying its own MRMP feature definition (and
labels), so a query .cg runs directly — no separate annotation
files.
# conda — installs the `methscope` binary, and `yame` with it: methscope
# declares yame as a runtime dependency, so one install covers both. It is a
# separate package because methscope bundles YAME as a LIBRARY and ships no
# `yame` binary of its own, and the Upscale and MRMP examples below call that
# binary to index, subset and view .cg stores. (Models and the CpG coordinate
# reference come from `methscope fetch`; yame is for the stores themselves.)
conda install -c zhou-lab -c conda-forge methscope
# optional, linux-64 only: adds a `methscope-cuda` binary with the CUDA
# backend for `upscale-train`. Everything else — upscale, classify,
# deconv — is pure C and needs no GPU, so most users can skip this.
conda install -c zhou-lab -c conda-forge methscope-cuda
# or build from source (bundles YAME). libxgboost is the one external
# dependency: the Makefile reads $CONDA_PREFIX, so activate the env first
# or pass XGB_PREFIX=/path/to/env.
conda create -n methscope -c conda-forge libxgboost
conda activate methscope
git clone --recurse-submodules https://github.com/zhou-lab/methscope-cli
cd methscope-cli && make # add CUDA=1 for the training backend
# the source build links libyame.a and emits no `yame` binary either; for the
# examples that call it, `make -C YAME` builds one at YAME/yame.
Data and models are fetched with methscope fetch; each workflow below
opens by fetching just the file it needs into the working directory
(methscope fetch -c, digest-checked). It never prompts, so it is safe
inside a container build. The catalogue and the model tag are compiled into this binary
(methscope --version prints both), so what a release documents is
exactly what it fetches; no other tool needs to be installed.
A name is a store path — what the browser shows and where the file lands, the
same string in both places. A directory takes that directory's own files, so
hg38 is the genome annotation and not the models beneath it; name
those explicitly — or reach them all at once with -R.
| form | what it does |
|---|---|
methscope fetch |
browse the catalogue; piped or redirected it dumps TSV instead, so a script never blocks on a keystroke |
methscope fetch -l |
the registry as TSV, with what the store already has |
methscope fetch -n hg38/models |
what that would fetch, and stop — exits 0, so a script can check first |
methscope fetch -y hg38/models |
the whole directory, unattended. A directory is confirmed first because a short name
can reach a lot: on a terminal, fetch hg38/models opens the
browser on that folder with the planned files checked (f fetches,
q leaves); off one, -y is that confirmation |
methscope fetch hg38/models/hg38_sex.clfx |
one file, no -y needed — naming the file IS the
confirmation, so a documented fetch line runs in a script as it stands |
methscope fetch -y -g celltype hg38/models |
only the files in that directory matching every term |
methscope fetch -y -R hg38 |
that directory and everything beneath it —
hg38, hg38/data and
hg38/models together, 16 files. Without
-R a bare hg38 is the one
coordinate file |
Atomic downloads. Downloads land on a .part
sibling and are renamed only once the bytes verify, so an interrupted
methscope fetch can never leave a short file that later reads as a
valid model. Without -c files go to a shared store
($METHSCOPE_DATA_HOME, else $YAME_DATA_HOME,
else ~/.local/share/yame) that other zhou-lab tools read too.
hg38 · 35 samples · 29,401,795 CpGs · 2,345,190 patterns · 155 MB header 136 B reference, sample names, binstring params, checksum patterns 16 B × 2,345,190 base-3 key + count, ranked 11111111111111111111111111111111111 10,736,116 P1 11111111111011111111111111111111111 2,859,193 P2 00000000000000000000000000000000000 2,214,235 P3 membership 4 B × 29,401,795 rank per CpG CpG 0 → 0 P1 CpG 7 → 4294967295 PNA
mrmp-build --top N for a single set (default 10,000, the
N patterns with the most CpGs, the rest folded into PNA), mrmp-pool
--pooled-top N when sets from a chain must compete for one budget. Ranking is
deterministic, so the build is byte-reproducible from the reference alone.An MRMP is one recurrent methylation pattern across a reference panel; the mask that
labels every genomic CpG with its pattern is what every model is built on.
mrmp-build constructs it natively from a discretized reference
.cg — no YAME + awk + sort pipeline, ~576 MB peak, ~50 s
for a whole-genome reference. The fetchable example below is chr20 only, so it finishes in
about a second. The default mode builds one set over every class, which is
what this demo, upscale, export and the violation rule use; --top
keeps the N patterns with the most CpGs and folds the rest into PNA (default 10,000). This
chr20 set has 261, so --top 500 cuts nothing here — it is the idiom
a whole-genome build uses.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# a 40-cell-type chr20 reference (39.5 MB)
methscope fetch -c hg38/data/human_hg38_40_celltypes_chr20.cg
methscope mrmp-build --top 500 human_hg38_40_celltypes_chr20.cg chr20_40celltypes.mrmp
mrmp-build: 40 classes, one MRMP over all of them (a tree of one level)
root 40 classes | 261 patterns | 3,787 CpGs
split threshold: none (flat) -> 1 group(s)
1 node(s), 16,792 bytes -> chr20_40celltypes.mrmp
40 samples is the ceiling: a pattern packs as a base-3 uint64 and 3^40 < 2^64 < 3^41. The whole-genome models use 35 Zhou pseudo-bulks.
What it is: dimensions, the binstring parameters, the PNA share.
methscope inspect chr20_40celltypes.mrmp
MRMP chr20_40celltypes.mrmp
format MRMPIDX1 v1, one set
name root
reference human_hg38_40_celltypes_chr20.cg
classes 40
patterns 261
covered 3,787 of 773,477 CpGs (0.49%)
PNA 769,690 CpGs (99.51%) match no pattern
block 16,792 bytes
membership 9,744 bytes, RLE + BGZF (318x vs dense)
thresholds present (per-pattern midpoints)
resolution --call-mindepth 1 --beta-threshold 0.500
--max-ambig-frac 1.000 --min-major-fold 10.000
checksum 8cf2026fa6de8d46 (over the per-CpG key stream)
--patterns adds the ranked head; --top sets
how many (default 20).
methscope inspect --patterns --top 5 chr20_40celltypes.mrmp
#pattern label count (top 5 of 261)
1111111101111111111111111111111111111111 P1 385
1111111111011111111111111111111111111111 P2 331
1111111110111111111111111111111111111111 P3 304
1111111111111101111111111111111111111111 P4 219
1111111111111111111111111111111111101111 P5 202
For a multi-set routing tree, --tree shows each MRMP
node and the classes it decides. It works on the tree bundle written by
mrmp-build.
methscope inspect --tree chr20_40celltypes.mrmp
TREE chr20_40celltypes.mrmp
1 node, 40 classes, 261 patterns over 3,787 CpGs
root 40 classes 261 patterns 3,787 CpGs
|-- GSM5652176_Adipocytes-Z000000T7
|-- GSM5652181_Saphenous-Vein-Endothel-Z000000RM
...
The runtime .cm mask is inspection / interop only, not a pipeline
step: upscale-train materializes it into the bundle from the
.mrmp itself.
methscope mrmp-export --top 1000 chr20_40celltypes.mrmp chr20_40celltypes.cm
what mrmp-export carries over, and what it leaves behind .mrmp .cm membership rank per CpG ──► state per CpG same RLE + BGZF non-top-K → Pna patterns base-3 key + count ✗ the label P7, never its rule thresholds float32 per pattern ✗ reference path + class names ✗ build params -c -b -m -M minseg ✗ PNA count pna_cpg ✗ Pna and background are one state checksum FNV-1a over keys ✗ set name inside the block ✗ moves out to the .cm.idx sidecar
.mrmp membership section
is a YAME format-2 RLE payload under BGZF, so the sizes track: on the chr20 set,
9,744 B of membership exports to a 9,729 B .cm; on the
178-set bank, 12,488,808 B becomes 12,390,627 B plus a 6,556 B
.idx. Folding the non-top-K patterns into the background is worth
about 4%. What a .cm cannot carry is the meaning of a column,
which is why classify --framework violation refuses an exported
mask: with no binstrings there is no rule to score. Read them out beside the mask with
mrmp-export --patterns.Ranking is deterministic (count desc, key asc), so a build is byte-reproducible from the
reference alone. The artifact is the build pipeline's currency:
upscale-featurize, upscale-set-units, and
upscale-train all read it, so the msur's group map and the mask a
model ships cannot drift apart.
Classify a methylome against a bundled classifier. Each step fetches the file it needs with
methscope fetch -c (digest-checked), so the block runs as-is on a clean
machine.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# 4 Loyfer cells vs the shipped 63-type bank-lite classifier
methscope fetch -c hg38/data/human_hg38_celltypes.cg
methscope fetch -c hg38/models/hg38_celltype_lite.clfx
methscope classify hg38_celltype_lite.clfx human_hg38_celltypes.cg
cell prediction_label confidence certainty
Oligodendrocyte ODC 0.999991 0.999971
Pancreas-Beta Beta 0.999998 0.999992
Blood-NK NK CD16 0.925882 0.935428
Blood-Monocytes Mono 0.949018 0.951083
Loyfer truth → Zhou2025 label: 4/4 concordant, and the 63-class bank names the
NK subtype. _lite is the smaller of the pair;
hg38_celltype_full.clfx is the same model with a resolver for every
class pair and the published accuracy — the
Bank classifiers table gives both sizes, generated from
the model rows so it cannot drift from the cards. The same call runs a sex /
age / any bundled model. Linear models store per-feature training means, so missing/sparse
betas are imputed at classify time — accurate even at <1% genome coverage.
Two scores, not one. confidence is P(called
class) — how likely the winner is. certainty is
1 - H(p)/log(K): the fraction of the initial ignorance the evidence
removed, 0 while the posterior is still flat over the K classes the node was deciding
between, 1 once it is settled. The two agree at the top, as all four rows above do, and
part company near the middle — a two-class call at P=0.6 reports confidence 0.600 and
certainty 0.029, because 1.96 of the 2 classes are still effectively in play. So read
certainty as how decisive the call was against its alternatives,
and threshold on it rather than on a posterior that has merely crossed 0.5.
methscope fetch -c hg38/models/hg38_sex.clfx
methscope inspect hg38_sex.clfx
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 30,931 bytes
bundle hg38_sex.clfx
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 30,653 YAME .cm mask, 3 states
1 kind 30,761 8 "logistic" framework mark
2 model 30,769 154 methscope-linear text spec
[0] mrmp MRMP feature definition (YAME .cm)
states 3
[1] kind bundle framework mark
value logistic
[2] model logistic linear classifier - run via `classify`
method logistic
labels Female, Male
bias -7.50339
scale 1
features 2
Xa_lo weight=-13.9125 mean=0.24253
Xa_hi weight=13.6301 mean=0.752643
hg38_celltype_full.clfx · 153.8 MB · 63 classes, one bank mrmp MRMP chain (MRMPIDX1) — the feature sets the nodes were fit on kind framework mark: tree | xgboost | logistic | threshold booster chain one xgboost booster per tree node, labels embedded
.clfx): one self-describing file
with the inner model, its MRMP feature definition, and the labels, so a query
.cg runs directly. The kind section records the framework, so
classify needs no flag to know how to run it. The shipped human model is a
tree: 13 sections, 11 modelled nodes over a 14-set chain, routing 33 classes
from the root. A flat single-booster model (xgboost) or a linear one
(logistic / threshold, which also stores per-feature means) has the
simpler shape — one model section instead of the chain, as in the
hg38_sex.clfx listing above.Three steps against a labelled reference store: build the bank (binstrings plus a
calibrated resolver per class pair), featurize a balanced training draw across the
coverage ladder, fit one pooled booster. These are the commands that built the shipped
hg38_celltype_full.clfx, copied from the driver that ran them
(labjournal/zhouw3/2026/20260904_bankfull_final.sh) rather than
simplified for the page — the flags are load-bearing and a reader who drops them gets a
different model. The whole difference between _lite and
_full is three flags: --min-pattern-cpgs 500
--resolvers 150 on the bank (the 150 class pairs with the thinnest hard-block
footprint get a resolver, instead of every pair) and -n 100 on
the booster. The driver writes bank.clfx; it is renamed when
staged.
## 20260904_bankfull_final.sh, SPECIES=human); needs the labelled reference store.
## Same recipe in CURRENT flags: that driver ran before 0.10 renamed them
## (--resolver-gate -1 -> --resolvers all, --pattern-floor -> --min-pattern-cpgs)
## and made the pooled booster the only mode (--pool-nodes is gone).
## The _lite footnotes are 20260914_lite_build.sh, which shipped the v10 lites.
# inputs: REF = 63 per-class pooled references, STORE = the 84,601 labelled cells,
# CLAB = cell -> label for every cell, TRIDS/TRLAB = a balanced training draw.
# The driver spreads the pairwise resolver calibration over 8 SLURM jobs with
# --resolver-stride k/8 into the same --resolver-cache; this is the one-job form.
LADDER=0,4096,16384,65536,262144; T=24; SEED=20260828
methscope mrmp-build --bank --resolvers all --min-pattern-cpgs 2000 \
--anneal-min-seg 10000,3000,1000 --max-depth 24 \
--resolver-cache $B/cache \
--cell-store $STORE --cell-labels $CLAB --calib-threads $T \
--force $REF $B/bank.mrmp # _lite: --min-pattern-cpgs 500 --resolvers 150
yame subset -l $TRIDS -o $B/train.cg $STORE
yame index -s $TRIDS $B/train.cg
methscope classify-featurize --threads $T --satellite-contrast replace \
--seed $SEED -b --sample $LADDER --reps 1 \
-l $TRLAB -o $B/train.msfm $B/train.cg $B/bank.mrmp
methscope classify-train --threads $T --data $B/train.msfm \
-o $B/bank.clfx # constrained defaults: max-depth 4, colsample 0.4
# _lite: add -n 100 (a 100-round booster)
Why it is not runnable here. It wants the 63-class pooled reference
and the 84,601 labelled single cells behind it, not a chr20 fixture — and the cost is the
resolver calibration: every class pair gets its own leave-one-cell-out grid over the cell
store, about 23 CPU-hours for the 1,953 human pairs, which the driver spreads across
SLURM jobs through the shared cache. $LADDER draws each training
cell at five coverages, which is what stops the booster reading sequencing depth as cell
identity; the booster itself trains on a balanced 400-cells-per-class draw, and its
accuracy is the ten-fold CV reported on the Published models tab, not a number measured
on the cells it trained on.
The same wrapping by hand: unpack a bundle into its parts, then re-wrap
(unbundle writes the inner model +
hg38_sex.mrmp; -k re-stamps the kind).
This round-trips the shipped sex model fetched above, and the result is
byte-identical to it. Flat and linear bundles round-trip this way; a
tree bundle carries a booster chain rather than one
model section, so unbundle does not take it.
methscope unbundle hg38_sex.clfx
methscope bundle -m hg38_sex.mrmp -k logistic -o rewrapped.clfx hg38_sex.clf
cmp hg38_sex.clfx rewrapped.clfx && echo identical
wrote inner model -> hg38_sex.clf
wrote bundled mrmp -> hg38_sex.mrmp
bundled hg38_sex.clf + hg38_sex.mrmp -> rewrapped.clfx
identical
Every model is positional over the genome's CpG rows and refuses a
query built on another row space. An array is another row space: one row
per probe, in the platform's ordering. mliftover
re-indexes the model onto it -- the booster, the weights and the
deconvolution trailer are copied untouched, only the CpG index moves -- so
array data is scored as it is, by a model built genome-wide. It takes a
.clfx, a loose .mrmp or
.cm, or a .msdref; it
refuses an upscaler, whose output is the genome (carry the data the other
way instead: sesame mliftover --to hg38 --simulated-depth 100).
Only cg probes are joined, decided by Probe_ID from
the ordering, never by where a probe maps: an rs
probe on a CpG site carries a genotype, not methylation. The per-set
summary says how much of each feature survives -- here 1,184 of 25,967 and
550 of 13,826 CpGs, enough to estimate the same means: on 168 MSA primary
tissues with reported sex this lifted model scores 97.0%, and the model
lifted here is byte-identical to one built by hand with
unbundle, sesame mliftover
and bundle -k.
# the platform's probe list and coordinates, and the genome's CpG rows
methscope fetch -c MSA/MSA.ordering.tsv.gz MSA/MSA.hg38.coord.tsv.gz hg38/cpg_nocontig.cr
methscope mliftover --to MSA --ordering MSA.ordering.tsv.gz \
--coord MSA.hg38.coord.tsv.gz --cr cpg_nocontig.cr \
hg38_sex.clfx -o hg38_sex.MSA.clfx
methscope inspect hg38_sex.MSA.clfx | head -3
mliftover: MSA: 274041 of 284309 probes sit on a CpG row of hg38 (29401795 rows); 9656 not lifted by type (not cg), 612 cg probes with no CpG row
mliftover: mask: 3 states over 284309 target rows; Pna 282575 Xa_lo 1184 Xa_hi 550
mliftover: hg38_sex.clfx -> hg38_sex.MSA.clfx (MSA row space, 284309 rows)
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 45,617 bytes
bundle hg38_sex.MSA.clfx
_full, _lite, and what they replacedA bank classifier does not route: its MRMP is a flat set of hierarchical LCA
binstrings plus redrawn, LOO-calibrated resolvers for close class pairs, all scored by
one pooled booster. _full and _lite are the
same build one flag apart — how many pairs get a resolver.
| sets | resolvers | size | |
|---|---|---|---|
hg38_celltype_full.clfx | 1,981 | 1,953 — every class pair | 153.8 MB |
hg38_celltype_lite.clfx | 178 | 150 — the thinnest-footprint pairs | 21.2 MB |
mm10_brain_full.clfx | 837 | 820 — every class pair | 64.5 MB |
mm10_brain_lite.clfx | 117 | 100 — the thinnest-footprint pairs | 13.3 MB |
Sizes are the catalogue's; sets and resolvers are read from each model's captured inspect output, the same block the cards below show, so the two cannot disagree.
Accuracy. Ten-fold CV of the recipe: human _full
class-balanced 0.9544 / micro 0.9569 over 84,601 cells; mouse 0.9746 / 0.9720 over 54,447.
The shipped mouse bytes were also scored on 49,113 Liu 2021 cells that never entered the
model — the cells beyond the 2,000-per-class training cap, so 18 of the 41 classes, the
large cortical ones included:
| micro | class-balanced (18) | |
|---|---|---|
mm10_brain_full.clfx | 0.9660 | 0.9707 |
mm10_brain_lite.clfx | 0.9482 | 0.9595 |
The lite model's cost sits in the cortical IT layers (IT-L23 0.946 → 0.924, IT-L5
0.951 → 0.915); elsewhere the two are within noise. Use _full when
the number matters, _lite when its size does, and validate on your
own data before relying on it.
What the banks replaced. Until v9 the classifiers were routing trees — a
tree of boosters over one shared MRMP chain, a query scored at the root and handed down to the
child that owns the call, with satellites lending columns for pairs that kept colliding.
A wrong turn high up could not be recovered lower down; a bank has one node and no routing
decision to get wrong, and covers 63 classes against the trees' 33. The trees
(hg38_celltype.clfx, mm10_celltype_brain.clfx)
are withdrawn at v10; earlier tags keep them for anything that cites them. Both kinds are
trained across a coverage ladder with reads binarised, so one model spans sparse single cells
and deep pseudobulks. A bank applies no sparsity gate: a record too thin to score is still
given a label, and as coverage falls the call converges on the model's prior instead of
coming back NA. certainty is what separates
the two — one Loyfer NK cell against hg38_celltype_lite.clfx
reads NK CD16 at certainty 0.988 with 20k covered CpGs, slips to
NK CD56 at 0.613 by 2k, is simply wrong at 1k
(B Mem, 0.199), and below ~50 returns one constant prior row
(B Mem, 0.239) that does not change when the query does. One cell
at one downsampling seed, so read the shape rather than the thresholds. Only the
violation rule gates on coverage, through
--min-patterns; see below.
The origin claim, on single-cell tumours. Bian 2018 colorectal cancer, scored with
hg38_celltype_full.clfx and split by site, because a colorectal cell
sitting in liver must still read as colon. The 63-class bank separates enterocyte from goblet,
so both the specific call and the lineage are shown:
| site | n | called colon enterocyte | GI lineage |
|---|---|---|---|
| primary tumour (PT) | 545 | 94.3% | 99.6% |
| lymph-node metastasis (LN) | 302 | 90.7% | 99.3% |
| liver metastasis (ML) | 135 | 88.1% | 99.3% |
| liver metastasis after chemotherapy (MP) | 112 | 92.0% | 100% |
| other metastasis (MO) | 12 | 91.7% | 100% |
| all tumour sites | 1,106 | 99.5% | |
| adjacent normal (NC) | 87 | 34.5% (goblet 60.9%) | 100% |
The five tumour cells off the GI lineage are three ileal enterocytes, one perivascular and
one keratinocyte-like call. Adjacent normal reads mostly goblet — a colon crypt, not a tumour.
Predictions: 20260904_coo_panel/bian_cells_final.tsv via
labjournal/zhouw3/2026/20260904_bian_coo_panels.sh.
--framework violation — classify with no trained modelNot a model but an option on classify. Every MRMP already carries a
binstring — the methylation state each reference class takes at that pattern — so a reference
is a per-class prediction already. The rule asks which class the query contradicts
least, and there is nothing to fit:
# The chr20 MRMP from the Build step above, scored against the very reference
# it was built from: same row space by construction, and every cell must come
# back as itself, so the answer is checkable without a held-out set.
methscope classify --framework violation \
chr20_40celltypes.mrmp human_hg38_40_celltypes_chr20.cg | head -4
cell prediction_label margin
GSM5652176_Adipocytes-Z000000T7 GSM5652176_Adipocytes-Z000000T7 0.050341
GSM5652181_Saphenous-Vein-Endothel-Z000000RM GSM5652181_Saphenous-Vein-Endothel-Z000000RM 0.023153
GSM5652187_Kidney-Tubular-Endothel-Z000000PX GSM5652187_Kidney-Tubular-Endothel-Z000000PX 0.028322
The .mrmp and the query must be built on the same
reference; a chr20 artifact against a whole-genome query is refused by name, with both CpG
counts. margin, not a confidence: the rule has no posterior.
E1 = patterns the class expects methylated; E0 = expects unmethylated
w_j = sqrt(n_cpg_j)
v1 = Σ_{j∈E1, observed} w_j·[β_j ≤ 0.5] / Σ_{j∈E1, observed} w_j
v0 = Σ_{j∈E0, observed} w_j·[β_j > 0.5] / Σ_{j∈E0, observed} w_j
predict argmin_class (v1 + v0), needing ≥20 observed patterns per side
Tune with --call-threshold, --pattern-weight,
--min-patterns and --top; nothing is fitted, so
changing one is a rescore, not a retrain. What it is for: you have a reference and no
training data — build an MRMP over it and you have a classifier the same minute, over any label
set. It is also the honest baseline for a trained model on the same features: whatever a
booster gains over it is what the training data bought.
confidence is the margin to the runner-up, not a
probability. It collapses toward zero when no class fits, which is how a query whose
true type is missing from the reference announces itself; records too sparse to clear
--min-patterns come back NA. Margins are not
comparable between references with different class counts.
ref.mrmp must be a single flat set. Do not point
it at a multi-set chain such as a bank with resolver blocks: the rule pools
every set it finds, and a chain's sets have different class memberships, so the binstrings are
not commensurable — it currently fails silently, with near-zero confidences and no
warning. Accuracy is coverage-dependent: on the 33-type reference it plateaued near 0.85
above ~130k covered CpGs against 0.163 below 8k. Do not use it below ~30k covered CpGs; for a
bulk tumour, deconvolve rather than classify.
Solve a mixture against a signature panel by non-negative least squares. The solver rebuilds its pattern set per query from the CpGs that query measured, so it is robust to extreme sparsity and works at cfDNA-scale coverage.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# nine known mixtures vs the shipped 62-type reference (1.5 GB)
methscope fetch -c hg38/data/human_hg38_immune_mixture.cg
methscope fetch -c hg38/models/hg38_62celltypes.msdref
methscope deconv hg38_62celltypes.msdref human_hg38_immune_mixture.cg
deconv: reference v3, 62 classes x 11798286 kept rows (1463 MB incl. depth), qfilter 0.30,0.70
deconv: confusion trailer present -- fold-held-out twin: Zhou single cells, folds 1-9 …
cell class fraction
mac70_mono30_2pow22 Macrophage 0.696336
mac70_mono30_2pow22 Mono 0.303664
…
cd4_70_cd8_30_2pow22 Tmem_CD4 0.498298
cd4_70_cd8_30_2pow22 Tmem_CD8 0.268361
…
The first mixture is 70/30 macrophage/monocyte and reads 0.70/0.30. The second is
70/30 CD4/CD8 T cells mixed from unsorted T cells, so the reference resolves each
into memory and naive, and the two Tmem rows shown are not the
whole answer: with the Tnaive rows below them, CD4 sums to 0.705
and CD8 to 0.295. Read a mixture against the reference's classes, not the labels it was
mixed under. Only the classes at or above --min-frac are listed, largest
first. --report gives the same thing to read rather than to
parse, and --wide the full 62-column matrix for a join.
methscope deconv --report hg38_62celltypes.msdref human_hg38_immune_mixture.cg
mac70_mono30_2pow22: Macrophage 69.6%; Mono 30.4%
mac70_mono30_2pow18: Macrophage 75.7%; Mono 23.5%; Tnaive_CD8 0.8%
mac70_mono30_2pow16: Macrophage 78.4%; Mono 20.9%; Delta 0.7%
cd4_70_cd8_30_2pow22: Tmem_CD4 49.8%; Tmem_CD8 26.8%; Tnaive_CD4 20.7%; Tnaive_CD8 2.6%
cd4_70_cd8_30_2pow18: Tmem_CD4 51.0%; Tmem_CD8 22.6%; Tnaive_CD4 20.1%; Tnaive_CD8 3.4%; B_Mem 2.9%
cd4_70_cd8_30_2pow16: Tmem_CD4 50.8%; Tmem_CD8 27.3%; Tnaive_CD4 21.3%; Macrophage 0.6%
tnaive_cd4_70_cd8_30_2pow22: Tnaive_CD4 69.9%; Tnaive_CD8 30.1%
tnaive_cd4_70_cd8_30_2pow18: Tnaive_CD4 73.6%; Tnaive_CD8 26.4%
tnaive_cd4_70_cd8_30_2pow16: Tnaive_CD4 66.7%; Tnaive_CD8 32.0%; Endo_Lym 0.7%; NP 0.6%
Every record has a known truth, so this checks an answer rather than merely producing one. Each name gives both the composition and the sparsity: three 70/30 pairs (macrophage/monocyte, CD4/CD8 T cell, naive CD4/naive CD8) at 222, 218 and 216 binarized CpGs. The pooled T-cell mixtures split across memory and naive sub-types, so sum those to score them. The 216 rung sits at the depth where a single draw becomes seed-dependent, which is why error grows down the ladder. The 62-type reference is built on the Zhou single-cell atlas plus Loyfer hepatocytes and BLUEPRINT granulocytes.
Per-cell panel, so each cell must deconvolve to itself.
methscope fetch -c hg38/data/human_hg38_celltypes.cg
methscope deconv-build-ref -o self.msdref human_hg38_celltypes.cg
methscope deconv self.msdref human_hg38_celltypes.cg
cell class fraction
Oligodendrocyte Oligodendrocyte 1.000000
Pancreas-Beta Pancreas-Beta 1.000000
Blood-NK Blood-NK 1.000000
Blood-Monocytes Blood-Monocytes 1.000000
.msdref, not a bundleDeconvolution takes a standalone reference; the solver rebuilds its pattern set per query from
the CpGs that query actually measured, so no pattern budget is baked in.
hg38_62celltypes.msdref is the primary reference (v8); the
single-cell backbone is what lets it read single-cell-derived immune mixtures and resolve
T-cell subtypes (Tmem/Tnaive × CD4/CD8). hg38_33celltypes.msdref is
the earlier bulk-pooled one; it replaced hg38_65celltypes.refx, a
withdrawn format the current binary can neither read nor produce.
For mouse brain, mm10_41celltypes.msdref (models v11) is the
first mouse reference: 41 Liu 2021 major types from binarised single cells. It is
brain only, like the mouse classifier, so route by tissue first. The
Published models tab carries its full card.
methscope deconv -o props.tsv hg38_62celltypes.msdref mixture.cg
Impute dense CpG methylation from a sparse methylome with a bundled decoder — pure-C
inference, no CUDA or BLAS at run time. The default output keeps the predicted
fraction per CpG (format 4); --binary thresholds it at 0.5
into 0/1 calls (format 6), and --probs writes the same values as
a TSV.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# what goes in: a sparse methylome
methscope fetch -c hg38/data/human_hg38_test.cg
yame summary human_hg38_test.cg
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
human_hg38_test.cg 1 NA global 29000 23857 29000 23857 NA 0.823 NA
# impute the whole genome from it — fractions, the default
methscope fetch -c hg38/models/hg38_10k1.updecx
methscope upscale -o human_hg38_test_reconstructed.cg hg38_10k1.updecx human_hg38_test.cg
yame summary human_hg38_test_reconstructed.cg
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
human_hg38_test_reconstructed.cg 1 NA global 29401795 10000 29401795 10000 NA 0.617 NA
# or 0/1 calls, for anything that wants a set rather than a level
methscope upscale --binary -o human_hg38_test_calls.cg hg38_10k1.updecx human_hg38_test.cg
yame summary human_hg38_test_calls.cg
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
human_hg38_test_calls.cg 1 NA global 10000 6239 10000 6239 NA 0.624 NA
This decoder writes one 10,000-CpG block. The two rows describe the same prediction from different sides: the fraction file spans the genome row space and reports 10,000 predicted CpGs at mean beta 0.617, while the binarized one is a set — 6,239 of the block's 10,000 CpGs called methylated, so its beta 0.624 is that fraction rather than a mean level. Keep the default unless a downstream tool demands format 6: thresholding is lossy, and it is what the accuracy check below needs.
Score against the truth, which is a set of 0/1 calls — so compare the
--binary output, not the fractions. Against the fraction file every
row differs and this prints 0/9794 (0.0%), which looks like a
broken model rather than a mismatched format.
methscope fetch -c hg38/data/human_hg38_test.truth.cg
paste <(yame rowsub -I 1_10000 human_hg38_test_calls.cg | yame unpack -a - | cut -f4) <(yame rowsub -I 1_10000 human_hg38_test.truth.cg | yame unpack -a - | cut -f4) | awk '$2!=2{t++;if($1==$2)c++} END{printf "%d/%d (%.1f%%)\n",c,t,100*c/t}'
9211/9794 (94.0%)
See the whole block first. The view is coordinate-aware — hprint
resolves the hg38 coordinate track from the store, window-averaging 10,000 CpGs
into 60 columns; -g shows deciles. The fetch below has no
-c, unlike the sample files above: this one is read from the
store, so putting it in the current directory leaves hprint
asking for -R.
cat human_hg38_test.truth.cg human_hg38_test.cg human_hg38_test_reconstructed.cg > three.cg
yame index -s <(printf 'truth\ninput\nreconstructed\n') three.cg
methscope fetch hg38/cpg_nocontig.cr
yame hprint -g -r chr1:921649-1151482 -w 60 three.cg
chr1:921649-1151482 (10000 CpGs, win=167)
|.........|.........|.........|.........|.........|.........
921649 958945 995856 1033422 1067326 1112421
truth 104523059802945587792005953275337999952335758999995168975689
input ...0.....5.0...........09...............9..0....9...9....9..
reconstructed 106644059812945577793006953275547999952536758999995168875689
0 unmethylated through 9 methylated, . no coverage, coloured blue to red. Each column averages 167 CpGs. The input row is mostly gaps, and where it does show a digit that digit is one or two observed CpGs standing in for the whole window — so it reads lower or higher than the truth because it is a far thinner sample of the same window, not because it disagrees with it. Truth and reconstruction, both dense, track each other across the block.
Zoom to single CpGs over the same block the score used — the same
rowsub -I as above. -c turns colour off,
which cut needs since it counts bytes, not columns.
The rows are named because the concatenation is written to a file and indexed first: a piped
stream carries no index, so hprint has no names to print.
-l sets the label column's width. Labels appear in the region
(-r) and whole-genome (-R) views only —
the plain dump has no label column — and a region this narrow gives one column per CpG
rather than a window average, which is what makes this single-CpG rather than binned.
# name the rows, then ask for a region small enough that one column is one CpG
cat human_hg38_test.truth.cg human_hg38_test.cg human_hg38_test_calls.cg > three_calls.cg
yame index -s <(printf 'truth\ninput\ncalls\n') three_calls.cg
yame hprint -c -r chr1:957686-959018 -l 8 three_calls.cg
chr1:957687-959017 (60 CpGs)
|.........|.........|.........|.........|.........|.........
957687 958034 958206 958496 958728 958945
truth ██████░█░██████████████░████░░░░░░█░░░░░░░░░░░░░░░░░░░░░░░░░
input .................█..........................░...............
calls ██████░█░██████████████░████░░░░░░█░░░░░░░░░░░░░░░░░░░░░░░░░
60 CpGs from the middle of the block. Two were observed, and the reconstruction matches the truth at all 60 — including the isolated 0 at position 7 and the lone 1 sitting in the unmethylated run.
Training the decoder is a separate, three-step pipeline, and every step reads the same
.mrmp.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# the whole pipeline from scratch, on the 40-cell-type chr20 reference
# 1. define the features
methscope fetch -c hg38/data/human_hg38_40_celltypes_chr20.cg
methscope mrmp-build --force human_hg38_40_celltypes_chr20.cg chr20.mrmp
mrmp-build: 40 classes, one MRMP over all of them (a tree of one level)
root 40 classes | 261 patterns | 3,787 CpGs
split threshold: none (flat) -> 1 group(s)
1 node(s), 16,792 bytes -> chr20.mrmp
# 2. featurize the truth atlas -> sampling/truth msur
methscope upscale-featurize --reps 20 --sample 8000 human_hg38_40_celltypes_chr20.cg chr20.mrmp chr20.msur
inflating 40 truth cells (1 worker, about 0.2 GB resident)
wrote chr20.msur and chr20.msur.tsv
# 3. group CpGs into processing units
methscope upscale-set-units --unit-cpgs 4096 human_hg38_40_celltypes_chr20.cg chr20.msui
CpGs=773477 real=657409 PNA=116068 memberships=116450
units=109 pure=13 mixed=67 PNA=29 oversized=7
unit CpGs min=1380 mean=7096.1 max=261816
wrote chr20.msui (5891452 bytes) and chr20.msui.tsv
inspect reads every build artifact, detected from its magic.
methscope inspect chr20.msui
format MSUIDX1 v1
cpgs 773477 (657409 real + 116068 PNA)
pattern_length 40
target_unit_cpgs 4096
units 109 (29 PNA)
real_memberships 116450
pattern_checksum acc25be7e28d1bba
file_bytes 5891452
methscope inspect chr20.msur
format MSURAW2 v2
cells 40
replicates 20
rows 800 (cells x replicates)
cpgs 773477
patterns 261
observed_beta continuous (full-depth truth)
observed_set sorted list (all replicates)
embedded_truth yes (trainable)
groups_bytes 1546954
truth_bytes 61878160
sampled_per_cell 8000
record_bytes 34088
records_bytes 27270400
hg38 · 207 cells × 100 reps · 29,401,795 CpGs · 15.0 GB offset section size 0 header 72 B 72 groups u16 × 29.4M 58.8 MB CpG → 1..1000, 0 = PNA 58,803,662 truth u16 × 207 × 29.4M 12.2 GB beta×65534, 65535 = missing (always embedded) 12,231,146,792 records 124,000 B × 20,700 2.6 GB 1000 beta + 1000 count + 29,000 CpG ids
.cg is never copied.hg38 · 16,384 CpGs/unit · 700 units · 2,345,190 memberships · 174 MB header 26,841,955 real CpGs + 2,559,840 PNA; checksum ties it to .mrmp units 24 B × 700 {offset, first_member, n_member, n_cpg, flags} u0 10,736,116 CpGs pure (P1, all-1) u1 2,859,193 CpGs pure (P2) u2 2,214,235 CpGs pure (P3, all-0) cpg 4 B × 29,401,795 output slot → genomic CpG index membership 24 B × 2,345,190 {key, offset, count, unit}
hg38 · 1000 MRMP in · 700 units · 29,401,795 CpGs out · 2.8 GB header patterns, input_dim, n_units, activation, section offsets mean/scale f32 × input_dim train-cell standardization units 32 B × 700 {out offset, param offset, bytes, n_cpg, rank} cpg 4 B × 29,401,795 output slot → genomic CpG index params f32 per unit: z = act(A·x + a); logit = E·z + b pure rank 16 · mixed/PNA rank 32
.cm mask.Training runs on CPU by default and threads over units, which are independent — that is what makes the run resumable. The 109 chr20 units take ~17 s on 8 threads.
methscope upscale-train -i chr20.msur --units chr20.msui --mrmp chr20.mrmp -o chr20.updecx --work-dir work/ --threads 8
P=261 features=beta+count input=522 trunk=none pure=16 mixed=factor/32 activation=linear
backend=cpu
CPU backend; cells=40 reps=20 patterns=261 input=522 units=109 split=28/6/6 threads=8
units 109/109 (resumed=0)
val MAE over 109 units: mean 0.159952, median 0.168192, worst 0.323179
wrote bare UPDEC2 work/model.updec2 (89885488 bytes), resumed=0
The validation MAE moves a little between runs — 0.171 and 0.194 on two passes over this same input — so treat it as a scale, not a checksum. The unit count, the pattern count and the model size are stable.
make CUDA=1 CUDA_HOME=/path/to/cuda CUDA_ARCH=sm_80 builds the GPU
backend, which is picked automatically when a device answers; checkpoints are interchangeable,
so you can start on CPU and finish on a GPU node. The chr20 demo is small; a whole-genome
decoder is memory- and GPU-heavy, so submit it to a GPU node:
sbatch -q gpu --gres=gpu:1 --mem=64G -c 8 -t 16:00:00 --wrap '\
methscope upscale-train -i $W/loyfer_100r_flat500_bank100.msur \
--units $W/units_16k.msui --mrmp $W/loyfer_train_flat500_bank100.mrmp \
-o $W/hg38_wg.updecx --work-dir $W/work --split $SPLIT \
--features scalar --activation leaky \
--pure-bottleneck 16 --mixed-bottleneck 32 --mixed-mode factor \
--min-steps 2000 --max-steps 60000 --patience 400 --eval-every 200 \
--eval-rows 24 --device 0'
--split rows: <cell_index>TAB<train|val|test>,
one per source cell; without it, cells are cut 70/15/15 by --seed.
--device cpu forces the portable backend even on a GPU node.
Classify, deconvolve and upscale each end in an answer. This tab stops one step earlier, at the matrix those answers are computed from: for every record, the mean methylation of each pattern in a set. It is the embedding to hand to a clustering, a projection or a heatmap when the question is what these samples look like rather than what they are. A pattern set built once serves any number of queries.
# a scratch dir for the outputs
mkdir -p ~/tmp/methscope && cd ~/tmp/methscope
# a 40-cell-type chr20 reference (39.5 MB); the same set the Classify tab
# builds, so --force lets either tab be run first
methscope fetch -c hg38/data/human_hg38_40_celltypes_chr20.cg
methscope mrmp-build --top 500 --force human_hg38_40_celltypes_chr20.cg chr20_40celltypes.mrmp
Any .mrmp works here: one built as above, a bank
with hundreds of sets, or the one inside a published bundle. See the Classify tab for
what a pattern is and how a set is inspected.
A pattern is a set of CpGs, so a record has one number per pattern: the mean of the betas
it covers there. That matrix is what classify-train fits on, and
mrmp-summary prints it as TSV — so it is also the embedding to
cluster, project or plot when the question is what these cells look like rather than what
they are. --min-cpgs N reads a thinly covered pattern as
NA instead of a confident mean.
# One mean methylation value per pattern per record, every set of the chain in
# one pass over the query. Query first, set second.
methscope mrmp-summary human_hg38_40_celltypes_chr20.cg chr20_40celltypes.mrmp \
-o chr20_patterns.tsv
head -4 chr20_patterns.tsv
sample set pattern beta
GSM5652176_Adipocytes-Z000000T7 root P1 0.8954
GSM5652176_Adipocytes-Z000000T7 root P2 0.9265
GSM5652176_Adipocytes-Z000000T7 root P3 0.9347
--wide gives the same numbers as a record × pattern matrix,
the shape a clustering or a heatmap wants. Columns are named
<set>.<pattern>, so a chain of sets stays unambiguous.
# the same numbers as one row per record, one column per pattern
methscope mrmp-summary --wide human_hg38_40_celltypes_chr20.cg chr20_40celltypes.mrmp \
-o chr20_patterns_wide.tsv
head -3 chr20_patterns_wide.tsv | cut -f1-4
sample root.P1 root.P2 root.P3
GSM5652176_Adipocytes-Z000000T7 0.8954 0.9265 0.9347
GSM5652181_Saphenous-Vein-Endothel-Z000000RM 0.9155 0.9413 0.9421
There is no Pna background column. It is the CpGs no pattern claims, which
measures coverage rather than biology, so the feature builder never gives it a column and
nothing downstream can embed sequencing depth as an axis. Its mean is the query's global
beta anyway — Pna is most of the genome — which
yame summary already prints. The numbers here agree with
yame summary -m over the exported
.cm to within 5×10-4, the width of yame's printed
precision.
Every model is a single self-contained bundle carrying its own MRMP feature definition
(and labels or signature), so a query .cg runs directly. The file
extension names the role — .clfx classifier,
.updecx upscale decoder, .msdref deconvolution
reference; bundles are detected by magic, never by name.
methscope fetch -y hg38/models takes the human set,
methscope fetch -y mm10/models the mouse
(a whole directory is gigabytes, so a bare fetch hg38/models on a
terminal opens the browser on that folder with the planned files checked rather than
starting, and refuses outright when nothing can answer — a script, a job
— which is what -y settles); methscope inspect
<model> prints the exact framework, label list and, for linear models, per-feature
weights. The models a given binary fetches are the ones it documents:
methscope --version prints the tag.
.clfx — run with classifyhg38_celltype_full.clfx — Cell-type classifier (human, full) · 153.8 MBWhat it is. Assigns a human methylome, from a single cell to a bulk sample, to one of 63 cell types. Use it when accuracy matters most; hg38_celltype_lite is a seventh of the size.
Architecture. One flat MRMP bank: hierarchical LCA binstrings plus a calibrated resolver for every pair of cell types (1,981 pattern sets, 1,953 resolvers), all scored by a single pooled xgboost booster. With no routing tree, there is no early decision whose mistake a later node cannot undo.
Training. The reference is 84,601 labelled cells: pseudobulks from the Zhou lab 2025 single-cell atlas, the Loyfer sorted-cell atlas and BLUEPRINT. Built 2026-09-06 with mrmp-build --bank --resolver-gate -1 and classify-train --pool-nodes (max depth 4, colsample 0.4), over a 24-rung coverage ladder with binarized reads. The final model was assembled from the per-fold resolver cache. Both flags have since been retired (--resolvers N replaces the gate, and the single-booster mode is gone), so the current recipe is the one on the methscope training page.
Accuracy. Class-balanced 0.9544 / micro 0.9569 in ten-fold cross-validation over all 84,601 cells.
History. It replaces the routing-tree classifier hg38_celltype (33 classes), withdrawn at model tag v10.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections
tree hg38_celltype_full.clfx
format 1 booster(s) over a 1981-set chain
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 136,146,848 MRMP chain (MRMPIDX1)
1 kind 136,146,956 4 "tree" framework mark
2- booster chain 136,146,960 25,176,288 xgboost boosters for tree nodes
Node details
node classes booster patterns parent
root 63 25,176,288 8 -- + 1953 soft, 77,819,311 CpGs
root.0 61 0 10 root
root.0.0 59 0 10 root.0
root.0.0.0 58 0 11 root.0.0
root.0.0.0.0 54 0 10 root.0.0.0
root.0.0.0.0.0 53 0 10 root.0.0.0.0
root.0.0.0.0.0.0 51 0 11 root.0.0.0.0.0
root.0.0.0.0.0.0.0 49 0 11 root.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0 47 0 11 root.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0 44 0 11 root.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0 43 0 11 root.0.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1 39 0 9 root.0.0.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1.0 3 0 6 root.0.0.0.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1 33 0 12 root.0.0.0.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0 31 0 10 root.0.0.0.0.0.0.0.0.0.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.0 9 0 7 root.0.0.0.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.0.2 4 0 3 root.0.0.0.0.0.0.0.0.0.0.1.1.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1 22 0 8 root.0.0.0.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1 21 0 8 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0 11 0 7 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0 8 0 5 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0 6 0 7 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0.1 3 0 3 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1 10 0 6 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1.0 4 0 5 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1.1 5 0 9 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1
root.0.0.0.0.0.0.0.0.1 3 0 4 root.0.0.0.0.0.0.0.0
root.0.0.0.1 4 0 2 root.0.0.0
1953 soft child(ren) lend columns to a node's booster and carry
no model of their own; `inspect --tree` shows them in place.
Score with: methscope classify hg38_celltype_full.clfx query.cg
hg38/models · tag v12 · 153.8 MB
hg38_celltype_lite.clfx — Cell-type classifier (human, lite) · 21.2 MBWhat it is. A compact version of hg38_celltype_full, about a seventh of its size (21 MB against 154 MB), covering the same 63 cell types. Use it when download size or memory matters.
Architecture. The same hard blocks as the full model, with calibrated resolvers only for the 150 cell-type pairs with the thinnest hard-block footprint (the fewest CpGs where the two class pseudobulks differ by more than 0.5), and a 100-round booster.
Training. As hg38_celltype_full, with mrmp-build --min-pattern-cpgs 500 --resolvers 150 and classify-train -n 100 (methscope-cli 60988b9).
Validation. Correctly calls the four Loyfer example cells (ODC, Beta, NK CD16, Mono).
Accuracy. Class-balanced 0.9403 / micro 0.9558 in ten-fold cross-validation at full coverage, against 0.9544 / 0.9569 for the full model. The gap widens on sparse data: at 4,096 observed CpGs, class-balanced accuracy is 0.6685 against 0.7612.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections
tree hg38_celltype_lite.clfx
format 1 booster(s) over a 178-set chain
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 12,488,808 MRMP chain (MRMPIDX1)
1 kind 12,488,916 4 "tree" framework mark
2- booster chain 12,488,920 9,716,376 xgboost boosters for tree nodes
Node details
node classes booster patterns parent
root 63 9,716,376 30 -- + 150 soft, 5,770,835 CpGs
root.0 61 0 36 root
root.0.0 59 0 36 root.0
root.0.0.0 58 0 33 root.0.0
root.0.0.0.0 54 0 34 root.0.0.0
root.0.0.0.0.0 53 0 34 root.0.0.0.0
root.0.0.0.0.0.0 51 0 36 root.0.0.0.0.0
root.0.0.0.0.0.0.0 49 0 35 root.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0 47 0 33 root.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0 44 0 32 root.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0 43 0 31 root.0.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1 39 0 28 root.0.0.0.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1.0 3 0 6 root.0.0.0.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1 33 0 23 root.0.0.0.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0 31 0 22 root.0.0.0.0.0.0.0.0.0.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.0 9 0 10 root.0.0.0.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.0.2 4 0 3 root.0.0.0.0.0.0.0.0.0.0.1.1.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1 22 0 17 root.0.0.0.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1 21 0 17 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0 11 0 12 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0 8 0 13 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0 6 0 12 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0.1 3 0 4 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.0.0.0
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1 10 0 17 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1.0 4 0 9 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1
root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1.1 5 0 11 root.0.0.0.0.0.0.0.0.0.0.1.1.0.1.1.1
root.0.0.0.0.0.0.0.0.1 3 0 4 root.0.0.0.0.0.0.0.0
root.0.0.0.1 4 0 3 root.0.0.0
150 soft child(ren) lend columns to a node's booster and carry
no model of their own; `inspect --tree` shows them in place.
Score with: methscope classify hg38_celltype_lite.clfx query.cg
hg38/models · tag v12 · 21.2 MB
hg38_sex.clfx — Sex classifier (human) · 30.2 KBWhat it is. Calls a sample female or male. At 30 KB it is the smallest model in the catalogue, and a quick way to check that an installation works end to end.
Architecture. A logistic model over two X-inactivation features, Xa_lo and Xa_hi, from a three-state MRMP.
Accuracy. About 95.8% on an independent cohort.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 30,931 bytes
bundle hg38_sex.clfx
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 30,653 YAME .cm mask, 3 states
1 kind 30,761 8 "logistic" framework mark
2 model 30,769 154 methscope-linear text spec
[0] mrmp MRMP feature definition (YAME .cm)
states 3
[1] kind bundle framework mark
value logistic
[2] model logistic linear classifier - run via `classify`
method logistic
labels Female, Male
bias -7.50339
scale 1
features 2
Xa_lo weight=-13.9125 mean=0.24253
Xa_hi weight=13.6301 mean=0.752643
hg38/models · tag v12 · 30.2 KB
mm10_brain_full.clfx — Cell-type classifier (mouse brain, full) · 64.5 MBWhat it is. Assigns a mouse-brain methylome to one of 41 brain cell types from the Liu 2021 snmC-seq2 atlas.
Architecture. One flat MRMP bank: LCA binstrings plus a calibrated resolver for every pair of cell types (837 pattern sets, 820 resolvers), scored by a single pooled booster with no routing.
Training. The reference is 41 all-cell binasum pools over the 54,447 cells of a 2,000-per-class draw. The booster was trained on a balanced 400-per-class draw (14,825 cells), over a 24-rung coverage ladder with binarized reads. The final model was assembled 2026-09-11 (8 calibration strides plus assembly, about 6.5 h).
Accuracy. Class-balanced 0.9746 / micro 0.9720 in ten-fold cross-validation. On the 49,113 Liu cells that never entered the model (18 of the 41 classes, the large cortical ones): micro 0.9660 / class-balanced 0.9707.
Caveats. It is brain only. With no non-brain classes, a cell from another tissue is assigned to the nearest brain type without warning, so confirm the tissue first.
History. It replaces the routing-tree classifier mm10_celltype_brain, withdrawn at model tag v10.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections
tree mm10_brain_full.clfx
format 1 booster(s) over a 837-set chain
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 54,445,640 MRMP chain (MRMPIDX1)
1 kind 54,445,748 4 "tree" framework mark
2- booster chain 54,445,752 13,168,045 xgboost boosters for tree nodes
Node details
node classes booster patterns parent
root 41 13,168,045 8 -- + 820 soft, 32,765,586 CpGs
root.0 39 0 9 root
root.0.0 36 0 7 root.0
root.0.0.0 35 0 7 root.0.0
root.0.0.0.0 33 0 5 root.0.0.0
root.0.0.0.0.0 31 0 8 root.0.0.0.0
root.0.0.0.0.0.0 27 0 8 root.0.0.0.0.0
root.0.0.0.0.0.0.0 26 0 12 root.0.0.0.0.0.0
root.0.0.0.0.0.0.0.1 24 0 12 root.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.1.1 22 0 11 root.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.1.1.0 21 0 11 root.0.0.0.0.0.0.0.1.1
root.0.0.0.0.0.0.0.1.1.0.2 8 0 9 root.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.1.1.0.2.1 4 0 3 root.0.0.0.0.0.0.0.1.1.0.2
root.0.0.0.0.0.0.0.1.1.0.2.1.0 3 0 2 root.0.0.0.0.0.0.0.1.1.0.2.1
root.0.0.0.0.0.0.0.1.1.0.3 8 0 7 root.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.1.1.0.3.0 5 0 5 root.0.0.0.0.0.0.0.1.1.0.3
root.0.0.0.0.0.2 3 0 6 root.0.0.0.0.0
820 soft child(ren) lend columns to a node's booster and carry
no model of their own; `inspect --tree` shows them in place.
Score with: methscope classify mm10_brain_full.clfx query.cg
mm10/models · tag v12 · 64.5 MB
mm10_brain_lite.clfx — Cell-type classifier (mouse brain, lite) · 13.3 MBWhat it is. A compact version of mm10_brain_full, about a fifth of its size, covering the same 41 cell types.
Architecture. The same hard blocks as the full model, with calibrated resolvers only for the 100 cell-type pairs with the thinnest hard-block footprint (the fewest CpGs where the two class pseudobulks differ by more than 0.5), and a 100-round booster.
Training. As mm10_brain_full, with mrmp-build --min-pattern-cpgs 500 --resolvers 100 and classify-train -n 100 (methscope-cli 60988b9).
Accuracy. On the 49,113 held-out Liu cells: micro 0.9583 / class-balanced 0.9638 (18 classes), within a point of the full model (0.9660 / 0.9707). On fold-0 cross-validation test cells, class-balanced accuracy is 0.9778 at full coverage (full model 0.9789) and 0.8195 at 4,096 observed CpGs (0.8778).
Caveats. Like every mouse classifier here, it is brain only.
History. It beats the v9 lite (micro 0.9482 / class-balanced 0.9595), which had 117 gate-picked resolvers and was twice this size.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections
tree mm10_brain_lite.clfx
format 1 booster(s) over a 117-set chain
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 7,835,048 MRMP chain (MRMPIDX1)
1 kind 7,835,156 4 "tree" framework mark
2- booster chain 7,835,160 6,082,814 xgboost boosters for tree nodes
Node details
node classes booster patterns parent
root 41 6,082,814 24 -- + 100 soft, 3,965,586 CpGs
root.0 39 0 23 root
root.0.0 36 0 24 root.0
root.0.0.0 35 0 27 root.0.0
root.0.0.0.0 33 0 25 root.0.0.0
root.0.0.0.0.0 31 0 28 root.0.0.0.0
root.0.0.0.0.0.0 27 0 20 root.0.0.0.0.0
root.0.0.0.0.0.0.0 26 0 21 root.0.0.0.0.0.0
root.0.0.0.0.0.0.0.1 24 0 23 root.0.0.0.0.0.0.0
root.0.0.0.0.0.0.0.1.1 22 0 22 root.0.0.0.0.0.0.0.1
root.0.0.0.0.0.0.0.1.1.0 21 0 22 root.0.0.0.0.0.0.0.1.1
root.0.0.0.0.0.0.0.1.1.0.2 8 0 12 root.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.1.1.0.2.1 4 0 7 root.0.0.0.0.0.0.0.1.1.0.2
root.0.0.0.0.0.0.0.1.1.0.2.1.0 3 0 5 root.0.0.0.0.0.0.0.1.1.0.2.1
root.0.0.0.0.0.0.0.1.1.0.3 8 0 16 root.0.0.0.0.0.0.0.1.1.0
root.0.0.0.0.0.0.0.1.1.0.3.0 5 0 11 root.0.0.0.0.0.0.0.1.1.0.3
root.0.0.0.0.0.2 3 0 6 root.0.0.0.0.0
100 soft child(ren) lend columns to a node's booster and carry
no model of their own; `inspect --tree` shows them in place.
Score with: methscope classify mm10_brain_lite.clfx query.cg
mm10/models · tag v12 · 13.3 MB
.msdref — run with deconvhg38_33celltypes.msdref — Deconvolution reference (human, bulk) · 528.4 MBWhat it is. Estimates the cell-type composition of a methylome against 33 human cell types. Prefer it when the query is itself sorted or bulk tissue; for single-cell-derived data use hg38_62celltypes, whose immune types come from single cells.
Architecture. Covers the 7.9 million CpGs (26.9% of hg38) that differ between at least some of the types. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart under one joined label.
Training. 286 normal bulk WGBS samples, grouped into 33 labels by a donor-resolved control sheet, pooled per class by read count (musum), then packed with deconv-build-ref at its default admission band 0.30,0.70.
Validation. A shipped build uses every donor, so the confusion matrix was measured on a donor-held-out twin: per label, every fourth donor by sample-ID hash was held out as mixture components, giving 480 mixtures of 2 to 5 classes pooled over 2^16 to 2^22 CpGs. (The control sheet's own split column is a cohort flag covering only 9 of the 33 classes, so it could not be used.) The closest pairs are CD4 and CD8 T cells (0.43), macrophage and monocyte (0.23), dendritic cell and monocyte (0.17), bladder and prostate epithelium (0.12), and NK and CD8 T cells (0.10).
Caveats. Its macrophage class is monocyte-derived in vitro, so tissue-resident macrophages have no matching reference.
History. It replaced hg38_65celltypes.refx, withdrawn at model tag v7; methscope no longer reads that format. Version 3 of the format since model tag v12 adds the confusion matrix; the M/U block is byte-identical to version 2.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
What inspect reports.
deconvolution reference MSDREF1/v3, 554,072,253 bytes
cell types 33
CpGs kept 7,915,244
CpG row space 29,401,795
qfilter 0.30,0.70
beta cut 0.50
min coverage 1
confusion 33 x 33, 1066 non-zero
most alike T.Cell.CD4 / T.Cell.CD8 0.430
Macrophage / Monocyte 0.229
Dendritic.Cell / Monocyte 0.169
measured donor-held-out twin: per label every 4th donor by sample-id hash held out as mixture components, the rest pooled into the reference; 480 mixtures of 2-5 classes, pooled over 2^16-2^22; 20260919_human33_deconv
hg38/models · tag v12 · 528.4 MB
hg38_62celltypes.msdref — Deconvolution reference (human) · 1.4 GBWhat it is. Estimates the cell-type composition of a methylome against 62 human cell types. Because its immune types come from single cells rather than sorted bulk, it resolves tissue-resident immune cells that the bulk-based hg38_33celltypes misreads.
Architecture. Sixty types are pseudobulks from the Zhou lab single-cell atlas (the early-embryonic syncytio- and villous trophoblast are excluded); Loyfer hepatocytes and BLUEPRINT granulocytes cover the two compartments single cells miss. It covers the 11.8 million CpGs (40.1% of hg38) that differ between at least some of the types, packed as uint16 M/U: a row is kept only where at least one class is unmethylated and at least one is methylated. No pattern budget is fixed; the solver selects its patterns for each query from the CpGs that query measured. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart under one joined label; a group total is identified even where its split is not.
Training. Per-class M/U pooled by read count (musum), then packed with deconv-build-ref at its default admission band 0.30,0.70.
Validation. A known 70/30 macrophage/monocyte mixture comes back within a few points of truth, where the bulk reference returns 0.14/0.56 and puts the missing mass on mast cells and stroma. T cells resolve to naive and memory CD4 and CD8. The confusion matrix was measured on a fold-held-out twin: pools rebuilt from Zhou cells in folds 1-9, fold 0 held out as mixture components, and Hepatocyte and Granulocyte held out by donor. The closest pairs are naive CD4 and CD8 T cells (0.29), colonic enterocyte and goblet cells (0.24), the CGE interneurons LAMP5, SNCG and VIP (0.22), muscle fibroblast and its adipogenic progenitor (0.20), NK CD16 and CD56 (0.15), adrenal ZG and ZR/ZF (0.14), cortical IT L4 and L5 (0.12), AT1 and AT2 (0.12), and Beta and Delta (0.11). A threshold of 0.15 joins five pairs, and 0.10 joins thirteen.
History. Version 3 of the format since model tag v12 adds the confusion matrix; the M/U block is byte-identical to version 2.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
What inspect reports.
deconvolution reference MSDREF1/v3, 1,510,196,803 bytes
cell types 62
CpGs kept 11,798,286
CpG row space 29,401,795
qfilter 0.30,0.70
beta cut 0.50
min coverage 1
confusion 62 x 62, 3785 non-zero
most alike Tnaive_CD4 / Tnaive_CD8 0.292
Ent_Co / Gob 0.242
LAMP5 / SNCG 0.219
measured fold-held-out twin: Zhou single cells, folds 1-9 pooled into the reference, fold 0 held out as mixture components (Hepatocyte/Granulocyte donor-held-out); 480 mixtures of 2-5 classes, pooled over 2^16-2^22; 20260919_human_deconv
hg38/models · tag v12 · 1.4 GB
mm10_41celltypes.msdref — Deconvolution reference (mouse brain) · 554.9 MBWhat it is. Estimates the cell-type composition of a mouse-brain methylome against 41 major cell types from the Liu 2021 atlas.
Architecture. Covers the 6,765,325 CpGs (30.9% of mm10) that differ between at least some of the types, packed as uint16 M/U. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart, such as CA3 and DG-po, under one joined label; a group total is identified even where its split is not.
Training. Each type is a pseudobulk of binarized single cells from the Liu 2021 snmC-seq2 atlas, pooled with yame rowop -o binasum over all 54,447 cells, so every cell counts once per CpG. The pools are concatenated in sorted label order and packed with methscope deconv-build-ref at qfilter 0.30,0.70.
Validation. An all-cells build has no held-out cells, so the confusion matrix was measured on a fold-held-out twin, each class alone at 2^16 to 2^22 CpGs, pooled across rungs. On P21 mouse-brain spatial data it agrees with an independent RNA-based deconvolution (RCTD): r = 0.74 for oligodendrocytes, 0.69 for CA1, 0.69 for dentate gyrus and 0.59 for CA3.
Caveats. Seven types rest on fewer than 300 cells, and grouping merges them into the neighbour that already absorbs them. The cortical IT ladder stays four separate labels.
History. The first version-3 .msdref, the format that carries a confusion matrix.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
What inspect reports.
deconvolution reference MSDREF1/v3, 581,825,095 bytes
cell types 41
CpGs kept 6,765,325
CpG row space 21,867,837
qfilter 0.30,0.70
beta cut 0.50
min coverage 1
confusion 41 x 41, 340 non-zero
most alike CA3 / DG-po 0.354
ANP / ASC 0.238
CGE-Lamp5 / Unc5c 0.203
measured pooled over 2^16, 2^18, 2^20, 2^22; each class alone, held-out cells, 20260918_mouse_msdref
mm10/models · tag v12 · 554.9 MB
.updecx — run with upscalehg38_10k1.updecx — Upscaling decoder (human, one-block demo) · 27.7 MBWhat it is. A small demonstration decoder that predicts one block of 10,000 hg38 CpGs from 101 MRMP pattern values. Use it to try `upscale` quickly before fetching the 2.7 GB whole-genome decoder; the documentation's upscale example runs it.
Architecture. UPDEC1: a single MLP decoder head for block 10k1.
Training. Trained on the same Loyfer sorted-cell atlas as hg38_wg.
Caveats. Its output is still a whole-genome .cg: the block is filled and every other CpG is NA.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 29,075,778 bytes
bundle hg38_10k1.updecx
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 8,331,342 YAME .cm mask, 101 states
1 outcpg 8,331,450 5,996 genome-wide output-CpG mask
2 model 8,337,446 20,738,324 UPDEC1 MLP decoder (101->512->10000)
[0] mrmp MRMP feature definition (YAME .cm)
states 101
[1] outcpg genome-wide output-CpG mask
imputed-CpG locations; makes `upscale` emit a whole-genome .cg
[2] model upscale decoder - run via `upscale`
n_in 101
n_hidden 512
n_out 10000
hg38/models · tag v12 · 27.7 MB
hg38_wg.updecx — Upscaling decoder (human, whole genome) · 2.7 GBWhat it is. Predicts methylation at all 29,401,795 hg38 CpGs from a sparse methylome, such as a single cell or a low-coverage sample. Inference runs in C, in about 2 s per sample.
Architecture. UPDEC2. The genome is divided into 700 units of about 16,000 CpGs, each with its own low-rank factor head (bottleneck 16 for pure units, 32 for mixed, leaky activation). The encoder input is 836 values: 835 MRMP pattern betas plus log1p of the number of CpGs observed.
Training. The patterns come from a 67-class Loyfer reference pooled from the 152 training samples only, cut to 500 patterns by CpG count and joined with the 100 resolver pairs of smallest footprint from the classifier bank (mrmp-pool --pooled-top 0) (835 columns in all). The model was trained on the 152/45/10 training/validation/test split of the 207-sample Loyfer atlas, against 100 binarized simulated replicates with coverage log-spaced from 71 to 2,940,180 observed CpGs (skew 0.5, dense-weighted). Each unit stops early on validation MAE (patience 400, 60,000-step cap); training took 7 h 38 min on one A100.
Caveats. At 2.7 GB this is the largest file in the catalogue; fetch it deliberately.
History. v10 (2026-09-16) changed where the patterns come from. Earlier references had no hepatocyte class, which caused an undercall at hepatocyte loci; at ELF3 with 29,000 observed CpGs the error fell from 0.309 to 0.047. On the 10 held-out samples v10 is more accurate at both ELF3 and MIR200C at every coverage tested (ELF3 0.113 to 0.107 at 29k, 0.098 to 0.089 at 290k, 0.093 to 0.089 at 2.9M; MIR200C 0.078 to 0.073, 0.072 to 0.069, 0.072 to 0.069). v9 and earlier took 500 patterns from a 35-class Zhou reference and trained at patience 20. v4 (2026-07-29) used a curated 117/45/45 split for a like-for-like comparison with the published MLP baseline.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 2,866,632,834 bytes
bundle hg38_wg.updecx
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 17,948,767 YAME .cm mask, 124 records, 501 states in the first
1 kind 17,948,875 7 "upscale" framework mark
2 model 17,948,882 2,848,683,944 UPDEC2 whole-genome unit decoder (700 units)
[0] mrmp MRMP feature definition (YAME .cm)
states 501
records 124 (one per set)
[1] kind bundle framework mark
value upscale
[2] model whole-genome upscale decoder - run via `upscale`
format UPDEC2/v3
patterns 835
input_dim 836
features standardized beta + missing indicator
missing NaN -> beta 0, indicator 1
shared trunk none
activation leaky_relu_0.01
units 700
pure/mixed/PNA 109 / 434 / 157
factor/direct 700 / 0
bottleneck_dim 16..32
memberships 2345190
CpGs 29,401,795
target/unit 16384
model bytes 2,848,683,944
index checksum a21ad74e190c803f
parameter sum 53ab3001c5a767c7
hg38/models · tag v12 · 2.7 GB
mm10_wg.updecx — Upscaling decoder (mouse, whole genome) · 1.9 GBWhat it is. Predicts methylation at all 21,867,837 mm10 CpGs from a sparse methylome. Inference runs in C.
Architecture. UPDEC2, the same architecture as hg38_wg: 472 units, each with its own low-rank factor head (bottleneck 16 for pure units, 32 for mixed, leaky activation). The encoder input is 501 values: 500 MRMP pattern betas plus log1p of the number of CpGs observed.
Training. The patterns come from 40 mouse-brain pseudobulk cell types of the Liu 2021 atlas. The atlas has 41, but MRMP packs a pattern as a base-3 uint64 and so caps at 40 samples; VLMC-Pia, a sub-split of the retained VLMC, was dropped. The training targets are 200 of the 258 Zhou 2018 mm10 methylomes with at least 80% genome coverage, against 100 binarized replicates log-spaced from 71 to 2,186,784 observed CpGs (skew 0.5). 60,000-step cap, patience 20; no unit reached the cap.
Accuracy. On held-out samples the mean absolute error is 0.098; with only 106 observed CpGs it is 0.134, at Pearson r 0.80.
Caveats. At 1.9 GB, fetch it deliberately.
History. Training across a coverage ladder is what closed the gap to the human model. A fixed-coverage build plateaued at 0.126 MAE; at 106 observed CpGs the ladder improved the error from 0.241 to 0.134 and Pearson r from 0.48 to 0.80.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8; Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
What inspect reports.
container MSBNDL1 (MethScope BuNDLe v1) - 3 sections, 2,057,125,237 bytes
bundle mm10_wg.updecx
# section offset size content
--- --------------- ----------- ----------- --------------------------------
0 mrmp 0 10,254,074 YAME .cm mask, 501 states
1 kind 10,254,182 7 "upscale" framework mark
2 model 10,254,189 2,046,871,040 UPDEC2 whole-genome unit decoder (472 units)
[0] mrmp MRMP feature definition (YAME .cm)
states 501
[1] kind bundle framework mark
value upscale
[2] model whole-genome upscale decoder - run via `upscale`
format UPDEC2/v3
patterns 500
input_dim 501
features standardized beta + missing indicator
missing NaN -> beta 0, indicator 1
shared trunk none
activation leaky_relu_0.01
units 472
pure/mixed/PNA 75 / 268 / 129
factor/direct 472 / 0
bottleneck_dim 16..32
memberships 1455530
CpGs 21,867,837
target/unit 16384
model bytes 2,046,871,040
index checksum ffb0a5ac1b36c70c
parameter sum 3f20e431aeeead67
mm10/models · tag v12 · 1.9 GB
Generated from the file table YAME/tools/registry/files.tsv at the submodule's tag by docs/build_models.py; the same rows describe each file in methscope fetch. Edit the TSV, not this page.
-h, as this build prints itGenerated from the binary at build time, so nothing here can drift from what
methscope <command> -h says. Each entry is folded — click to open.
In every example on the other tabs the highlighted subcommand links to its entry below and
opens it. Grouped as the methscope banner groups them.
methscope fetch — Download models and example data into the store (or -c: here)
Usage:
methscope fetch browse the catalogue
methscope fetch [options] <name> ...
methscope fetch [options] -u <url> -s <sha256> -o <dest>
Naming:
A name is a store directory, as the browser shows it: hg38,
hg38/KYCG, hg38/data, EPIC. It takes that directory's own files --
`hg38` is the genome annotation, not the knowledgebase and models
beneath it; name those explicitly. Narrow within one with -g.
A file resolves too, best written out: `hg38/data/test.cg`. The
bare name works when only one directory publishes it.
Several names may be given, separated by spaces or by commas --
commas so a list fits an option that takes one argument. A
directory named twice is taken once, and naming it whole absorbs a
file picked out of it.
Browsing:
With no target on a terminal, opens a tree browser: species, then
platform or genome build, then its knowledgebase and files. Arrows
move, right/left open and close a row, space or x selects (a folder
takes everything under it), f fetches what is selected, h lists every
key, q leaves. `/` filters the tree -- by name, source, collection
or title -- and enter keeps the filter so you can then select.
What is already in the store shows as present and
cannot be selected.
Piped or redirected it dumps the registry as TSV instead (same as
-l), so a script never blocks on a keystroke.
Purpose:
Download reference assets into the shared store that every tool in the
suite reads, verifying each file against a digest this build pins.
-c puts them in the current directory instead, for a one-off or a
demo: no SHA256SUMS is written beside them, but each file is checked
against the same digest.
Store:
Resolved in order: -d, $METHSCOPE_DATA_HOME, $YAME_DATA_HOME,
${XDG_DATA_HOME:-~/.local/share}/yame
METHSCOPE_DATA_HOME: /home/user/.local/share/yame
Mirror:
$YAME_ASSETS_MIRROR=<scheme://host[:port]> downloads from a site that
mirrors the public repositories, keeping each URL's path: the file at
https://raw.githubusercontent.com/zhou-lab/X/v1/f is fetched from <mirror>/zhou-lab/X/v1/f.
Every byte is still checked against the compiled-in digest.
Options:
-d <dir> Store root, overriding the environment.
-f Re-download what is present, and replace a file the store's
manifest records at a different digest than this build pins.
-R A directory name also takes every directory beneath it:
`-R hg38` is hg38, hg38/KYCG, hg38/data and hg38/models.
Without it a name is one directory's own files. Works
with -l and -n, so `-n -R hg38` says what that reaches.
-l Dump the registry as TSV and exit: one row per file, with
its size, digest, description and whether the store has it.
Takes the same <name> and -g a fetch does, so `-l -g
methscope hg38` is the dry run for fetching exactly that.
-g <a,b> Only files matching every term: name, source, collection,
title or upstream database. `-g chromatin` inside a
knowledgebase, `-g celltype` across a whole genome.
-n Say what would be fetched and stop, successfully. The same
plan a fetch prints before asking, but it exits 0, so a
script can check first. -l gives the same set as TSV.
-y Fetch a whole folder without asking. A folder is confirmed
first, since a short name can reach a lot -- `hg38` is 3.5
GB. On a terminal the browser is the confirmation: it opens
on the folder with exactly those files checked, f fetches,
q leaves. Off one, a folder needs -y. A name that picks out
ONE file needs neither: naming the file is the confirmation,
so a documented fetch line runs in a script as it stands.
-q No progress output.
-u <url> Single-file form: what to download.
-s <sha256> Single-file form: the digest it must have.
-o <dest> Single-file form: where it goes (a path, not a dir).
-c Into the current directory rather than the store.
-h This help.
Notes:
* Each file is verified against the digest this build carries for it;
the SHA256SUMS a directory keeps records what was verified there.
A file recorded at another digest is stale and needs -f to replace.
* `shasum -a 256 -c SHA256SUMS` in any store directory re-verifies it
by hand, with none of this code involved.
* Built with libcurl: fetch available.
methscope mrmp-build — Build an MRMP set (--bank: the classifier's chain)Usage: methscope mrmp-build [options] REF.cg OUT.mrmp
methscope mrmp-build --bank [options] REF.cg OUT.mrmp
Build an MRMP from a labelled reference store: one record per class,
record name = class name. The default is ONE set over every class,
for upscale, the violation rule and export; --bank is the classifier
artifact (the release models), a chain of blocks from an annealed
split hierarchy plus calibrated 2-class resolvers.
Workflow
methscope mrmp-build --top 500 REF.cg OUT.mrmp
methscope mrmp-build --bank --anneal-min-seg 10000,3000,1000 \
--resolvers 100 --cell-store CELLS.cg --cell-labels CELLS.tsv \
REF.cg OUT.mrmp
Split discovery (the hierarchy behind a bank's binstring blocks)
--anneal-min-seg HI,LO[,STEP]
Anneal the split threshold PER NODE:
try HI, and when nothing separates,
drop exactly to the first disconnection of the
single-linkage graph (the largest threshold
that splits anything -- the slowest possible
peel), stopping at LO. A node unsplittable at
LO is a leaf. Each node re-anneals from HI, so
the root peels only its best-separated groups
while deep leaves can still split fine
structure. A STEP quantizes the drop to the
grid HI, HI-STEP, ...: near-tied
disconnections land in one split instead of a
cascade of near-duplicate levels, and the
recorded thresholds are round numbers. Every
node records the threshold it used in its
header (inspect --tree shows it). Required by
--bank; a fixed threshold is HI,HI.
--max-depth N Maximum split depth. Default: 16.
Modes
--top N Default mode only: keep the N patterns with the most
CpGs, the rest folded into PNA. Default:
10000; 0 keeps every pattern. The build ranks
every candidate; this is the consumer's cut,
applied here because a flat set has one
consumer -- the same prune as mrmp-pool
--pooled-top, for chains that must compete.
--bank The classifier artifact: a FLAT chain of two
feature types and no routing. Type 1: each split
node's binstrings, pruned to patterns with
>= --min-pattern-cpgs CpGs (tree-named for
provenance only). Type 2: calibrated 2-class
rebuilds named root@A__B -- every terminal
pair, every pair inside an unsplittable leaf
(which emits no type-1 of its own), and every
cross-split pair --resolvers picks.
Needs --anneal-min-seg and the calibration
inputs. Featurize the result with
--satellite-contrast replace and train with
classify-train --data.
--min-pattern-cpgs N A type-1 pattern must carry at least N CpGs
to stay in its block; smaller ones fold into
the background. Discovery is unaffected (it
runs on full evidence), so the tree and the
resolvers are the same at any N; only the
hard blocks get fatter. Default 2000; 500
is the release models' value.
--resolvers N|all How many pairs get a calibrated resolver.
Default 0: none, hard blocks only (no
calibration, so no cell store needed).
all = every pair, the bank-full model.
N = the N thinnest pairs by hard-
block FOOTPRINT -- the hard-block CpGs at
which the two class pseudobulks are both
covered and > 0.5 apart in beta -- the
pairs a sparse cell cannot tell apart from
the type-1 columns alone. Pairs the tree
never parts rank at the top by themselves;
the log warns if N leaves one out. 100 is
the bank-lite release model. The build log
lists every chosen pair with its footprint.
--resolver-cache DIR Reuse resolver blocks across runs: before
building root@A__B, look for DIR/A__B.mrmp
(sanitized tags) and splice it in; else build
it and save it there. A block embodies its
calibration, so bank-full and bank-lite on
the same reference + cells share one cache,
and a rerun rebuilds nothing. The cache is
valid ONLY for one (reference, cell store,
labels, selection) combination -- point
different builds at different directories.
--resolver-stride k/N Populate the cache in parallel: build only
the resolver pairs with ordinal == k mod N
(0-based); an uncached out-of-stride pair is
skipped and the artifact is PARTIAL. Run N
array jobs with the same command and
k = 0..N-1, then one final run without the
stride to assemble everything from cache.
Requires --resolver-cache.
Per-pair calibration
--cell-store F.cg Per-cell fmt3 store (with F.cg.idx) holding
the training cells. With --cell-labels, every
2-class resolver gets its own
(shrink-pseudocnt, gap) by leave-one-cell-out:
selection pools drop the held-out cell (no
leak, and the inner pool depth matches the
final build), the anchored sign feature
scores it, and the LEAST RESTRICTIVE grid
cell within --calib-eps of the best LOO
macro wins -- looser admission keeps more
CpGs for sparse queries. Grid: pseudocount
1-14, gap 0.28-0.53. Validated on fold-0:
13 pairs, mean 2.8 points off the per-pair
oracle, worst 10.7.
--cell-labels F.tsv cell<TAB>class, class names as in REF.cg.
Cells missing from the store index are
skipped with a note.
--calib-eps E Tolerance below the best LOO macro when
picking the loosest grid cell. Default: 0.02.
--calib-threads N Threads per calibration pass. Default: 8.
Binstring: which pattern a CpG belongs to
--beta-threshold B Call beta > B methylated. Default: 0.5.
--call-mindepth N Reads a class needs at a CpG to be CALLED
there; below it the class is ambiguous (see
below). Per class per CpG, not a class
filter. Default: 1. (was --mincov)
--include-all-0 Keep the all-unmethylated pattern.
--include-all-1 Keep the all-methylated pattern.
Ambiguous classes. A class is ambiguous at a CpG when it is below
--call-mindepth or sits exactly at --beta-threshold. Such a class
is imputed, by default from the majority call of the classes that
ARE confident at that CpG. --impute-ambiguous chooses the strategy,
`none` making the CpG PNA instead -- the strict reading, since a
binstring otherwise states something about a class with no
evidence behind it. The rest of this block tunes imputation and
means nothing under `none`.
--impute-ambiguous S How to fill an ambiguous class:
none the CpG is PNA
majority the majority call of the
confidently called classes AT
THAT CpG (more 1s -> 1, more
0s -> 0, exact tie -> 0)
[default]
zero always unmethylated
one always methylated
random an unbiased coin per (CpG,
class), from --impute-seed
A nearest-neighbour strategy -- fill from
the CpGs around it rather than the classes
beside it -- is not implemented.
--impute-seed N Seed for `random` (default 1). The draw is
a hash of the seed, the CpG and the class,
so it does not depend on row order or
thread count. NOT stored in the artifact:
keep the build command.
--min-major-fold F `majority` only: trust the fill just where
the majority sweeps, the larger side being
>= F times the smaller (unanimous always
qualifies). Otherwise the CpG is PNA after
all, because filling from a near-even split
fabricates calls. Default: 10. 0 fills
whatever the majority is.
--max-ambig-frac F PNA a CpG whose ambiguous classes exceed
this fraction, whatever the strategy says.
Default: 1.0 (off).
Feature selection: which CpGs back a pattern
--call-band LO,HI Keep a CpG only if every 0-class reads <= LO
and every 1-class >= HI, on shrunk betas.
Default: 0.30,0.70. A class with reads is
tested even where its digit was imputed: an
intermediate measurement is a real answer,
and a band failure is never tolerated. `none`
(LO == HI == --beta-threshold) filters
nothing, which is the way to see a set before
and after the band. (was --qfilter)
--shrink-pseudocnt A Shrink the per-class beta to (M+A)/(M+U+2A)
for BOTH the admission test and the ranking.
In Bayesian terms this is the posterior mean
under a Beta(A,A) prior: A is the PRIOR
INERTIA, A pseudo-observations on each side
pulling the estimate toward 0.5 until real
votes outweigh them. Default: 3 -- a
unanimous site needs depth 4 to clear the
0.70 leg, one dissenting read pushes that to
~8. Without it a single read scores beta
exactly 0 or 1, passing admission more easily
than well-measured evidence and taking the
maximum rank. 0 restores raw betas.
--feature-mindepth N Reads a class needs at a CpG for its beta to
count as EVIDENCE there. A class below it is
not tested. Default: --call-mindepth; a class
that could not be called is not evidence.
(was --min-cg-depth)
--max-lowdepth-frac F Fraction of classes allowed below
--feature-mindepth before the CpG is dropped
as a feature. Default: 0. This is the only
tolerance: untested classes may be waived,
tested-and-failed ones may not.
(was --max-frac-na)
--delta-mean-top N Keep at most N CpGs per binstring, ranked by
mean class gap. Default: 20000; 0 keeps all.
Output
--name NAME Root name. Default: root.
--force Overwrite OUT.mrmp.
Inspect the result with: inspect --tree OUT.mrmp
methscope mrmp-export — Emit the runtime .cm mask (and pattern / count tables)Usage: methscope mrmp-export [options] IN.mrmp OUT.cm
IN a .mrmp: one set, or a chain of several
OUT.cm per-CpG P1..PK/Pna labels as a YAME format-2 mask
OUT.cm OPTIONAL when --patterns/--counts is given
OUT.cm a multi-set input writes ONE .cm holding every set as
a record, plus OUT.cm.idx naming them -- YAME's own
store convention, so `yame subset -s <set>` works.
This command is an INTERFACE TO YAME, not part of the pipeline.
classify-featurize reads a .mrmp directly and a bundle carries one, so
nothing in methscope consumes a .cm any more; export exists to hand
sets to yame and should speak yame's idiom rather than ours.
--set NAME export just this set of a container, to OUT.cm
--top K rank cut for the mask and --patterns (default 1000)
--patterns TSV top-K patterns: string<tab>P<rank><tab>count (- = stdout)
--counts TSV every pattern (incl. PNA): count<tab>string (- = stdout)
--pna-label NAME background label in the mask (default Pna)
methscope mrmp-pool — Combine MRMP sets and cut them to a shared column budgetUsage: methscope mrmp-pool [options] -o OUT.mrmp IN.mrmp [IN.mrmp ...]
Pool several MRMP sets into one chain and cut them to a
shared pattern budget. Needs no store and no reference -- pattern CpG
counts are already in each artifact -- so re-pooling at a different
budget costs seconds. Plain concatenation is just `cat`; this is the
step that SELECTS.
A set is a set: the inputs may come from any generator and from
DIFFERENT stores, as long as they share a row space. Nothing here
distinguishes a global from a satellite.
Each input is a CHAIN of one or more sets and expands into all of
them under their own names, so a satellite-bearing chain competes for
slots as the many 2-class sets it is, not as one blob. A set carries
its name in its own block (mrmp-build --name), so pooling needs no
NAME: prefix and cannot mislabel anything.
--pooled-top N total pattern budget across every input, ranked by
CpG count (default 1000). Sets COMPETE for these
slots rather than being reserved any, so a set that
cannot field well-covered patterns loses -- the right
verdict, since a pattern too thin to rank is too thin
to trust. The cut PRUNES: the output holds exactly N
patterns, with the CpGs of the rest folded into PNA
and any set winning nothing dropped, so the file IS
what it claims and can go to a model unchanged. To
re-pool at another budget, re-run this on the same
generator outputs -- they are untouched. 0 disables
the cut, leaving a plain `cat` plus the row check.
-o OUT output container
--min-cpgs N drop any pattern carrying fewer than N CpGs,
BEFORE the budget ranks what is left. A pattern is
a feature, and a feature averages the reads a cell
happens to have in it -- so a 1-CpG pattern is one
binarized read, which is noise wearing a column.
The tail is nearly free to cut: on the 7-set Zhou
tree, --min-cpgs 10 drops 82% of the patterns and
1.6% of the CpGs (4,626 of 9,965 are singletons).
Gate and budget compose: --min-cpgs alone prunes to
everything that clears the floor.
--include-all-0 keep patterns no class calls 1
--include-all-1 keep patterns no class calls 0
Both are folded into PNA by default: a
binstring with no class on one side
separates nothing however many CpGs it
carries, so it should not consume budget.
mrmp-build --call-band also excludes them,
but only when a filter is given, and pool
takes blocks from any generator.
methscope classify — Classify a methylome -> labels + confidence
Usage:
methscope classify [options] <model.clfx> <query.cg>
methscope classify --data <in.msfm> [options] <model.clfx>
Purpose:
Predict a label (cell type, sex, ... — whatever the model was trained on)
and a confidence score for each query record, by featurizing the query
against the MRMP reference and running the booster.
Arguments:
<model.clfx> A self-contained bundle of the model + its MRMP (from
`classify-train -o model.clfx` or `bundle`) — the recommended
single-file form.
<query.cg> Query methylome(s); '-' reads a .cg stream from stdin
(cells are then named 1,2,3,... as a stream has no index).
Query records may be YAME format 3 (M/U counts) or
format 6 (universe bit plus binary 0/1 call); positions
outside the format-6 universe are treated as missing.
Options:
-o <out.tsv> Write output to a file instead of stdout.
--framework violation Score directly from a .mrmp with the violation
rule. Give <ref.mrmp> <query.cg>. The rule is an
unfitted function of the artifact and the three
options below.
--call-threshold T Violation beta cutoff. Default: 0.5.
--pattern-weight W Violation weighting: sqrt, log1p, linear, or flat.
Default: sqrt.
--min-patterns N Patterns required on each side. Default: 20.
--top K Patterns to use by rank. Default: 1000.
--probs Append one predicted-probability column per class.
This is unavailable for routing trees because each
node scores a different class subset.
--no-header Suppress the output header.
--threads T Featurize with T workers. Default: 1. Workers seek
indexed query records; streams use the serial path.
--data <in.msfm> Score prebuilt features from classify-featurize.
Omit <query.cg> when this option is used.
-h Show this help message.
Output columns:
cell prediction_label confidence certainty [<class1> ... with --probs]
confidence = P(called class); certainty = 1 - H(p)/log(K), 0 at a
uniform posterior and 1 at a decided one. `violation` has no posterior
and reports a margin instead.
methscope classify-train — Fit a label classifier (xgboost / threshold / logistic)
Usage:
methscope classify-train -l <labels.txt> -o <out.clfx> [options] <query.cg> <ref.cm>
methscope classify-train --data <in.msfm> -o <out.clfx> [options] <ref.cm>
Purpose:
Train a multiclass classifier for any per-record label (cell type, sex,
...; fixed hyperparameters, no grid search) and write a self-describing
model with the class labels embedded.
Arguments:
<query.cg> Training methylome(s), one record per sample.
<ref.cm> MRMP pattern definition. A bundle (.clfx/.updecx) also
works; its MRMP is used. Features are the MRMP states; the
'Pna' NA-background state is always excluded.
NOTE the default framework (xgboost) does not take this form
at all -- it needs --data TRAIN.msfm, which carries its own
chain. This positional serves the unfitted frameworks.
Options:
-l <labels.txt> One label per query record, in query order (required).
-o <out.clfx> Output model path (required). A '.clfx' name writes a
self-contained bundle (model + MRMP) that `classify` can
run directly; a plain '.ubj' writes just the loose booster.
The bundled MRMP is TRIMMED to exactly the patterns used
(others folded into 'Pna'), and `classify` uses that same set.
-p <npattern> Use the first N patterns, in the artifact's own order and
after the 'Pna' backgrounds have been excluded. For an
auto MRMP that order
is recurrence rank (P1,P2,...), so N means the N most
recurrent; for curated named markers it is definition
order, and leaving it unset is what you want.
Note this cuts across a FUSED multi-set artifact rather
than within each set, so a value below the total will
drop whole trailing sets. Cut with mrmp-pool --pooled-top
instead if you want a per-set budget.
Default: every non-'Pna' state.
--framework <f> Model framework (default: xgboost):
xgboost gradient-boosted trees (multiclass).
logistic binary L2-regularized logistic regression;
needs a .clfx out.
The violation rule is unfitted, so it is a scoring mode:
`classify --framework violation ref.mrmp query.cg`.
-n <nrounds> Boosting rounds, xgboost only (default: round(sqrt(n_cells))).
--eval-every <N> Report training mlogloss every N rounds, as a rolling
window of the last 5. Default: nrounds/20, about 5%%
overhead -- each evaluation predicts over the whole
training matrix. 0 disables, 1 evaluates every round.
--threads <N> xgboost nthread (default: the CPUs this process may
actually use). Left to xgboost it sizes its pool from
the machine, which on a shared batch node is not what
the job was allocated -- so the same matrix and the
same rounds have timed 76 s on one node and 46 min on
another. Pinning it makes runs comparable.
--max-depth <d> Cap tree depth (xgboost default 6). Lower it when the
training set is redundant -- e.g. one pseudobulk repeated
across a coverage ladder -- since the repeats inflate
split confidence and let one feature decide a class.
--min-child-weight <w> Minimum child weight (xgboost default 1).
--colsample <f> Per-tree feature subsample fraction, 0-1 (default 1).
--data <in.msfm> Train from a prebuilt feature artifact (classify-featurize)
instead of featurizing <query.cg> here. Only <ref.cm> is then
positional, and labels come from the artifact unless -l is
given. Featurization is single-threaded, so this is how to
train repeatedly on the same cells without repeating it.
-h Show this help message.
methscope classify-featurize — Prebuild the .msfm feature matrix (parallel, reusable)
Usage:
methscope classify-featurize [options] -o <out.msfm> <query.cg> <ref.mrmp>
methscope classify-featurize [options] -o <out.msfm> <query.cg> <ref.cm> [ref2.cm ...]
Purpose:
Summarize each query record against the MRMP once and store the result,
so training and scoring never repeat the featurization. `classify-train`
and `classify` both accept the artifact with --data.
--threads T partitions the records across workers, each seeking its own
through the .cg index.
Arguments:
<query.cg> Query methylome(s), one record per sample or cell.
<ref.mrmp> A pooled MRMP. A .mrmp is a chain, so ONE argument carries
every set, and set names come from the blocks -- no loose
.cm files and no order.txt to keep in step with it. This is
the normal input.
<ref.cm> Exported masks, one per set; a bundle also works. Equivalent,
but the caller owns keeping them together and in order.
Options:
-o <out> Output .msfm (required).
-l <labels> One label per query record, in query order. Embedded in the
artifact, so a downstream train cannot be handed a
mismatched label file. Omit for an unlabeled query.
--sample LIST Comma list of target covered-CpG counts, one replicate each
(0 = keep native coverage). Every cell is inflated ONCE and
all replicates are drawn from that copy, so a coverage ladder
costs one pass rather than one pass per level.
--reps N Replicates per --sample level (default 1).
--patterns N Feature patterns P1..PN. Default: every non-Pna mask state.
-b, --binarize Draw one read per sampled CpG, so the call is 0/1 not a
fraction -- what a real sparse methylome delivers.
--continuous-features
Emit each pattern's mean beta instead of a call. By DEFAULT
a pattern beta is cut at 0.5, so a feature says "this cell
is methylated across this pattern" rather than "this cell
reads 0.918". The call is absolute, so it means the same
thing on a cohort this reference never saw -- which a raw
beta does not once a global shift, mitotic or otherwise,
moves every value at once.
continuous either way; they are not a contrast.
--set NAME featurize only this set of a chain, so one node of an
mrmp-tree can be scored straight out of the tree it
lives in -- a block inside a chain is byte-identical to
the standalone .mrmp, so nothing need be kept beside it.
--satellite-contrast off|add|replace (default off)
A 2-class satellite's two patterns have opposite polarity,
so "is P1 above P2" asks the pair's question RELATIVELY --
anything shifting both patterns together cancels, which is
the violation rule's argmin cancellation as one feature.
replace: emit the contrast INSTEAD of the two patterns.
add: emit it alongside them -- but note the model then has
no reason to prefer it, since in training the absolute
columns work just as well; that is why replace is the one
that removes the fragility rather than offering an
alternative to it.
--thresh-pattern
Cut at each pattern's OWN midpoint (stored in the .mrmp)
rather than 0.5. Separates a close pair better, since
inside a satellite both classes can sit above 0.5 -- but it
is fitted to this reference and travels worse.
--seed S Sampling seed (default 1). The draw is a pure function of it
and is NOT affected by --threads.
--side-floor N Pairwise features only: a side backed by fewer than N
observed CpGs contributes the neutral 0.5 instead of a
one-read coin flip; both sides under N -> NA. Default 3.
The 0.30/0.70 admission band is what makes 0.5 a
calibrated anchor between the reference poles. Score
with the same value the model was trained with.
--threads T Worker threads (default 1). Cells are partitioned across
workers, each seeking its own records via the .cg index.
--counts <N> Require N measured CpGs behind a pattern's beta; below that
the beta is recorded MISSING rather than kept. A beta from a
single CpG can only be 0 or 1, so after the 0.5 call it is
always maximally confident and never borderline -- yet it
carries the same weight as one backed by hundreds of CpGs.
In mouse single cells 46.5% of observed pattern-betas rest
on one or two CpGs. Default 1 (keep every observed pattern).
-h, --help Show this help message.
Notes:
Every pattern is stored, never a -p prefix, so one artifact serves any
`classify-train --patterns N`. Betas are u16 fixed point (code/65534,
65535 = NA), the same encoding the upscale msur uses for truth.
methscope mrmp-summary — Per-pattern mean methylation as TSV (the MRMP average)Usage: methscope mrmp-summary [options] <query.cg> <ref.mrmp>
Purpose:
Per-pattern mean methylation for every query record: one pass over the
query, every set of the chain at once. This is the feature matrix a
classifier would train on, as text -- an embedding for clustering, a
projection, or a heatmap.
Arguments:
<query.cg> Query methylome(s), one record per sample or cell.
<ref.mrmp> A .mrmp (one set or a chain); a bundle carrying one works.
Options:
-o <out> Write here instead of stdout.
--wide One row per record, one column per pattern (the embedding
matrix). Default is long: sample, set, pattern, beta.
--threads T Records per worker, each seeking through the .cg index.
--min-cpgs N A pattern seen at fewer than N CpGs in a record reads NA
rather than a confident mean (default 1).
methscope deconv-build-ref — Pack a cell-type store into the .msdref deconvolution referenceUsage:
methscope deconv-build-ref -o <out.msdref> [options] <celltypes.cg>
Purpose:
Pack a per-cell-type M/U store into the uint16 M/U reference that `deconv`
rebuilds its pattern set from, keeping only the CpGs that could ever
contribute to some class subset.
A class is CALLABLE at a CpG when it is covered and outside the ambiguous
band (beta <= LO or beta >= HI). A row is kept only when at least one
callable class is unmethylated AND at least one is methylated -- otherwise
no subset of classes can form a non-constant admissible binstring there,
whatever the query, so the row is dead weight. The test is a necessary
condition under EVERY scope, so dropping a row cannot remove a pattern a
later rebuild would have found.
Arguments:
<celltypes.cg> One record per cell type, YAME format 3 (M/U) or format 6
(universe 0/1), with a matching <celltypes.cg>.idx whose
index order IS the binstring digit order. Format-6
positions outside the universe are treated as missing.
Options:
-o <out.msdref> Write the reference here. Required.
--qfilter LO,HI The admission band baked into the row test. Default:
0.30,0.70. Must match the solver's.
--beta-threshold B Call a class methylated above B. Default: 0.5.
--call-mindepth N A class is covered at N reads or more. Default: 1.
(was --mincov, still accepted.)
--confusion FILE Embed a validation confusion matrix, which is what
`deconv --group-threshold` reads to decide which classes
are reported under one label. Long-form TSV,
true<TAB>predicted<TAB>value, unnamed pairs zero, rows
normalised on write; the first '#' line is kept as the
provenance note. Measuring it is NOT this command's job
-- run the validation, write the TSV, pass it here. With
this flag the artifact is version 3; without it, 2.
--keep-all Skip the never-useful row test, keeping every row so a
consumer can re-apply any band in memory. Costs 1.94 GB
for 33 classes x 29.4M rows.
--force Overwrite an existing output.
-h Show this help message.
methscope deconv — Estimate cell-type proportions, rebuilding the MRMP per queryUsage:
methscope deconv [options] <ref.msdref> <query.cg>
Purpose:
Estimate the cell-type composition of each query record by non-negative
least squares against a .msdref reference. The pattern set is rebuilt per
query from the CpGs that query actually measured, so no pattern budget is
baked in and a sparse query is scored on its own evidence.
The fit is ADAPTIVE. Round 1 solves over every class. Each round after it
keeps the classes holding at least --mass-floor, then pulls back in any
class the panel separates from a survivor by --max-segregating measured
CpGs or fewer -- repeating until nothing new enters, so a scope is a union
of connected components rather than a mass cutoff. The pattern set is then
rebuilt over that scope and refit: dropping classes relaxes the admission
conjunction, so a narrower scope admits CpGs the full one could not. The
scope only ever narrows, and rounds stop once it stops changing or at
--max-round. --no-adaptive keeps the round-1 fit over every class.
Arguments:
<ref.msdref> Deconvolution reference from `deconv-build-ref`. Its index
order is the binstring digit order.
<query.cg> Query methylome(s), YAME format 3 (M/U counts) or format 6
(universe bit plus binary call). Record names come from
<query.cg>.idx; without one they are 1,2,3,...
Options:
-o <out.tsv> Write output to a file instead of stdout.
--wide Emit the full composition matrix instead: one row
per record, one column per reference class, in the
reference's index order. Every class appears, so
--min-frac does not apply.
--report Emit a readable one-line summary per record --
"cell: Class 69.6%%; Class 30.4%%" -- instead of a
TSV.
--group-threshold F Report confusable classes under one joined label,
"CA3/DG-po", summing their fractions. A class
handing F or more of its mass to another in the
reference's confusion trailer joins it; groups are
the connected components of that relation. Default:
0, off. Needs a reference built with
`deconv-build-ref --confusion`, and applies to every
output shape but --wide, which stays per-class so
the full-resolution answer is never lost.
F IS A TUNING PARAMETER, NOT AN ERROR RATE: the
matrix is pooled over sparsity levels and measured
on the reference atlas's own held-out cells, so it
names pairs that were not separable THERE and is a
lower bound on what another protocol will confuse.
--min-frac F Hide classes below this fraction. Default: 0.005.
Display only: it never changes what is fitted. 0
shows every non-zero class.
--threads N Deconvolve N records in parallel. Default: 1.
The reference is shared, so
memory grows by the per-thread workspace (~16 MB),
not linearly. Needs a <query.cg>.idx to seek by; -v
and the dump options force a single thread.
--max-round N Cap on rebuild rounds. Default: 8. Rounds stop early
once the class set stops changing.
-v Report per-record panel size and measured-row count.
-h Show this help message.
Tuning:
Every default below was chosen on the 200-mixture benchmarks, where moving it
scored worse or made no measurable difference. Change one to reproduce an
experiment, not to improve an answer.
--min-cpg N Drop a pattern seen on fewer than N measured CpGs.
Default: 5.
--mass-floor F A class seeds the scope at this mass or above.
Default: 0.005. Not 0, because NNLS hands out
arbitrary exact zeros among collinear columns.
--max-segregating N Two classes the panel separates by this many
measured CpGs or fewer are inseparable for this
query, so both enter the scope when either carries
mass. Absolute, because an absolute CpG count is
what reaches the solve. Default: 200.
--weight-exponent E Row weight is n^E, so influence is n^2E. Default:
0.5, the precision weight. One value per round is
accepted as E1,E2,... with the last repeating.
--delta-mean-top N Per binstring, keep only the N measured CpGs with
the cleanest class gap, bounding how much one
pattern can weigh. Default: 0, keep all.
--global-ref Compute each pattern's reference profile over every
reference CpG carrying that binstring, not only the
measured ones. Costs a second pass over kept rows.
--rescue-below N A pair the rebuilt panel separates by fewer than N
measured CpGs is weak; the panel is then rebuilt
admitting CpGs that split a weak pair under a wider
rule, every class still getting a digit. Default:
2000. 0 turns the rescue off, which was the worst
configuration measured.
--rescue-gap G Rescue on |beta_a - beta_b| >= G instead of on the
band, exempting the pair's own classes. Default:
0.20. 0 uses the band rule.
--rescue-target N Fit each weak pair its own gap, loosening until it
has N measured CpGs or hits the floor. Default: 0,
use the flat --rescue-gap.
--rescue-floor G Never relax a pair past this gap. Default: 0.15.
--rescue-min-depth N Reads required on both classes of a pair before its
gap is believed. Default: 10. Needs a v2 .msdref,
which carries M/U rather than a bare beta.
--rescue-qfilter LO,HI The wider admission band. Default: 0.40,0.60.
Diagnostics:
--no-adaptive Stop after round 1: fit the frozen global panel
over every class, skipping the per-query rebuild.
It is the CONTROL that says what narrowing is
worth, and the one way to score a cohort on a
single panel. Not a tuning knob: it measured 9-12x
worse TVD than the default on the 41-class mouse
benchmark, at every coverage rung.
--panel-out F Dump the rebuilt panel: binstring, measured CpGs and
observed beta, one row per pattern per record.
--scope-out F Write the settled scope, 1/0 per class per record.
--design-out F Dump the exact system handed to the solve -- one row
per pattern, its weight, the observed value and every
scope class's profile -- so the fit can be redone
outside.
--force-scope C[,C..] Use exactly these classes, skipping discovery. It
answers whether a narrow panel is weak because the
scope is wrong or because it is narrow.
--pair-count A,B Report how many measured CpGs segregate classes A and
B on their own, and how many survive the all-classes
admission conjunction.
--eval-x SPEC Print the objective at hand-given compositions beside
the fitted one, as "Cls=frac,Cls=frac;Cls=frac,...".
Answers whether a wrong answer means the solver
missed the optimum or the panel scores the wrong
composition better. Needs -v.
Output:
By default a tidy TSV -- cell, class, fraction -- carrying only the classes
at or above --min-frac, largest first. --wide gives the full matrix and
--report a one-line summary per record.
A record's fractions are proportions of the mass the reference could
explain: the fit is renormalized to sum to 1, so a query whose true
composition lies outside the reference still sums to 1, spread over its
nearest classes. --report is the one shape that does NOT rescale -- what
--min-frac hides is named as "Others", so a printed percentage always
means the same thing.
methscope upscale — Impute genome-wide CpG methylation from a sparse methylome
Usage:
methscope upscale [options] <model.updec|.updecx> [input]
Purpose:
Impute CpG methylation from a sparse methylome. Unified UPDEC2 models
predict the whole genome; legacy UPDEC1 block models remain readable.
Both inference paths are pure C and require neither CUDA nor BLAS.
Arguments:
<model> A bare decoder (.updec/.updec2) OR a bundle with the MRMP attached
(.updecx, from `bundle`). The form selects the input below.
[input] For .updec : a feature TSV (one row per sample, n_in numeric cols).
For .updecx: a query .cg, featurized internally against the bundled
MRMP. '-' or omitted reads from stdin; for .updec a header row is
auto-detected and empty/NA fields are imputed with the model means.
Options:
-o <out> Write output to a file instead of stdout. A .cg is binary, so
without -o it is refused on a terminal; a pipe or redirect is fine.
--binary Threshold predictions at 0.5 and write 0/1 calls (format 6).
Lossy -- the default format 4 keeps the fraction itself.
--probs Emit per-CpG probabilities as TSV instead of a .cg.
-h Show this help message.
Output:
A YAME .cg (format 4: continuous methylation fraction), one record per input
sample. UPDEC2 output is always in whole-genome CpG order. For legacy UPDEC1,
if the bundle carries an outcpg.cm (the imputed-CpG locations), the .cg spans
the whole genome with the block's CpGs predicted (no NA) and the rest NA;
otherwise a dense block of n_out CpGs. With --binary, format 6 of 0/1 calls;
with --probs, a TSV of per-CpG probabilities (one row per sample).
methscope upscale-featurize — Build the MSURAW2/3 training msur from a truth .cgUsage: methscope upscale-featurize [options] TRUTH.cg IN.mrmp OUT.msur
Create a compact exact-YAME sampling msur for global upscale training.
The original TRUTH.cg remains the truth store and is never copied.
TRUTH.cg continuous format-3 YAME .cg truth store
IN.mrmp MRMPIDX1 artifact (preferred) or an exported .cm
OUT.msur output msur
Options:
--reps N deterministic simulations per --sample level (default 100)
--sample N[,N..] CpGs retained per cell/replicate (default 29000). Give a
comma list to mix coverages: --reps applies to EACH level,
so --reps 100 --sample 14701,29000,147009 writes 300
replicates. Each level stores exactly its own record size.
--sample-logrange MIN,MAX
continuous ladder instead of --sample: --reps is the TOTAL
replicate count, each getting its own sample size spread
in log space over [MIN,MAX]
--sample-skew A bend that draw (default 0.5). A<1 puts more replicates at
the DENSE end, where finer differences must be resolved;
A=1 is plain log-uniform
--binarize one read per observed CpG: replace each sampled beta with a
Bernoulli(beta) draw, as `yame dsample -b` does
--threads N inflate and scan N samples at once (default 1). Needs a
<store>.idx, since each worker opens its own handle and
seeks to its sample. Output is byte-identical at any N.
--manifest PATH write provenance TSV (default OUT.msur.tsv)
-h, --help show this help
EVERY pattern in IN.mrmp becomes a column -- there is no width knob here.
Pattern SELECTION happens when the MRMP is made: `mrmp-build --top N`
for a single set, or `mrmp-pool --pooled-top N` when several sets must
compete for one budget. Both rank on CpG count and fold the rest into PNA.
A width flag here was worse than redundant: it RESERVED slots per set, so a
chain handed its root the first N columns whether or not the root's Nth
pattern outweighed a child's 1st. Cut first, then featurize what survives.
One msur serves many models. What is stored per (cell, replicate) is the raw
summary -- beta and covered-CpG count for every pattern, plus the observed
set -- NOT the encoder input. Which projection the encoder sees is chosen at
TRAINING time, so one msur covers every `upscale-train --features` mode
(missing / count / beta / scalar). Only the *simulation* is fixed here: the
cells, the replicate count, the coverage ladder and --binarize.
methscope upscale-set-units — Build the MSUIDX1 processing-unit index from the reference storeUsage: methscope upscale-set-units [options] REF.cg OUT.msui
Build the whole-genome processing-unit index used by UPDEC2. Each CpG's
binstring is derived from the reference store itself. Memberships are
size-ranked and never split; PNA CpGs are implicit singleton memberships
packed after all real memberships.
Units are an OUTPUT partition, so this reads the STORE and not a .mrmp.
A built .mrmp is a SELECTION -- mrmp-build drops the constant binstrings
because a CpG every class calls the same way separates nothing, which is
right for a classifier and fatal here: those CpGs are 54% of the genome
and they still need reconstructing. Reading the store keeps them, so
every CpG has a real membership and coverage is 100% by construction.
REF.cg reference store
OUT.msui output MSUIDX1 index
--call-mindepth N min per-class coverage (default 1)
(was --mincov, still accepted.)
--beta-threshold B call a class methylated above B (default 0.5).
Must match mrmp-build's.
--unit-cpgs N target CpGs per unit (default 16384)
-h, --help show this help
methscope upscale-train — Train the upscale decoder (CPU, CUDA optional)Usage: methscope upscale-train -i DATA.msur --units UNITS.msui
--mrmp TOP1000.mrmp -o MODEL.updecx --work-dir DIR [options]
Train whole-genome UPDEC2 processing units. CPU by default; --device
selects a CUDA device in a CUDA build. Each MRMP contributes
beta plus log1p(observed-CpG count); count zero represents missingness.
Required:
-i, --data PATH embedded-truth MSURAW2/3 training msur
--units PATH whole-genome MSUIDX1 processing-unit index
(from upscale-set-units)
--mrmp PATH MRMPIDX1 artifact (preferred) or exported .cm;
the runtime mask is bundled into the output
-o PATH self-contained output .updecx
--work-dir DIR resumable unit checkpoints and bare UPDEC2
Architecture:
--features MODE beta, count, missing, or scalar (default count).
scalar = beta + ONE log1p(total covered CpGs), so
input is P+1 wide instead of 2P.
Both of these are projections of what the msur
already stores, so one msur trains every mode and
width -- retrain to vary them rather than paying
for another upscale-featurize
--pure-bottleneck N one-membership unit dimension (default 16)
--mixed-bottleneck N mixed/PNA unit dimension (default 32)
--mixed-mode MODE factor or direct (default factor)
--activation MODE linear or leaky (default linear)
Optimization:
--min-steps N earliest stopping point (default 2000)
--max-steps N maximum updates per unit (default 8000)
--eval-every N validation interval (default 200)
--patience N non-improving checks (default 8)
--batch N target CpGs per update (default 8192)
--eval-rows N fixed validation rows per unit (default 8)
--learning-rate X AdamW rate (default 0.001)
--weight-decay X AdamW decay (default 0.00001)
--seed N deterministic seed (default 1)
--split FILE curated cell split, rows <cell_index>TAB<train|
val|test> (default: seeded 70/15/15 shuffle)
--device N|cpu CUDA device (default 0), or cpu to force the
portable backend
--threads N CPU backend worker threads (default 1)
--pilot-units FILE train listed unit IDs only. The bundle is still
assembled, holding just those units, so it can be
run with `upscale` -- CpGs outside them come back
NA. That is what makes a locus-sized experiment
possible without training all 700 units.
--force replace final output and manifest
--dry-run validate options and print configuration
-h, --help show this help
methscope bundle — Wrap a model + its MRMP into a self-contained bundle
Usage:
methscope bundle -m <ref.mrmp> -o <out> <model>
Purpose:
Bundle a model with the MRMP definition it needs into one self-contained
file, so a query .cg can be featurized + run without a separate .mrmp.
Works for any model; by convention the output extension names the ROLE:
a classifier -> model.clfx (classify)
an upscale decoder -> model.updecx (upscale)
Both are the same MSBNDL1 container and are detected by magic, not by
name; '.ubjx' is the former classifier extension and is still accepted.
Options:
-m <ref.mrmp> the MRMP definition (a YAME .cm) to bundle (required).
-k <kind> framework mark to record (xgboost/violation/logistic);
required for a classifier .clfx that `classify` will run (it
rejects an unmarked bundle). Also re-stamps an existing model.
-O <outcpg.cm> output-CpG locations (upscale only): a YAME mask marking the
CpGs the model imputes. With it, `upscale` writes a full-genome
.cg; without it, a dense block .cg.
-o <out> output bundle path (required).
-h Show this help message.
methscope relabel — Rename a class label in a trained model, no retrainingUsage:
methscope relabel [options] OLD=NEW <in.clfx> -o <out.clfx>
Purpose:
Rename a class label inside a trained model. The label lives in one XGBoost
booster attribute ("methscope_labels", comma-separated in class-index
order), so this rewrites that string and nothing else -- every prediction is
bit-identical and only the emitted name changes.
A class name is a claim about what the reference CONTAINS, and it can be
wrong while the model is right. Retraining to correct a name would perturb
numbers that were never at fault.
On a TREE bundle every node carrying the class has its own copy of the label
list, so all of them are walked; renaming the root alone would leave the
leaves disagreeing with it.
Arguments:
OLD=NEW The existing label and its replacement.
<in.clfx> The model to read. It is never modified.
Options:
-o <out.clfx> Write the relabelled model here. Required.
--force Overwrite an existing output.
-h Show this help message.
Notes:
The bundled MRMP's reference sample names are left ALONE on purpose: they
record what was built, not what it is.
Example:
methscope relabel -o hg38_celltype_v2.clfx \
'Macrophage=Macrophage.(Monocyte.Derived)' hg38_celltype_full.clfx
methscope unbundle — Unpack a bundle into its model, MRMP, and outcpg mask
Usage:
methscope unbundle [-o <model_out>] [--mrmp <mrmp_out>] <bundle>
Purpose:
Unpack a bundle (.clfx / .updecx) into its inner model, its MRMP,
and (if present) the output-CpG mask.
With no -o/--mrmp, output names are derived from the bundle path (the
original .mrmp filename is not stored): drop the 'x' from the extension for
the model, and add sibling suffixes for the rest --
foo.clfx -> foo.model + foo.mrmp (+ foo.outcpg.cm if present)
foo.updecx -> foo.updec + foo.mrmp
Options:
-o <model_out> write the inner model here (default: derived, see above).
--mrmp <mrmp_out> write the bundled MRMP (.cm) here (default: derived).
-h Show this help message.
methscope inspect — Describe any artifact: bundle, .mrmp, .msui, .msur, or .msfm
Usage:
methscope inspect <FILE>
Purpose:
Describe any methscope artifact without running it. The format is detected
from its magic:
.clfx/.updecx bundle: kind mark, section layout, model breakdown
.mrmp MRMPIDX1 pattern set: dimensions, binstring parameters, top ranks
.msui MSUIDX1 processing-unit index: units, memberships, CpG split
.msur MSURAW2/3 training msur: cells, replicates, embedded truth
.msfm MSFMAT1 feature matrix: records, patterns, labels, coverage
.msdref MSDREF1 deconvolution reference: cell types and CpG rows
Options:
--tree render an mrmp-tree: its nodes, and which class each one
decides. Works on a chain or a tree bundle; the structure
is DERIVED, so there is no tree file to keep in step.
--patterns .mrmp only: list the top-ranked patterns
--top K .mrmp only: how many to list (default 20)
-h Show this help message.
methscope mliftover — Re-index a model onto an array platform's row spaceUsage:
methscope mliftover --to <platform> [options] <model> -o <out>
Purpose:
Re-index a model onto an array platform's row space, so array data is
scored in its own space by a model built genome-wide: "lift the model,
use the data as it is". The model itself -- boosters, weights, the
deconvolution trailer -- is copied untouched; only the CpG index moves.
Reads any .clfx (the bundled .mrmp chain or .cm mask), a loose .mrmp or
.cm, or a .msdref deconvolution reference.
It refuses an upscaler (.updecx): its output is the genome, so it does
not lift. Carry the DATA the other way instead --
`sesame mliftover --to hg38 --simulated-depth 100 beta.cg beta.hg38.cg`
-- and run the original model on that.
Options:
--to <platform> Target platform (MSA, EPIC, HM450, ...). Its ordering
(<platform>.ordering.tsv.gz: the probe IDs, in row
order), its coordinate table (<platform>.<genome>
.coord.tsv.gz) and the genome's cpg_nocontig.cr come
from the store; `methscope fetch` names all three.
--genome <g> The model's genome (default hg38).
--coord <file> Coordinate table to use instead of the store's.
--ordering <file> Ordering to use instead of the store's.
--cr <file> Row-coordinate track to use instead of the store's.
-o <out> Output path (required).
--pna-label <s> The background key of a .cm mask (default Pna).
--min-retained N Refuse if any pattern keeps fewer than N CpGs on the
platform. Default 0: report and continue -- an emptied
pattern stays a column and reads as missing.
--force Overwrite an existing output.
-h Show this help message.
Output:
The same kind of artifact, in the platform's row space (one row per
probe, ordering order), with its reference named after the platform.
Only cg probes are joined, decided by Probe_ID: an rs, nv or ch probe
that maps onto a CpG site carries a genotype or a non-CpG call, not
methylation, so it takes the background state whatever its coordinate
says. Those are counted separately from cg probes with no CpG row.
Per-set retained-CpG counts go to stderr: a feature that keeps 4% of
its CpGs is still an estimate of the same mean, but the reader should
see that number before trusting a call made on it.