Bit-packing for DNA methylation. Arrays and whole genomes in one binary format family, where analysis is bitwise — and so stays fast across a thousandfold of scale, from 28K probes to 29M CpGs.
No reference data is bundled: no coordinates, no feature sets, no example
methylomes. yame fetch opens a browser over the
catalogue — species, then platform or genome build, then its files.
yame fetch
YAME_DATA_HOME (default): ~/.local/share/yame
TARGET KIND ROWS TAG IN STORE
▪ HUMAN
❯ ▾ [*] hg38 genome 29,401,795 mixed ▰▱▱▱▱ 5/47
├── ● cpg_nocontig.cr v2 28.9 MB
├── ▸ [ ] KYCG sets v2 ▱▱▱▱▱ 0/32
├── ▸ ✓ data sets v3 ▰▰▰▰▰ 5/5
────────────────────────────────────────
hg38 genome build
Arrows move, / filters by name, source,
collection or title, x takes a row (a folder takes
everything under it) and f fetches. What the store
already holds shows as present and cannot be re-taken.
d points the browser at another store without
leaving it.
Opened over a store that holds files from an earlier release than this
build pins, the browser first lists them under their directories and asks
whether to replace them now; declined, it marks each such file
stale beside its size and offers it again.
Files land in a digest-verified store shared by every tool in the suite
— $YAME_DATA_HOME, default
~/.local/share/yame — so what you fetch here
is what KnowYourCG and SeSAMe
find. Everything else on this page runs against it.
The store is laid out the way the tree reads. What you see under a
name is what lands under that name on disk, so hg38/data
in the browser is $YAME_DATA_HOME/hg38/data, and a
genome's knowledgebase simply nests inside it. Which upstream repo published
a directory is a column in fetch -l, not a path
component — one name, whether you are browsing, fetching or looking at
the disk.
What it carries are the row spaces everything else is measured in. Each length is distinct, so the count is the identity — which is what lets -R and -m take a name instead of a path.
row space rows row space rows ───────────────────────────────────────────── hg38 29,401,795 HM450 486,427 mm10 21,867,837 MM285 287,692 mm39 21,889,506 MSA 284,309 EPICv2 937,690 Mammal40 38,607 EPIC 866,553 HM27 27,722
Piped or redirected, fetch prints the catalogue
as TSV instead, so a script never blocks on a keystroke. A folder is
confirmed before it transfers. On a terminal that is the browser, opened
on the folder with exactly the files it would fetch already checked
(f to fetch, q to leave);
off one it means -y, since a short name can reach a
lot — but
a name that picks out one file is its own confirmation. A
filtered directory is still a directory, though, even when the filter
leaves exactly one file, so the two -g lines below
carry -y; with those, every fetch line on this page
runs in a script as it stands.
A name is a store path. A directory name is that directory's
own files; a subdirectory is named to add it, or -R
takes everything beneath. One file needs no -y.
These are the forms, and the only spellings there are:
yame fetch -l # the catalogue as TSV, with what the store holds
yame fetch -n hg38/models # the plan for that directory; downloads nothing
yame fetch -n -R hg38 # -R: hg38 and every directory beneath it
yame fetch -y hg38/data # a directory, in a script
yame fetch -y EPICv2 EPICv2/KYCG # two directories; names compose
yame fetch hg38/data/human_hg38_test.cg # one file: its own confirmation
yame fetch -y -g CGI hg38/KYCG # a directory, filtered
Everything else — -c for the current
directory, -f to replace a stale file,
-d for another store: yame fetch
-h.
Almost everything in practice is one of three kinds of data, and they land in two of the seven formats. One table covers how each is stored and what pack reads to build it.
the data format on disk pack one line per CpG ──────────────────────────────────────────────────────────────────────────────── methylation, M/U counts 3 M/U counts 1-8 B/site -fm M ⇥ U methylation, binarized 6 set+universe 2 bits/row -fd call ⇥ covered a CpG set 6 set+universe 2 bits/row -fd in set ⇥ in universe
pack is the only way in. Input is one row per CpG already in the reference's order — nothing is sorted or matched for you. Columns are tab-separated, and under -fd the first column must be exactly 1 to count as in-set while the second is any nonzero int.
yame fetch -c hg38/data/human_hg38_test.cg
yame fetch hg38/cpg_nocontig.cr # the coordinates, into the store
yame unpack -R hg38 human_hg38_test.cg | awk '$4!=2' | head -5
chr1 659296 659298 1
chr1 779099 779101 0
chr1 795737 795739 1
chr1 933194 933196 0
chr1 958160 958162 1
Format 6 prints one column by default; -f -1 gives set and universe separately, which is what pack -f6 reads back:
yame unpack -f -1 human_hg38_test.cg | awk '$1!="NA"' | head -3
1 1
0 1
1 1
yame unpack -f -1 human_hg38_test.cg > rt.txt
yame pack -f6 rt.txt rt.cg # and back again
Verified byte-identical on the round trip.
Rows are samples, columns are CpGs or windows. The reference is inferred from the row count, so -r alone is enough.
yame fetch -c hg38/data/human_hg38_celltypes.cg
yame fetch hg38/cpg_nocontig.cr # the coordinates, into the store
yame hprint -r chr20:30000000-30001000 human_hg38_celltypes.cg
chr20:30000011-30000862 (10 CpGs)
|.........
30000011
Oligodendrocyte █...██..░█
Pancreas-Beta █.....░█..
Blood-NK ....██░█░█
Blood-Monocytes ....███.░█
-w caps the width — wider regions are window-averaged rather than truncated — and -R alone gives the whole-genome view. With neither, every row is dumped as one glyph, which is how a whole reconstruction is compared against a whole truth: M/U counts and fractions render as H/M/L (-g for deciles), binary calls as 1/0/2.
Records concatenate, so a cohort is one file. The .idx beside it maps sample names to offsets, which is what makes them addressable by name.
yame fetch -c hg38/data/human_hg38_celltypes.cg
yame info human_hg38_celltypes.cg | cut -f2-7
Sample NSample Nrow Format UnitBytes Keys
Oligodendrocyte 4 29401795 6 2 NA
Pancreas-Beta 4 29401795 6 2 NA
Blood-NK 4 29401795 6 2 NA
Blood-Monocytes 4 29401795 6 2 NA
yame chunk -s 500000 human_hg38_celltypes.cg chunks/ # fixed row blocks
yame subset -o two.cg human_hg38_celltypes.cg Blood-NK Blood-Monocytes
yame info human_hg38_celltypes.cg | cut -f2 | tail -n +2 > names.txt
yame split -s names.txt human_hg38_celltypes.cg out_ # one file each
yame index -s names.txt human_hg38_celltypes.cg # rebuild the index
Sample Nrow Format
Blood-NK 29401795 6
Blood-Monocytes 29401795 6
info -1 reports the first record of each file — the cheap way to see what a directory holds:
File Sample Nrow Format
human_hg38_40_celltypes_chr20.cg GSM5652176_Adipocytes-Z000000T7 773477 3
human_hg38_celltypes.cg Oligodendrocyte 29401795 6
human_hg38_immune_mixture.cg mac70_mono30_2pow22 29401795 3
human_hg38_test.cg 1 29401795 6
human_hg38_test.truth.cg 1 29401795 6
With no mask, summary describes the file. With one it gives the overlap table — universe, query, mask, overlap, log2 odds ratio, beta, depth. The mask is named, not spelled out: the query's row count says which row space to resolve it in.
yame fetch -c hg38/data/human_hg38_test.cg
yame fetch -y -g CGI hg38/KYCG # the mask, into the store
yame summary -m CGI human_hg38_test.cg
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
human_hg38_test.cg 1 CGI.20220904.cm OpenSea 29000 23857 23534 21149 3.17 0.899 NA
human_hg38_test.cg 1 CGI.20220904.cm Shelf 29000 23857 1224 1076 0.67 0.879 NA
human_hg38_test.cg 1 CGI.20220904.cm Shore 29000 23857 2074 1263 -1.74 0.609 NA
human_hg38_test.cg 1 CGI.20220904.cm Island 29000 23857 2168 369 -5.10 0.170 NA
Read the last row first: CpG islands are depleted in this methylome (log2 odds −5.10) and what little overlaps is unmethylated (beta 0.170), against 0.899 in the open sea. That contrast is the whole point of the table, and it is one -m away.
Two spellings, two namespaces. fetch takes
the BROWSER path, hg38/KYCG/CGI, because that is
where the file goes in the store. -m takes the BARE
name, CGI, and resolves it inside the query's own row
space. The two are not interchangeable: -m
hg38/KYCG/CGI fails, because the slashes make it look for a file of
that literal name. The block above uses both, one line apart, and that is the
asymmetry most likely to cost you an afternoon.
Many masks cost little more than one. The query is read once and every
mask is measured against it 64 rows at a time, so a whole knowledgebase is one
pass rather than one per set: the 1359-set TFBS collection against a 29.4M-row
methylome takes 6 s where it used to take 193, with the same numbers and the
same 45 MB. YAME_SUMMARY_PATH=1 reports on stderr how
many masks took that path, since the output cannot tell you. With many
records against many masks, -I inverts the mask file
into an index keyed by row instead: a build worth about four record walks and
855 MB on TFBS, then almost nothing per record. The walk stays the default
because the break-even, about four records, moves with the machine, and the
wrong index costs a gigabyte where the wrong walk costs seconds.
A format 2 mask yields one row per state, so -m ChromHMM tests every chromatin state at once. -b picks the set in the browser instead.
yame fetch -c hg38/data/human_hg38_immune_mixture.cg
yame fetch -c hg38/data/human_hg38_40_celltypes_chr20.cg
yame binarize -t 0.5 -c 1 -o calls.cg human_hg38_immune_mixture.cg # M/U -> set + universe (one read per site here, so -c 1)
yame subset -o a.cg human_hg38_40_celltypes_chr20.cg GSM5652176_Adipocytes-Z000000T7
yame subset -o b.cg human_hg38_40_celltypes_chr20.cg GSM5652181_Saphenous-Vein-Endothel-Z000000RM
yame pairwise -H 1 -c 5 -d 0.2 a.cg b.cg # differential, fmt3 -> fmt6
yame pairwise -S a.cg b.cg # or score the agreement instead: one TSV line
yame rowop -o binasum human_hg38_40_celltypes_chr20.cg pseudobulk.cg # collapse
Sample Nrow Format
1 773477 3
That last one is how pseudobulks are made: subset the cells of a cluster, pipe into rowop.
pairwise -S answers the other question about the same two samples — not where they differ but how much they agree — as one line of n mae rmse acc pearson over the sites both cover. Either side may be beta (format 4) rather than counts, since comparing a model’s output to its ground truth is inherently cross-format. Give it one sample and a multi-sample store and it scores every record against that sample, inflating the sample once while the store streams:
yame pairwise -S -1 GSM5652176_Adipocytes-Z000000T7 human_hg38_40_celltypes_chr20.cg predictions.cg
To score imputation rather than recall, blank the sites the model was shown before comparing — mask with the query as the mask leaves exactly the sites it had to guess:
yame mask truth.cg query.cg | yame pairwise -S -1 <cell> - predictions.cg
Two different jobs. mask is deterministic and file-driven — it decides which sites are valid. dsample and perturb are stochastic and seed-driven — they decide how much signal is left.
yame fetch -c hg38/data/human_hg38_test.cg
yame dsample -N 5000 -s 42 -o ds.cg human_hg38_test.cg
yame summary ds.cg # dsample itself prints nothing
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
ds.cg 0 NA global 5000 4087 5000 4087 NA 0.817 NA
yame perturb -p 0.05 -s 7 -o noisy.cg human_hg38_test.cg
yame summary noisy.cg
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
noisy.cg 1 NA global 29000 22953 29000 22953 NA 0.791 NA
Set -s. Without a fixed seed neither is reproducible, and the same seed reproduces a run only on the same build and platform — these counts come from Linux, and a macOS build of the same version draws different sites. mask takes its mask as a positional argument, not -m, and both files must have the same row count.
yame fetch -y -c -g blacklist hg38/KYCG
yame mask -o clean.cg human_hg38_test.cg Blacklist.20220304.cm
mask blanks the sites the mask covers and leaves the rest: the line above removes the blacklisted sites, and -v would keep only them. The mask may be a set (.cm) or another methylome, whose covered sites are the mask — which is what the imputation recipe above relies on. A bare name (Blacklist) resolves in the store the way summary -m does.
The four extensions. They are all the same container and differ
only in what they hold, so any command reads any of them:
.cg a methylome, one value per CpG;
.cm a feature set or mask;
.cr the coordinates of a row space;
.cx the generic name, used where the content is not
fixed — split writes it for that reason.
Nothing enforces them; they are a convention for readers.
A YAME file is values in order and nothing else. Row i means whatever row i of the reference means, so intersecting a methylome with a feature set is a bitwise AND rather than a genomic join — fast, and indifferent to how many rows there are.
cpg_nocontig.cr the reference: every CpG in the genome, in order row 1 2 3 4 5 6 chr1: 10,468 10,470 10,483 10,489 10,493 10,497 .cg methylome 1 0 1 1 0 1 .cm feature set 0 1 1 0 0 1 ─────────────────────────────────────────────────────────────── AND 0 0 1 0 0 1 ▲ row 3 is chr1:10,483 in every file, so an overlap is a bitwise AND, not a genomic join
-R to be inferred.Two files are comparable only if they share a row space, and since each has a distinct length, a row count is an identity. That is what lets -R and -m take a name instead of a path. An Infinium manifest and a whole-genome CpG set are the same kind of object here, differing only in length.
one .cg file = BGZF frames; any record reachable without the rest ┌────────┬─────┬────────┬───────────────────────────────┐ │ sig │ fmt │ n │ data │ │ 8 B │ 1 B │ 8 B │ packed per the format code │ └────────┴─────┴────────┴───────────────────────────────┘ └─────────────── one record = one sample ───────────────┘ [ sample 1 ][ sample 2 ][ sample 3 ][ … ] concatenated; the .idx beside it stores where each begins
n, then the data packed
according to that code. Records are concatenated, so a multi-sample file is
just more of them, and the .idx beside it stores where each one
begins — which is why yame subset can pull one sample out of
hundreds without decompressing the others.The three kinds above are the common cases. The rest exist because a row sometimes has to carry something else: a fraction, a label, a plain yes/no, or the coordinates themselves.
the data format one line per CpG on disk NA pack what to know ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── methylation, M/U counts 3 M/U counts M ⇥ U 1-8 B/site M=U=0 -fm a tagged record; only 2^31+ counts lose bits methylation, binarized 6 set+universe call ⇥ covered 2 bits/row 00 -fd the 2nd bit says the site was covered at all a CpG set 6 set+universe in set ⇥ in universe 2 bits/row 00 -fd query and background travel in one file a beta value 4 beta a float, or NA 4 B/value negative -fn NA runs collapse, so sparse is cheap a plain yes/no 0 binary 0 or 1 1 bit/row none -fb any first character but 0 reads as 1 a state label 2 states the whole line dict + runs a label -fs keys stored once in a table, then runs a small int 1 byte one char per line 3 B/run none -fc only the first character of a line coordinates (.cr) 7 coordinates BED: chrom ⇥ start delta varint n/a -fr only chrom and start are read Format 5 is obsolete: still decodable, no longer packable.
Every format lays the same row space down at a different density, which is the whole reason there is a family rather than one format. Two of them are fixed width and cost the same per CpG whatever the data says; the rest widen with the values they carry, so a format 3 store scales with sequencing depth rather than with the row count. These layouts are the ones in src/format*.c.
one byte, and how many CpGs it holds fmt 0 binary │1│0│0│1│1│0│1│0│ 8 CpGs · 1 bit each fmt 6 set+universe │11 │10 │00 │11 │ 4 CpGs · 2 bits each fmt 1 small int │ 37 │ 1 CpG · one whole byte hg38 is 29,401,795 rows: 3.5 MB at one bit, 7.0 MB at two.
Format 5 is obsolete: still decodable, no longer packable.
inflated c->unit bytes a site, default 8 (M and U 32 bits each);
pack -u takes 1..8, and 1 leaves only 4 bits apiece
on disk the low two bits tag the record
01 1 byte [ M(3) | U(3) | 01 ] M,U ≤ 6
10 2 bytes [ M(7) | U(7) | 10 ] M,U ≤ 126
11 8 bytes [ M(31) | U(31) | 11 ] M,U ≤ 2^31-1
00 2 bytes [ run length (14) | 00 ] a run of M = U = 0
beyond 2^31-1 fitMU right-shifts M and U together until both
fit, so the ratio survives and the absolute depth does not
inflated byte s[i>>2], offset = (i&3)*2 (n+3)>>2 bytes
set bit = (b >> offset) & 1
universe bit = (b >> (offset+1)) & 1
codes 00 outside the universe · 10 in universe, 0 · 11 in universe, 1
on disk identical; "compression" is only a flag
Two bits is the interesting one: the universe bit separates not measured from measured and zero — a distinction one bit cannot make and every enrichment test needs, which is why a query and its background travel in one file.
two bits per CpG: the universe bit, then the set bit ┌──────────────┬──────────────┬──────────────┐ │ 0 0 │ 1 0 │ 1 1 │ │ NA │ value 0 │ value 1 │ ├──────────────┼──────────────┼──────────────┤ │ outside │ measured, │ measured, │ │ the universe │ not in set │ in the set │ └──────────────┴──────────────┴──────────────┘ The universe bit separates absent from present-and-zero — a distinction one bit cannot make, and every enrichment test needs.
summary can compute an odds ratio from a single
record.inflated bit (i&7) of s[i>>3] (n+7)>>3 bytes on disk identical; format 0 defines no coding of its own
inflated s[i], any value 0-255
on disk 3-byte records [ value (1 B) | length (uint16 LE) ]
length 1..32767; a longer run is split across records
layout [ key section ][ '\0' ][ data section ]
keys "state0\0state1\0...\0", stored once; row value v means keys[v]
inflated data section is one integer a row, c->unit bytes
on disk [ unit ][ (value, length) ... ], unit in {1,2,3,8},
length uint16 LE 1..65535
inflated float array, 4 bytes a row; NA is a negative float (-1.0)
on disk 32-bit words, the MSB tags each
1 [ 1 | run length (31) ] that many consecutive NAs
0 [ 0 | float bits (31) ] the value, sign bit dropped
stream [ chrom '\0' ][ delta ][ delta ] ... [ 0xff ][ chrom '\0' ] ...
deltas 0bbbbbbb 1 byte, 7-bit delta
10xxxxxx yyyyyyyy 2 bytes, 14-bit delta
11...... + 7 more 8 bytes, 62-bit delta, big-endian
each delta is loc(i) - loc(i-1); 0xff ends a chromosome block
yame fetch can downloadGenerated from tools/registry/files.tsv, the
table the fetch registry is compiled from, so this list is what this
release offers. Every card is one GitHub or HuggingFace repository at a
pinned tag, and every file is verified against its SHA-256 on arrival.
A drawer lists one store directory; click a file for its description,
source and citation. yame fetch -l prints the
same table with what your store already holds.
EPIC/ — 8 files, 34.7 MBEPIC.ordering.tsv.gzProbe ordering7.4 MBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
EPIC.hg38.coord.tsv.gzProbe coordinates5.3 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPIC.hg38.snp.tsv.gzSNP overlap1.7 MBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPIC.typeI_ext.tsv.gzType I extension89.3 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
EPIC.hg38.mask.cmRecommended probe mask357.2 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPIC.hg38.mask.cm.idxRecommended probe mask733 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPIC.cnvnormals.cgCNV normal reference panel19.9 MBWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Five normals from Bioconductor sesameData `EPIC.5.SigDF.normal`, the same default set as R sesame's cnSegmentation(): PrEC prostate epithelial cells (2 replicates) and normal adjacent fibroblasts NAF.A/B/C. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
EPIC.cnvnormals.cg.idxCNV normal reference panel93 BWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Five normals from Bioconductor sesameData `EPIC.5.SigDF.normal`, the same default set as R sesame's cnSegmentation(): PrEC prostate epithelial cells (2 replicates) and normal adjacent fibroblasts NAF.A/B/C. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
EPIC/KYCG/ — 26 files, 41.7 MBABCompartment.20220911.cmA/B compartments392.0 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
Blacklist.20220304.cmBlacklisted regions1.0 KBWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands312.8 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220303.cmChromatin states461.6 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
CTCFbind.20220911.cmCTCF binding sites37.4 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
HM.20221013.cmHistone modifications2.3 MBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs2.5 KBWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry166.3 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position1.8 MBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cmFlanking CpG density644.9 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cmPartially methylated domains191.1 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeType.cmProbe type99 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap459.9 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
rmsk1.20220307.cmRepeats, class191.0 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family272.2 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Tetranuc2.20220321.cmTetranucleotide context276.7 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TFBSrm.20221005.cmTranscription factor binding sites, ReMap34.3 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
ABCompartment.20220911.cm.idxA/B compartments79 BWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cm.idxCpG islands64 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
HM.20221013.cm.idxHistone modifications1.7 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position762 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cm.idxFlanking CpG density504 BWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cm.idxPartially methylated domains33 BWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
rmsk1.20220307.cm.idxRepeats, class352 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1.1 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap22.9 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Cites: Rao 2014 · Amemiya 2019 · Gardiner-Garden 1987 · Ernst 2012 · Wang 2012 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Smit · Hammal 2022
EPICv2/ — 8 files, 24.3 MBEPICv2.ordering.tsv.gzProbe ordering8.2 MBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
EPICv2.hg38.coord.tsv.gzProbe coordinates5.5 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPICv2.hg38.snp.tsv.gzSNP overlap1.6 MBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPICv2.typeI_ext.tsv.gzType I extension87.8 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
EPICv2.hg38.mask.cmRecommended probe mask294.0 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPICv2.hg38.mask.cm.idxRecommended probe mask760 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
EPICv2.cnvnormals.cgCNV normal reference panel8.8 MBWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Two normals from sesameData `EPICv2.8.SigDF`, the same default set as R sesame's cnSegmentation(): GM12878 lymphoblastoid replicates 206909630042_R08C01 and 206909630040_R03C01. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
EPICv2.cnvnormals.cg.idxCNV normal reference panel71 BWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Two normals from sesameData `EPICv2.8.SigDF`, the same default set as R sesame's cnSegmentation(): GM12878 lymphoblastoid replicates 206909630042_R08C01 and 206909630040_R03C01. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
EPICv2/KYCG/ — 24 files, 43.6 MBABCompartment.20220911.cmA/B compartments419.6 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cmCpG islands326.6 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220303.cmChromatin states485.4 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
CTCFbind.20220911.cmCTCF binding sites42.7 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
HM.20221013.cmHistone modifications2.3 MBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs2.4 KBWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry157.7 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position1.9 MBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cmFlanking CpG density628.0 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cmPartially methylated domains206.4 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeType.cmProbe type105 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap484.4 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
rmsk1.20220307.cmRepeats, class206.7 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family295.5 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Tetranuc2.20220321.cmTetranucleotide context300.7 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TFBSrm.20221005.cmTranscription factor binding sites, ReMap35.9 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
ABCompartment.20220911.cm.idxA/B compartments79 BWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cm.idxCpG islands64 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
HM.20221013.cm.idxHistone modifications1.7 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position764 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
PMD.20220911.cm.idxPartially methylated domains33 BWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
rmsk1.20220307.cm.idxRepeats, class352 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1.1 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap23.0 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Cites: Rao 2014 · Gardiner-Garden 1987 · Ernst 2012 · Wang 2012 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Smit · Hammal 2022
HM27/ — 6 files, 829.9 KBHM27.ordering.tsv.gzProbe ordering298.4 KBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
HM27.hg38.coord.tsv.gzProbe coordinates169.6 KBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM27.hg38.snp.tsv.gzSNP overlap350.3 KBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM27.typeI_ext.tsv.gzType I extension3.4 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
HM27.hg38.mask.cmRecommended probe mask7.8 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM27.hg38.mask.cm.idxRecommended probe mask466 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Cites: Zhou 2017
HM27/KYCG/ — 20 files, 318.0 KBABCompartment.20220911.cmA/B compartments12.0 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cmCpG islands10.4 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220303.cmChromatin states14.4 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
CTCFbind.20220911.cmCTCF binding sites1.6 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
HM.20221013.cmHistone modifications139.9 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs407 BWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry8.3 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position62.6 KBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cmFlanking CpG density21.5 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cmPartially methylated domains8.5 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeType.cmProbe type83 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap13.1 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
rmsk1.20220307.cmRepeats, class5.0 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family8.3 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Tetranuc2.20220321.cmTetranucleotide context8.7 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
CGI.20220904.cm.idxCpG islands59 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
HM.20221013.cm.idxHistone modifications1.5 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position724 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
rmsk1.20220307.cm.idxRepeats, class229 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family682 BWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Cites: Rao 2014 · Gardiner-Garden 1987 · Ernst 2012 · Wang 2012 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Smit
HM450/ — 8 files, 32.8 MBHM450.ordering.tsv.gzProbe ordering4.4 MBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
HM450.hg38.coord.tsv.gzProbe coordinates2.9 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM450.hg38.snp.tsv.gzSNP overlap1.6 MBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM450.typeI_ext.tsv.gzType I extension71.2 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
HM450.hg38.mask.cmRecommended probe mask208.7 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM450.hg38.mask.cm.idxRecommended probe mask727 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
HM450.cnvnormals.cgCNV normal reference panel23.6 MBWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Ten normals from sesameData `HM450.10.SigDF`: TCGA-BLCA solid-tissue normals (barcode sample type 11A; TCGA-BL-A13J, -BT-A20J/N/P/R/U/V/W/X, -BT-A2LA). R sesame has no HM450 default panel; this is the set sesame cnv uses. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
HM450.cnvnormals.cg.idxCNV normal reference panel412 BWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Ten normals from sesameData `HM450.10.SigDF`: TCGA-BLCA solid-tissue normals (barcode sample type 11A; TCGA-BL-A13J, -BT-A20J/N/P/R/U/V/W/X, -BT-A2LA). R sesame has no HM450 default panel; this is the set sesame cnv uses. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
HM450/KYCG/ — 27 files, 28.2 MBABCompartment.20220911.cmA/B compartments216.8 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
Blacklist.20220304.cmBlacklisted regions730 BWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands191.1 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220303.cmChromatin states283.0 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
Chromosome.20221129.cmChromosome366.3 KBWhat it is. Which chromosome each CpG is on. Mostly a sanity check, and the way to spot chromosome-scale effects such as aneuploidy. One set per chromosome, contigs and alts excluded.
Source. Reference genome index
CTCFbind.20220911.cmCTCF binding sites21.8 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
HM.20221013.cmHistone modifications1.6 MBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs2.3 KBWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry134.8 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position1.0 MBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cmFlanking CpG density414.0 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cmPartially methylated domains106.0 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeType.cmProbe type108 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap273.3 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
rmsk1.20220307.cmRepeats, class92.7 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family127.1 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Tetranuc2.20220321.cmTetranucleotide context151.0 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TFBSrm.20221005.cmTranscription factor binding sites, ReMap23.2 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
ABCompartment.20220911.cm.idxA/B compartments77 BWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cm.idxCpG islands62 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
HM.20221013.cm.idxHistone modifications1.6 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position755 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cm.idxFlanking CpG density484 BWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cm.idxPartially methylated domains33 BWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
rmsk1.20220307.cm.idxRepeats, class341 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1021 BWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap22.7 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Cites: Rao 2014 · Amemiya 2019 · Gardiner-Garden 1987 · Ernst 2012 · Wang 2012 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Smit · Hammal 2022
Mammal40/ — 6 files, 617.1 KBMammal40.ordering.tsv.gzProbe ordering337.2 KBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
Mammal40.hg38.coord.tsv.gzProbe coordinates229.2 KBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Mammal40.hg38.snp.tsv.gzSNP overlap36.2 KBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Mammal40.typeI_ext.tsv.gzType I extension2.5 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
Mammal40.hg38.mask.cmRecommended probe mask11.4 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Mammal40.hg38.mask.cm.idxRecommended probe mask629 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Cites: Zhou 2017
Mammal40/KYCG/ — 18 files, 296.5 KBABCompartment.20220911.cmA/B compartments17.9 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cmCpG islands12.7 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220303.cmChromatin states20.5 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
CTCFbind.20220911.cmCTCF binding sites1.1 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
HM.20221013.cmHistone modifications80.1 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs214 BWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry4.1 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position81.2 KBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
nFlankCG.20220321.cmFlanking CpG density25.5 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
PMD.20220911.cmPartially methylated domains12.6 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeType.cmProbe type94 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap21.0 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
rmsk1.20220307.cmRepeats, class2.2 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family2.5 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Tetranuc2.20220321.cmTetranucleotide context12.3 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
CGI.20220904.cm.idxCpG islands59 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
HM.20221013.cm.idxHistone modifications1.5 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position723 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
Cites: Rao 2014 · Gardiner-Garden 1987 · Ernst 2012 · Wang 2012 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Smit
MM285/ — 9 files, 6.1 MBMM285.ordering.tsv.gzProbe ordering2.7 MBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
MM285.mm10.coord.tsv.gzProbe coordinates1.2 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.mm39.coord.tsv.gzProbe coordinates1.2 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.mm10.snp.tsv.gzSNP overlap748.1 KBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.typeI_ext.tsv.gzType I extension31.8 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
MM285.mm10.mask.cmRecommended probe mask165.3 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.mm39.mask.cmRecommended probe mask10.6 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.mm10.mask.cm.idxRecommended probe mask3.0 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MM285.mm39.mask.cm.idxRecommended probe mask52 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
Cites: Zhou 2017
MM285/KYCG/ — 24 files, 7.4 MBBlacklist.20220304.cmBlacklisted regions157 BWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands60.8 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220318.cmChromatin states103.1 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
Chromosome.20221129.cmChromosome3.7 KBWhat it is. Which chromosome each CpG is on. Mostly a sanity check, and the way to spot chromosome-scale effects such as aneuploidy. One set per chromosome, contigs and alts excluded.
Source. Reference genome index
EnsRegBuild.20220710.cmEnsembl Regulatory Build78.6 KBWhat it is. Ensembl's integrated regulatory annotation: promoters, promoter flanks, enhancers, CTCF sites and open chromatin, called across many cell types. Feature type per CpG.
Source. Ensembl Regulatory Build GFF
Citation. Zerbino et al. 2015, Genome Biol, doi:10.1186/s13059-015-0621-5
HM.20221013.cmHistone modifications666.8 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
InfiniumChemistry.cmInfinium chemistry68.4 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
MetagenePC.20220911.cmMetagene position395.6 KBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
PMD.20220911.cmPartially methylated domains11.0 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeCGnum.cmCpGs per probe106.4 KBWhat it is. How many CpGs a probe body covers, which affects its sensitivity to local methylation state. Technical, like ProbeType: rule it out before believing a surprising array enrichment. CpG count over the probe body, read from the manifest; the exact counting rule is not recorded in the lab journal.
Source. Infinium platform manifest
ProbeType.cmProbe type92 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
TFBSrm.20221005.cmTranscription factor binding sites, ReMap5.4 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Tetranuc2.20220321.cmTetranucleotide context90.4 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
Tetranuc4.20220321.cmTetranucleotide context, uncollapsed171.3 KBWhat it is. The same context before reverse-complement collapsing. N[CG]N context as read.
Source. Derived from the reference genome
nFlankCG.20220321.cmFlanking CpG density166.3 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
rmsk1.20220321.cmRepeats, class44.1 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family65.2 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
CGI.20220904.cm.idxCpG islands62 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
EnsRegBuild.20220710.cm.idxEnsembl Regulatory Build142 BWhat it is. Ensembl's integrated regulatory annotation: promoters, promoter flanks, enhancers, CTCF sites and open chromatin, called across many cell types. Feature type per CpG.
Source. Ensembl Regulatory Build GFF
Citation. Zerbino et al. 2015, Genome Biol, doi:10.1186/s13059-015-0621-5
HM.20221013.cm.idxHistone modifications1.5 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position747 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap11.6 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
rmsk1.20220321.cm.idxRepeats, class351 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family983 BWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Cites: Amemiya 2019 · Gardiner-Garden 1987 · Ernst 2012 · Zerbino 2015 · Zheng 2019 · Frankish 2019 · Zhou 2018 · Hammal 2022 · Smit
MSA/ — 8 files, 9.8 MBMSA.ordering.tsv.gzProbe ordering2.8 MBWhat it is. The canonical probe order for the platform -- the row space every mask, coordinate table and knowledgebase set for that platform is written against. Not an annotation in itself: it is what makes the others addressable, since a .cm carries bits and no probe names. One probe ID per row, in the fixed order every other file for this platform uses. Comes with anything else fetched from the platform, because a file indexed against it cannot be read without it.
Source. zhou-lab/InfiniumAnnotation, derived from the Illumina manifest
MSA.hg38.coord.tsv.gzProbe coordinates1.5 MBWhat it is. Where each probe's target sits in a given genome build. What turns a probe-space result into genomic coordinates, and what a genome-space annotation is projected through to reach the array. One row per probe in platform ordering, one file per genome build, so the hg38 and mm10 views of a platform coexist.
Source. zhou-lab/InfiniumAnnotation, from the probe realignment
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MSA.hg38.snp.tsv.gzSNP overlap1.0 MBWhat it is. Common variants falling on a probe's target or its extension base, where the signal reports genotype rather than methylation. The basis for dropping or flagging those probes, and for the rs-probe genotyping that identifies a sample. Per-probe variant overlap in platform ordering. The variant database build is not recorded in this table.
Source. zhou-lab/InfiniumAnnotation, variants intersected with probe coordinates
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MSA.typeI_ext.tsv.gzType I extension29.8 KBWhat it is. Infinium type I probes read both alleles in one colour channel, leaving the other channel carrying out-of-band signal. This holds the type I extension information that colour-channel inference and out-of-band intensity are computed from. One row per type I probe. Only platforms with type I probes publish it.
Source. zhou-lab/InfiniumAnnotation
MSA.hg38.mask.cmRecommended probe mask152.9 KBWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MSA.hg38.mask.cm.idxRecommended probe mask774 BWhat it is. Which probes to drop before analysis: cross-hybridizing probes, probes overlapping common variants, and probes that do not map uniquely. Applying it removes a large and reproducible source of false signal, and using the same mask is what makes results comparable across studies. A .cm over the platform's probe ordering, one bit per probe, one file per genome build. Ships with a .cm.idx companion.
Source. zhou-lab/InfiniumAnnotation
Citation. Zhou, Laird & Shen 2017, Nucleic Acids Res, doi:10.1093/nar/gkw967
MSA.cnvnormals.cgCNV normal reference panel4.3 MBWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Three normals from sesameData `MSA.9.SigDF`: GM12878 lymphoblastoid replicates 207760740037_R09C01/R11C01/R13C01 (the HCT116 and LNCaP samples in that set are cancer lines and excluded). R sesame has no MSA default panel; this is the set sesame cnv uses. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
MSA.cnvnormals.cg.idxCNV normal reference panel111 BWhat it is. The normal reference panel that `sesame cnv` regresses a query's per-probe total intensity on (OLS) before taking log2(query/fitted), binning along the genome and segmenting by CBS. Raw (un-normalized) signal, so the query must be extracted with `preprocess --raw-signal --output total_intensity`. Three normals from sesameData `MSA.9.SigDF`: GM12878 lymphoblastoid replicates 207760740037_R09C01/R11C01/R13C01 (the HCT116 and LNCaP samples in that set are cancer lines and excluded). R sesame has no MSA default panel; this is the set sesame cnv uses. A format 3 .cg (M/U) over the platform's probe ordering, one record per normal sample. Ships with a .cg.idx companion.
Source. zhou-lab/InfiniumAnnotation, built by sesame's make cnv-normals (tools/export_cnvnormals.R + mu2cg), positional on the v8.2 ordering
Citation. Zhou, Triche, Laird & Shen 2018, Nucleic Acids Res, doi:10.1093/nar/gky691
MSA/KYCG/ — 44 files, 14.0 MBABCompartment.20220911.cmA/B compartments117.0 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
Blacklist.20220304.cmBlacklisted regions663 BWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands92.9 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
CTCFbind.20220911.cmCTCF binding sites9.4 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
Centromere.20221129.cmCentromeres2.6 KBWhat it is. Centromeric regions -- repeat-rich, poorly mappable, and usually excluded rather than interpreted. Like Blacklist, a filter rather than biology. Build not recorded in the lab journal.
Source. UCSC centromere annotation (inferred from the set name; not recorded in the lab journal)
ChromHMM.20220303.cmChromatin states134.2 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
ChromHMMfullStack.20230515.cmChromatin states, full stack399.6 KBWhat it is. One universal 100-state annotation learned jointly across many cell types, so a state means the same thing everywhere -- at the cost of cell-type specificity. hg38: hg38lift_genome_100_segments intersected with the CpG reference. mm10: mm10_100_segments, first state per CpG.
Source. Ernst lab full-stack annotation
Citation. Vu & Ernst 2022, Genome Biol, doi:10.1186/s13059-022-02648-4 (mouse model: Vu & Ernst 2023, doi:10.1186/s13059-023-02994-x)
Chromosome.20221129.cmChromosome171.5 KBWhat it is. Which chromosome each CpG is on. Mostly a sanity check, and the way to spot chromosome-scale effects such as aneuploidy. One set per chromosome, contigs and alts excluded.
Source. Reference genome index
CoRSIV.20220309.cmCorrelated regions of systemic interindividual variation4.6 KBWhat it is. Regions where individuals differ reproducibly across every tissue, implying establishment in early development rather than tissue identity. Region coordinates with per-region CpG counts.
Source. CoRSIV coordinates (upstream download undocumented)
Citation. Gunasekara et al. 2019, Genome Biol, doi:10.1186/s13059-019-1708-1 (follow-up: 2023, doi:10.1186/s13059-022-02827-3)
EvoCons.20220314.cmEvolutionary conservation29.9 KBWhat it is. Cross-species conservation. Conserved unmethylated regions are usually developmental regulators, so conservation separates functional from incidental CpGs. phastCons score thresholded at 0.5.
Source. UCSC phastCons 60-way
Citation. Siepel et al. 2005, Genome Res, doi:10.1101/gr.3715005
G4peaksHighConfidence.20220719.cmG-quadruplexes5.2 KBWhat it is. Sequences that fold into four-stranded G-quadruplexes, which affect replication and methylation maintenance. G4 formation and methylation are mutually antagonistic -- a folded G4 inhibits DNMT1 -- which is one proposed reason some CpG islands stay constitutively unmethylated. The high-confidence peak calls only; the exact upstream file is not recorded.
Source. G4-seq whole-genome G4 maps (K+ and PDS conditions). Identification inferred from the set name, not recorded in the lab journal.
Citation. Marsico et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gkz179 (inferred)
HM.20221013.cmHistone modifications601.3 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs1.6 KBWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
InfiniumChemistry.cmInfinium chemistry76.6 KBWhat it is. Finer design-chemistry partition than probe type, including colour channel. Generated by the manifest build pipeline; the rule is not in the lab journal.
Source. Infinium platform manifest
IntergenicCpGs.20220811.cmIntergenic CpGs37.3 KBWhat it is. CpGs outside genes and their promoters. The complement of the extended gene set.
Source. GENCODE
IntermediateMeth.20221121.cmIntermediate methylation6.6 KBWhat it is. CpGs persistently between 25% and 75%. Intermediate values usually mean a mixed cell population or allele-specific methylation rather than a genuinely half-methylated cell. Carries the number of datasets in which the CpG was intermediate.
Source. WGBS compendium, MSA design work
MetagenePC.20220911.cmMetagene position552.7 KBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
PMD.20220911.cmPartially methylated domains81.8 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
ProbeCGnum.cmCpGs per probe121.7 KBWhat it is. How many CpGs a probe body covers, which affects its sensitivity to local methylation state. Technical, like ProbeType: rule it out before believing a surprising array enrichment. CpG count over the probe body, read from the manifest; the exact counting rule is not recorded in the lab journal.
Source. Infinium platform manifest
ProbeType.cmProbe type98 BWhat it is. Infinium type I versus type II chemistry. Technical, not biological -- and the first thing to rule out when an array enrichment looks surprising, since the two types differ in dynamic range and CpG density. Read from the probe ID prefix in the manifest.
Source. Infinium platform manifest
REMCChromHMM.20220911.cmChromatin states, Roadmap134.9 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
RoadMapNegGeneExpCpG.20220814.cmCpGs negatively correlated with expression16.8 KBWhat it is. The classical direction: methylation up, expression down. Mostly promoter-proximal. Spearman rho < -0.5 at p < 0.01.
Source. Roadmap methylation-expression correlation
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
RoadMapPosGeneExpCpG.20220814.cmCpGs positively correlated with expression12.2 KBWhat it is. CpGs whose methylation rises with expression of a nearby gene -- typically gene-body rather than promoter, the direction that surprises people. Spearman rho > 0.5 at p < 0.01.
Source. Roadmap methylation-expression correlation
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
TFBSrm.20221005.cmTranscription factor binding sites, ReMap8.5 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Tetranuc2.20220321.cmTetranucleotide context88.2 KBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TiSigBLUEPRINT.20221209.cmTissue signatures, blood790.1 KBWhat it is. Haematopoietic lineage signatures -- the reference panel for deconvolving blood, where most human EWAS material comes from. One set per blood cell type, same pairwise-signature form.
Source. BLUEPRINT Epigenome WGBS
Citation. Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007
TiSigBrain.20221209.cmTissue signatures, brain1.1 MBWhat it is. Brain cell-type signatures -- neurons, glia and subtypes -- from single-cell methylation. One set per cell type. Brain is the one tissue with abundant non-CpG methylation in neurons, and neuron-versus-glia is the largest cell-type methylation contrast in any human tissue, so this doubles as a cell-composition control.
Source. Combined brain single-cell DNAm signatures. The upstream atlas is not recorded; the likely source is Luo et al. 2017, Science, doi:10.1126/science.aan3351 (single-neuron methylomes, 21 human and 16 mouse subtypes) -- treat that as unconfirmed.
TiSigLoyfer.20221209.cmTissue signatures, Loyfer atlas589.6 KBWhat it is. Cell-type-specific methylation markers from a deeply sequenced atlas of sorted human cell types. The set to reach for when asking what tissue or cell type a methylome resembles. One set per pairwise cell-type comparison, named celltype.in.background.
Source. Loyfer human methylome atlas, sorted-cell WGBS
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
XCILinkedWGBSSorted.20221121.cmX-inactivation linked, sorted cells2.5 KBWhat it is. The same, from sorted-cell WGBS rather than bulk. As above, from sorted cells.
Source. Sorted-cell WGBS, MSA design work
nFlankCG.20220321.cmFlanking CpG density192.1 KBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
rmsk1.20220307.cmRepeats, class60.0 KBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family87.5 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
ABCompartment.20220911.cm.idxA/B compartments75 BWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
CGI.20220904.cm.idxCpG islands62 BWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMMfullStack.20230515.cm.idxChromatin states, full stack2.0 KBWhat it is. One universal 100-state annotation learned jointly across many cell types, so a state means the same thing everywhere -- at the cost of cell-type specificity. hg38: hg38lift_genome_100_segments intersected with the CpG reference. mm10: mm10_100_segments, first state per CpG.
Source. Ernst lab full-stack annotation
Citation. Vu & Ernst 2022, Genome Biol, doi:10.1186/s13059-022-02648-4 (mouse model: Vu & Ernst 2023, doi:10.1186/s13059-023-02994-x)
HM.20221013.cm.idxHistone modifications1.6 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cm.idxMetagene position750 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap22.0 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
TiSigBLUEPRINT.20221209.cm.idxTissue signatures, blood19.3 KBWhat it is. Haematopoietic lineage signatures -- the reference panel for deconvolving blood, where most human EWAS material comes from. One set per blood cell type, same pairwise-signature form.
Source. BLUEPRINT Epigenome WGBS
Citation. Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007
TiSigBrain.20221209.cm.idxTissue signatures, brain12.5 KBWhat it is. Brain cell-type signatures -- neurons, glia and subtypes -- from single-cell methylation. One set per cell type. Brain is the one tissue with abundant non-CpG methylation in neurons, and neuron-versus-glia is the largest cell-type methylation contrast in any human tissue, so this doubles as a cell-composition control.
Source. Combined brain single-cell DNAm signatures. The upstream atlas is not recorded; the likely source is Luo et al. 2017, Science, doi:10.1126/science.aan3351 (single-neuron methylomes, 21 human and 16 mouse subtypes) -- treat that as unconfirmed.
TiSigLoyfer.20221209.cm.idxTissue signatures, Loyfer atlas14.5 KBWhat it is. Cell-type-specific methylation markers from a deeply sequenced atlas of sorted human cell types. The set to reach for when asking what tissue or cell type a methylome resembles. One set per pairwise cell-type comparison, named celltype.in.background.
Source. Loyfer human methylome atlas, sorted-cell WGBS
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
nFlankCG.20220321.cm.idxFlanking CpG density420 BWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
rmsk1.20220307.cm.idxRepeats, class324 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1.0 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Cites: Rao 2014 · Amemiya 2019 · Gardiner-Garden 1987 · Wang 2012 · Ernst 2012 · Vu 2022 · Gunasekara 2019 · Siepel 2005 · Marsico 2019 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Hammal 2022 · Stunnenberg 2016 · Loyfer 2023 · Smit
hg38/KYCG/ — 41 files, 334.3 MBABCompartment.20220911.cmA/B compartments9.5 KBWhat it is. The two large-scale nuclear compartments from Hi-C: A is open and gene-rich, B is closed, late-replicating and the substrate on which partially methylated domains form. Subcompartment labels per CpG. A second build takes the most frequent call across 95 tumour and normal Hi-C maps, lifted hg19 to hg38.
Source. GM12878 subcompartments, GSE63525
Citation. Rao et al. 2014, Cell, doi:10.1016/j.cell.2014.11.021
Blacklist.20220304.cmBlacklisted regions3.1 KBWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands201.9 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
CTCFbind.20220911.cmCTCF binding sites159.2 KBWhat it is. CTCF anchors chromatin loops and TAD boundaries, and its binding is methylation-sensitive -- the mechanism by which methylation can rewire 3D contacts. Lifted hg19 to hg38.
Source. CTCF sites from Wang et al. 2012
Citation. Wang et al. 2012, Genome Res, doi:10.1101/gr.136101.111
Centromere.20221129.cmCentromeres347 BWhat it is. Centromeric regions -- repeat-rich, poorly mappable, and usually excluded rather than interpreted. Like Blacklist, a filter rather than biology. Build not recorded in the lab journal.
Source. UCSC centromere annotation (inferred from the set name; not recorded in the lab journal)
ChromHMM.20220303.cmChromatin states517.0 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
ChromHMMfullStack.20230515.cmChromatin states, full stack6.8 MBWhat it is. One universal 100-state annotation learned jointly across many cell types, so a state means the same thing everywhere -- at the cost of cell-type specificity. hg38: hg38lift_genome_100_segments intersected with the CpG reference. mm10: mm10_100_segments, first state per CpG.
Source. Ernst lab full-stack annotation
Citation. Vu & Ernst 2022, Genome Biol, doi:10.1186/s13059-022-02648-4 (mouse model: Vu & Ernst 2023, doi:10.1186/s13059-023-02994-x)
Chromosome.20221129.cmChromosome298 BWhat it is. Which chromosome each CpG is on. Mostly a sanity check, and the way to spot chromosome-scale effects such as aneuploidy. One set per chromosome, contigs and alts excluded.
Source. Reference genome index
ChromosomeXY.20230901.cmSex chromosomes99 BWhat it is. X and Y only -- the first check when a result might be sex-driven. Build not recorded in the lab journal.
Source. Reference genome index
HM.20221013.cmHistone modifications7.2 MBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
HM.20221013.cm.idxHistone modifications1.7 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
ImprintingDMR.20220818.cmImprinted DMRs907 BWhat it is. Methylation inherited parent-of-origin-specifically and held near 50%. Small, well characterised, and a good positive control: hitting them means finding real allele-specific biology. 450K sliding-window DMR calls, lifted hg19 to hg38.
Source. Court et al. 2014 supplementary table 1
Citation. Court et al. 2014, Genome Res, doi:10.1101/gr.164913.113 (PMID 24402520)
IntermediateMeth.20221121.cmIntermediate methylation26.1 KBWhat it is. CpGs persistently between 25% and 75%. Intermediate values usually mean a mixed cell population or allele-specific methylation rather than a genuinely half-methylated cell. Carries the number of datasets in which the CpG was intermediate.
Source. WGBS compendium, MSA design work
IntermediateMethS.20221121.cmIntermediate methylation, sorted cells18.1 KBWhat it is. The subset that stays intermediate in sorted cells, so mixture is excluded and allele-specificity is the likely cause. As above, from sorted cells.
Source. Sorted-cell WGBS, MSA design work
MetagenePC.20220911.cmMetagene position2.8 MBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
MetagenePC.20220911.cm.idxMetagene position712 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
PMD.20220911.cmPartially methylated domains17.9 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
REMCChromHMM.20220911.cmChromatin states, Roadmap479.3 KBWhat it is. Roadmap Epigenomics core 15-state chromatin model across 129 reference human epigenomes -- the same idea as ChromHMM, tied to the Roadmap tissue panel. Per-CpG state from all 129 epigenomes, reduced to a consensus call.
Source. NIH Roadmap Epigenomics, coreMarks hg38lift
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
RoadMapNegGeneExpCpG.20220814.cmCpGs negatively correlated with expression452.4 KBWhat it is. The classical direction: methylation up, expression down. Mostly promoter-proximal. Spearman rho < -0.5 at p < 0.01.
Source. Roadmap methylation-expression correlation
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
RoadMapPosGeneExpCpG.20220814.cmCpGs positively correlated with expression556.7 KBWhat it is. CpGs whose methylation rises with expression of a nearby gene -- typically gene-body rather than promoter, the direction that surprises people. Spearman rho > 0.5 at p < 0.01.
Source. Roadmap methylation-expression correlation
Citation. Roadmap Epigenomics Consortium 2015, Nature, doi:10.1038/nature14248
TFBS.20220921.cmTranscription factor binding sites86.8 MBWhat it is. Where transcription factors bind. Methylation blocks binding for some factors and is shaped by others, so TF overlap is the usual route from a CpG set to a candidate mechanism. One set per factor, capped at the top 500,000 CpGs per factor by peak-overlap count. hg38 carries about 1,359 factors.
Source. Cistrome Data Browser TF ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
TFBS.20220921.cm.idxTranscription factor binding sites26.4 KBWhat it is. Where transcription factors bind. Methylation blocks binding for some factors and is shaped by others, so TF overlap is the usual route from a CpG set to a candidate mechanism. One set per factor, capped at the top 500,000 CpGs per factor by peak-overlap count. hg38 carries about 1,359 factors.
Source. Cistrome Data Browser TF ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
TFBSrm.20221005.cmTranscription factor binding sites, ReMap78.4 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap23.2 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Tetranuc2.20220321.cmTetranucleotide context8.7 MBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TiSigBLUEPRINT.20221209.cmTissue signatures, blood25.7 MBWhat it is. Haematopoietic lineage signatures -- the reference panel for deconvolving blood, where most human EWAS material comes from. One set per blood cell type, same pairwise-signature form.
Source. BLUEPRINT Epigenome WGBS
Citation. Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007
TiSigBLUEPRINT.20221209.cm.idxTissue signatures, blood21.9 KBWhat it is. Haematopoietic lineage signatures -- the reference panel for deconvolving blood, where most human EWAS material comes from. One set per blood cell type, same pairwise-signature form.
Source. BLUEPRINT Epigenome WGBS
Citation. Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007
TiSigBrain.20221209.cmTissue signatures, brain50.0 MBWhat it is. Brain cell-type signatures -- neurons, glia and subtypes -- from single-cell methylation. One set per cell type. Brain is the one tissue with abundant non-CpG methylation in neurons, and neuron-versus-glia is the largest cell-type methylation contrast in any human tissue, so this doubles as a cell-composition control.
Source. Combined brain single-cell DNAm signatures. The upstream atlas is not recorded; the likely source is Luo et al. 2017, Science, doi:10.1126/science.aan3351 (single-neuron methylomes, 21 human and 16 mouse subtypes) -- treat that as unconfirmed.
TiSigBrain.20221209.cm.idxTissue signatures, brain13.1 KBWhat it is. Brain cell-type signatures -- neurons, glia and subtypes -- from single-cell methylation. One set per cell type. Brain is the one tissue with abundant non-CpG methylation in neurons, and neuron-versus-glia is the largest cell-type methylation contrast in any human tissue, so this doubles as a cell-composition control.
Source. Combined brain single-cell DNAm signatures. The upstream atlas is not recorded; the likely source is Luo et al. 2017, Science, doi:10.1126/science.aan3351 (single-neuron methylomes, 21 human and 16 mouse subtypes) -- treat that as unconfirmed.
TiSigLoyfer.20221209.cmTissue signatures, Loyfer atlas27.5 MBWhat it is. Cell-type-specific methylation markers from a deeply sequenced atlas of sorted human cell types. The set to reach for when asking what tissue or cell type a methylome resembles. One set per pairwise cell-type comparison, named celltype.in.background.
Source. Loyfer human methylome atlas, sorted-cell WGBS
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
TiSigLoyfer.20221209.cm.idxTissue signatures, Loyfer atlas16.0 KBWhat it is. Cell-type-specific methylation markers from a deeply sequenced atlas of sorted human cell types. The set to reach for when asking what tissue or cell type a methylome resembles. One set per pairwise cell-type comparison, named celltype.in.background.
Source. Loyfer human methylome atlas, sorted-cell WGBS
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
Win100k.20220228.cm100 kb windows165.1 KBWhat it is. Fixed 100 kb tiles, for locating signal that follows no annotation. bedtools makewindows, contigs excluded.
Source. Reference genome index
XCILinkedWGBS.20221121.cmX-inactivation linked CpGs12.3 KBWhat it is. CpGs whose methylation tracks X-chromosome inactivation. On the X, methylation reports allelic silencing rather than regulation, so these should usually be handled separately. Derived during MSA array design; carries counts of intermediate and male unmethylated/methylated observations.
Source. WGBS X-inactivation analysis, MSA design work
XCILinkedWGBSSorted.20221121.cmX-inactivation linked, sorted cells16.4 KBWhat it is. The same, from sorted-cell WGBS rather than bulk. As above, from sorted cells.
Source. Sorted-cell WGBS, MSA design work
nFlankCG.20220321.cmFlanking CpG density11.0 MBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
nFlankCG100.20231025.cmFlanking CpG density, 100 bp9.7 MBWhat it is. Count of other CpGs within 50 bp either side. bedtools intersect count in a +/-50 bp window, minus the CpG itself.
Source. Derived from the CpG reference
nFlankCG50.20231025.cmFlanking CpG density, 50 bp7.2 MBWhat it is. Count of other CpGs within 25 bp either side. bedtools intersect count in a +/-25 bp window, minus the CpG itself.
Source. Derived from the CpG reference
rmsk1.20220307.cmRepeats, class4.4 MBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk1.20220307.cm.idxRepeats, class361 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family5.5 MBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1.2 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Cites: Rao 2014 · Amemiya 2019 · Gardiner-Garden 1987 · Wang 2012 · Ernst 2012 · Vu 2022 · Zheng 2019 · Court 2014 · Frankish 2019 · Zhou 2018 · Roadmap 2015 · Hammal 2022 · Stunnenberg 2016 · Loyfer 2023 · Smit
mm10/KYCG/ — 36 files, 242.0 MBBlacklist.20220304.cmBlacklisted regions636 BWhat it is. Regions that produce artefactual signal in short-read assays. Not biology: a filter, and a check that an enrichment is not an artefact. Published intervals, unmodified.
Source. ENCODE blacklist v2 (hg38 ENCFF356LFX, mm10 ENCFF547MET)
Citation. Amemiya et al. 2019, Sci Rep, doi:10.1038/s41598-019-45839-z
CGI.20220904.cmCpG islands120.0 KBWhat it is. CpG-dense, usually promoter-associated and usually unmethylated; the single most informative partition of the methylome. Shores and shelves flank them and carry most tissue-specific variation. Island as published; Shore = +/-2 kb around it; Shelf = +/-4 kb minus the shore; OpenSea = the complement. Nearly disjoint rather than strictly so: on EPIC 49 probes fall in two states where flanks of neighbouring islands meet, and 716 (mostly control and SNP probes) fall in none.
Source. UCSC Genome Browser cpgIslandExt track
Citation. Gardiner-Garden & Frommer 1987, J Mol Biol, doi:10.1016/0022-2836(87)90689-9
ChromHMM.20220414.cmChromatin states857.4 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
ChromHMMfullStack.20231222.cmChromatin states, full stack6.1 MBWhat it is. One universal 100-state annotation learned jointly across many cell types, so a state means the same thing everywhere -- at the cost of cell-type specificity. hg38: hg38lift_genome_100_segments intersected with the CpG reference. mm10: mm10_100_segments, first state per CpG.
Source. Ernst lab full-stack annotation
Citation. Vu & Ernst 2022, Genome Biol, doi:10.1186/s13059-022-02648-4 (mouse model: Vu & Ernst 2023, doi:10.1186/s13059-023-02994-x)
Chromosome.20240119.cmChromosome269 BWhat it is. Which chromosome each CpG is on. Mostly a sanity check, and the way to spot chromosome-scale effects such as aneuploidy. One set per chromosome, contigs and alts excluded.
Source. Reference genome index
ChromosomeXY.20230901.cmSex chromosomes95 BWhat it is. X and Y only -- the first check when a result might be sex-driven. Build not recorded in the lab journal.
Source. Reference genome index
EnsRegBuild.20220710.cmEnsembl Regulatory Build995.3 KBWhat it is. Ensembl's integrated regulatory annotation: promoters, promoter flanks, enhancers, CTCF sites and open chromatin, called across many cell types. Feature type per CpG.
Source. Ensembl Regulatory Build GFF
Citation. Zerbino et al. 2015, Genome Biol, doi:10.1186/s13059-015-0621-5
EnsRegBuild.20220710.cm.idxEnsembl Regulatory Build146 BWhat it is. Ensembl's integrated regulatory annotation: promoters, promoter flanks, enhancers, CTCF sites and open chromatin, called across many cell types. Feature type per CpG.
Source. Ensembl Regulatory Build GFF
Citation. Zerbino et al. 2015, Genome Biol, doi:10.1186/s13059-015-0621-5
EvoCons.20220314.cmEvolutionary conservation113.0 KBWhat it is. Cross-species conservation. Conserved unmethylated regions are usually developmental regulators, so conservation separates functional from incidental CpGs. phastCons score thresholded at 0.5.
Source. UCSC phastCons 60-way
Citation. Siepel et al. 2005, Genome Res, doi:10.1101/gr.3715005
GCfrac.20220323.cmGC fraction19.1 MBWhat it is. Local GC content -- the broader correlate of CpG density, and a common confounder. Fraction of G+C in a 152 bp window centred on the CpG.
Source. Derived from the reference genome
GeneFeatures.20220710.cmGene features1.4 MBWhat it is. Gene-structural partition: promoter, gene body, exon, intron, transcript. The most direct answer to whether a CpG set is regulatory or transcribed. Promoter is TSS +/-1.5 kb; gene body is the gene minus that promoter; introns are exons subtracted from the body.
Source. GENCODE vM25 (mm10) / v39 (hg38)
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
GeneFeatures.20220710.cm.idxGene features123 BWhat it is. Gene-structural partition: promoter, gene body, exon, intron, transcript. The most direct answer to whether a CpG set is regulatory or transcribed. Promoter is TSS +/-1.5 kb; gene body is the gene minus that promoter; introns are exons subtracted from the body.
Source. GENCODE vM25 (mm10) / v39 (hg38)
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
HM.20221013.cmHistone modifications5.6 MBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
HM.20221013.cm.idxHistone modifications1.6 KBWhat it is. ChIP-seq peaks for marks such as H3K4me3, H3K27ac, H3K27me3 and H3K9me3. Active marks anti-correlate with promoter methylation; repressive marks pick out Polycomb and heterochromatin. Per-mark CpG peak counts, keeping CpGs above the 300,000th-ranked count for that mark. Later builds add QC filters and keep marks with >5,000 overlapping CpGs.
Source. Cistrome Data Browser histone ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
MetagenePC.20220911.cmMetagene position2.4 MBWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
MetagenePC.20220911.cm.idxMetagene position709 BWhat it is. Where a CpG sits along an averaged protein-coding gene: upstream flank, body bins, downstream flank. The axis along which methylation's meaning flips -- promoter methylation silences, gene-body methylation tracks transcription. wzmetagene with 10 internal bins and 10 flanking steps either side.
Source. GENCODE v36 (hg38) / vM25 (mm10), protein-coding transcripts >1 kb
Citation. Frankish et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky955
PMD.20220911.cmPartially methylated domains16.3 KBWhat it is. Large late-replicating, lamina-associated domains that lose methylation with cell division. The dominant source of global hypomethylation in cancer and ageing, so a query enriched here is usually reporting proliferation history rather than regulation. PMD and the complementary commonHMD; CpGs called Neither are dropped.
Source. Zhou lab PMD coordinates
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
TFBS.20220922.cmTranscription factor binding sites54.0 MBWhat it is. Where transcription factors bind. Methylation blocks binding for some factors and is shaped by others, so TF overlap is the usual route from a CpG set to a candidate mechanism. One set per factor, capped at the top 500,000 CpGs per factor by peak-overlap count. hg38 carries about 1,359 factors.
Source. Cistrome Data Browser TF ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
TFBS.20220922.cm.idxTranscription factor binding sites13.4 KBWhat it is. Where transcription factors bind. Methylation blocks binding for some factors and is shaped by others, so TF overlap is the usual route from a CpG set to a candidate mechanism. One set per factor, capped at the top 500,000 CpGs per factor by peak-overlap count. hg38 carries about 1,359 factors.
Source. Cistrome Data Browser TF ChIP-seq
Citation. Zheng et al. 2019, Nucleic Acids Res, doi:10.1093/nar/gky1094
TFBSFreq.20220922.cmTFBS density3.7 MBWhat it is. How many distinct factors bind near each CpG, as a graded measure of regulatory density rather than per-factor membership. Per-CpG factor count from the TFBS peaks, binned at 1, 3, 10, 30, 100 and 300; CpGs with no factor form their own bin.
Source. Derived from TFBS
TFBSrm.20221005.cmTranscription factor binding sites, ReMap50.9 MBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
TFBSrm.20221005.cm.idxTranscription factor binding sites, ReMap12.2 KBWhat it is. A larger, uniformly reprocessed TF binding compendium -- around 1,188 factors -- built from published ChIP-seq rather than one lab's panel. The rm suffix is ReMap, not repeat-masked. remap2022_nr_macs2 peaks, one set per factor.
Source. ReMap 2022, non-redundant MACS2 peaks
Citation. Hammal et al. 2022, Nucleic Acids Res, doi:10.1093/nar/gkab996
Tetranuc2.20220321.cmTetranucleotide context6.5 MBWhat it is. The base either side of the CpG -- the WCGW / SCGS distinction. Solo-WCGW CpGs are the ones that lose methylation fastest with cell division. N[CG]N context, reverse-complement collapsed to 10 alphabetical tetranucleotides.
Source. Derived from the reference genome
Citation. Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
Tetranuc4.20220321.cmTetranucleotide context, uncollapsed12.5 MBWhat it is. The same context before reverse-complement collapsing. N[CG]N context as read.
Source. Derived from the reference genome
UCSCEnsemblRegulatory.20220727.cmRegulatory features via UCSC624.7 KBWhat it is. Regulatory features -- promoter, enhancer, CTCF, TFBS -- as served through UCSC. Largely redundant with EnsRegBuild; enrichment in both is one finding, not two. Feature type per CpG. The Ensembl release behind the UCSC mirror is not recorded in the lab journal.
Source. Ensembl Regulatory Build as mirrored by the UCSC track
Citation. Zerbino et al. 2015, Genome Biol, doi:10.1186/s13059-015-0621-5
Win100k.20220228.cm100 kb windows151.1 KBWhat it is. Fixed 100 kb tiles, for locating signal that follows no annotation. bedtools makewindows, contigs excluded.
Source. Reference genome index
Win1m.20230709.cm1 Mb windows16.6 KBWhat it is. Fixed 1 Mb tiles, for chromosome-scale structure. CpG position divided into 1 Mb bins.
Source. Reference genome index
Win30k.20220228.cm30 kb windows501.4 KBWhat it is. Fixed 30 kb tiles. bedtools makewindows.
Source. Reference genome index
kmer10.20231109.cm10-mer context50.0 MBWhat it is. The 10-base sequence around each CpG, for sequence-level effects finer than the immediate flanks. 10-mer per CpG.
Source. Derived from the reference genome
nFlankCG.20220321.cmFlanking CpG density7.4 MBWhat it is. How many CpGs surround this one, together with whether the immediate flanks are A/T or G/C. Local density drives methylation more strongly than almost any annotation, so this is the covariate to check before believing a subtler enrichment. Built from the WCGW / SCGS / SCGW context partition.
Source. Derived from the CpG reference and genome
nFlankCG100.20231025.cmFlanking CpG density, 100 bp6.5 MBWhat it is. Count of other CpGs within 50 bp either side. bedtools intersect count in a +/-50 bp window, minus the CpG itself.
Source. Derived from the CpG reference
nFlankCG50.20231025.cmFlanking CpG density, 50 bp4.8 MBWhat it is. Count of other CpGs within 25 bp either side. bedtools intersect count in a +/-25 bp window, minus the CpG itself.
Source. Derived from the CpG reference
rmsk1.20220321.cmRepeats, class3.2 MBWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk1.20220321.cm.idxRepeats, class374 BWhat it is. RepeatMasker classes -- SINE, LINE, LTR, DNA, satellite. Repeats are usually heavily methylated, and their demethylation is a hallmark of cancer and ageing. Verified: 21 classes in mm10. Repeat class per CpG as a categorical set.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cmRepeats, family4.4 MBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
rmsk2.20220321.cm.idxRepeats, family1.1 KBWhat it is. The same annotation one level finer -- families such as Alu, MIR, L2 and ERVL-MaLR. Verified: 59 families in mm10. Repeat family per CpG.
Source. UCSC RepeatMasker track
Citation. Smit, Hubley & Green, RepeatMasker (repeatmasker.org)
Cites: Amemiya 2019 · Gardiner-Garden 1987 · Ernst 2012 · Vu 2022 · Zerbino 2015 · Siepel 2005 · Frankish 2019 · Zheng 2019 · Zhou 2018 · Hammal 2022 · Smit
mm39/KYCG/ — 1 files, 861.6 KBChromHMM.20250102.cmChromatin states861.6 KBWhat it is. Chromatin segmentation into states such as active TSS, enhancer, transcribed, quiescent and Polycomb-repressed. Methylation means different things in each, so this is the usual first pass at what kind of region a CpG set sits in. hg38: consensus state per CpG across 833 ENCODE samples. mm10: ENCODE v5 mouse segmentations, consensus over 200 bp windows.
Source. ENCODE chromatin-state segmentations
Citation. Ernst & Kellis 2012, Nat Methods, doi:10.1038/nmeth.1906
Cites: Ernst 2012
hg38/ — 6 files, 39.3 MBseqinfo.tsv.gzAssembly sequence info221 BWhat it is. Chromosome names and lengths for the assembly. What bounds a coordinate, orders the chromosomes, and separates a primary chromosome from a scaffold. One row per sequence. The upstream assembly report it was taken from is not recorded in this table.
Source. zhou-lab/genomes
gaps.tsv.gzAssembly gaps5.1 KBWhat it is. Where the assembly has no sequence -- telomeres, centromeres, heterochromatin and inter-scaffold gaps. Regions to exclude before calling anything a coverage or a methylation difference. An interval list per assembly, conventionally the UCSC gap table; the exact download is not recorded in this table.
Source. zhou-lab/genomes
cytoband.tsv.gzCytogenetic bands9.4 KBWhat it is. Giemsa band coordinates -- the chromosome-arm coordinate system (8q24 and the like) that cytogenetic and clinical work uses, and the ideogram under a genome-wide plot. A band table per assembly, conventionally UCSC cytoBand; the exact download is not recorded in this table.
Source. zhou-lab/genomes
genes.bed.gzGene models10.2 MBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
genes.bed.gz.tbiGene models187.5 KBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
cpg_nocontig.crCpG coordinate index28.9 MBWhat it is. Every CpG in the assembly on the primary chromosomes, in coordinate order -- the row space every genome-wide .cm set and every .cg methylome is written against. A genome-wide set is a bit vector with no coordinates of its own; this is what gives its bits positions. A YAME format-7 coordinate stream (.cr). Unplaced contigs and alternate haplotypes are excluded -- what 'nocontig' means, and why a set built against the full assembly will not line up. Comes with anything else fetched from the genome. Published in zhou-lab/genomes as of tag v4, and byte-identically in each whole-genome knowledgebase repo, which is unreadable without it. Fetched with the rest of the genome annotation (`yame fetch genomes/hg38`), or on its own by name (`yame fetch hg38/cpg_nocontig.cr`) when coordinates are all a tool needs -- both write the one copy at <store>/<build>/.
Source. zhou-lab/genomes
Citation. Archived at Zenodo: 10.5281/zenodo.18175837 (hg38), 10.5281/zenodo.18175655 (mm10)
Cites: Frankish 2021
mm10/ — 6 files, 29.4 MBseqinfo.tsv.gzAssembly sequence info203 BWhat it is. Chromosome names and lengths for the assembly. What bounds a coordinate, orders the chromosomes, and separates a primary chromosome from a scaffold. One row per sequence. The upstream assembly report it was taken from is not recorded in this table.
Source. zhou-lab/genomes
gaps.tsv.gzAssembly gaps4.6 KBWhat it is. Where the assembly has no sequence -- telomeres, centromeres, heterochromatin and inter-scaffold gaps. Regions to exclude before calling anything a coverage or a methylation difference. An interval list per assembly, conventionally the UCSC gap table; the exact download is not recorded in this table.
Source. zhou-lab/genomes
cytoband.tsv.gzCytogenetic bands4.0 KBWhat it is. Giemsa band coordinates -- the chromosome-arm coordinate system (8q24 and the like) that cytogenetic and clinical work uses, and the ideogram under a genome-wide plot. A band table per assembly, conventionally UCSC cytoBand; the exact download is not recorded in this table.
Source. zhou-lab/genomes
genes.bed.gzGene models6.5 MBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
genes.bed.gz.tbiGene models197.3 KBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
cpg_nocontig.crCpG coordinate index22.7 MBWhat it is. Every CpG in the assembly on the primary chromosomes, in coordinate order -- the row space every genome-wide .cm set and every .cg methylome is written against. A genome-wide set is a bit vector with no coordinates of its own; this is what gives its bits positions. A YAME format-7 coordinate stream (.cr). Unplaced contigs and alternate haplotypes are excluded -- what 'nocontig' means, and why a set built against the full assembly will not line up. Comes with anything else fetched from the genome. Published in zhou-lab/genomes as of tag v4, and byte-identically in each whole-genome knowledgebase repo, which is unreadable without it. Fetched with the rest of the genome annotation (`yame fetch genomes/hg38`), or on its own by name (`yame fetch hg38/cpg_nocontig.cr`) when coordinates are all a tool needs -- both write the one copy at <store>/<build>/.
Source. zhou-lab/genomes
Citation. Archived at Zenodo: 10.5281/zenodo.18175837 (hg38), 10.5281/zenodo.18175655 (mm10)
Cites: Frankish 2021
mm39/ — 6 files, 29.7 MBseqinfo.tsv.gzAssembly sequence info203 BWhat it is. Chromosome names and lengths for the assembly. What bounds a coordinate, orders the chromosomes, and separates a primary chromosome from a scaffold. One row per sequence. The upstream assembly report it was taken from is not recorded in this table.
Source. zhou-lab/genomes
gaps.tsv.gzAssembly gaps2.0 KBWhat it is. Where the assembly has no sequence -- telomeres, centromeres, heterochromatin and inter-scaffold gaps. Regions to exclude before calling anything a coverage or a methylation difference. An interval list per assembly, conventionally the UCSC gap table; the exact download is not recorded in this table.
Source. zhou-lab/genomes
cytoband.tsv.gzCytogenetic bands757 BWhat it is. Giemsa band coordinates -- the chromosome-arm coordinate system (8q24 and the like) that cytogenetic and clinical work uses, and the ideogram under a genome-wide plot. A band table per assembly, conventionally UCSC cytoBand; the exact download is not recorded in this table.
Source. zhou-lab/genomes
genes.bed.gzGene models6.7 MBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
genes.bed.gz.tbiGene models198.9 KBWhat it is. Transcript models -- exons, CDS, strand -- as BED12 with a tabix index. What turns a coordinate into "in the third exon of X", and what a promoter or gene-body partition of the methylome is cut from. One BED12 line per transcript, bgzipped and tabix-indexed, so a region query reads only the blocks it needs. The per-build release above is the whole of the provenance: it used to live in a genes.bed.gz.source file beside the data, which no manifest covered and so nothing could fetch.
Source. GENCODE: human release 36 (hg38), mouse release M25 (mm10), mouse release M31 (mm39), converted from the published annotation GTF
Citation. Frankish et al. 2021, Nucleic Acids Res, doi:10.1093/nar/gkaa1087
cpg_nocontig.crCpG coordinate index22.8 MBWhat it is. Every CpG in the assembly on the primary chromosomes, in coordinate order -- the row space every genome-wide .cm set and every .cg methylome is written against. A genome-wide set is a bit vector with no coordinates of its own; this is what gives its bits positions. A YAME format-7 coordinate stream (.cr). Unplaced contigs and alternate haplotypes are excluded -- what 'nocontig' means, and why a set built against the full assembly will not line up. Comes with anything else fetched from the genome. Published in zhou-lab/genomes as of tag v4, and byte-identically in each whole-genome knowledgebase repo, which is unreadable without it. Fetched with the rest of the genome annotation (`yame fetch genomes/hg38`), or on its own by name (`yame fetch hg38/cpg_nocontig.cr`) when coordinates are all a tool needs -- both write the one copy at <store>/<build>/.
Source. zhou-lab/genomes
Citation. Archived at Zenodo: 10.5281/zenodo.18175837 (hg38), 10.5281/zenodo.18175655 (mm10)
Cites: Frankish 2021
hg38/data/ — 8 files, 59.5 MBhuman_hg38_40_celltypes_chr20.cgForty cell types, chr20 only39.5 MBWhat it is. One sample for each of 40 Loyfer cell types -- neurons, hepatocytes, pancreas, immune, epithelia, endothelium, muscle -- restricted to chr20 (773,477 CpGs). Broad enough to be a real reference panel, small enough to run in seconds -- mrmp-build turns it into 116,450 patterns in about a second. Forty is the hard ceiling, not a round number: a pattern packs as a base-3 uint64 and 3^40 < 2^64 < 3^41. Its .cg.idx IS the sample list and has to travel with it. Sorted-cell WGBS subset to chr20, one record per cell type, packed to format 3 (M/U).
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
human_hg38_40_celltypes_chr20.cg.idxForty cell types, chr20 only2.1 KBWhat it is. One sample for each of 40 Loyfer cell types -- neurons, hepatocytes, pancreas, immune, epithelia, endothelium, muscle -- restricted to chr20 (773,477 CpGs). Broad enough to be a real reference panel, small enough to run in seconds -- mrmp-build turns it into 116,450 patterns in about a second. Forty is the hard ceiling, not a round number: a pattern packs as a base-3 uint64 and 3^40 < 2^64 < 3^41. Its .cg.idx IS the sample list and has to travel with it. Sorted-cell WGBS subset to chr20, one record per cell type, packed to format 3 (M/U).
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
human_hg38_celltypes.cgFour sorted cell types6.1 MBWhat it is. Oligodendrocyte, pancreatic beta, NK and monocyte in one 4-record file, 23.6M CpGs covered at mean beta 0.847. Four known, unmixed cell types in one file make it the natural input for anything that compares samples to each other. Sample names live in the .cg.idx fetched alongside, so the file is unusable without it. Sorted-cell WGBS binarized to format 6, concatenated into one multi-sample .cg with an index of the four names.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
human_hg38_celltypes.cg.idxFour sorted cell types96 BWhat it is. Oligodendrocyte, pancreatic beta, NK and monocyte in one 4-record file, 23.6M CpGs covered at mean beta 0.847. Four known, unmixed cell types in one file make it the natural input for anything that compares samples to each other. Sample names live in the .cg.idx fetched alongside, so the file is unusable without it. Sorted-cell WGBS binarized to format 6, concatenated into one multi-sample .cg with an index of the four names.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
human_hg38_immune_mixture.cgSimulated immune mixtures12.0 MBWhat it is. Nine deconvolution mixtures with EXACT known truths, in one 9-record file. Three 70/30 pairs -- macrophage/monocyte, CD4/CD8 T cell (memory and naive pooled, so a reference must split each across its sub-types), and naive CD4/naive CD8, the hardest same-lineage pair -- each drawn at three sparsities. Record names carry both: mac70_mono30_2pow22 is macrophage 0.70 / monocyte 0.30 over 2^22 binarized CpGs. Because the mixing fractions are known, each record checks an answer rather than merely producing one; the 2^16 rung sits deliberately at the depth where a single draw becomes seed-dependent, so it shows where the method stops being trustworthy. Sample names live in the .cg.idx fetched alongside, so the file is unusable without it. Per-CpG beta blended 0.7*A + 0.3*B on CpGs both class pools cover, repacked at depth 10,000 (beta to 1e-4), then one seeded binarized draw per rung -- 1 read per sampled CpG, kept as format 3 (M/U). Components are the same pooled profiles a deconvolution reference is packed from, so the truth is exact by construction rather than nominal.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Zhou lab single-cell atlas
Citation. Zhou lab single-cell atlas 2025
human_hg38_immune_mixture.cg.idxSimulated immune mixtures313 BWhat it is. Nine deconvolution mixtures with EXACT known truths, in one 9-record file. Three 70/30 pairs -- macrophage/monocyte, CD4/CD8 T cell (memory and naive pooled, so a reference must split each across its sub-types), and naive CD4/naive CD8, the hardest same-lineage pair -- each drawn at three sparsities. Record names carry both: mac70_mono30_2pow22 is macrophage 0.70 / monocyte 0.30 over 2^22 binarized CpGs. Because the mixing fractions are known, each record checks an answer rather than merely producing one; the 2^16 rung sits deliberately at the depth where a single draw becomes seed-dependent, so it shows where the method stops being trustworthy. Sample names live in the .cg.idx fetched alongside, so the file is unusable without it. Per-CpG beta blended 0.7*A + 0.3*B on CpGs both class pools cover, repacked at depth 10,000 (beta to 1e-4), then one seeded binarized draw per rung -- 1 read per sampled CpG, kept as format 3 (M/U). Components are the same pooled profiles a deconvolution reference is packed from, so the truth is exact by construction rather than nominal.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Zhou lab single-cell atlas
Citation. Zhou lab single-cell atlas 2025
human_hg38_test.cgSparse example methylome101.3 KBWhat it is. A single deeply-undersequenced cell: 23,857 of 29.4M CpGs covered, about 0.1%, at mean beta 0.823. The sparsity is the point -- it is what a real single cell looks like, and the smallest honest input to run a command against. Its dense counterpart human_hg38_test.truth.cg is the same cell sequenced deeply (22.9M CpGs, mean beta 0.821), so the pair scores an imputation. Binarized to format 6 over the cpg_nocontig.cr row space: the universe bit marks covered CpGs, the set bit the methylated ones.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
human_hg38_test.truth.cgSparse example methylome1.9 MBWhat it is. A single deeply-undersequenced cell: 23,857 of 29.4M CpGs covered, about 0.1%, at mean beta 0.823. The sparsity is the point -- it is what a real single cell looks like, and the smallest honest input to run a command against. Its dense counterpart human_hg38_test.truth.cg is the same cell sequenced deeply (22.9M CpGs, mean beta 0.821), so the pair scores an imputation. Binarized to format 6 over the cpg_nocontig.cr row space: the universe bit marks covered CpGs, the set bit the methylated ones.
Source. zhou-lab/methscope_data -- example .cg files published alongside methscope-cli, derived from the Loyfer sorted-cell WGBS atlas
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
Cites: Loyfer 2023 · Zhou 2025
hg38/models/ — 7 files, 4.8 GBhg38_10k1.updecxUpscaling decoder (human, one-block demo)27.7 MBWhat it is. A small demonstration decoder that predicts one block of 10,000 hg38 CpGs from 101 MRMP pattern values. Use it to try `upscale` quickly before fetching the 2.7 GB whole-genome decoder; the documentation's upscale example runs it.
Architecture. UPDEC1: a single MLP decoder head for block 10k1.
Training. Trained on the same Loyfer sorted-cell atlas as hg38_wg.
Caveats. Its output is still a whole-genome .cg: the block is filled and every other CpG is NA.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
hg38_wg.updecxUpscaling decoder (human, whole genome)2.7 GBWhat it is. Predicts methylation at all 29,401,795 hg38 CpGs from a sparse methylome, such as a single cell or a low-coverage sample. Inference runs in C, in about 2 s per sample.
Architecture. UPDEC2. The genome is divided into 700 units of about 16,000 CpGs, each with its own low-rank factor head (bottleneck 16 for pure units, 32 for mixed, leaky activation). The encoder input is 836 values: 835 MRMP pattern betas plus log1p of the number of CpGs observed.
Training. The patterns come from a 67-class Loyfer reference pooled from the 152 training samples only, cut to 500 patterns by CpG count and joined with the 100 resolver pairs of smallest footprint from the classifier bank (mrmp-pool --pooled-top 0) (835 columns in all). The model was trained on the 152/45/10 training/validation/test split of the 207-sample Loyfer atlas, against 100 binarized simulated replicates with coverage log-spaced from 71 to 2,940,180 observed CpGs (skew 0.5, dense-weighted). Each unit stops early on validation MAE (patience 400, 60,000-step cap); training took 7 h 38 min on one A100.
Caveats. At 2.7 GB this is the largest file in the catalogue; fetch it deliberately.
History. v10 (2026-09-16) changed where the patterns come from. Earlier references had no hepatocyte class, which caused an undercall at hepatocyte loci; at ELF3 with 29,000 observed CpGs the error fell from 0.309 to 0.047. On the 10 held-out samples v10 is more accurate at both ELF3 and MIR200C at every coverage tested (ELF3 0.113 to 0.107 at 29k, 0.098 to 0.089 at 290k, 0.093 to 0.089 at 2.9M; MIR200C 0.078 to 0.073, 0.072 to 0.069, 0.072 to 0.069). v9 and earlier took 500 patterns from a 35-class Zhou reference and trained at patience 20. v4 (2026-07-29) used a curated 117/45/45 split for a like-for-like comparison with the published MLP baseline.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
hg38_celltype_full.clfxCell-type classifier (human, full)153.8 MBWhat it is. Assigns a human methylome, from a single cell to a bulk sample, to one of 63 cell types. Use it when accuracy matters most; hg38_celltype_lite is a seventh of the size.
Architecture. One flat MRMP bank: hierarchical LCA binstrings plus a calibrated resolver for every pair of cell types (1,981 pattern sets, 1,953 resolvers), all scored by a single pooled xgboost booster. With no routing tree, there is no early decision whose mistake a later node cannot undo.
Training. The reference is 84,601 labelled cells: pseudobulks from the Zhou lab 2025 single-cell atlas, the Loyfer sorted-cell atlas and BLUEPRINT. Built 2026-09-06 with mrmp-build --bank --resolver-gate -1 and classify-train --pool-nodes (max depth 4, colsample 0.4), over a 24-rung coverage ladder with binarized reads. The final model was assembled from the per-fold resolver cache. Both flags have since been retired (--resolvers N replaces the gate, and the single-booster mode is gone), so the current recipe is the one on the methscope training page.
Accuracy. Class-balanced 0.9544 / micro 0.9569 in ten-fold cross-validation over all 84,601 cells.
History. It replaces the routing-tree classifier hg38_celltype (33 classes), withdrawn at model tag v10.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
hg38_celltype_lite.clfxCell-type classifier (human, lite)21.2 MBWhat it is. A compact version of hg38_celltype_full, about a seventh of its size (21 MB against 154 MB), covering the same 63 cell types. Use it when download size or memory matters.
Architecture. The same hard blocks as the full model, with calibrated resolvers only for the 150 cell-type pairs with the thinnest hard-block footprint (the fewest CpGs where the two class pseudobulks differ by more than 0.5), and a 100-round booster.
Training. As hg38_celltype_full, with mrmp-build --min-pattern-cpgs 500 --resolvers 150 and classify-train -n 100 (methscope-cli 60988b9).
Validation. Correctly calls the four Loyfer example cells (ODC, Beta, NK CD16, Mono).
Accuracy. Class-balanced 0.9403 / micro 0.9558 in ten-fold cross-validation at full coverage, against 0.9544 / 0.9569 for the full model. The gap widens on sparse data: at 4,096 observed CpGs, class-balanced accuracy is 0.6685 against 0.7612.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
hg38_sex.clfxSex classifier (human)30.2 KBWhat it is. Calls a sample female or male. At 30 KB it is the smallest model in the catalogue, and a quick way to check that an installation works end to end.
Architecture. A logistic model over two X-inactivation features, Xa_lo and Xa_hi, from a three-state MRMP.
Accuracy. About 95.8% on an independent cohort.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6
hg38_33celltypes.msdrefDeconvolution reference (human, bulk)528.4 MBWhat it is. Estimates the cell-type composition of a methylome against 33 human cell types. Prefer it when the query is itself sorted or bulk tissue; for single-cell-derived data use hg38_62celltypes, whose immune types come from single cells.
Architecture. Covers the 7.9 million CpGs (26.9% of hg38) that differ between at least some of the types. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart under one joined label.
Training. 286 normal bulk WGBS samples, grouped into 33 labels by a donor-resolved control sheet, pooled per class by read count (musum), then packed with deconv-build-ref at its default admission band 0.30,0.70.
Validation. A shipped build uses every donor, so the confusion matrix was measured on a donor-held-out twin: per label, every fourth donor by sample-ID hash was held out as mixture components, giving 480 mixtures of 2 to 5 classes pooled over 2^16 to 2^22 CpGs. (The control sheet's own split column is a cohort flag covering only 9 of the 33 classes, so it could not be used.) The closest pairs are CD4 and CD8 T cells (0.43), macrophage and monocyte (0.23), dendritic cell and monocyte (0.17), bladder and prostate epithelium (0.12), and NK and CD8 T cells (0.10).
Caveats. Its macrophage class is monocyte-derived in vitro, so tissue-resident macrophages have no matching reference.
History. It replaced hg38_65celltypes.refx, withdrawn at model tag v7; methscope no longer reads that format. Version 3 of the format since model tag v12 adds the confusion matrix; the M/U block is byte-identical to version 2.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
hg38_62celltypes.msdrefDeconvolution reference (human)1.4 GBWhat it is. Estimates the cell-type composition of a methylome against 62 human cell types. Because its immune types come from single cells rather than sorted bulk, it resolves tissue-resident immune cells that the bulk-based hg38_33celltypes misreads.
Architecture. Sixty types are pseudobulks from the Zhou lab single-cell atlas (the early-embryonic syncytio- and villous trophoblast are excluded); Loyfer hepatocytes and BLUEPRINT granulocytes cover the two compartments single cells miss. It covers the 11.8 million CpGs (40.1% of hg38) that differ between at least some of the types, packed as uint16 M/U: a row is kept only where at least one class is unmethylated and at least one is methylated. No pattern budget is fixed; the solver selects its patterns for each query from the CpGs that query measured. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart under one joined label; a group total is identified even where its split is not.
Training. Per-class M/U pooled by read count (musum), then packed with deconv-build-ref at its default admission band 0.30,0.70.
Validation. A known 70/30 macrophage/monocyte mixture comes back within a few points of truth, where the bulk reference returns 0.14/0.56 and puts the missing mass on mast cells and stroma. T cells resolve to naive and memory CD4 and CD8. The confusion matrix was measured on a fold-held-out twin: pools rebuilt from Zhou cells in folds 1-9, fold 0 held out as mixture components, and Hepatocyte and Granulocyte held out by donor. The closest pairs are naive CD4 and CD8 T cells (0.29), colonic enterocyte and goblet cells (0.24), the CGE interneurons LAMP5, SNCG and VIP (0.22), muscle fibroblast and its adipogenic progenitor (0.20), NK CD16 and CD56 (0.15), adrenal ZG and ZR/ZF (0.14), cortical IT L4 and L5 (0.12), AT1 and AT2 (0.12), and Beta and Delta (0.11). A threshold of 0.15 joins five pairs, and 0.10 joins thirteen.
History. Version 3 of the format since model tag v12 adds the confusion matrix; the M/U block is byte-identical to version 2.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Loyfer et al. 2023, Nature, doi:10.1038/s41586-022-05580-6; Stunnenberg et al. 2016, Cell, doi:10.1016/j.cell.2016.11.007 (BLUEPRINT); Zhou lab single-cell atlas 2025
Cites: Loyfer 2023 · Stunnenberg 2016 · Zhou 2025
mm10/models/ — 4 files, 2.5 GBmm10_wg.updecxUpscaling decoder (mouse, whole genome)1.9 GBWhat it is. Predicts methylation at all 21,867,837 mm10 CpGs from a sparse methylome. Inference runs in C.
Architecture. UPDEC2, the same architecture as hg38_wg: 472 units, each with its own low-rank factor head (bottleneck 16 for pure units, 32 for mixed, leaky activation). The encoder input is 501 values: 500 MRMP pattern betas plus log1p of the number of CpGs observed.
Training. The patterns come from 40 mouse-brain pseudobulk cell types of the Liu 2021 atlas. The atlas has 41, but MRMP packs a pattern as a base-3 uint64 and so caps at 40 samples; VLMC-Pia, a sub-split of the retained VLMC, was dropped. The training targets are 200 of the 258 Zhou 2018 mm10 methylomes with at least 80% genome coverage, against 100 binarized replicates log-spaced from 71 to 2,186,784 observed CpGs (skew 0.5). 60,000-step cap, patience 20; no unit reached the cap.
Accuracy. On held-out samples the mean absolute error is 0.098; with only 106 observed CpGs it is 0.134, at Pearson r 0.80.
Caveats. At 1.9 GB, fetch it deliberately.
History. Training across a coverage ladder is what closed the gap to the human model. A fixed-coverage build plateaued at 0.126 MAE; at 106 observed CpGs the ladder improved the error from 0.241 to 0.134 and Pearson r from 0.48 to 0.80.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8; Zhou et al. 2018, Nat Genet, doi:10.1038/s41588-018-0073-4
mm10_brain_full.clfxCell-type classifier (mouse brain, full)64.5 MBWhat it is. Assigns a mouse-brain methylome to one of 41 brain cell types from the Liu 2021 snmC-seq2 atlas.
Architecture. One flat MRMP bank: LCA binstrings plus a calibrated resolver for every pair of cell types (837 pattern sets, 820 resolvers), scored by a single pooled booster with no routing.
Training. The reference is 41 all-cell binasum pools over the 54,447 cells of a 2,000-per-class draw. The booster was trained on a balanced 400-per-class draw (14,825 cells), over a 24-rung coverage ladder with binarized reads. The final model was assembled 2026-09-11 (8 calibration strides plus assembly, about 6.5 h).
Accuracy. Class-balanced 0.9746 / micro 0.9720 in ten-fold cross-validation. On the 49,113 Liu cells that never entered the model (18 of the 41 classes, the large cortical ones): micro 0.9660 / class-balanced 0.9707.
Caveats. It is brain only. With no non-brain classes, a cell from another tissue is assigned to the nearest brain type without warning, so confirm the tissue first.
History. It replaces the routing-tree classifier mm10_celltype_brain, withdrawn at model tag v10.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
mm10_brain_lite.clfxCell-type classifier (mouse brain, lite)13.3 MBWhat it is. A compact version of mm10_brain_full, about a fifth of its size, covering the same 41 cell types.
Architecture. The same hard blocks as the full model, with calibrated resolvers only for the 100 cell-type pairs with the thinnest hard-block footprint (the fewest CpGs where the two class pseudobulks differ by more than 0.5), and a 100-round booster.
Training. As mm10_brain_full, with mrmp-build --min-pattern-cpgs 500 --resolvers 100 and classify-train -n 100 (methscope-cli 60988b9).
Accuracy. On the 49,113 held-out Liu cells: micro 0.9583 / class-balanced 0.9638 (18 classes), within a point of the full model (0.9660 / 0.9707). On fold-0 cross-validation test cells, class-balanced accuracy is 0.9778 at full coverage (full model 0.9789) and 0.8195 at 4,096 observed CpGs (0.8778).
Caveats. Like every mouse classifier here, it is brain only.
History. It beats the v9 lite (micro 0.9482 / class-balanced 0.9595), which had 117 gate-picked resolvers and was twice this size.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
mm10_41celltypes.msdrefDeconvolution reference (mouse brain)554.9 MBWhat it is. Estimates the cell-type composition of a mouse-brain methylome against 41 major cell types from the Liu 2021 atlas.
Architecture. Covers the 6,765,325 CpGs (30.9% of mm10) that differ between at least some of the types, packed as uint16 M/U. The file carries a confusion matrix, which `deconv --group-threshold` uses to report types that cannot be told apart, such as CA3 and DG-po, under one joined label; a group total is identified even where its split is not.
Training. Each type is a pseudobulk of binarized single cells from the Liu 2021 snmC-seq2 atlas, pooled with yame rowop -o binasum over all 54,447 cells, so every cell counts once per CpG. The pools are concatenated in sorted label order and packed with methscope deconv-build-ref at qfilter 0.30,0.70.
Validation. An all-cells build has no held-out cells, so the confusion matrix was measured on a fold-held-out twin, each class alone at 2^16 to 2^22 CpGs, pooled across rungs. On P21 mouse-brain spatial data it agrees with an independent RNA-based deconvolution (RCTD): r = 0.74 for oligodendrocytes, 0.69 for CA1, 0.69 for dentate gyrus and 0.59 for CA3.
Caveats. Seven types rest on fewer than 300 cells, and grouping merges them into the neighbour that already absorbs them. The cortical IT ladder stays four separate labels.
History. The first version-3 .msdref, the format that carries a confusion matrix.
Source. zhou-lab/methscope on HuggingFace, published with methscope-cli
Citation. Liu et al. 2021, Nature, doi:10.1038/s41586-020-03182-8
-h, as this build prints itGenerated from the binary, so nothing here can drift from what
yame <command> -h says. Each entry is
folded; click to open. Every highlighted command in an example on the
Usage tab links to its entry here. Grouped as the bare
yame banner groups them.
yame pack — Pack text/bed-like inputs into a .cx streamUsage:
yame pack [options] <in.txt> <out.cx>
Pack tab-delimited text into a compressed cx file.
The input file must have one row per CpG and match the
dimension and order of the reference CpG BED file.
Options:
-f [char] Format specification (one of b,c,s,m,d,n,r):
(b) Binary data (format 0).
Each entry is 0 or 1.
Example (single-sample, one column):
0
1
1
(c) Character data (format 1).
ONE character per line -- the byte is stored as typed,
not parsed as a number, so a longer value is refused
rather than truncated to its first character.
Example:
0
5
9
(s) State data (format 2).
Categorical strings compressed via an index + RLE.
Best for chromatin states or other labels.
Example:
quies
quies
enhA
(m) Sequencing MU data (format 3).
Input is 2-column text: M and U counts per CpG.
M=U=0 is treated as missing.
Example (M U):
10 5
20 0
13 17
(d) Differential / mask data (format 6).
2-bit boolean for S (set) and U (universe).
Input is 2-column text: S and U, each 0 or 1.
Example (S U):
1 1
0 1
0 0
(n) Fraction / beta data (format 4).
Floating-point fraction in [0,1] or NA.
Example:
0.250
NA
1.000
(r) Reference coordinates (format 7).
Compressed BED records for CpG coordinates.
Input is 4-column BED: chrom, start, end, name.
Example:
chr1 100 101 CpG_1
chr1 200 201 CpG_2
chr1 300 301 CpG_3
The examples above show single-sample input.
Multi-sample input can be provided as additional
columns per row, following the same conventions.
-u [int] Number of bytes per unit when inflated (1-8).
Lower values are more memory efficient but may be lossier.
0 - infer from data.
-v Verbose mode.
-h Display this help message.
yame unpack — Unpack a .cx stream back to textUsage:
yame unpack [options] <in.cx> [sample1 sample2 ...]
Purpose:
Print selected records from a .cx file as a tab-delimited table.
Each output row is a genomic row index; each output column is a selected sample/record.
Sample selection (default: first record):
-a Output all records in the file, one COLUMN per record (a matrix,
not stacked rows); -C names the columns.
-l <list> Sample list file (one name per line).
Ignored if sample names are provided as trailing arguments.
-H <N> Output the first N samples.
-T <N> Output the last N samples (requires index).
Row coordinates (optional first column):
-R <rows.cx|name> Row coordinate dataset (CX; typically format 7).
A name works: -R hg38 finds it in the store. Not
inferred, since omitting it means no coordinate column.
-r <mode> Coordinate print mode (default: 0):
0: chrm<tab>beg0<tab>end1 (cg-style)
1: chrm<tab>beg0<tab>end0 (allc-style)
else: chrm_beg1
Output formatting:
-C Print a header line (column names).
-u <bytes> Inflated unit-size override (0=auto; allowed: 1,2,4,6,8).
Value printing (-f):
-f <N> Print mode for certain formats (default: 0 for format 6;
REQUIRED for format 3, which has no default):
For format 3 (MU):
N == 0 : print packed MU (uint64) -- raw storage, not a beta
N < 0 : print M<tab>U (two columns)
N > 0 : print beta; print NA if cov < N or cov==0
For format 6 (set+universe):
N == 0 : print 0/1, NA coded as '2'
N < 0 : print value<tab>universe (e.g., 1<tab>1, 0<tab>1, NA<tab>0)
N > 0 : print raw 2-bit code (FMT6_2BIT)
Chunked printing:
-c Enable chunked printing (reduces peak memory).
-s <rows> Chunk size in rows (default: 1000000).
Other:
-h Show this help message.
Notes:
* Selecting by sample name or using -T requires an index (.cxi) unless reading from stdin.
* Chunking does not support format 7 datasets.
yame hprint — Horizontal printing (primarily format 6)Usage:
yame hprint [options] <in.cx>
Modes:
-R <ref|name> Whole-genome view: one column per CpG window across all chroms.
-r <reg> Region view: rows=samples, columns=CpG sites in region.
-R is OPTIONAL here: the file's row count identifies its
reference, and the store is searched for it. So `-r chr16`
alone is usually enough.
(neither) Full-dataset dump: every row, no windowing. fmt0/3/4/6.
Options:
-c Never colour. Colour is on only when stdout is a terminal (and TERM is not dumb), so a redirect or a pipe is plain text already.
-g Granular output: 0-9 deciles instead of H/M/L
-s <file> Record names for the input, one per line, in order,
the same file `index -s` takes. Names live in the .idx
sidecar, which a PIPE does not carry, so this is how a
streamed view gets labels. A count that disagrees with the
records is an error.
-R <ref.cr|name> Reference coordinates (format 7).
OPTIONAL with -r -- inferred from the row count.
A name works too: -R hg38 finds it in the store.
-r <region> Genomic region: chr16 or chr16:10000000-10100000
-l <int> Sample label column width (default: 20)
-t <int> Ruler tick interval in columns (default: 10)
-w <int> Max data columns; wider views are window-averaged (default: 80)
-h This help
Symbols (per-site): fmt6: █ meth ░ unmeth . NA fmt3/4: H/M/L/. fmt0: 1/0
Symbols (windowed): H >0.67 M 0.33-0.67 L <0.33 . no coverage
Annotation: a sample whose name ends with '_' is shown without the '_' and its whole row is underlined
yame index — Create/refresh a sample index for a .cx fileUsage:
yame index [options] <in.cx>
The index file name default to <in.cx>.idx
Options:
-s [file path] tab-delimited sample name list (use first column)
-1 [sample name] add one sample to the end of the index
-c output index to console
-h This help
yame split — Split a multi-sample .cx into single-sample filesUsage:
yame split [options] <in.cx> out_prefix
Options:
-v verbose; also says which path ran
-s sample name list
-h This help
Notes:
* Each output holds one record unchanged, so when the input is indexed
and its records start on BGZF block boundaries, the compressed bytes
are copied out verbatim -- no decompression, no re-compression. An
unindexed or unaligned input falls back to decode/encode.
yame info — Show basic metadata/parameters of a .cx fileUsage:
yame info [options] <in.cx>
Options:
-1 Report one record per file.
-h This help
Columns:
File, Sample, NSample (samples in the FILE, from its .idx --
NA without one), Nrow, Format, UnitBytes, Keys.
yame subset — Subset samples from a .cx (or terms from format 2 with -s)Usage:
yame subset [options] <in.cx> [sample1 sample2 ...] > out.cx
Purpose:
Subset a multi-sample .cx by sample names (requires an index), or
(with -s) convert a format-2 state track into one binary track per state.
Modes:
(A) Sample subsetting (default):
Select named samples from <in.cx> and emit them in the given order.
Requires <in.cx>.cxi index.
(B) Subset format-2 states (-s):
Interpret <in.cx> as a single format-2 dataset (must be fmt2).
For each requested state name, emit one format-0 bitset where
bit=1 iff row state == that term.
Input sample list:
Provide sample names either:
* as trailing arguments on the command line, OR
* via -l <list.txt> (one name per line).
Options:
-o <out.cx> Write output to a file. If provided, an output index (.cxi)
is also generated. If omitted, writes to stdout (no index).
-l <list> Path to sample/state list. Ignored if names are provided as
trailing command-line arguments.
-s Format-2 state filtering mode (output format 0; one record per term).
-H <N> If no names are provided, take the first N samples from the input index.
-T <N> If no names are provided, take the last N samples from the input index.
-z <0-9> Re-encode the output at this zlib level, turning off the raw
copy below. Only useful when the copy cannot run, or when you
want a different compression than the input has: -z1 is much
faster than the default 6, -z0 stores. Any level reads back
identically.
-v Say which path ran (raw copy or re-encode).
-h Show this help message.
Notes:
* A subset copies whole records unchanged, so when every requested record
starts on a BGZF block boundary its compressed bytes are moved verbatim
-- no decompression, no re-compression. Records written to their own
file and concatenated are aligned; records appended to a shared writer
are not, and those fall back to decode/encode. -v says which ran.
* -H/-T only apply when you did NOT provide an explicit name list.
* -T requires an index (same as default sample subsetting).
* In -s mode, the input is expected to be a single fmt2 record; the output
contains one fmt0 record per requested term/state.
yame rowsub — Subset rows by index list / mask / coordinates / block rangeUsage:
yame rowsub [options] in.cx >out.cx
Purpose:
Subset (slice) rows from each dataset (record) in a CX stream.
Output is always written to stdout.
Row selection modes (choose one):
(A) Explicit row indices (1-based list):
-l <idx.txt> One [index1] per line (1-based). Order preserved; no sorting required.
(B) Explicit genomic coordinates via row coordinate table (format 7):
-R <rows.cx|name> Row coordinates (format 7; BED-like).
OPTIONAL with -L or -1: inferred from the input's row
count. A name works too: -R hg38 finds it in the store.
-L <coord.txt> One [chrm]_[beg1] per line (1-based beg). Needs coordinates,
which are inferred when -R is not given.
Order preserved; no sorting required.
-1 If -R is provided, emit the subsetted row coordinates as the FIRST dataset.
Neighbourhood and row map (with -l or -L):
-w <N> Keep rows [i-N, i+N] around each selected row i, clipped to
its chromosome; one window per query, in query order, overlaps
kept. Needs coordinates (inferred like -L); not for array rows.
-M <map.tsv> Write one line per output row: query, offset from the
query's own row (0 without -w), 1-based row.
(C) Mask-based filtering (binary mask):
-m <mask.cx> Mask file (format 0/1 only). Rows with bit=1 are kept.
(D) Contiguous block by absolute row range (0-based):
-B <beg0>[_<end1>]
Keep rows in [beg0, end0] where end0 = end1-1.
If <end1> is omitted, keep a single row at beg0.
(E) Contiguous block by block index and size (0-based):
-I <blockIndex0>[_<blockSize>]
Keep rows:
beg0 = blockIndex0 * blockSize
end0 = (blockIndex0+1)*blockSize - 1
If <blockSize> is omitted, default blockSize=1000000.
Other options:
-h Show this help message.
Index conventions:
- '0' suffix means 0-based (beg0, blockIndex0).
- '1' suffix means 1-based (index1, beg1, end1).
- For -B, end is provided as end1 (exclusive, 1-based), internally converted to end0.
Notes:
* For format 2 (state data), the key section is preserved when slicing.
* Format 7 (row coordinates) is sliced with fmt7_* helpers.
* If multiple selection options are given, the effective precedence is:
-l/-L > -m > -B/-I > default.
yame chunk — Chunk binary CX into smaller fragmentsUsage:
yame chunk [options] <in.cx> <outdir>
Options:
-v verbose
-s chunk size
-h This help
yame chunkchar — Chunk text data into smaller fragmentsUsage:
yame chunkchar [options] <in.txt>
Options:
-v verbose
-s chunk size
-h This help
yame summary — Summarize query features, optionally against masksUsage:
yame summary [options] <query.cx> [query2.cx ...]
Purpose:
Summarize a query feature set (or per-state composition) and optionally
its overlap/enrichment against one or more masks.
Input:
<query.cx> may contain one or multiple samples (records). Supported query
formats: 0/1 (binary), 2 (state), 3 (MU counts), 4 (float),
6 (set+universe), 7 (genomic coordinates).
Masking:
-m <mask.cx|name> Optional mask feature file (can be multi-sample).
A name is looked up in the store for this query's row
space: -m ChromHMM finds the newest ChromHMM set.
-b Browse the catalogue instead, choose masks there, and
summarize against each. Fetches what is missing. The tree
opens at the query's own row space; with -m, those set
names arrive checked (`-b -m CGI,ChromHMM`). d changes the store.
If provided, every query sample is summarized against every
mask sample (cartesian product).
-M Load all masks into memory, for a mask file on slow IO.
It no longer affects SPEED: a query is read once whatever
this says, and the masks stream either way.
Also auto-enabled when the mask stream is unseekable.
-I Invert the mask file into an index keyed by row, built once
and reused for every query record. Pays a build about as
long as four record walks, then almost nothing per record,
and holds 4 bytes per row plus 2 per membership (855 MB for
TFBS). Worth it for many records against many masks; the
default walk is right for a few. Format 3 queries and
binary masks; anything else walks. Declines and walks past
YAME_SUMMARY_INDEX_MB (default 2048), saying what it needed.
Naming / output formatting:
-H Suppress the header line.
-F Use full paths in QFile/MFile (default: basename only).
-T Always include section/state names in output labels when
summarizing format-2 (state) data.
-s <list.txt> Override query sample names using a plain-text list.
Only applies to the first query file.
Stdin helpers:
-q <name> Backup query file name used only when <query.cx> is '-'.
Format-6 view (the 2 bits are used with more than one meaning):
-V <view> set (default) universe = background, set = feature member
meth universe = covered, set = methylated
2bit count the 4 quaternary states separately
-6 Deprecated alias for -V 2bit.
Other:
-h Show this help message.
Output columns (-V set, the default; also -V 2bit):
QFile Query MFile Mask N_univ N_query N_mask N_overlap Log2OddsRatio Beta Depth
Output columns (-V meth):
QFile Query MFile Mask N_covered N_meth N_covered_in_mask N_meth_in_mask Log2OddsRatio Beta Beta_bg
Notes:
* For state masks (format 2), summary is emitted per state key (one row per key).
* When no mask is given, Mask is reported as 'global'.
* -V meth requires every query record to be format 6; Beta is then the
methylated fraction inside the mask (whole sample when no mask is given)
and Beta_bg the same fraction outside the mask.
yame pairwise — Call pairwise differential methylation (fmt3 -> fmt6)Usage:
yame pairwise [options] <MU1.cx> [MU2.cx] > out.cx
Purpose:
Compare two samples site-by-site on the same row space. By default, write the
differential-methylation set as one format-6 track (set + universe). With -S,
print agreement statistics instead (one TSV line per pairing).
Inputs:
<MU1.cx> Format 3 (M/U) or format 4 (beta). Sample 1 is the record named by -1,
or the first record.
[MU2.cx] Optional second file. Sample 2 is the record named by -2, or the first
record; with -S and no -2, EVERY record of MU2.cx is scored against
sample 1. If omitted, sample 2 is the SECOND record of MU1.cx.
Output:
Default: one format-6 record of length N (same as the inputs).
Universe: site i is in-universe only if BOTH samples cover it (-c for format 3;
a value that is not NA for format 4).
Set: site i is set if it passes the direction rule (-H) and effect threshold (-d).
-S: TSV with a header: sample1 sample2 n mae rmse acc pearson, over the
universe. acc = share of sites where both betas fall on the same side of -t.
Options:
-1 <name> Sample to take from MU1.cx (needs MU1.cx.idx).
-2 <name> Sample to take from MU2.cx (needs MU2.cx.idx).
-S Print agreement statistics instead of writing a set.
-t <beta> Threshold for the acc column (default: 0.5). -S only.
-o <out.cx> Write output to file (default: stdout).
-c <cov> Minimum coverage (M+U) in a format-3 sample to count a site (default: 1).
-d <delta> Minimum absolute beta difference required to call a site differential (default: 0).
-H <mode> Direction mode (default: 1):
1 beta1 > beta2 (hypermethylated in sample 1)
2 beta1 < beta2 (hypomethylated in sample 1)
3 beta1 != beta2 (any difference; with -d uses |beta1-beta2|>delta)
-h Show this help message.
Notes:
* If you omit MU2.cx, MU1.cx must contain at least two records.
* The set output is binary; it does not store the delta magnitude.
* To score imputation, first blank the sites the model was shown:
yame mask truth.cg query.cg | yame pairwise -S - prediction.cg
yame binarize — Convert fmt3 (M/U) to fmt6 (set+universe) by beta/M thresholdUsage:
yame binarize [options] <mu.cx>
Purpose:
Convert per-site M/U counts (format 3) into a packed binary-with-universe
track (format 6).
Input / Output:
Input : format 3 (.cx) with per-site (M,U) stored as uint64.
Output: format 6 (.cx), where each site stores two bits:
- universe bit: 1 if depth>=min_cov, else 0 (NA/outside-universe)
- set bit: 1 if methylated by rule, else 0
Binarization rules:
Default: set=1 if beta > T (beta = M/(M+U)), set=0 otherwise.
If -m is provided (>0): set=1 if M >= Mmin, else 0 (overrides -t).
Universe is always defined by coverage: (M+U) >= min_cov.
Options:
-t <Tmin> Beta threshold (default: 0.5).
-m <Mmin> M-count threshold (default: 0; if >0 overrides -t).
-c <cov> Minimum coverage (M+U) to include a site in universe (default: 1).
-o <out.cx> Write output to file (default: stdout).
-h Show this help message.
Notes:
* Sites with depth < min_cov remain NA in format 6 (universe bit = 0).
* If the input has a sample index and -o is used, an output index is written.
yame mask — Invalidate specific CpG sites using an external mask file (deterministic, file-driven: controls *which* sites are valid)Usage:
yame mask [options] <in.cg> <mask>
Blank the sites the mask covers; every other site is left as it is.
`mask x.cg Blacklist.cm` REMOVES the blacklisted sites; -v keeps only
them. The mask may be format 0, 1, 3 (covered = M+U > 0) or 6 (covered
= in the universe), so a methylome can mask another: `mask truth.cg
query.cg` blanks the sites the query was shown. A bare name (CGI,
Blacklist) resolves in the store for the input's row space, as -m does
for summary. Row counts must match.
Options:
-o output cx file name. if missing, output to stdout without index.
-c contextualize binary input to format 6 using '1's in mask.
implicit for format 6 input (output is always format 6).
-v invert the mask: blank the sites it does NOT cover.
-h This help
yame dsample — Randomly subsample N covered CpG sites, masking the rest (stochastic, rate-driven: reduces site count and coverage)Usage:
yame dsample [options] <in.cx> [out.cx]
Downsample methylation data for format 3 or 6.
- For format 3, downsampling masks by setting M=U=0.
- For format 6, downsampling masks by clearing the universe bit.
Options:
-o [PATH] output .cx file name.
If missing, write to stdout (no index will be written).
-s [int] seed for random sampling (default: current time).
Fix it to make a downsampling reproducible: the same seed
and input give the same sites every run.
-b After sampling, randomly binarize sampled format 3 (MU)
rows. Output is still format 3.
-N [int] number of records to sample/keep per sample (default: 100).
If N >= available records, all available records are kept.
-r [int] number of downsampled replicates per input sample (default: 1).
Each replicate is independently re-sampled.
-p [str] replicate sample name prefix [default: None].
If given, the out sample name is: [sname]-[pre]-[rep_id].
-h this help.
yame perturb — Randomly flip 0/1 bits in fmt0 or fmt6 (noise injection)Usage:
yame perturb [options] <in.cx>
Randomly flip 0/1 bits for format 0 and format 6.
- For format 0, each set bit (1) or unset bit (0) is independently
flipped with probability p.
- For format 6, only in-universe sites are eligible; their set bit
is flipped with probability p. NA sites are left unchanged.
Options:
-s [int] random seed (default: current time).
-p [float] fraction of CpGs to flip, in [0,1] (default: 0.05).
-o [PATH] output .cx file (default: stdout, no index written).
-h this help.
yame rowop — Row-wise operations (e.g., sum / combine binary tracks)Usage:
yame rowop [options] <in.cx> [out]
Purpose:
Perform row-wise operations across multiple records (samples) in a CX file.
Depending on the operation, output is either a new CX file or plain text.
Operation:
-o <op> Operation name (default: binasum)
CX-output operations:
binasum Convert per-sample values into per-row sample counts (M/U) as format 3.
Input: fmt0, fmt1, or fmt3.
For fmt3, beta thresholds (-p/-q) define methylated vs unmethylated calls.
musum Sum MU sequencing counts across samples.
Input: fmt3 only. Output: one fmt3 record.
Text-output operations:
stat Per-row summary statistics across samples.
Input: fmt3 only.
Output columns:
count mean_beta sd_beta delta_beta min_n delta_mean q95_0 q05_1 delta_q
q95_0 = 95th pct of beta<0.5; q05_1 = 5th pct of beta>0.5.
delta_q = q05_1 - q95_0: delta_beta's worst-case idea, but tolerant of
an outlier per side. Both quantiles are reported, not just the gap,
because a filter usually constrains ONE side -- "no expected-0 sample
creeps toward 0.5" is q95_0, which the difference cannot express.
delta_beta = min(beta>0.5) - max(beta<0.5) (worst-case margin).
min_n = min(#beta<0.5, #beta>0.5).
delta_mean = mean(beta>0.5) - mean(beta<0.5) (group-center separation).
binstring Convert per-sample beta values into row-wise binary strings.
Input: fmt3 only. Uses -b as the beta threshold, -c as min coverage.
Ambiguous cells (mu==0, cov<mincov, or beta==threshold) are filled
with the CpG's majority state (deterministic; -s does not apply).
cometh Neighbor co-methylation summary within a window.
Input: fmt3 only.
Output: packed 4-way counts (UU, UM, MU, MM) per neighbor offset.
Use -v to print unpacked lanes.
Common filters:
-c <mincov> Minimum coverage (M+U) for a sample/row to contribute (default: 1).
-d <0-9> Decimals printed for stat's fractions (default 6). The old
default of 3 put many values on a rounding boundary: a beta is a
ratio of small integers, so a mean like 9/400 = 0.0225 sits exactly
between 0.022 and 0.023 and any change to the arithmetic flips it.
It also matters downstream, since a threshold applied to a printed
value inherits half the last digit as slop.
Threads:
-t <N> Split the records over N threads (default 1). Applies to
binasum, musum and stat, whose accumulators combine; the other
ops stay serial. An indexed file lets each thread seek to its
own run, so the decompression parallelises too; a stream or an
unindexed file is read once and dispatched to the threads, which
still parallelises everything after the inflate. Each thread
holds its own accumulator, so memory is N x (8 bytes/row) for
binasum/musum and N x (84 bytes/row) for stat -- at hg38 scale,
235 MB and 2.2 GB per thread (stat also carries a 16-bin beta
histogram per row for delta_q). Output is byte-identical at every
N for all three ops.
binasum (fmt3 input) thresholds:
-p <beta0> Call unmethylated if beta < beta0 (default: 0.4).
-q <beta1> Call methylated if beta > beta1 (default: 0.6).
Betas in [beta0, beta1] are ignored.
binstring options:
-b <beta> Call methylated if beta > threshold (default: 0.5).
-m <frac> Max ambiguous fraction per CpG; above this the line is emitted
as an all-'2' sentinel (default: 1.0 = no filtering).
-M <fold> Min majority fold (hi/lo of confident calls) to trust the fill
(default: 10). If a CpG has ambiguous cells but the majority is
below this fold, its line is emitted as the all-'2' sentinel.
cometh options:
-w <W> Neighbor window size (default: 5).
-v Verbose output (print UU-UM-MU-MM instead of packed uint64).
Other:
-h Show this help message.
yame fetch — Download reference assets into the shared store ('yame fetch' to browse, 'yame fetch -l' to dump the registry)Usage:
yame fetch browse the catalogue
yame fetch [options] <name> ...
yame fetch [options] -u <url> -s <sha256> -o <dest>
Naming:
A name is a store directory, as the browser shows it: hg38,
hg38/KYCG, hg38/data, EPIC. It takes that directory's own files --
`hg38` is the genome annotation, not the knowledgebase and models
beneath it; name those explicitly. Narrow within one with -g.
A file resolves too, best written out: `hg38/data/test.cg`. The
bare name works when only one directory publishes it.
Several names may be given, separated by spaces or by commas --
commas so a list fits an option that takes one argument. A
directory named twice is taken once, and naming it whole absorbs a
file picked out of it.
Browsing:
With no target on a terminal, opens a tree browser: species, then
platform or genome build, then its knowledgebase and files. Arrows
move, right/left open and close a row, space or x selects (a folder
takes everything under it), f fetches what is selected, h lists every
key, q leaves. `/` filters the tree -- by name, source, collection
or title -- and enter keeps the filter so you can then select.
What is already in the store shows as present and
cannot be selected.
Piped or redirected it dumps the registry as TSV instead (same as
-l), so a script never blocks on a keystroke.
Purpose:
Download reference assets into the shared store that every tool in the
suite reads, verifying each file against a digest this build pins.
-c puts them in the current directory instead, for a one-off or a
demo: no SHA256SUMS is written beside them, but each file is checked
against the same digest.
Store:
Resolved in order: -d, $YAME_DATA_HOME, $METHSCOPE_DATA_HOME,
${XDG_DATA_HOME:-~/.local/share}/yame
YAME_DATA_HOME: /home/user/.local/share/yame
Mirror:
$YAME_ASSETS_MIRROR=<scheme://host[:port]> downloads from a site that
mirrors the public repositories, keeping each URL's path: the file at
https://raw.githubusercontent.com/zhou-lab/X/v1/f is fetched from <mirror>/zhou-lab/X/v1/f.
Every byte is still checked against the compiled-in digest.
Options:
-d <dir> Store root, overriding the environment.
-f Re-download what is present, and replace a file the store's
manifest records at a different digest than this build pins.
-R A directory name also takes every directory beneath it:
`-R hg38` is hg38, hg38/KYCG, hg38/data and hg38/models.
Without it a name is one directory's own files. Works
with -l and -n, so `-n -R hg38` says what that reaches.
-l Dump the registry as TSV and exit: one row per file, with
its size, digest, description and whether the store has it.
Takes the same <name> and -g a fetch does, so `-l -g
methscope hg38` is the dry run for fetching exactly that.
-g <a,b> Only files matching every term: name, source, collection,
title or upstream database. `-g chromatin` inside a
knowledgebase, `-g celltype` across a whole genome.
-n Say what would be fetched and stop, successfully. The same
plan a fetch prints before asking, but it exits 0, so a
script can check first. -l gives the same set as TSV.
-y Fetch a whole folder without asking. A folder is confirmed
first, since a short name can reach a lot -- `hg38` is 3.5
GB. On a terminal the browser is the confirmation: it opens
on the folder with exactly those files checked, f fetches,
q leaves. Off one, a folder needs -y. A name that picks out
ONE file needs neither: naming the file is the confirmation,
so a documented fetch line runs in a script as it stands.
-q No progress output.
-u <url> Single-file form: what to download.
-s <sha256> Single-file form: the digest it must have.
-o <dest> Single-file form: where it goes (a path, not a dir).
-c Into the current directory rather than the store.
-h This help.
Notes:
* Each file is verified against the digest this build carries for it;
the SHA256SUMS a directory keeps records what was verified there.
A file recorded at another digest is stale and needs -f to replace.
* `shasum -a 256 -c SHA256SUMS` in any store directory re-verifies it
by hand, with none of this code involved.
* Built with libcurl: fetch available.