misha

CRAN status R-CMD-check

The misha package is a toolkit for analysis of genomic data. it implements an efficient data structure for storing genomic data, and provides a set of functions for data extraction, manipulation and analysis.

Installation

You can install the released version of misha from CRAN with:

install.packages("misha")

Or from conda:

conda install -c aviezerl r-misha

And the development version from GitHub with:

remotes::install_github("tanaylab/misha")

Quick start

The package ships a small example database, so there is nothing to download before the first query:

library(misha)
gdb.init_examples() # a tiny example genome, unpacked into tempdir()
gtrack.ls() # what is in it
#> [1] "array_track"         "dense_track"         "rects_track"        
#> [4] "sparse_track"        "subdir.dense_track2"
gextract("dense_track", gintervals(1, 0, 500), iterator = 100) # signal in 100 bp bins
#>   chrom start end dense_track intervalID
#> 1  chr1     0 100   0.1688889          1
#> 2  chr1   100 200   0.1700000          1
#> 3  chr1   200 300   0.1800000          1
#> 4  chr1   300 400   0.1600000          1
#> 5  chr1   400 500   0.1100000          1
head(gscreen("dense_track > 0.2", gintervals(1, 0, 50000), iterator = 100)) # bins above a threshold
#>   chrom start   end
#> 1  chr1 17200 17300
#> 2  chr1 20000 20100
#> 3  chr1 23300 23400
#> 4  chr1 26200 26300
#> 5  chr1 32600 32800
#> 6  chr1 32900 33000

Every misha analysis is that shape: a scope (where to look), an iterator (in what chunks), and a track expression evaluated over it.

Usage

Start with the Misha Basics short guide.

See the Genomes vignette for instructions on how to create a misha database for common genomes.

See the user manual for more usage details.

Using misha with an LLM agent

Drop-in prompt (no clone needed). Paste the block below into your agent at the start of a misha task. It points the agent at the raw files on GitHub, so it works without a local checkout:

Before writing any misha code, fetch and read:

- https://raw.githubusercontent.com/tanaylab/misha/master/agent-guides/misha-core.md  (mandatory: concepts + everyday recipes)
- https://raw.githubusercontent.com/tanaylab/misha/master/agent-guides/misha-anti-patterns.md  (silent footguns; cross-referenced from core)
- https://raw.githubusercontent.com/tanaylab/misha/master/agent-guides/misha-advanced.md  (consult on demand: 2D / Hi-C, PWM, import/export, new genomes)

Follow the conventions in those files. When you hit a recipe with an "Avoid:" block, treat it as a hard rule.

For agents (Claude Code, Copilot, Cursor, etc.) writing misha analysis code in a downstream project, point them at the maintained agent guides in this repo:

The core guide is ~4k words and targets a system-prompt-sized context. For Claude Code-style setups, dropping misha-core.md (or all three) into the project’s CLAUDE.md / AGENTS.md is the intended use.

Running scripts from old versions of misha (< 4.2.0)

Starting in misha 4.2.0, the package no longer stores global variables such as ALLGENOME or GROOT. Instead, these variables are stored in a special environment called .misha. This means that scripts written for older versions of misha will no longer work. To run such scripts, either add a prefix of .misha$ to all those variables (.misha$ALLGENOME instead of ALLGENOME), or run the following command before running the script:

ALLGENOME <<- .misha$ALLGENOME
GROOT <<- .misha$GROOT
ALLGENOME <<- .misha$ALLGENOME
GINTERVID <<- .misha$GINTERVID
GITERATOR.INTERVALS <<- .misha$GITERATOR.INTERVALS
GROOT <<- .misha$GROOT
GWD <<- .misha$GWD
GTRACKS <<- .misha$GTRACKS