High-performance error correction for Illumina resequencing data
Fast and precise comparison of genomes and metagenomes (in the order of terabytes) on a typical personal laptop
Implementation of Neural Distance Embeddings for Biological Sequences (NeuroSEED) in PyTorch (NeurIPS 2021)
📛 A Python package for using ontologies, terminologies, and biomedical nomenclatures
Efficient implementations of Needleman-Wunsch and other sequence alignment algorithms written in Rust with Python bindings via PyO3.
🦠🧬🧑💻📇 Microbial genomes-to-report pipeline
A WGS de novo assembler based on the FMD-index for large genomes
Analysis pipelines for genomic sequencing data
Bayesian MCMC matrix factorization algorithm
The MPI Bioinformatics Toolkit
Lightweight converter between hic and cool contact matrices.
dcHiC: Differential compartment analysis for Hi-C datasets
A Nextflow workflow to generate lift over files for any pair of genomes
Tibanna helps you run your genomic pipelines on Amazon cloud (AWS). It is used by the 4DN DCIC (4D Nucleome Data Coordination and Integration Center)...
An R package for studying mutational signatures and structural variant signatures along clonal evolution in cancer.
Ontology Tutorial
Python3 scripts to manipulate FASTA and FASTQ files
Monitor computational workflows in real time
ClassifyCNV: a tool for clinical annotation of copy-number variants
An ultra-high-performance protein-protein docking for heterogeneous supercomputers
Allelic decomposition and exact genotyping of highly polymorphic and structurally variant genes
Statistical approach for removing unwanted variation from multiple single-cell datasets
ICML 2026 autonomous AI agent for end-to-end spatial proteomics analysis, with SP-Bench for agentic multiplexed-imaging workflows.
Standalone C library for assembling Illumina short reads in small regions
A tutorial on methods of 16S analysis with QIIME 1
Incremental construction of FM-index for DNA sequences
Universal Transcript Archive: comprehensive genome-transcript alignments; multiple transcript sources, versions, and alignment methods; available as a...
:golf: GO-terms Semantic Similarity Measures
ElasticBLAST is a cloud-based tool to perform your BLAST searches faster and make you more effective
zol (& fai): large-scale targeted detection and evolutionary investigation of gene clusters (i.e. BGCs, phages, etc.)
𝐠𝐠𝐯𝐨𝐥𝐜 effortlessly translates differential expression datasets and RNAseq data into informative volcano plots. Highlight genes of interest with unpre...
Unsupervised Deep Disentangled Representation of Single-Cell Omics
EpiViz is a scientific information visualization tool for genetic and epigenetic data, used to aid in the exploration and understanding of correlation...
Thermo MSFileReader Python bindings
Gfapy: a flexible and extensible software library for handling sequence graphs in Python
Genomics data visualization in Python by using matplotlib.
Horizontal gene transfer (HGT) identification pipeline
Exon is an OLAP query engine specifically for biology and life science applications.
Nucleic Acids Research 2024:RNA-MSM model is an unsupervised RNA language model based on multiple sequences that outputs both embedding and attention...
NEAT (NExt-generation Analysis Toolkit) simulates next-gen sequencing reads and can learn simulation parameters from real data.
A modern multiple sequence alignment browser - built for the terminal.
Node.js module for working with the NCBI API (aka e-utils).
Butler is a framework for running scientific workflows on public and academic clouds.
Bonsai: Fast, flexible taxonomic analysis and classification
Performs memory-efficient reservoir sampling on very large input files delimited by newlines
G-BLASTN is a GPU-accelerated nucleotide alignment tool based on the widely used NCBI-BLAST.
A modern genomics framework for julia
Arioc: GPU-accelerated DNA short-read alignment
BISulfite-seq CUI Toolkit
Rust implementation of a fast, easy, interval tree library nim-lapper