Download sequencing data and metadata from GSA, SRA, ENA, and DDBJ databases.
🐙 KrakenUniq: Metagenomics classifier with unique k-mer counting for more specific results
Tools for manipulating sequence graphs in the GFA and rGFA formats
Pegasus Workflow Management System - Automate, recover, and debug scientific computations.
A cool place to store your Hi-C
📚 Tools and databases for analyzing HLA and VDJ genes.
A curated list of resources for machine learning for small-molecule drug discovery
Next generation sequencing reads de novo assembler.
A universal toolkit for upstream processing of long RNA reads
Simple simulation of single-cell RNA sequencing data
Metadata and website for the Open Bio Ontologies Foundry Ontology Registry
A framework for state-of-the-art pre-trained bio foundation models on genomics and transcriptomics modalities.
A collection of scripts and notes related to genomics and bioinformatics
Clone with Python! Data structures for double stranded DNA & simulation of homologous recombination, Gibson assembly, cut & paste cloning.
goleft is a collection of bioinformatics tools distributed under MIT license in a single static binary
Fast alignment and preprocessing of chromatin profiles
:rocket: A sequencing simulator
R package for the analysis of massive SNP arrays.
Ten Quick Tips for Deep Learning in Biology
CodonTransformer (2M+ Downloads); The tool for codon optimization, optimizing DNA for protein expression
:package: :whale: Dockerfiles and documentation on tools for public health bioinformatics
Earl Grey: A fully automated TE curation and annotation pipeline
Rapid phylogenetic analysis of large samples of recombinant bacterial whole genome sequences using Gubbins
INDRA (Integrated Network and Dynamical Reasoning Assembler) is an automated model assembly system interfacing with NLP systems and databases to colle...
Nature Communications | BASALT (Binning Across a Series of Assemblies Toolkit) for binning and refinement of short- and long-read sequencing data
Application and Python module for average nucleotide identity analyses of microbes.
PhysiCell: Scientist end users should use latest release! Developers please fork the development branch and submit PRs to the dev branch. Thanks!
Fast indexing and search of discontinuous motifs in protein structures
A bioinformatics workflow engine built on top of the Workflow Description Language (WDL).
A fully reproducible and state-of-the-art ancient DNA analysis pipeline
Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data
Clustering scRNAseq by genotypes
Bioinformatics Workbook repository
Constructing a pangenome gene graph
课题组每周研讨会
TOGA (Tool to infer Orthologs from Genome Alignments): implements a novel paradigm to infer orthologous genes. TOGA integrates gene annotation, inferr...
A visualization grammar and GPU-accelerated toolkit for genomic data
MetaEuk - sensitive, high-throughput gene discovery and annotation for large-scale eukaryotic metagenomics
A structural variation pipeline for short-read sequencing
Scans genome contigs against the ResFinder, PlasmidFinder, and PointFinder databases.
Aligns short reads using dynamic seed size with strobemers
C-library for calculating Solvent Accessible Surface Areas
🔬 Bioinformatics Notebook. Scripts for bioinformatics pipelines, with quick start guides for programs and video demonstrations.
Tutorials on machine learning, artificial intelligence, data science with math explanation and reusable code (in python and R)
Make Picrust2 Output Analysis and Visualization Easier
Differential expression analysis for single-cell RNA-seq data.
Population-scale genotyping using pangenome graphs
Blazing-Fast Bioinformatic Operations on Python DataFrames
MSA(Multiple Sequence Alignment) visualization python package for sequence analysis
tools for working with Bisulfite Sequencing data while preserving reads intrinsic dependencies