S4Vectors - Foundation of vector-like and list-like containers in Bioconductor
The S4Vectors package defines the Vector and List virtual classes and a set of generic functions that extend the semantic of ordinary vectors and lists in R. Package developers can easily implement vector-like or list-like objects as concrete subclasses of Vector or List. In addition, a few low-level concrete subclasses of general interest (e.g. DataFrame, Rle, Factor, and Hits) are implemented in the S4Vectors package itself (many more are implemented in the IRanges package and in other Bioconductor infrastructure packages).
Last updated
infrastructuredatarepresentationbioconductor-packagecore-package
16.70 score 18 stars 2.1k dependents 2.1k scripts 150k downloadsDelayedArray - A unified framework for working transparently with on-disk and in-memory array-like datasets
Wrapping an array-like object (typically an on-disk object) in a DelayedArray object allows one to perform common array operations on it without loading the object in memory. In order to reduce memory usage and optimize performance, operations on the object are either delayed or executed using a block processing mechanism. Note that this also works on in-memory array-like objects like DataFrame objects (typically with Rle columns), Matrix objects, ordinary arrays and, data frames.
Last updated
infrastructuredatarepresentationannotationgenomeannotationbioconductor-packagecore-packageu24ca289073
15.71 score 29 stars 1.4k dependents 932 scripts 100k downloadsKEGGREST - Client-side REST access to the Kyoto Encyclopedia of Genes and Genomes (KEGG)
A package that provides a client interface to the Kyoto Encyclopedia of Genes and Genomes (KEGG) REST API. Only for academic use by academic users belonging to academic institutions (see <https://www.kegg.jp/kegg/rest/>). Note that KEGGREST is based on KEGGSOAP by J. Zhang, R. Gentleman, and Marc Carlson, and KEGG (python package) by Aurelien Mazurie.
Last updated
annotationpathwaysthirdpartyclientkeggbioconductor-packagecore-package
15.24 score 15 stars 798 dependents 1.7k scripts 77k downloadsGOSemSim - GO-terms Semantic Similarity Measures
The semantic comparisons of Gene Ontology (GO) annotations provide quantitative ways to compute similarities between genes and gene groups, and have became important basis for many bioinformatics analysis approaches. GOSemSim is an R package for semantic similarity computation among GO terms, sets of GO terms, gene products and gene clusters. GOSemSim implemented five methods proposed by Resnik, Schlicker, Jiang, Lin and Wang respectively.
Last updated
annotationgoclusteringpathwaysnetworksoftwarebioinformaticsgene-ontologysemantic-similaritycpp
14.72 score 72 stars 77 dependents 1.0k scripts 38k downloads
Spectra - Spectra Infrastructure for Mass Spectrometry Data
The Spectra package defines an efficient infrastructure for storing and handling mass spectrometry spectra and functionality to subset, process, visualize and compare spectra data. It provides different implementations (backends) to store mass spectrometry data. These comprise backends tuned for fast data access and processing and backends for very large data sets ensuring a small memory footprint.
Last updated
infrastructureproteomicsmassspectrometrymetabolomicsbioconductorhacktoberfestmass-spectrometry
14.11 score 46 stars 69 dependents 880 scripts 5.5k downloadsAnnotationHub - Client to access AnnotationHub resources
This package provides a client for the Bioconductor AnnotationHub web resource. The AnnotationHub web resource provides a central location where genomic files (e.g., VCF, bed, wig) and other resources from standard locations (e.g., UCSC, Ensembl) can be discovered. The resource includes metadata about each resource, e.g., a textual description, tags, and date of modification. The client creates and manages a local cache of files retrieved by the user, helping with quick and reproducible access.
Last updated
infrastructuredataimportguithirdpartyclientcore-packageu24ca289073
14.10 score 19 stars 122 dependents 4.6k scripts 25k downloadsscDblFinder - scDblFinder
The scDblFinder package gathers various methods for the detection and handling of doublets/multiplets in single-cell sequencing data (i.e. multiple cells captured within the same droplet or reaction volume). It includes methods formerly found in the scran package, the new fast and comprehensive scDblFinder method, and a reimplementation of the Amulet detection method for single-cell ATAC-seq.
Last updated
preprocessingsinglecellrnaseqatacseqdoubletssingle-cell
12.73 score 258 stars 2 dependents 2.4k scriptsSeqArray - Data management of large-scale whole-genome sequence variant calls using GDS files
Data management of large-scale whole-genome sequencing variant calls with thousands of individuals: genotypic data (e.g., SNVs, indels and structural variation calls) and annotations in SeqArray GDS files are stored in an array-oriented and compressed manner, with efficient data access using the R programming language.
Last updated
infrastructuredatarepresentationsequencinggeneticsbioinformaticsgds-formatsnpsnvweswgscpp
12.23 score 48 stars 10 dependents 1.1k scripts 1.8k downloadssparseMatrixStats - Summary Statistics for Rows and Columns of Sparse Matrices
High performance functions for row and column operations on sparse matrices. For example: col / rowMeans2, col / rowMedians, col / rowVars etc. Currently, the optimizations are limited to data in the column sparse format. This package is inspired by the matrixStats package by Henrik Bengtsson.
Last updated
infrastructuresoftwaredatarepresentationcpp
12.16 score 55 stars 161 dependents 373 scripts 29k downloadsMOFA2 - Multi-Omics Factor Analysis v2
The MOFA2 package contains a collection of tools for training and analysing multi-omic factor analysis (MOFA). MOFA is a probabilistic factor model that aims to identify principal axes of variation from data sets that can comprise multiple omic layers and/or groups of samples. Additional time or space information on the samples can be incorporated using the MEFISTO framework, which is part of MOFA2. Downstream analysis functions to inspect molecular features underlying each factor, visualisation, imputation etc are available.
Last updated
dimensionreductionbayesianvisualizationfactor-analysismofamulti-omics
12.04 score 414 stars 2 dependents 920 scripts 1.9k downloadsdestiny - Creates diffusion maps
Create and plot diffusion maps.
Last updated
cellbiologycellbasedassaysclusteringsoftwarevisualizationdiffusion-mapsdimensionality-reductioncpp
11.87 score 108 stars 1 dependents 1.0k scripts 2.3k downloadsUCell - Rank-based signature enrichment analysis for single-cell data
UCell is a package for evaluating gene signatures in single-cell datasets. UCell signature scores, based on the Mann-Whitney U statistic, are robust to dataset size and heterogeneity, and their calculation demands less computing time and memory than other available methods, enabling the processing of large datasets in a few minutes even on machines with limited computing power. UCell can be applied to any single-cell data matrix, and includes functions to directly interact with SingleCellExperiment and Seurat objects.
Last updated
singlecellgenesetenrichmenttranscriptomicsgeneexpressioncellbasedassays
11.49 score 203 stars 3 dependents 1.2k scripts 3.0k downloadsscRepertoire - A toolkit for single-cell immune receptor profiling
scRepertoire is a toolkit for processing and analyzing single-cell T-cell receptor (TCR) and immunoglobulin (Ig). The scRepertoire framework supports use of 10x, AIRR, BD, MiXCR, TRUST4, and WAT3R single-cell formats. The functionality includes basic clonal analyses, repertoire summaries, distance-based clustering and interaction with the popular Seurat and SingleCellExperiment/Bioconductor R single-cell workflows.
Last updated
softwareimmunooncologysinglecellclassificationannotationsequencingcpp
11.46 score 373 stars 1 dependents 640 scripts 1.6k downloadsscater - Single-Cell Analysis Toolkit for Gene Expression Data in R
A collection of tools for doing various analyses of single-cell RNA-seq gene expression data, with a focus on quality control and visualization.
Last updated
immunooncologysinglecellrnaseqqualitycontrolpreprocessingnormalizationvisualizationdimensionreductiontranscriptomicsgeneexpressionsequencingsoftwaredataimportdatarepresentationinfrastructurecoverage
11.40 score 56 dependents 14k scripts 17k downloadsgdsfmt - R Interface to CoreArray Genomic Data Structure (GDS) Files
Provides a high-level R interface to CoreArray Genomic Data Structure (GDS) data files. GDS is portable across platforms with hierarchical structure to store multiple scalable array-oriented data sets with metadata information. It is suited for large-scale datasets, especially for data which are much larger than the available random-access memory. The gdsfmt package offers the efficient operations specifically designed for integers of less than 8 bits, since a diploid genotype, like single-nucleotide polymorphism (SNP), usually occupies fewer bits than a byte. Data compression and decompression are available with relatively efficient random access. It is also allowed to read a GDS file in parallel with multiple R processes supported by the package parallel.
Last updated
infrastructuredataimportbioinformaticsgds-formatgenomicscpp
11.34 score 20 stars 32 dependents 1.1k scripts 4.3k downloadsS4Arrays - Foundation of array-like containers in Bioconductor
The S4Arrays package defines the Array virtual class to be extended by other S4 classes that wish to implement a container with an array-like semantic. It also provides: (1) low-level functionality meant to help the developer of such container to implement basic operations like display, subsetting, or coercion of their array-like objects to an ordinary matrix or array, and (2) a framework that facilitates block processing of array-like objects (typically on-disk objects).
Last updated
infrastructuredatarepresentationbioconductor-packagecore-packageu24ca289073
11.01 score 7 stars 1.4k dependents 16 scripts 84k downloadsGlimma - Interactive visualizations for gene expression analysis
This package produces interactive visualizations for RNA-seq data analysis, utilizing output from limma, edgeR, or DESeq2. It produces interactive htmlwidgets versions of popular RNA-seq analysis plots to enhance the exploration of analysis results by overlaying interactive features. The plots can be viewed in a web browser or embedded in notebook documents.
Last updated
differentialexpressiongeneexpressionmicroarrayreportwritingrnaseqsequencingvisualizationdifferential-expressioninteractive-visualizations
10.84 score 35 stars 2 dependents 808 scripts 1.8k downloadsggtreeExtra - An R Package To Add Geometric Layers On Circular Or Other Layout Tree Of "ggtree"
'ggtreeExtra' extends the method for mapping and visualizing associated data on phylogenetic tree using 'ggtree'. These associated data can be presented on the external panels to circular layout, fan layout, or other rectangular layout tree built by 'ggtree' with the grammar of 'ggplot2'.
Last updated
softwarevisualizationphylogeneticsannotation
10.79 score 98 stars 2 dependents 838 scripts 3.1k downloadssesame - SEnsible Step-wise Analysis of DNA MEthylation BeadChips
Tools For analyzing Illumina Infinium DNA methylation arrays. SeSAMe provides utilities to support analyses of multiple generations of Infinium DNA methylation BeadChips, including preprocessing, quality control, visualization and inference. SeSAMe features accurate detection calling, intelligent inference of ethnicity, sex and advanced quality control routines.
Last updated
dnamethylationmethylationarraypreprocessingqualitycontrolbioinformaticsdna-methylationmicroarray
10.58 score 84 stars 1 dependents 389 scriptstxdbmaker - Tools for making TxDb objects from genomic annotations
A set of tools for making TxDb objects from genomic annotations from various sources (e.g. UCSC, Ensembl, and GFF files). These tools allow the user to download the genomic locations of transcripts, exons, and CDS, for a given assembly, and to import them in a TxDb object. TxDb objects are implemented in the GenomicFeatures package, together with flexible methods for extracting the desired features in convenient formats.
Last updated
infrastructuredataimportannotationgenomeannotationgenomeassemblygeneticssequencingbioconductor-packagecore-package
10.44 score 5 stars 63 dependents 378 scripts 7.7k downloadsBiocCheck - Bioconductor-specific package checks
BiocCheck guides maintainers through Bioconductor best practicies. It runs Bioconductor-specific package checks by searching through package code, examples, and vignettes. Maintainers are required to address all errors, warnings, and most notes produced.
Last updated
infrastructurebioconductor-packagecore-services
10.22 score 10 stars 6 dependents 125 scripts 4.3k downloadsmaaslin3 - "Refining and extending generalized multivariate linear models for meta-omic association discovery"
MaAsLin 3 refines and extends generalized multivariate linear models for meta-omicron association discovery. It finds abundance and prevalence associations between microbiome meta-omics features and complex metadata in population-scale epidemiological studies. The software includes multiple analysis methods (including support for multiple covariates, repeated measures, and ordered predictors), filtering, normalization, and transform options to customize analysis for your specific study.
Last updated
metagenomicssoftwaremicrobiomenormalizationmultiplecomparisonr-tools
10.14 score 84 stars 4 dependents 189 scripts 643 downloadsSpatialFeatureExperiment - Integrating SpatialExperiment with Simple Features in sf
A new S4 class integrating Simple Features with the R package sf to bring geospatial data analysis methods based on vector data to spatial transcriptomics. Also implements management of spatial neighborhood graphs and geometric operations. This pakage builds upon SpatialExperiment and SingleCellExperiment, hence methods for these parent classes can still be used.
Last updated
datarepresentationtranscriptomicsspatial
10.02 score 57 stars 2 dependents 688 scripts 786 downloadssingleCellTK - Comprehensive and Interactive Analysis of Single Cell RNA-Seq Data
The Single Cell Toolkit (SCTK) in the singleCellTK package provides an interface to popular tools for importing, quality control, analysis, and visualization of single cell RNA-seq data. SCTK allows users to seamlessly integrate tools from various packages at different stages of the analysis workflow. A general "a la carte" workflow gives users the ability access to multiple methods for data importing, calculation of general QC metrics, doublet detection, ambient RNA estimation and removal, filtering, normalization, batch correction or integration, dimensionality reduction, 2-D embedding, clustering, marker detection, differential expression, cell type labeling, pathway analysis, and data exporting. Curated workflows can be used to run Seurat and Celda. Streamlined quality control can be performed on the command line using the SCTK-QC pipeline. Users can analyze their data using commands in the R console or by using an interactive Shiny Graphical User Interface (GUI). Specific analyses or entire workflows can be summarized and shared with comprehensive HTML reports generated by Rmarkdown. Additional documentation and vignettes can be found at camplab.net/sctk.
Last updated
singlecellgeneexpressiondifferentialexpressionalignmentclusteringimmunooncologybatcheffectnormalizationqualitycontroldataimportgui
9.78 score 189 stars 335 scripts 763 downloadsPSMatch - Handling and Managing Peptide Spectrum Matches
The PSMatch package helps proteomics practitioners to load, handle and manage Peptide Spectrum Matches. It provides functions to model peptide-protein relations as adjacency matrices and connected components, visualise these as graphs and make informed decision about shared peptide filtering. The package also provides functions to calculate and visualise MS2 fragment ions.
Last updated
infrastructureproteomicsmassspectrometrymass-spectrometrypeptide-spectrum-matches
9.68 score 6 stars 40 dependents 58 scripts 3.8k downloadsBayesSpace - Clustering and Resolution Enhancement of Spatial Transcriptomes
Tools for clustering and enhancing the resolution of spatial gene expression experiments. BayesSpace clusters a low-dimensional representation of the gene expression matrix, incorporating a spatial prior to encourage neighboring spots to cluster together. The method can enhance the resolution of the low-dimensional representation into "sub-spots", for which features such as gene expression or cell type composition can be imputed.
Last updated
softwareclusteringtranscriptomicsgeneexpressionsinglecellimmunooncologydataimportopenblascppopenmp
9.65 score 180 stars 1 dependents 684 scripts 956 downloadsBiocBaseUtils - Utility and internal functions for Bioconductor packages
The package coalesces typical helper functions that are scattered throughout the Bioconductor ecosystem. It aims to reduce code redundancy by formalizing functions often used by Bioconductor developers. These functions include operations such as replacing slots in an object, selecting observations for show methods, labeling function life cycles, and more.
Last updated
softwareinfrastructurebioconductor-packagecore-package
9.49 score 4 stars 923 dependents 11 scripts 17k downloadsUCSC.utils - Low-level utilities to retrieve data from the UCSC Genome Browser
A set of low-level utilities to retrieve data from the UCSC Genome Browser. Most functions in the package access the data via the UCSC REST API but some of them query the UCSC MySQL server directly. Note that the primary purpose of the package is to support higher-level functionalities implemented in downstream packages like GenomeInfoDb or txdbmaker.
Last updated
infrastructuregenomeassemblyannotationgenomeannotationdataimportbioconductor-packagecore-package
9.43 score 1 stars 341 dependents 14 scripts 47k downloadsbambu - Context-Aware Transcript Quantification from Long Read RNA-Seq data
bambu is a R package for multi-sample transcript discovery and quantification using long read RNA-Seq data. You can use bambu after read alignment to obtain expression estimates for known and novel transcripts and genes. The output from bambu can directly be used for visualisation and downstream analysis such as differential gene expression or transcript usage.
Last updated
alignmentcoveragedifferentialexpressionfeatureextractiongeneexpressiongenomeannotationgenomeassemblyimmunooncologylongreadmultiplecomparisonnormalizationrnaseqregressionsequencingsoftwaretranscriptiontranscriptomicsbambubioconductorlong-readsnanoporenanopore-sequencingrna-seqrna-seq-analysistranscript-quantificationtranscript-reconstructioncpp
9.40 score 251 stars 1 dependents 246 scripts 884 downloadsscp - Mass Spectrometry-Based Single-Cell Proteomics Data Analysis
Utility functions for manipulating, processing, and analyzing mass spectrometry-based single-cell proteomics data. The package is an extension to the 'QFeatures' package and relies on 'SingleCellExpirement' to enable single-cell proteomics analyses. The package offers the user the functionality to process quantitative table (as generated by MaxQuant, Proteome Discoverer, and more) into data tables ready for downstream analysis and data visualization.
Last updated
geneexpressionproteomicssinglecellmassspectrometrypreprocessingcellbasedassaysbioconductormass-spectrometrysingle-cellsoftware
9.31 score 33 stars 340 scripts 424 downloadsInteractiveComplexHeatmap - Make Interactive Complex Heatmaps
This package can easily make heatmaps which are produced by the ComplexHeatmap package into interactive applications. It provides two types of interactivities: 1. on the interactive graphics device, and 2. on a Shiny app. It also provides functions for integrating the interactive heatmap widgets for more complex Shiny app development.
Last updated
softwarevisualizationsequencinginteractive-heatmaps
8.94 score 142 stars 4 dependents 253 scripts 986 downloadstidySummarizedExperiment - Brings SummarizedExperiment to the Tidyverse
The tidySummarizedExperiment package provides a set of tools for creating and manipulating tidy data representations of SummarizedExperiment objects. SummarizedExperiment is a widely used data structure in bioinformatics for storing high-throughput genomic data, such as gene expression or DNA sequencing data. The tidySummarizedExperiment package introduces a tidy framework for working with SummarizedExperiment objects. It allows users to convert their data into a tidy format, where each observation is a row and each variable is a column. This tidy representation simplifies data manipulation, integration with other tidyverse packages, and enables seamless integration with the broader ecosystem of tidy tools for data analysis.
Last updated
assaydomaininfrastructurernaseqdifferentialexpressiongeneexpressionnormalizationclusteringqualitycontrolsequencingtranscriptiontranscriptomicsbioconductorgenomicssummarizedexperimenttidyverse
8.92 score 30 stars 1 dependents 368 scripts 601 downloadstidySingleCellExperiment - Brings SingleCellExperiment to the Tidyverse
'tidySingleCellExperiment' is an adapter that abstracts the 'SingleCellExperiment' container in the form of a 'tibble'. This allows *tidy* data manipulation, nesting, and plotting. For example, a 'tidySingleCellExperiment' is directly compatible with functions from 'tidyverse' packages `dplyr` and `tidyr`, as well as plotting with `ggplot2` and `plotly`. In addition, the package provides various utility functions specific to single-cell omics data analysis (e.g., aggregation of cell-level data to pseudobulks).
Last updated
assaydomaininfrastructurernaseqdifferentialexpressionsinglecellgeneexpressionnormalizationclusteringqualitycontrolsequencingbioconductordplyrggplot2plotlysingle-cell-rna-seqsingle-cell-sequencingsinglecellexperimenttibbletidyrtidyverse
8.89 score 37 stars 2 dependents 212 scripts 660 downloadscrisprDesign - Comprehensive design of CRISPR gRNAs for nucleases and base editors
Provides a comprehensive suite of functions to design and annotate CRISPR guide RNA (gRNAs) sequences. This includes on- and off-target search, on-target efficiency scoring, off-target scoring, full gene and TSS contextual annotations, and SNP annotation (human only). It currently support five types of CRISPR modalities (modes of perturbations): CRISPR knockout, CRISPR activation, CRISPR inhibition, CRISPR base editing, and CRISPR knockdown. All types of CRISPR nucleases are supported, including DNA- and RNA-target nucleases such as Cas9, Cas12a, and Cas13d. All types of base editors are also supported. gRNA design can be performed on reference genomes, transcriptomes, and custom DNA and RNA sequences. Both unpaired and paired gRNA designs are enabled.
Last updated
crisprfunctionalgenomicsgenetargetbioconductorbioconductor-packagecrispr-cas9crispr-designcrispr-targetgenomics-analysisgrnagrna-sequencegrna-sequencessgrnasgrna-design
8.87 score 32 stars 3 dependents 123 scripts 488 downloadsVoyager - From geospatial to spatial omics
SpatialFeatureExperiment (SFE) is a new S4 class for working with spatial single-cell genomics data. The voyager package implements basic exploratory spatial data analysis (ESDA) methods for SFE. Univariate methods include univariate global spatial ESDA methods such as Moran's I, permutation testing for Moran's I, and correlograms. Bivariate methods include Lee's L and cross variogram. Multivariate methods include MULTISPATI PCA and multivariate local Geary's C recently developed by Anselin. The Voyager package also implements plotting functions to plot SFE data and ESDA results.
Last updated
geneexpressionspatialtranscriptomicsvisualizationbioconductoredaesdaexploratory-data-analysisomicsspatial-statisticsspatial-transcriptomics
8.84 score 103 stars 415 scripts 570 downloads
CompoundDb - Creating and Using (Chemical) Compound Annotation Databases
CompoundDb provides functionality to create and use (chemical) compound annotation databases from a variety of different sources such as LipidMaps, HMDB, ChEBI or MassBank. The database format allows to store in addition MS/MS spectra along with compound information. The package provides also a backend for Bioconductor's Spectra package and allows thus to match experimetal MS/MS spectra against MS/MS spectra in the database. Databases can be stored in SQLite format and are thus portable.
Last updated
massspectrometrymetabolomicsannotationdatabasesmass-spectrometry
8.82 score 19 stars 3 dependents 97 scripts 870 downloadspwalign - Perform pairwise sequence alignments
The two main functions in the package are pairwiseAlignment() and stringDist(). The former solves (Needleman-Wunsch) global alignment, (Smith-Waterman) local alignment, and (ends-free) overlap alignment problems. The latter computes the Levenshtein edit distance or pairwise alignment score matrix for a set of strings.
Last updated
alignmentsequencematchingsequencinggeneticsbioconductor-package
8.80 score 1 stars 112 dependents 146 scripts 13k downloadsscDesign3 - A unified framework of realistic in silico data generation and statistical model inference for single-cell and spatial omics
We present a statistical simulator, scDesign3, to generate realistic single-cell and spatial omics data, including various cell states, experimental designs, and feature modalities, by learning interpretable parameters from real data. Using a unified probabilistic model for single-cell and spatial omics data, scDesign3 infers biologically meaningful parameters; assesses the goodness-of-fit of inferred cell clusters, trajectories, and spatial locations; and generates in silico negative and positive controls for benchmarking computational tools.
Last updated
softwaresinglecellsequencinggeneexpressionspatial
8.71 score 121 stars 1 dependents 94 scripts 314 downloads
MsExperiment - Infrastructure for Mass Spectrometry Experiments
Infrastructure to store and manage all aspects related to a complete proteomics or metabolomics mass spectrometry (MS) experiment. The MsExperiment package provides light-weight and flexible containers for MS experiments building on the new MS infrastructure provided by the Spectra, QFeatures and related packages. Along with raw data representations, links to original data files and sample annotations, additional metadata or annotations can also be stored within the MsExperiment container. To guarantee maximum flexibility only minimal constraints are put on the type and content of the data within the containers.
Last updated
infrastructureproteomicsmassspectrometrymetabolomicsexperimentaldesigndataimport
8.68 score 5 stars 18 dependents 243 scripts 1.8k downloadsMSstatsPTM - Statistical Characterization of Post-translational Modifications
MSstatsPTM provides general statistical methods for quantitative characterization of post-translational modifications (PTMs). Supports DDA, DIA, SRM, and tandem mass tag (TMT) labeling. Typically, the analysis involves the quantification of PTM sites (i.e., modified residues) and their corresponding proteins, as well as the integration of the quantification results. MSstatsPTM provides functions for summarization, estimation of PTM site abundance, and detection of changes in PTMs across experimental conditions.
Last updated
immunooncologymassspectrometryproteomicssoftwaredifferentialexpressiononechanneltwochannelnormalizationqualitycontrolpost-translational-modificationcpp
8.62 score 15 stars 2 dependents 92 scripts 577 downloads
dreamlet - Scalable differential expression analysis of single cell transcriptomics datasets with complex study designs
Recent advances in single cell/nucleus transcriptomic technology has enabled collection of cohort-scale datasets to study cell type specific gene expression differences associated disease state, stimulus, and genetic regulation. The scale of these data, complex study designs, and low read count per cell mean that characterizing cell type specific molecular mechanisms requires a user-frieldly, purpose-build analytical framework. We have developed the dreamlet package that applies a pseudobulk approach and fits a regression model for each gene and cell cluster to test differential expression across individuals associated with a trait of interest. Use of precision-weighted linear mixed models enables accounting for repeated measures study designs, high dimensional batch effects, and varying sequencing depth or observed cells per biosample.
Last updated
rnaseqgeneexpressiondifferentialexpressionbatcheffectqualitycontrolregressiongenesetenrichmentgeneregulationepigeneticsfunctionalgenomicstranscriptomicsnormalizationsinglecellpreprocessingsequencingimmunooncologysoftwarecpp
8.61 score 23 stars 327 scripts 512 downloadsalabaster.base - Save Bioconductor Objects to File
Save Bioconductor data structures into file artifacts, and load them back into memory. This is a more robust and portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
datarepresentationdataimportcurlopensslzlibcpp
8.56 score 4 stars 20 dependents 67 scripts 6.5k downloadslefser - R implementation of the LEfSE method for microbiome biomarker discovery
lefser is the R implementation of the popular microbiome biomarker discovery too, LEfSe. It uses the Kruskal-Wallis test, Wilcoxon-Rank Sum test, and Linear Discriminant Analysis to find biomarkers from two-level classes (and optional sub-classes).
Last updated
softwaresequencingdifferentialexpressionmicrobiomestatisticalmethodclassificationbioconductor-packager01ca230551
8.56 score 67 stars 114 scripts 1.3k downloadssccomp - Differential Composition and Variability Analysis for Single-Cell Data
Comprehensive R package for differential composition and variability analysis in single-cell RNA sequencing, CyTOF, and microbiome data. Provides robust Bayesian modeling with outlier detection, random effects, and advanced statistical methods for cell type proportion analysis. Features include probabilistic outlier identification, mixed-effect modeling, differential variability testing, and comprehensive visualization tools. Perfect for cancer research, immunology, developmental biology, and single-cell genomics applications.
Last updated
bayesianregressiondifferentialexpressionsinglecellmetagenomicsflowcytometryspatialbatch-correctioncompositioncytofdifferential-proportionmicrobiomemultilevelproportionsrandom-effectssingle-cellunwanted-variation
8.55 score 132 stars 194 scripts 474 downloadsScaledMatrix - Creating a DelayedMatrix of Scaled and Centered Values
Provides delayed computation of a matrix of scaled and centered values. The result is equivalent to using the scale() function but avoids explicit realization of a dense matrix during block processing. This permits greater efficiency in common operations, most notably matrix multiplication.
Last updated
softwaredatarepresentation
8.51 score 126 dependents 12 scripts 24k downloadsSPIAT - Spatial Image Analysis of Tissues
SPIAT (**Sp**atial **I**mage **A**nalysis of **T**issues) is an R package with a suite of data processing, quality control, visualization and data analysis tools. SPIAT is compatible with data generated from single-cell spatial proteomics platforms (e.g. OPAL, CODEX, MIBI, cellprofiler). SPIAT reads spatial data in the form of X and Y coordinates of cells, marker intensities and cell phenotypes. SPIAT includes six analysis modules that allow visualization, calculation of cell colocalization, categorization of the immune microenvironment relative to tumor areas, analysis of cellular neighborhoods, and the quantification of spatial heterogeneity, providing a comprehensive toolkit for spatial data analysis.
Last updated
biomedicalinformaticscellbiologyspatialclusteringdataimportimmunooncologyqualitycontrolsinglecellsoftwarevisualization
8.50 score 29 stars 76 scripts 452 downloads
recount3 - Explore and download data from the recount3 project
The recount3 package enables access to a large amount of uniformly processed RNA-seq data from human and mouse. You can download RangedSummarizedExperiment objects at the gene, exon or exon-exon junctions level with sample metadata and QC statistics. In addition we provide access to sample coverage BigWig files.
Last updated
coveragedifferentialexpressiongeneexpressionrnaseqsequencingsoftwaredataimportannotation-agnosticbioconductorcountderfinderexongenehumanilluminajunctionmouserecountrecount3
8.47 score 39 stars 317 scripts 983 downloadsmetapod - Meta-Analyses on P-Values of Differential Analyses
Implements a variety of methods for combining p-values in differential analyses of genome-scale datasets. Functions can combine p-values across different tests in the same analysis (e.g., genomic windows in ChIP-seq, exons in RNA-seq) or for corresponding tests across separate analyses (e.g., replicated comparisons, effect of different treatment conditions). Support is provided for handling log-transformed input p-values, missing values and weighting where appropriate.
Last updated
multiplecomparisondifferentialpeakcallingcpp
8.34 score 2 stars 55 dependents 15 scripts 11k downloadsggkegg - Analyzing and visualizing KEGG information using the grammar of graphics
This package aims to import, parse, and analyze KEGG data such as KEGG PATHWAY and KEGG MODULE. The package supports visualizing KEGG information using ggplot2 and ggraph through using the grammar of graphics. The package enables the direct visualization of the results from various omics analysis packages.
Last updated
pathwaysdataimportkeggggplot2ggraphpathwaytidygraphvisualization
8.24 score 247 stars 2 dependents 58 scripts 788 downloadsrBLAST - R Interface for the Basic Local Alignment Search Tool
Seamlessly interfaces the Basic Local Alignment Search Tool (BLAST) running locally to search genetic sequence data bases. This work was partially supported by grant no. R21HG005912 from the National Human Genome Research Institute.
Last updated
geneticssequencingsequencematchingalignmentdataimportbioconductorbioinformaticsblast-search
8.21 score 114 stars 1 dependents 157 scriptsRBioFormats - R interface to Bio-Formats
An R package which interfaces the OME Bio-Formats Java library to allow reading of proprietary microscopy image data and metadata.
Last updated
dataimportbio-formatsbioconductorimage-processingopenjdk
8.06 score 27 stars 4 dependents 119 scripts 714 downloadsEBSeq - An R package for gene and isoform differential expression analysis of RNA-seq data
Differential Expression analysis at both gene and isoform level using RNA-seq data
Last updated
immunooncologystatisticalmethoddifferentialexpressionmultiplecomparisonrnaseqsequencingcpp
7.87 score 6 dependents 205 scripts 990 downloads
hermes - Preprocessing, analyzing, and reporting of RNA-seq data
Provides classes and functions for quality control, filtering, normalization and differential expression analysis of pre-processed `RNA-seq` data. Data can be imported from `SummarizedExperiment` as well as `matrix` objects and can be annotated from `BioMart`. Filtering for genes without too low expression or containing required annotations, as well as filtering for samples with sufficient correlation to other samples or total number of reads is supported. The standard normalization methods including cpm, rpkm and tpm can be used, and 'DESeq2` as well as voom differential expression analyses are available.
Last updated
rnaseqdifferentialexpressionnormalizationpreprocessingqualitycontrolrna-seqstatistical-engineering
7.81 score 12 stars 1 dependents 46 scripts 489 downloadscrisprScore - On-Target and Off-Target Scoring Algorithms for CRISPR gRNAs
Provides R wrappers of several on-target and off-target scoring methods for CRISPR guide RNAs (gRNAs). The following nucleases are supported: SpCas9, AsCas12a, enAsCas12a, and RfxCas13d (CasRx). The available on-target cutting efficiency scoring methods are RuleSet1, RuleSet3, DeepHF, enPAM+GB, and CRISPRscan. Both the CFD and MIT scoring methods are available for off-target specificity prediction. The package also provides a Lindel-derived score to predict the probability of a gRNA to produce indels inducing a frameshift for the Cas9 nuclease. Note that DeepHF and enPAM+GB are not available on Windows machines.
Last updated
crisprfunctionalgenomicsfunctionalpredictionbioconductorbioconductor-packagecrispr-cas9crispr-designcrispr-targetgenomicsgrnagrna-sequencegrna-sequencesscoring-algorithmsgrnasgrna-design
7.81 score 29 stars 4 dependents 23 scripts 490 downloadsspacexr - SpatialeXpressionR: Cell Type Identification in Spatial Transcriptomics
Spatial-eXpression-R (spacexr) is a package for analyzing cell types in spatial transcriptomics data. This implementation is a fork of the spacexr GitHub repo (https://github.com/dmcable/spacexr), adapted to work with Bioconductor objects. The original package implements two statistical methods: RCTD for learning cell types and CSIDE for inferring cell type-specific differential expression. Currently, this fork only implements RCTD, which learns cell type profiles from annotated RNA sequencing (RNA-seq) reference data and uses these profiles to identify cell types in spatial transcriptomic pixels while accounting for platform-specific effects. Future releases will include an implementation of CSIDE.
Last updated
geneexpressiondifferentialexpressionsinglecellrnaseqsoftwarespatialtranscriptomics
7.81 score 6 stars 888 scripts 676 downloads
MetaboAnnotation - Utilities for Annotation of Metabolomics Data
High level functions to assist in annotation of (metabolomics) data sets. These include functions to perform simple tentative annotations based on mass matching but also functions to consider m/z and retention times for annotation of LC-MS features given that respective reference values are available. In addition, the function provides high-level functions to simplify matching of LC-MS/MS spectra against spectral libraries and objects and functionality to represent and manage such matched data.
Last updated
infrastructuremetabolomicsmassspectrometryannotationmass-spectromtry
7.77 score 20 stars 1 dependents 65 scripts 506 downloadsSpatialDecon - Deconvolution of mixed cells from spatial and/or bulk gene expression data
Using spatial or bulk gene expression data, estimates abundance of mixed cell types within each observation. Based on "Advances in mixed cell deconvolution enable quantification of cell types in spatial transcriptomic data", Danaher (2022). Designed for use with the NanoString GeoMx platform, but applicable to any gene expression data.
Last updated
immunooncologyfeatureextractiongeneexpressiontranscriptomicsspatial
7.76 score 43 stars 112 scripts 636 downloadsSimBu - Simulate Bulk RNA-seq Datasets from Single-Cell Datasets
SimBu can be used to simulate bulk RNA-seq datasets with known cell type fractions. You can either use your own single-cell study for the simulation or the sfaira database. Different pre-defined simulation scenarios exist, as are options to run custom simulations. Additionally, expression values can be adapted by adding an mRNA bias, which produces more biologically relevant simulations.
Last updated
softwarernaseqsinglecell
7.76 score 19 stars 1 dependents 63 scripts 316 downloadsAlpsNMR - Automated spectraL Processing System for NMR
Reads Bruker NMR data directories both zipped and unzipped. It provides automated and efficient signal processing for untargeted NMR metabolomics. It is able to interpolate the samples, detect outliers, exclude regions, normalize, detect peaks, align the spectra, integrate peaks, manage metadata and visualize the spectra. After spectra proccessing, it can apply multivariate analysis on extracted data. Efficient plotting with 1-D data is also available. Basic reading of 1D ACD/Labs exported JDX samples is also available.
Last updated
softwarepreprocessingvisualizationclassificationcheminformaticsmetabolomicsdataimport
7.74 score 17 stars 1 dependents 20 scripts 546 downloads
velociraptor - Toolkit for Single-Cell Velocity
This package provides Bioconductor-friendly wrappers for RNA velocity calculations in single-cell RNA-seq data. We use the basilisk package to manage Conda environments, and the zellkonverter package to convert data structures between SingleCellExperiment (R) and AnnData (Python). The information produced by the velocity methods is stored in the various components of the SingleCellExperiment class.
Last updated
singlecellgeneexpressionsequencingcoveragerna-velocity
7.70 score 61 stars 61 scripts 454 downloadsgDRimport - Package for handling the import of dose-response data
The package is a part of the gDR suite. It helps to prepare raw drug response data for downstream processing. It mainly contains helper functions for importing/loading/validating dose-response data provided in different file formats.
Last updated
softwareinfrastructuredataimport
7.67 score 3 stars 1 dependents 26 scripts 345 downloadsEpiCompare - Comparison, Benchmarking & QC of Epigenomic Datasets
EpiCompare is used to compare and analyse epigenetic datasets for quality control and benchmarking purposes. The package outputs an HTML report consisting of three sections: (1. General metrics) Metrics on peaks (percentage of blacklisted and non-standard peaks, and peak widths) and fragments (duplication rate) of samples, (2. Peak overlap) Percentage and statistical significance of overlapping and non-overlapping peaks. Also includes upset plot and (3. Functional annotation) functional annotation (ChromHMM, ChIPseeker and enrichment analysis) of peaks. Also includes peak enrichment around TSS.
Last updated
epigeneticsgeneticsqualitycontrolchipseqmultiplecomparisonfunctionalgenomicsatacseqdnaseseqbenchmarkbenchmarkingbioconductorbioconductor-packagecomparisonhtmlinteractive-reporting
7.66 score 19 stars 46 scripts 352 downloadsgDRutils - A package with helper functions for processing drug response data
This package contains utility functions used throughout the gDR platform to fit data, manipulate data, and convert and validate data structures. This package also has the necessary default constants for gDR platform. Many of the functions are utilized by the gDRcore package.
Last updated
softwareinfrastructure
7.65 score 2 stars 4 dependents 34 scripts 312 downloadsbiocthis - Automate package and project setup for Bioconductor packages
This package expands the usethis package with the goal of helping automate the process of creating R packages for Bioconductor or making them Bioconductor-friendly.
Last updated
softwarereportwritingactionsbioconductorbiocthisgithubstylerusethis
7.65 score 56 stars 1 dependents 5 scripts 1.0k downloadsepiregulon - Gene regulatory network inference from single cell epigenomic data
Gene regulatory networks model the underlying gene regulation hierarchies that drive gene expression and observed phenotypes. Epiregulon infers TF activity in single cells by constructing a gene regulatory network (regulons). This is achieved through integration of scATAC-seq and scRNA-seq data and incorporation of public bulk TF ChIP-seq data. Links between regulatory elements and their target genes are established by computing correlations between chromatin accessibility and gene expressions.
Last updated
singlecellgeneregulationnetworkinferencenetworkgeneexpressiontranscriptiongenetargetcpp
7.61 score 28 stars 1 dependents 32 scripts 313 downloadsstandR - Spatial transcriptome analyses of Nanostring's DSP data in R
standR is an user-friendly R package providing functions to assist conducting good-practice analysis of Nanostring's GeoMX DSP data. All functions in the package are built based on the SpatialExperiment object, allowing integration into various spatial transcriptomics-related packages from Bioconductor. standR allows data inspection, quality control, normalization, batch correction and evaluation with informative visualizations.
Last updated
spatialtranscriptomicsgeneexpressiondifferentialexpressionqualitycontrolnormalizationexperimenthubsoftware
7.58 score 26 stars 1 dependents 69 scripts 464 downloadsMGnifyR - R interface to EBI MGnify metagenomics resource
Utility package to facilitate integration and analysis of EBI MGnify data in R. The package can be used to import microbial data for instance into TreeSummarizedExperiment (TreeSE). In TreeSE format, the data is directly compatible with miaverse framework.
Last updated
infrastructuredataimportmetagenomicsmicrobiomemicrobiomedata
7.57 score 23 stars 45 scripts 342 downloadsscMultiSim - Simulation of Multi-Modality Single Cell Data Guided By Gene Regulatory Networks and Cell-Cell Interactions
scMultiSim simulates paired single cell RNA-seq, single cell ATAC-seq and RNA velocity data, while incorporating mechanisms of gene regulatory networks, chromatin accessibility and cell-cell interactions. It allows users to tune various parameters controlling the amount of each biological factor, variation of gene-expression levels, the influence of chromatin accessibility on RNA sequence data, and so on. It can be used to benchmark various computational methods for single cell multi-omics data, and to assist in experimental design of wet-lab experiments.
Last updated
singlecelltranscriptomicsgeneexpressionsequencingexperimentaldesign
7.57 score 68 stars 34 scripts 304 downloadsbiocmake - CMake for Bioconductor
Manages the installation of CMake for building Bioconductor packages. This avoids the need for end-users to manually install CMake on their system. No action is performed if a suitable version of CMake is already available.
Last updated
infrastructure
7.55 score 1 stars 361 dependents 5 scripts 1.6k downloadsCHETAH - Fast and accurate scRNA-seq cell type identification
CHETAH (CHaracterization of cEll Types Aided by Hierarchical classification) is an accurate, selective and fast scRNA-seq classifier. Classification is guided by a reference dataset, preferentially also a scRNA-seq dataset. By hierarchical clustering of the reference data, CHETAH creates a classification tree that enables a step-wise, top-to-bottom classification. Using a novel stopping rule, CHETAH classifies the input cells to the cell types of the references and to "intermediate types": more general classifications that ended in an intermediate node of the tree.
Last updated
classificationrnaseqsinglecellclusteringgeneexpressionimmunooncology
7.53 score 44 stars 85 scripts 404 downloadsMSstatsShiny - MSstats GUI for Statistical Anaylsis of Proteomics Experiments
MSstatsShiny is an R-Shiny graphical user interface (GUI) integrated with the R packages MSstats, MSstatsTMT, and MSstatsPTM. It provides a point and click end-to-end analysis pipeline applicable to a wide variety of experimental designs. These include data-dependedent acquisitions (DDA) which are label-free or tandem mass tag (TMT)-based, as well as DIA, SRM, and PRM acquisitions and those targeting post-translational modifications (PTMs). The application automatically saves users selections and builds an R script that recreates their analysis, supporting reproducible data analysis.
Last updated
immunooncologymassspectrometryproteomicssoftwareshinyappsdifferentialexpressiononechanneltwochannelnormalizationqualitycontrolgui
7.46 score 20 stars 11 scripts 375 downloadsbaySeq - Empirical Bayesian analysis of patterns of differential expression in count data
This package identifies differential expression in high-throughput 'count' data, such as that derived from next-generation sequencing machines, calculating estimated posterior likelihoods of differential expression (or more complex hypotheses) via empirical Bayesian methods.
Last updated
sequencingdifferentialexpressionmultiplecomparisonsagebayesiancoverage
7.42 score 3 dependents 91 scripts 920 downloadsimmApex - Tools for Adaptive Immune Receptor Sequence-Based Machine and Deep Learning
A set of tools to for machine and deep learning in R from amino acid and nucleotide sequences focusing on adaptive immune receptors. The package includes pre-processing of sequences, unifying gene nomenclature usage, encoding sequences, and combining models. This package will serve as the basis of future immune receptor sequence functions/packages/models compatible with the scRepertoire ecosystem.
Last updated
softwareimmunooncologysinglecellclassificationannotationsequencingmotifannotationcppopenmp
7.40 score 14 stars 4 dependents 15 scripts 794 downloadsGenomicDistributions - GenomicDistributions: fast analysis of genomic intervals with Bioconductor
If you have a set of genomic ranges, this package can help you with visualization and comparison. It produces several kinds of plots, for example: Chromosome distribution plots, which visualize how your regions are distributed over chromosomes; feature distance distribution plots, which visualizes how your regions are distributed relative to a feature of interest, like Transcription Start Sites (TSSs); genomic partition plots, which visualize how your regions overlap given genomic features such as promoters, introns, exons, or intergenic regions. It also makes it easy to compare one set of ranges to another.
Last updated
softwaregenomeannotationgenomeassemblydatarepresentationsequencingcoveragefunctionalgenomicsvisualization
7.40 score 27 stars 33 scriptspsichomics - Graphical Interface for Alternative Splicing Quantification, Analysis and Visualisation
Interactive R package with an intuitive Shiny-based graphical interface for alternative splicing quantification and integrative analyses of alternative splicing and gene expression based on The Cancer Genome Atlas (TCGA), the Genotype-Tissue Expression project (GTEx), Sequence Read Archive (SRA) and user-provided data. The tool interactively performs survival, dimensionality reduction and median- and variance-based differential splicing and gene expression analyses that benefit from the incorporation of clinical and molecular sample-associated features (such as tumour stage or survival). Interactive visual access to genomic mapping and functional annotation of selected alternative splicing events is also included.
Last updated
sequencingrnaseqalternativesplicingdifferentialsplicingtranscriptionguiprincipalcomponentsurvivalbiomedicalinformaticstranscriptomicsimmunooncologyvisualizationmultiplecomparisongeneexpressiondifferentialexpressionalternative-splicingbioconductordata-analysesdifferential-gene-expressiondifferential-splicing-analysisgene-expressiongtexrecount2rna-seq-datasplicing-quantificationsratcgavast-toolscpp
7.36 score 37 stars 39 scripts 474 downloadscrisprBase - Base functions and classes for CRISPR gRNA design
Provides S4 classes for general nucleases, CRISPR nucleases, CRISPR nickases, and base editors.Several CRISPR-specific genome arithmetic functions are implemented to help extract genomic coordinates of spacer and protospacer sequences. Commonly-used CRISPR nuclease objects are provided that can be readily used in other packages. Both DNA- and RNA-targeting nucleases are supported.
Last updated
crisprfunctionalgenomicsbioconductorbioconductor-packagecrispr-cas9crispr-designcrispr-targetgrnagrna-sequencegrna-sequences
7.33 score 5 stars 6 dependents 79 scripts 472 downloadsHiCExperiment - Bioconductor class for interacting with Hi-C files in R
R generic interface to Hi-C contact matrices in `.(m)cool`, `.hic` or HiC-Pro derived formats, as well as other Hi-C processed file formats. Contact matrices can be partially parsed using a random access method, allowing a memory-efficient representation of Hi-C data in R. The `HiCExperiment` class stores the Hi-C contacts parsed from local contact matrix files. `HiCExperiment` instances can be further investigated in R using the `HiContacts` analysis package.
Last updated
hicdna3dstructuredataimport
7.32 score 13 stars 3 dependents 40 scripts 454 downloads
gemma.R - A wrapper for Gemma's Restful API to access curated gene expression data and differential expression analyses
Low- and high-level wrappers for Gemma's RESTful API. They enable access to curated expression and differential expression data from over 10,000 published studies. Gemma is a web site, database and a set of tools for the meta-analysis, re-use and sharing of genomics data, currently primarily targeted at the analysis of gene expression profiles.
Last updated
softwaredataimportmicroarraysinglecellthirdpartyclientdifferentialexpressiongeneexpressionbayesianannotationexperimentaldesignnormalizationbatcheffectpreprocessingbioinformaticsgemmagenomicstranscriptomics
7.31 score 10 stars 55 scripts 384 downloadsgDRcore - Processing functions and interface to process and analyze drug dose-response data
This package contains core functions to process and analyze drug response data. The package provides tools for normalizing, averaging, and calculation of gDR metrics data. All core functions are wrapped into the pipeline function allowing analyzing the data in a straightforward way.
Last updated
softwareshinyappscpp
7.31 score 2 stars 1 dependents 23 scripts 320 downloadsmariner - Mariner: Explore the Hi-Cs
Tools for manipulating paired ranges and working with Hi-C data in R. Functionality includes manipulating/merging paired regions, generating paired ranges, extracting/aggregating interactions from `.hic` files, and visualizing the results. Designed for compatibility with plotgardener for visualization.
Last updated
functionalgenomicsvisualizationhic
7.27 score 12 stars 260 scripts 339 downloadssimona - Semantic Similarity on Bio-Ontologies
This package implements infrastructures for ontology analysis by offering efficient data structures, fast ontology traversal methods, and elegant visualizations. It provides a robust toolbox supporting over 70 methods for semantic similarity analysis.
Last updated
softwareannotationgobiomedicalinformaticscpp
7.27 score 18 stars 2 dependents 44 scripts 1.3k downloads
DeconvoBuddies - Helper Functions for LIBD Deconvolution
Functions helpful for LIBD deconvolution project. Includes tools for marker finding with mean ratio, expression plotting, and plotting deconvolution results. Working to include DLPFC datasets.
Last updated
softwaresinglecellrnaseqgeneexpressiontranscriptomicsexperimenthubsoftwarebioconductordeconvolution
7.25 score 10 stars 49 scripts 284 downloadsCytoPipeline - Automation and visualization of flow cytometry data analysis pipelines
This package provides support for automation and visualization of flow cytometry data analysis pipelines. In the current state, the package focuses on the preprocessing and quality control part. The framework is based on two main S4 classes, i.e. CytoPipeline and CytoProcessingStep. The pipeline steps are linked to corresponding R functions - that are either provided in the CytoPipeline package itself, or exported from a third party package, or coded by the user her/himself. The processing steps need to be specified centrally and explicitly using either a json input file or through step by step creation of a CytoPipeline object with dedicated methods. After having run the pipeline, obtained results at all steps can be retrieved and visualized thanks to file caching (the running facility uses a BiocFileCache implementation). The package provides also specific visualization tools like pipeline workflow summary display, and 1D/2D comparison plots of obtained flowFrames at various steps of the pipeline.
Last updated
flowcytometrypreprocessingqualitycontrolworkflowstepimmunooncologysoftwarevisualization
7.22 score 7 stars 3 dependents 19 scriptslfa - Logistic Factor Analysis for Categorical Data
Logistic Factor Analysis is a method for a PCA analogue on Binomial data via estimation of latent structure in the natural parameter. The main method estimates genetic population structure from genotype data. There are also methods for estimating individual-specific allele frequencies using the population structure. Lastly, a structured Hardy-Weinberg equilibrium (HWE) test is developed, which quantifies the goodness of fit of the genotype data to the estimated population structure, via the estimated individual-specific allele frequencies (all of which generalizes traditional HWE tests).
Last updated
snpdimensionreductionprincipalcomponentregressionopenblas
7.22 score 16 stars 1 dependents 58 scripts 686 downloadskoinar - KoinaR - Remote machine learning inference using Koina
A client to simplify fetching predictions from the Koina web service. Koina is a model repository enabling the remote execution of models. Predictions are generated as a response to HTTP/S requests, the standard protocol used for nearly all web traffic.
Last updated
massspectrometryproteomicsinfrastructuresoftwarebioinformaticsdeep-learningmachine-learningmass-spectrometrypython
7.19 score 57 stars 9 scripts 268 downloadsTnT - Interactive Visualization for Genomic Features
A R interface to the TnT javascript library (https://github.com/ tntvis) to provide interactive and flexible visualization of track-based genomic data.
Last updated
infrastructurevisualizationbioconductorgenome-browserhtmlwidgetsshiny
7.18 score 15 stars 17 scripts 490 downloads
BulkSignalR - Infer Ligand-Receptor Interactions from bulk expression (transcriptomics/proteomics) data, or spatial transcriptomics
Inference of ligand-receptor (LR) interactions from bulk expression (transcriptomics/proteomics) data, or spatial transcriptomics. BulkSignalR bases its inferences on the LRdb database included in our other package, SingleCellSignalR available from Bioconductor. It relies on a statistical model that is specific to bulk data sets. Different visualization and data summary functions are proposed to help navigating prediction results.
Last updated
networkrnaseqsoftwareproteomicstranscriptomicsnetworkinferencespatial
7.18 score 28 stars 1 dependents 17 scripts 456 downloadssystemPipeShiny - systemPipeShiny: An Interactive Framework for Workflow Management and Visualization
systemPipeShiny (SPS) extends the widely used systemPipeR (SPR) workflow environment with a versatile graphical user interface provided by a Shiny App. This allows non-R users, such as experimentalists, to run many systemPipeR’s workflow designs, control, and visualization functionalities interactively without requiring knowledge of R. Most importantly, SPS has been designed as a general purpose framework for interacting with other R packages in an intuitive manner. Like most Shiny Apps, SPS can be used on both local computers as well as centralized server-based deployments that can be accessed remotely as a public web service for using SPR’s functionalities with community and/or private data. The framework can integrate many core packages from the R/Bioconductor ecosystem. Examples of SPS’ current functionalities include: (a) interactive creation of experimental designs and metadata using an easy to use tabular editor or file uploader; (b) visualization of workflow topologies combined with auto-generation of R Markdown preview for interactively designed workflows; (d) access to a wide range of data processing routines; (e) and an extendable set of visualization functionalities. Complex visual results can be managed on a 'Canvas Workbench’ allowing users to organize and to compare plots in an efficient manner combined with a session snapshot feature to continue work at a later time. The present suite of pre-configured visualization examples. The modular design of SPR makes it easy to design custom functions without any knowledge of Shiny, as well as extending the environment in the future with contributions from the community.
Last updated
shinyappsinfrastructuredataimportsequencingqualitycontrolreportwritingexperimentaldesignclusteringbioconductorbioconductor-packagedata-visualizationshinysystempiper
7.15 score 36 stars 44 scripts 356 downloads
escheR - Unified multi-dimensional visualizations with Gestalt principles
The creation of effective visualizations is a fundamental component of data analysis. In biomedical research, new challenges are emerging to visualize multi-dimensional data in a 2D space, but current data visualization tools have limited capabilities. To address this problem, we leverage Gestalt principles to improve the design and interpretability of multi-dimensional data in 2D data visualizations, layering aesthetics to display multiple variables. The proposed visualization can be applied to spatially-resolved transcriptomics data, but also broadly to data visualized in 2D space, such as embedding visualizations. We provide this open source R package escheR, which is built off of the state-of-the-art ggplot2 visualization framework and can be seamlessly integrated into genomics toolboxes and workflows.
Last updated
spatialsinglecelltranscriptomicsvisualizationsoftwaremultidimensionalsingle-cellspatial-omics
7.15 score 8 stars 1 dependents 296 scripts 451 downloadstidytof - Analyze High-dimensional Cytometry Data Using Tidy Data Principles
This package implements an interactive, scientific analysis pipeline for high-dimensional cytometry data built using tidy data principles. It is specifically designed to play well with both the tidyverse and Bioconductor software ecosystems, with functionality for reading/writing data files, data cleaning, preprocessing, clustering, visualization, modeling, and other quality-of-life functions. tidytof implements a "grammar" of high-dimensional cytometry data analysis.
Last updated
singlecellflowcytometrybioinformaticscytometrydata-sciencesingle-celltidyversecpp
7.09 score 20 stars 37 scripts 62 downloads
tidyomics - Easily install and load the tidyomics ecosystem
The tidyomics ecosystem is a set of packages for ’omic data analysis that work together in harmony; they share common data representations and API design, consistent with the tidyverse ecosystem. The tidyomics package is designed to make it easy to install and load core packages from the tidyomics ecosystem with a single command.
Last updated
assaydomaininfrastructurernaseqdifferentialexpressiongeneexpressionnormalizationclusteringqualitycontrolsequencingtranscriptiontranscriptomicscytometrygenomicstidyverse
7.08 score 74 stars 27 scripts 260 downloadscardelino - Clone Identification from Single Cell Data
Methods to infer clonal tree configuration for a population of cells using single-cell RNA-seq data (scRNA-seq), and possibly other data modalities. Methods are also provided to assign cells to inferred clones and explore differences in gene expression between clones. These methods can flexibly integrate information from imperfect clonal trees inferred based on bulk exome-seq data, and sparse variant alleles expressed in scRNA-seq data. A flexible beta-binomial error model that accounts for stochastic dropout events as well as systematic allelic imbalance is used.
Last updated
singlecellrnaseqvisualizationtranscriptomicsgeneexpressionsequencingsoftwareexomeseqclonal-clusteringgibbs-samplingscrna-seqsingle-cellsomatic-mutations
7.07 score 65 stars 60 scripts 382 downloadsmegadepth - megadepth: BigWig and BAM related utilities
This package provides an R interface to Megadepth by Christopher Wilks available at https://github.com/ChristopherWilks/megadepth. It is particularly useful for computing the coverage of a set of genomic regions across bigWig or BAM files. With this package, you can build base-pair coverage matrices for regions or annotations of your choice from BigWig files. Megadepth was used to create the raw files provided by https://bioconductor.org/packages/recount3.
Last updated
softwarecoveragedataimporttranscriptomicsrnaseqpreprocessingbambigwigdasptermegadepthrecount2recount3
7.03 score 14 stars 3 dependents 19 scripts 508 downloadsconcordexR - Identify Spatial Homogeneous Regions with concordex
Spatial homogeneous regions (SHRs) in tissues are domains that are homogenous with respect to cell type composition. We present a method for identifying SHRs using spatial transcriptomics data, and demonstrate that it is efficient and effective at finding SHRs for a wide variety of tissue types. concordex relies on analysis of k-nearest-neighbor (kNN) graphs. The tool is also useful for analysis of non-spatial transcriptomics data, and can elucidate the extent of concordance between partitions of cells derived from clustering algorithms, and transcriptomic similarity as represented in kNN graphs.
Last updated
singlecellclusteringspatialtranscriptomics
7.03 score 15 stars 178 scripts 302 downloadsAlphaMissenseR - Accessing AlphaMissense Data Resources in R
The AlphaMissense publication <https://www.science.org/doi/epdf/10.1126/science.adg7492> outlines how a variant of AlphaFold / DeepMind was used to predict missense variant pathogenicity. Supporting data on Zenodo <https://zenodo.org/record/10813168> include, for instance, 71M variants across hg19 and hg38 genome builds. The 'AlphaMissenseR' package allows ready access to the data, downloading individual files to DuckDB databases for exploration and integration into *R* and *Bioconductor* workflows.
Last updated
snpannotationfunctionalgenomicsstructuralpredictiontranscriptomicsvariantannotationgenepredictionimmunooncology
7.01 score 13 stars 13 scripts 332 downloads
syntenet - Inference And Analysis Of Synteny Networks
syntenet can be used to infer synteny networks from whole-genome protein sequences and analyze them. Anchor pairs are detected with the MCScanX algorithm, which was ported to this package with the Rcpp framework for R and C++ integration. Anchor pairs from synteny analyses are treated as an undirected unweighted graph (i.e., a synteny network), and users can perform: i. network clustering; ii. phylogenomic profiling (by identifying which species contain which clusters) and; iii. microsynteny-based phylogeny reconstruction with maximum likelihood.
Last updated
softwarenetworkinferencefunctionalgenomicscomparativegenomicsphylogeneticssystemsbiologygraphandnetworkwholegenomenetworkcomparative-genomicsevolutionary-genomicsnetwork-sciencephylogenomicssyntenysynteny-networkcpp
6.99 score 43 stars 1 dependents 19 scripts 450 downloadsGloScope - Population-level Representation on scRNA-Seq data
This package aims at representing and summarizing the entire single-cell profile of a sample. It allows researchers to perform important bioinformatic analyses at the sample-level such as visualization and quality control. The main functions Estimate sample distribution and calculate statistical divergence among samples, and visualize the distance matrix through MDS plots.
Last updated
datarepresentationqualitycontrolrnaseqsequencingsoftwaresinglecell
6.98 score 8 stars 96 scripts 334 downloadsSpotSweeper - Spatially-aware quality control for spatial transcriptomics
Spatially-aware quality control (QC) software for both spot-level and artifact-level QC in spot-based spatial transcripomics, such as 10x Visium. These methods calculate local (nearest-neighbors) mean and variance of standard QC metrics (library size, unique genes, and mitochondrial percentage) to identify outliers spot and large technical artifacts.
Last updated
softwarespatialtranscriptomicsqualitycontrolgeneexpressionbioconductorquality-controlspatial-transcriptomics
6.93 score 16 stars 177 scripts 474 downloadsRCX - R package implementing the Cytoscape Exchange (CX) format
Create, handle, validate, visualize and convert networks in the Cytoscape exchange (CX) format to standard data types and objects. The package also provides conversion to and from objects of iGraph and graphNEL. The CX format is also used by the NDEx platform, a online commons for biological networks, and the network visualization software Cytocape.
Last updated
pathwaysdataimportnetwork
6.91 score 8 stars 1 dependents 17 scripts 362 downloadsNetPathMiner - NetPathMiner for Biological Network Construction, Path Mining and Visualization
NetPathMiner is a general framework for network path mining using genome-scale networks. It constructs networks from KGML, SBML and BioPAX files, providing three network representations, metabolic, reaction and gene representations. NetPathMiner finds active paths and applies machine learning methods to summarize found paths for easy interpretation. It also provides static and interactive visualizations of networks and paths to aid manual investigation.
Last updated
graphandnetworkpathwaysnetworkclusteringclassificationlibsbmllibxml2openblascpp
6.89 score 9 stars 1 dependents 18 scriptsNanoMethViz - Visualise methylation data from Oxford Nanopore sequencing
NanoMethViz is a toolkit for visualising methylation data from Oxford Nanopore sequencing. It can be used to explore methylation patterns from reads derived from Oxford Nanopore direct DNA sequencing with methylation called by callers including nanopolish, f5c and megalodon. The plots in this package allow the visualisation of methylation profiles aggregated over experimental groups and across classes of genomic features.
Last updated
softwarelongreadvisualizationdifferentialmethylationdnamethylationepigeneticsdataimportzlibcpp
6.89 score 37 stars 20 scripts 482 downloadsspatialFDA - A Tool for Spatial Multi-sample Comparisons
spatialFDA is a package to calculate spatial statistics metrics. The package takes a SpatialExperiment object and calculates spatial statistics metrics using the package spatstat. Then it compares the resulting functions across samples/conditions using functional additive models as implemented in the package refund. Furthermore, it provides exploratory visualisations using functional principal component analysis, as well implemented in refund.
Last updated
softwarespatialtranscriptomics
6.89 score 8 stars 30 scriptsTileDBArray - Using TileDB as a DelayedArray Backend
Implements a DelayedArray backend for reading and writing dense or sparse arrays in the TileDB format. The resulting TileDBArrays are compatible with all Bioconductor pipelines that can accept DelayedArray instances.
Last updated
datarepresentationinfrastructuresoftware
6.88 score 11 stars 1 dependents 29 scripts 404 downloadsmiaSim - Microbiome Data Simulation
Microbiome time series simulation with generalized Lotka-Volterra model, Self-Organized Instability (SOI), and other models. Hubbell's Neutral model is used to determine the abundance matrix. The resulting abundance matrix is applied to (Tree)SummarizedExperiment objects.
Last updated
microbiomesoftwaresequencingdnaseqatacseqcoveragenetwork
6.84 score 22 stars 35 scripts 354 downloadsStatial - A package to identify changes in cell state relative to spatial associations
Statial is a suite of functions for identifying changes in cell state. The functionality provided by Statial provides robust quantification of cell type localisation which are invariant to changes in tissue structure. In addition to this Statial uncovers changes in marker expression associated with varying levels of localisation. These features can be used to explore how the structure and function of different cell types may be altered by the agents they are surrounded with.
Last updated
singlecellspatialclassificationsingle-cell
6.83 score 6 stars 36 scriptsVisiumIO - Import Visium data from the 10X Space Ranger pipeline
The package allows users to readily import spatial data obtained from either the 10X website or from the Space Ranger pipeline. Supported formats include tar.gz, h5, and mtx files. Multiple files can be imported at once with *List type of functions. The package represents data mainly as SpatialExperiment objects.
Last updated
softwareinfrastructuredataimportsinglecellspatialbioconductor-packagegenomicsu24ca289073
6.82 score 3 stars 1 dependents 106 scripts 504 downloadsCaDrA - Candidate Driver Analysis
Performs both stepwise and backward heuristic search for candidate (epi)genetic drivers based on a binary multi-omics dataset. CaDrA's main objective is to identify features which, together, are significantly skewed or enriched pertaining to a given vector of continuous scores (e.g. sample-specific scores representing a phenotypic readout of interest, such as protein expression, pathway activity, etc.), based on the union occurence (i.e. logical OR) of the events.
Last updated
microarrayrnaseqgeneexpressionsoftwarefeatureextraction
6.81 score 24 stars 10 scripts 272 downloadsextraChIPs - Additional functions for working with ChIP-Seq data
This package builds on existing tools and adds some simple but extremely useful capabilities for working wth ChIP-Seq data. The focus is on detecting differential binding windows/regions. One set of functions focusses on set-operations retaining mcols for GRanges objects, whilst another group of functions are to aid visualisation of results. Coercion to tibble objects is also implemented.
Last updated
chipseqhicsequencingcoverage
6.79 score 7 stars 37 scripts 475 downloadsSpotClean - SpotClean adjusts for spot swapping in spatial transcriptomics data
SpotClean is a computational method to adjust for spot swapping in spatial transcriptomics data. Recent spatial transcriptomics experiments utilize slides containing thousands of spots with spot-specific barcodes that bind mRNA. Ideally, unique molecular identifiers at a spot measure spot-specific expression, but this is often not the case due to bleed from nearby spots, an artifact we refer to as spot swapping. SpotClean is able to estimate the contamination rate in observed data and decontaminate the spot swapping effect, thus increase the sensitivity and precision of downstream analyses.
Last updated
dataimportrnaseqsequencinggeneexpressionspatialsinglecelltranscriptomicspreprocessingrna-seqspatial-transcriptomics
6.75 score 39 stars 48 scripts 389 downloadsBumpyMatrix - Bumpy Matrix of Non-Scalar Objects
Implements the BumpyMatrix class and several subclasses for holding non-scalar objects in each entry of the matrix. This is akin to a ragged array but the raggedness is in the third dimension, much like a bumpy surface - hence the name. Of particular interest is the BumpyDataFrameMatrix, where each entry is a Bioconductor data frame. This allows us to naturally represent multivariate data in a format that is compatible with two-dimensional containers like the SummarizedExperiment and MultiAssayExperiment objects.
Last updated
softwareinfrastructuredatarepresentation
6.73 score 1 stars 16 dependents 52 scripts 1.1k downloads
MsBackendMsp - Mass Spectrometry Data Backend for NIST msp Files
Mass spectrometry (MS) data backend supporting import and handling of MS/MS spectra from NIST MSP Format (msp) files. Import of data from files with different MSP *flavours* is supported. Objects from this package add support for MSP files to Bioconductor's Spectra package. This package is thus not supposed to be used without the Spectra package that provides a complete infrastructure for MS data handling.
Last updated
infrastructureproteomicsmassspectrometrymetabolomicsdataimportmass-spectrometry
6.72 score 5 stars 2 dependents 44 scripts 709 downloadsSPONGE - Sparse Partial Correlations On Gene Expression
This package provides methods to efficiently detect competitive endogeneous RNA interactions between two genes. Such interactions are mediated by one or several miRNAs such that both gene and miRNA expression data for a larger number of samples is needed as input. The SPONGE package now also includes spongEffects: ceRNA modules offer patient-specific insights into the miRNA regulatory landscape.
Last updated
geneexpressiontranscriptiongeneregulationnetworkinferencetranscriptomicssystemsbiologyregressionrandomforestmachinelearning
6.72 score 1 dependents 58 scripts 452 downloadsfastreeR - Phylogenetic, Distance and Other Calculations on VCF and Fasta Files
Calculate distances, build phylogenetic trees or perform hierarchical clustering between the samples of a VCF or FASTA file. Functions are implemented in Java-11 and called via rJava. Parallel implementation that operates directly on the VCF or FASTA file for fast execution.
Last updated
phylogeneticsmetagenomicsclusteringopenjdk
6.72 score 31 stars 28 scripts 322 downloadsMuData - Serialization for MultiAssayExperiment Objects
Save MultiAssayExperiments to h5mu files supported by muon and mudata. Muon is a Python framework for multimodal omics data analysis. It uses an HDF5-based format for data storage.
Last updated
dataimportanndatabioconductormudatamulti-omicsmultimodal-omicsscrna-seq
6.67 score 10 stars 39 scripts 377 downloadsBSgenomeForge - Forge your own BSgenome data package
A set of tools to forge BSgenome data packages. Supersedes the old seed-based tools from the BSgenome software package. This package allows the user to create a BSgenome data package in one function call, simplifying the old seed-based process.
Last updated
infrastructuredatarepresentationgenomeassemblyannotationgenomeannotationsequencingalignmentdataimportsequencematchingbioconductor-packagecore-package
6.67 score 5 stars 31 scripts 537 downloadsSpatialExperimentIO - Read in Xenium, CosMx, MERSCOPE or STARmapPLUS data as SpatialExperiment object
Read in imaging-based spatial transcriptomics technology data. Current available modules are for Xenium by 10X Genomics, CosMx by Nanostring, MERSCOPE by Vizgen, or STARmapPLUS from Broad Institute. You can choose to read the data in as a SpatialExperiment or a SingleCellExperiment object.
Last updated
datarepresentationdataimportinfrastructuretranscriptomicssinglecellspatialgeneexpression
6.64 score 19 stars 1 dependents 77 scripts 360 downloadsENmix - Quality control and analysis tools for Illumina DNA methylation BeadChip
Tools for quanlity control, analysis and visulization of Illumina DNA methylation array data.
Last updated
dnamethylationpreprocessingqualitycontroltwochannelmicroarrayonechannelmethylationarraybatcheffectnormalizationdataimportregressionprincipalcomponentepigeneticsmultichanneldifferentialmethylationimmunooncology
6.64 score 1 dependents 182 scripts 904 downloads
tidySpatialExperiment - SpatialExperiment with tidy principles
tidySpatialExperiment provides a bridge between the SpatialExperiment package and the tidyverse ecosystem. It creates an invisible layer that allows you to interact with a SpatialExperiment object as if it were a tibble; enabling the use of functions from dplyr, tidyr, ggplot2 and plotly. But, underneath, your data remains a SpatialExperiment object.
Last updated
infrastructurernaseqgeneexpressionsequencingspatialtranscriptomicssinglecell
6.64 score 8 stars 1 dependents 24 scripts 382 downloads
dar - Differential Abundance Analysis by Consensus
Differential abundance testing in microbiome data challenges both parametric and non-parametric statistical methods, due to its sparsity, high variability and compositional nature. Microbiome-specific statistical methods often assume classical distribution models or take into account compositional specifics. These produce results that range within the specificity vs sensitivity space in such a way that type I and type II error that are difficult to ascertain in real microbiome data when a single method is used. Recently, a consensus approach based on multiple differential abundance (DA) methods was recently suggested in order to increase robustness. With dar, you can use dplyr-like pipeable sequences of DA methods and then apply different consensus strategies. In this way we can obtain more reliable results in a fast, consistent and reproducible way.
Last updated
softwaresequencingmicrobiomemetagenomicsmultiplecomparisonnormalizationbioconductorbiomarker-discoverydifferential-abundance-analysisfeature-selectionmicrobiologyphyloseq
6.62 score 6 stars 11 scripts 270 downloadszenith - Gene set analysis following differential expression using linear (mixed) modeling with dream
Zenith performs gene set analysis on the result of differential expression using linear (mixed) modeling with dream by considering the correlation between gene expression traits. This package implements the camera method from the limma package proposed by Wu and Smyth (2012). Zenith is a simple extension of camera to be compatible with linear mixed models implemented in variancePartition::dream().
Last updated
rnaseqgeneexpressiongenesetenrichmentdifferentialexpressionbatcheffectqualitycontrolregressionepigeneticsfunctionalgenomicstranscriptomicsnormalizationpreprocessingmicroarrayimmunooncologysoftware
6.59 score 1 dependents 217 scripts 508 downloads
cogeqc - Systematic quality checks on comparative genomics analyses
cogeqc aims to facilitate systematic quality checks on standard comparative genomics analyses to help researchers detect issues and select the most suitable parameters for each data set. cogeqc can be used to asses: i. genome assembly and annotation quality with BUSCOs and comparisons of statistics with publicly available genomes on the NCBI; ii. orthogroup inference using a protein domain-based approach and; iii. synteny detection using synteny network properties. There are also data visualization functions to explore QC summary statistics.
Last updated
softwaregenomeassemblycomparativegenomicsfunctionalgenomicsphylogeneticsqualitycontrolnetworkcomparative-genomicsevolutionary-genomics
6.59 score 12 stars 36 scripts 388 downloadsSpaNorm - Spatially-aware normalisation for spatial transcriptomics data
This package implements the spatially aware library size normalisation algorithm, SpaNorm. SpaNorm normalises out library size effects while retaining biology through the modelling of smooth functions for each effect. Normalisation is performed in a gene- and cell-/spot- specific manner, yielding library size adjusted data.
Last updated
softwaregeneexpressiontranscriptomicsspatialcellbiology
6.56 score 19 stars 32 scripts 335 downloadscellxgenedp - Discover and Access Single Cell Data Sets in the CELLxGENE Data Portal
The cellxgene data portal (https://cellxgene.cziscience.com/) provides a graphical user interface to collections of single-cell sequence data processed in standard ways to 'count matrix' summaries. The cellxgenedp package provides an alternative, R-based inteface, allowind data discovery, viewing, and downloading.
Last updated
singlecelldataimportthirdpartyclient
6.53 score 9 stars 47 scripts 372 downloadsalabaster.schemas - Schemas for the Alabaster Framework
Stores all schemas required by various alabaster.* packages. No computation should be performed by this package, as that is handled by alabaster.base. We use a separate package instead of storing the schemas in alabaster.base itself, to avoid conflating management of the schemas with code maintenence.
Last updated
datarepresentationdataimport
6.52 score 21 dependents 2 scripts 5.3k downloadsHiContacts - Analysing cool files in R with HiContacts
HiContacts provides a collection of tools to analyse and visualize Hi-C datasets imported in R by HiCExperiment.
Last updated
hicdna3dstructure
6.52 score 16 stars 69 scripts 444 downloadscliqueMS - Annotation of Isotopes, Adducts and Fragmentation Adducts for in-Source LC/MS Metabolomics Data
Annotates data from liquid chromatography coupled to mass spectrometry (LC/MS) metabolomics experiments. Based on a network algorithm (O.Senan, A. Aguilar- Mogas, M. Navarro, O. Yanes, R.Guimerà and M. Sales-Pardo, Bioinformatics, 35(20), 2019), 'CliqueMS' builds a weighted similarity network where nodes are features and edges are weighted according to the similarity of this features. Then it searches for the most plausible division of the similarity network into cliques (fully connected components). Finally it annotates metabolites within each clique, obtaining for each annotated metabolite the neutral mass and their features, corresponding to isotopes, ionization adducts and fragmentation adducts of that metabolite.
Last updated
metabolomicsmassspectrometrynetworknetworkinferencecpp
6.52 score 13 stars 28 scripts 420 downloadscoMethDMR - Accurate identification of co-methylated and differentially methylated regions in epigenome-wide association studies
coMethDMR identifies genomic regions associated with continuous phenotypes by optimally leverages covariations among CpGs within predefined genomic regions. Instead of testing all CpGs within a genomic region, coMethDMR carries out an additional step that selects co-methylated sub-regions first without using any outcome information. Next, coMethDMR tests association between methylation within the sub-region and continuous phenotype using a random coefficient mixed effects model, which models both variations between CpG sites within the region and differential methylation simultaneously.
Last updated
dnamethylationepigeneticsmethylationarraydifferentialmethylationgenomewideassociation
6.51 score 7 stars 46 scriptsscFeatures - scFeatures: Multi-view representations of single-cell and spatial data for disease outcome prediction
scFeatures constructs multi-view representations of single-cell and spatial data. scFeatures is a tool that generates multi-view representations of single-cell and spatial data through the construction of a total of 17 feature types. These features can then be used for a variety of analyses using other software in Biocondutor.
Last updated
cellbasedassayssinglecellspatialsoftwaretranscriptomics
6.50 score 15 stars 21 scripts 312 downloadssosta - A package for the analysis of anatomical tissue structures in spatial omics data
sosta (Spatial Omics STructure Analysis) is a package for analyzing spatial omics data to explore tissue organization at the anatomical structure level. It reconstructs anatomically relevant structures based on molecular features or cell types. It further calculates a range of metrics at the structure level to quantitatively describe tissue architecture. The package is designed to integrate with other packages for the analysis of spatial omics data.
Last updated
softwarespatialtranscriptomicsvisualization
6.48 score 8 stars 21 scripts 299 downloadspoem - POpulation-based Evaluation Metrics
This package provides a comprehensive set of external and internal evaluation metrics. It includes metrics for assessing partitions or fuzzy partitions derived from clustering results, as well as for evaluating subpopulation identification results within embeddings or graph representations. Additionally, it provides metrics for comparing spatial domain detection results against ground truth labels, and tools for visualizing spatial errors.
Last updated
dimensionreductionclusteringgraphandnetworkspatialatacseqsinglecellrnaseqsoftwarevisualization
6.46 score 11 stars 29 scripts 270 downloadsDESpace - DESpace: a framework to discover spatially variable genes and differential spatial patterns across conditions
Intuitive framework for identifying spatially variable genes (SVGs) and differential spatial variable pattern (DSP) between conditions via edgeR, a popular method for performing differential expression analyses. Based on pre-annotated spatial clusters as summarized spatial information, DESpace models gene expression using a negative binomial (NB), via edgeR, with spatial clusters as covariates. SVGs are then identified by testing the significance of spatial clusters. For multi-sample, multi-condition datasets, we again fit a NB model via edgeR, incorporating spatial clusters, conditions and their interactions as covariates. DSP genes-representing differences in spatial gene expression patterns across experimental conditions-are identified by testing the interaction between spatial clusters and conditions.
Last updated
spatialsinglecellrnaseqtranscriptomicsgeneexpressionsequencingdifferentialexpressionstatisticalmethodvisualization
6.44 score 8 stars 58 scripts 496 downloadsDFplyr - A `DataFrame` (`S4Vectors`) backend for `dplyr`
Provides `dplyr` verbs (`mutate`, `select`, `filter`, etc...) supporting `S4Vectors::DataFrame` objects. Importantly, this is achieved without conversion to an intermediate `tibble`. Adds grouping infrastructure to `DataFrame` which is respected by the transformation verbs.
Last updated
datarepresentationinfrastructuresoftware
6.42 score 21 stars 1 dependents 14 scripts 262 downloadsRAIDS - Robust Ancestry Inference using Data Synthesis
This package implements specialized algorithms that enable genetic ancestry inference from various cancer sequences sources (RNA, Exome and Whole-Genome sequences). This package also implements a simulation algorithm that generates synthetic cancer-derived data. This code and analysis pipeline was designed and developed for the following publication: Belleau, P et al. Genetic Ancestry Inference from Cancer-Derived Molecular Data across Genomic and Transcriptomic Platforms. Cancer Res 1 January 2023; 83 (1): 49–58.
Last updated
geneticssoftwaresequencingwholegenomeprincipalcomponentgeneticvariabilitydimensionreductionbiocviewsancestrycancer-genomicsexome-sequencinggenomicsinferencer-languagerna-seqrna-sequencingwhole-genome-sequencing
6.42 score 7 stars 21 scripts 252 downloadsMoleculeExperiment - Prioritising a molecule-level storage of Spatial Transcriptomics Data
MoleculeExperiment contains functions to create and work with objects from the new MoleculeExperiment class. We introduce this class for analysing molecule-based spatial transcriptomics data (e.g., Xenium by 10X, Cosmx SMI by Nanostring, and Merscope by Vizgen). This allows researchers to analyse spatial transcriptomics data at the molecule level, and to have standardised data formats accross vendors.
Last updated
dataimportdatarepresentationinfrastructuresoftwarespatialtranscriptomics
6.42 score 12 stars 55 scripts 316 downloadsalabaster.matrix - Load and Save Artifacts from File
Save matrices, arrays and similar objects into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentationcpp
6.42 score 11 dependents 13 scripts 6.1k downloadsLimROTS - LimROTS: A Hybrid Method Integrating Empirical Bayes and Reproducibility-Optimized Statistics for Robust Differential Expression Analysis
Differential expression analysis is commonly used to study diverse biological datasets. The reproducibility-optimized test statistic (ROTS) (Elo et al., 2008, <doi:10.1109/tcbb.2007.1078>) uses a modified t-statistic to prioritise features that differ between two or more groups. However, the ROTS Bioconductor implementation (Suomi et al., 2017, <doi:10.1371/journal.pcbi.1005562>) did not accommodate technical or biological covariates. LimROTS (Anwar et al., 2025, <doi:10.1093/bioinformatics/btaf570>) addressed this limitation by combining a reproducibility-optimized test statistic with the limma empirical Bayes approach (Ritchie et al., 2015, <doi:10.1093/nar/gkv007>). This enables the analysis of more complex experimental designs and the incorporation of covariates.
Last updated
softwaregeneexpressiondifferentialexpressionmicroarrayrnaseqproteomicsimmunooncologymetabolomicsmrnamicroarray
6.41 score 4 stars 27 scripts 254 downloadspathlinkR - Analyze and interpret RNA-Seq results
pathlinkR is an R package designed to facilitate analysis of RNA-Seq results. Specifically, our aim with pathlinkR was to provide a number of tools which take a list of DE genes and perform different analyses on them, aiding with the interpretation of results. Functions are included to perform pathway enrichment, with muliplte databases supported, and tools for visualizing these results. Genes can also be used to create and plot protein-protein interaction networks, all from inside of R.
Last updated
genesetenrichmentnetworkpathwaysreactomernaseqnetworkenrichmentbioinformaticsnetworkspathway-enrichment-analysisvisualization
6.41 score 32 stars 4 scripts 314 downloadsalabaster.ranges - Load and Save Ranges-related Artifacts from File
Save GenomicRanges, IRanges and related data structures into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
6.41 score 11 dependents 13 scripts 6.0k downloadsAPL - Association Plots
APL is a package developed for computation of Association Plots (AP), a method for visualization and analysis of single cell transcriptomics data. The main focus of APL is the identification of genes characteristic for individual clusters of cells from input data. The package performs correspondence analysis (CA) and allows to identify cluster-specific genes using Association Plots. Additionally, APL computes the cluster-specificity scores for all genes which allows to rank the genes by their specificity for a selected cell cluster of interest.
Last updated
statisticalmethoddimensionreductionsinglecellsequencingrnaseqgeneexpression
6.41 score 17 stars 25 scripts 362 downloadsClustIRR - Clustering of Immune Receptor Repertoires
ClustIRR analyzes repertoires of B- and T-cell receptors. It starts by identifying communities of immune receptors with similar specificities, based on the sequences of their complementarity-determining regions (CDRs). Next, it employs a Bayesian probabilistic models to quantify differential community occupancy (DCO) between repertoires, allowing the identification of expanding or contracting communities in response to e.g. infection or cancer treatment.
Last updated
clusteringimmunooncologysinglecellsoftwareclassificationbayesianbiomedicalinformaticsmathematicalbiologyb-cell-receptorbioinformaticsimmunoinformaticsimmunologyquantitative-methodsrep-seqrepertoire-analysist-cell-receptoronetbbcpp
6.40 score 5 stars 13 scripts 312 downloadsalabaster.se - Load and Save SummarizedExperiments from File
Save SummarizedExperiments into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
6.40 score 10 dependents 14 scripts 6.0k downloadstreeclimbR - An algorithm to find optimal signal levels in a tree
The arrangement of hypotheses in a hierarchical structure appears in many research fields and often indicates different resolutions at which data can be viewed. This raises the question of which resolution level the signal should best be interpreted on. treeclimbR provides a flexible method to select optimal resolution levels (potentially different levels in different parts of the tree), rather than cutting the tree at an arbitrary level. treeclimbR uses a tuning parameter to generate candidate resolutions and from these selects the optimal one.
Last updated
statisticalmethodcellbasedassays
6.40 score 21 stars 60 scripts 282 downloadsscider - Spatial cell-type inter-correlation by density in R
scider is an user-friendly R package providing functions to model the global density of cells in a slide of spatial transcriptomics data. All functions in the package are built based on the SpatialExperiment object, allowing integration into various spatial transcriptomics-related packages from Bioconductor. After modelling density, the package allows for several downstream analysis, including colocalization analysis, boundary detection analysis and differential density analysis.
Last updated
spatialtranscriptomicscppopenjdk
6.40 score 12 stars 15 scripts 309 downloadsSpaceMarkers - Spatial Interaction Markers
Spatial transcriptomic technologies have helped to resolve the connection between gene expression and the 2D orientation of tissues relative to each other. However, the limited single-cell resolution makes it difficult to highlight the most important molecular interactions in these tissues. SpaceMarkers, R/Bioconductor software, can help to find molecular interactions, by identifying genes associated with latent space interactions in spatial transcriptomics.
Last updated
singlecellgeneexpressionsoftwarespatialtranscriptomics
6.39 score 8 stars 41 scripts 256 downloads
doubletrouble - Identification and classification of duplicated genes
doubletrouble aims to identify duplicated genes from whole-genome protein sequences and classify them based on their modes of duplication. The duplication modes are i. segmental duplication (SD); ii. tandem duplication (TD); iii. proximal duplication (PD); iv. transposed duplication (TRD) and; v. dispersed duplication (DD). Transposon-derived duplicates (TRD) can be further subdivided into rTRD (retrotransposon-derived duplication) and dTRD (DNA transposon-derived duplication). If users want a simpler classification scheme, duplicates can also be classified into SD- and SSD-derived (small-scale duplication) gene pairs. Besides classifying gene pairs, users can also classify genes, so that each gene is assigned a unique mode of duplication. Users can also calculate substitution rates per substitution site (i.e., Ka and Ks) from duplicate pairs, find peaks in Ks distributions with Gaussian Mixture Models (GMMs), and classify gene pairs into age groups based on Ks peaks.
Last updated
softwarewholegenomecomparativegenomicsfunctionalgenomicsphylogeneticsnetworkclassificationbioinformaticscomparative-genomicsgene-duplicationmolecular-evolutionwhole-genome-duplication
6.39 score 37 stars 22 scripts 412 downloadsRiboCrypt - Interactive visualization in genomics
R Package for interactive visualization and browsing NGS data. It contains a browser for both transcript and genomic coordinate view. In addition a QC and general metaplots are included, among others differential translation plots and gene expression plots. The package is still under development.
Last updated
softwaresequencingriboseqrnaseq
6.37 score 6 stars 26 scripts 416 downloadsChromSCape - Analysis of single-cell epigenomics datasets with a Shiny App
ChromSCape - Chromatin landscape profiling for Single Cells - is a ready-to-launch user-friendly Shiny Application for the analysis of single-cell epigenomics datasets (scChIP-seq, scATAC-seq, scCUT&Tag, ...) from aligned data to differential analysis & gene set enrichment analysis. It is highly interactive, enables users to save their analysis and covers a wide range of analytical steps: QC, preprocessing, filtering, batch correction, dimensionality reduction, vizualisation, clustering, differential analysis and gene set analysis.
Last updated
shinyappssoftwaresinglecellchipseqatacseqmethylseqclassificationclusteringepigeneticsprincipalcomponentannotationbatcheffectmultiplecomparisonnormalizationpathwayspreprocessingqualitycontrolreportwritingvisualizationgenesetenrichmentdifferentialpeakcallingepigenomicsshinysingle-cellcpp
6.36 score 14 stars 18 scripts 375 downloadsscMET - Bayesian modelling of cell-to-cell DNA methylation heterogeneity
High-throughput single-cell measurements of DNA methylomes can quantify methylation heterogeneity and uncover its role in gene regulation. However, technical limitations and sparse coverage can preclude this task. scMET is a hierarchical Bayesian model which overcomes sparsity, sharing information across cells and genomic features to robustly quantify genuine biological heterogeneity. scMET can identify highly variable features that drive epigenetic heterogeneity, and perform differential methylation and variability analyses. We illustrate how scMET facilitates the characterization of epigenetically distinct cell populations and how it enables the formulation of novel hypotheses on the epigenetic regulation of gene expression.
Last updated
immunooncologydnamethylationdifferentialmethylationdifferentialexpressiongeneexpressiongeneregulationepigeneticsgeneticsclusteringfeatureextractionregressionbayesiansequencingcoveragesinglecellbayesian-inferencegeneralised-linear-modelsheterogeneityhierarchical-modelsmethylation-analysissingle-cellonetbbcpp
6.31 score 25 stars 41 scripts 358 downloadsiNETgrate - Integrates DNA methylation data with gene expression in a single gene network
The iNETgrate package provides functions to build a correlation network in which nodes are genes. DNA methylation and gene expression data are integrated to define the connections between genes. This network is used to identify modules (clusters) of genes. The biological information in each of the resulting modules is represented by an eigengene. These biological signatures can be used as features e.g., for classification of patients into risk categories. The resulting biological signatures are very robust and give a holistic view of the underlying molecular changes.
Last updated
geneexpressionrnaseqdnamethylationnetworkinferencenetworkgraphandnetworkbiomedicalinformaticssystemsbiologytranscriptomicsclassificationclusteringdimensionreductionprincipalcomponentmrnamicroarraynormalizationgenepredictionkeggsurvivalcore-services
6.30 score 76 stars 1 scriptsdeconvR - Simulation and Deconvolution of Omic Profiles
This package provides a collection of functions designed for analyzing deconvolution of the bulk sample(s) using an atlas of reference omic signature profiles and a user-selected model. Users are given the option to create or extend a reference atlas and,also simulate the desired size of the bulk signature profile of the reference cell types.The package includes the cell-type-specific methylation atlas and, Illumina Epic B5 probe ids that can be used in deconvolution. Additionally,we included BSmeth2Probe, to make mapping WGBS data to their probe IDs easier.
Last updated
dnamethylationregressiongeneexpressionrnaseqsinglecellstatisticalmethodtranscriptomicsbioconductor-packagedeconvolutiondna-methylationomics
6.29 score 10 stars 26 scripts 420 downloads
CuratedAtlasQueryR - Queries the Human Cell Atlas
Provides access to a copy of the Human Cell Atlas, but with harmonised metadata. This allows for uniform querying across numerous datasets within the Atlas using common fields such as cell type, tissue type, and patient ethnicity. Usage involves first querying the metadata table for cells of interest, and then downloading the corresponding cells into a SingleCellExperiment object.
Last updated
assaydomaininfrastructurernaseqdifferentialexpressiongeneexpressionnormalizationclusteringqualitycontrolsequencingtranscriptiontranscriptomicsdatabaseduckdbhdf5human-cell-atlassingle-cellsinglecellexperimenttidyverse
6.27 score 93 stars 45 scripts 293 downloadsPLSDAbatch - PLSDA-batch
A novel framework to correct for batch effects prior to any downstream analysis in microbiome data based on Projection to Latent Structures Discriminant Analysis. The main method is named “PLSDA-batch”. It first estimates treatment and batch variation with latent components, then subtracts batch-associated components from the data whilst preserving biological variation of interest. PLSDA-batch is highly suitable for microbiome data as it is non-parametric, multivariate and allows for ordination and data visualisation. Combined with centered log-ratio transformation for addressing uneven library sizes and compositional structure, PLSDA-batch addresses all characteristics of microbiome data that existing correction methods have ignored so far. Two other variants are proposed for 1/ unbalanced batch x treatment designs that are commonly encountered in studies with small sample sizes, and for 2/ selection of discriminative variables amongst treatment groups to avoid overfitting in classification problems. These two variants have widened the scope of applicability of PLSDA-batch to different data settings.
Last updated
statisticalmethoddimensionreductionprincipalcomponentclassificationmicrobiomebatcheffectnormalizationvisualization
6.26 score 16 stars 57 scripts 288 downloadsMsQuality - MsQuality - Quality metric calculation from Spectra, MsExperiment and Chromatograms objects
The MsQuality provides functionality to calculate quality metrics for mass spectrometry-derived, spectral data at the per-sample level. MsQuality relies on the mzQC framework of quality metrics defined by the Human Proteom Organization-Proteomics Standards Initiative (HUPO-PSI). These metrics quantify the quality of spectral raw files using a controlled vocabulary. The package is especially addressed towards users that acquire mass spectrometry data on a large scale (e.g. data sets from clinical settings consisting of several thousands of samples). The MsQuality package allows to calculate low-level quality metrics that require minimum information on mass spectrometry data: retention time, m/z values, and associated intensities. MsQuality relies on the Spectra package, or alternatively the MsExperiment package, and its infrastructure to store spectral data. Additionally, MsQuality supports Chromatograms objects from the Chromatograms package for chromatographic quality metrics.
Last updated
metabolomicsproteomicsmassspectrometryqualitycontrolmass-spectrometryqc
6.26 score 9 stars 10 scripts 332 downloadsOSTA.data - OSTA book data
'OSTA.data' is a companion package for the "Orchestrating Spatial Transcriptomics Analysis" (OSTA) with Bioconductor online book. Throughout OSTA, we rely on a set of publicly available datasets that cover different sequencing- and imaging-based platforms, such as Visium, Visium HD, Xenium (10x Genomics) and CosMx (NanoString). In addition, we rely on scRNA-seq (Chromium) data for tasks, e.g., spot deconvolution and label transfer (i.e., supervised clustering). These data been deposited in an Open Storage Framework (OSF) repository, and can be queried and downloaded using functions from the 'osfr' package. For convenience, we have implemented 'OSTA.data' to query and retrieve data from our OSF node, and cache retrieved Zip archives using 'BiocFileCache'.
Last updated
dataimportdatarepresentationexperimenthubsoftwareinfrastructureimmunooncologygeneexpressiontranscriptomicssinglecellspatial
6.25 score 2 stars 148 scripts 379 downloadsCatsCradle - This package provides methods for analysing spatial transcriptomics data and for discovering gene clusters
This package addresses two broad areas. It allows for in-depth analysis of spatial transcriptomic data by identifying tissue neighbourhoods. These are contiguous regions of tissue surrounding individual cells. 'CatsCradle' allows for the categorisation of neighbourhoods by the cell types contained in them and the genes expressed in them. In particular, it produces Seurat objects whose individual elements are neighbourhoods rather than cells. In addition, it enables the categorisation and annotation of genes by producing Seurat objects whose elements are genes.
Last updated
biologicalquestionstatisticalmethodgeneexpressionsinglecelltranscriptomicsspatial
6.24 score 7 stars 7 scripts 311 downloadsPRONE - The PROteomics Normalization Evaluator
High-throughput omics data are often affected by systematic biases introduced throughout all the steps of a clinical study, from sample collection to quantification. Normalization methods aim to adjust for these biases to make the actual biological signal more prominent. However, selecting an appropriate normalization method is challenging due to the wide range of available approaches. Therefore, a comparative evaluation of unnormalized and normalized data is essential in identifying an appropriate normalization strategy for a specific data set. This R package provides different functions for preprocessing, normalizing, and evaluating different normalization approaches. Furthermore, normalization methods can be evaluated on downstream steps, such as differential expression analysis and statistical enrichment analysis. Spike-in data sets with known ground truth and real-world data sets of biological experiments acquired by either tandem mass tag (TMT) or label-free quantification (LFQ) can be analyzed.
Last updated
proteomicspreprocessingnormalizationdifferentialexpressionvisualizationdata-analysisevaluation
6.24 score 8 stars 12 scripts 299 downloadsRigraphlib - igraph library as an R package
Vendors the igraph C source code and builds it into a static library. Other Bioconductor packages can link to libigraph.a in their own C/C++ code. This is intended for packages wrapping C/C++ libraries that depend on the igraph C library and cannot be easily adapted to use the igraph R package.
Last updated
clusteringgraphandnetwork
6.22 score 1 stars 15 dependents 1.9k downloadsAnVILBase - Generic functions for interacting with the AnVIL ecosystem
Provides generic functions for interacting with the AnVIL ecosystem. Packages that use either GCP or Azure in AnVIL are built on top of AnVILBase. Extension packages will provide methods for interacting with other cloud providers.
Last updated
softwareinfrastructureu24hg010263
6.21 score 17 dependents 79 scripts 739 downloadsfaers - R interface for FDA Adverse Event Reporting System
The FDA Adverse Event Reporting System (FAERS) is a database used for the spontaneous reporting of adverse events and medication errors related to human drugs and therapeutic biological products. faers pacakge serves as the interface between the FAERS database and R. Furthermore, faers pacakge offers a standardized approach for performing pharmacovigilance analysis.
Last updated
softwaredataimportbiomedicalinformaticspharmacogenomicsadverse-eventsdrug-safetyfaersfaers-procedurepharmacovigilancesignal-detection
6.20 score 48 stars 11 scripts 346 downloadsxCell2 - A Tool for Generic Cell Type Enrichment Analysis
xCell2 provides methods for cell type enrichment analysis using cell type signatures. It includes three main functions - 1. xCell2Train for training custom references objects from bulk or single-cell RNA-seq datasets. 2. xCell2Analysis for conducting the cell type enrichment analysis using the custom reference. 3. xCell2GetLineage for identifying dependencies between different cell types using ontology.
Last updated
geneexpressiontranscriptomicsmicroarrayrnaseqsinglecelldifferentialexpressionimmunooncologygenesetenrichment
6.19 score 23 stars 34 scripts 440 downloadsscECODA - Single-Cell Exploratory Compositional Data Analysis
The scECODA R package provides a complete workflow for the analysis and visualization of compositional data, primarily focusing on cell type proportions derived from single-cell data. It implements specialized methods, such as the Centered Log-Ratio (CLR) transformation, to properly analyze proportional data while avoiding the biases introduced by the compositional constraint. The package encapsulates data management, transformation, and analysis into a single SummarizedExperiment object, offering downstream tools for dimensionality reduction via PCA, calculating critical metrics like the Adjusted Rand Index (ARI) and Modularity to quantify sample grouping quality, and generating high-quality visualizations like heatmaps and scatter plots.
Last updated
softwaresinglecelltranscriptomicscellbasedassaysnormalizationpreprocessingvisualizationclusteringdimensionreductionfeatureextractionprincipalcomponent
6.18 score 10 stars 5 scripts 319 downloadsSpliceWiz - interactive analysis and visualization of alternative splicing in R
The analysis and visualization of alternative splicing (AS) events from RNA sequencing data remains challenging. SpliceWiz is a user-friendly and performance-optimized R package for AS analysis, by processing alignment BAM files to quantify read counts across splice junctions, IRFinder-based intron retention quantitation, and supports novel splicing event identification. We introduce a novel visualization for AS using normalized coverage, thereby allowing visualization of differential AS across conditions. SpliceWiz features a shiny-based GUI facilitating interactive data exploration of results including gene ontology enrichment. It is performance optimized with multi-threaded processing of BAM files and a new COV file format for fast recall of sequencing coverage. Overall, SpliceWiz streamlines AS analysis, enabling reliable identification of functionally relevant AS events for further characterization.
Last updated
softwaretranscriptomicsrnaseqalternativesplicingcoveragedifferentialsplicingdifferentialexpressionguisequencingcppopenmp
6.16 score 24 stars 15 scripts 408 downloadslimpa - Quantification and Differential Analysis of Proteomics Data
Quantification and differential analysis of mass-spectrometry proteomics data, with probabilistic recovery of information from missing values. Avoids the need for imputation. Estimates the detection probability curve (DPC), which relates the probability of successful detection to the underlying log-intensity of each precursor ion, and uses it to incorporate missing values into protein quantification and into subsequent differential expression analyses. The package produces objects suitable for downstream analysis in limma. The package accepts precursor (or peptide) intensities including missing values and produces complete protein quantifications without the need for imputation. The uncertainty of the protein quantifications is propagated through to the limma analyses using variance modeling and precision weights, ensuring accurate error rate control. The analysis pipeline can alternatively work with PTM or protein level data. The package name "limpa" is an acronym for "Linear Models for Proteomics Data".
Last updated
bayesianbiologicalquestiondataimportdifferentialexpressiongeneexpressionmassspectrometrypreprocessingproteomicsregressionsoftwaredifferential-expressionmass-spectrometry
6.15 score 22 stars 40 scripts 399 downloadsMSstatsLiP - LiP Significance Analysis in shotgun mass spectrometry-based proteomic experiments
Tools for LiP peptide and protein significance analysis. Provides functions for summarization, estimation of LiP peptide abundance, and detection of changes across conditions. Utilizes functionality across the MSstats family of packages.
Last updated
immunooncologymassspectrometryproteomicssoftwaredifferentialexpressiononechanneltwochannelnormalizationqualitycontrolcpp
6.15 score 7 stars 8 scripts 412 downloadsrawDiag - Brings Orbitrap Mass Spectrometry Data to Life; Fast and Colorful
Optimizing methods for liquid chromatography coupled to mass spectrometry (LC-MS) poses a nontrivial challenge. The rawDiag package facilitates rational method optimization by generating MS operator-tailored diagnostic plots of scan-level metadata. The package is designed for use on the R shell or as a Shiny application on the Orbitrap instrument PC.
Last updated
massspectrometryproteomicsmetabolomicsinfrastructuresoftwareshinyappsfastmass-spectrometrymultiplatformorbitrapvisualization
6.14 score 37 stars 25 scripts 273 downloadsnipalsMCIA - Multiple Co-Inertia Analysis via the NIPALS Method
Computes Multiple Co-Inertia Analysis (MCIA), a dimensionality reduction (jDR) algorithm, for a multi-block dataset using a modification to the Nonlinear Iterative Partial Least Squares method (NIPALS) proposed in (Hanafi et. al, 2010). Allows multiple options for row- and table-level preprocessing, and speeds up computation of variance explained. Vignettes detail application to bulk- and single cell- multi-omics studies.
Last updated
softwareclusteringclassificationmultiplecomparisonnormalizationpreprocessingsinglecell
6.14 score 7 stars 11 scripts 335 downloadsCBNplot - plot bayesian network inferred from gene expression data based on enrichment analysis results
This package provides the visualization of bayesian network inferred from gene expression data. The networks are based on enrichment analysis results inferred from packages including clusterProfiler and ReactomePA. The networks between pathways and genes inside the pathways can be inferred and visualized.
Last updated
visualizationbayesiangeneexpressionnetworkinferencepathwaysreactomenetworknetworkenrichmentgenesetenrichment
6.13 score 68 stars 10 scripts 370 downloadsTENxIO - Import methods for 10X Genomics files
Provides a structured S4 approach to importing data files from the 10X pipelines. It mainly supports Single Cell Multiome ATAC + Gene Expression data among other data types. The main Bioconductor data representations used are SingleCellExperiment and RaggedExperiment.
Last updated
softwareinfrastructuredataimportsinglecellbioconductor-packageu24ca289073
6.08 score 1 stars 5 dependents 16 scripts 626 downloadsCCPlotR - Plots For Visualising Cell-Cell Interactions
CCPlotR is an R package for visualising results from tools that predict cell-cell interactions from single-cell RNA-seq data. These plots are generic and can be used to visualise results from multiple tools such as Liana, CellPhoneDB, NATMI etc.
Last updated
singlecellnetworkvisualizationcellbiologysystemsbiology
6.08 score 47 stars 17 scripts 329 downloadscytoviewer - An interactive multi-channel image viewer for R
This R package supports interactive visualization of multi-channel images and segmentation masks generated by imaging mass cytometry and other highly multiplexed imaging techniques using shiny. The cytoviewer interface is divided into image-level (Composite and Channels) and cell-level visualization (Masks). It allows users to overlay individual images with segmentation masks, integrates well with SingleCellExperiment and SpatialExperiment objects for metadata visualization and supports image downloads.
Last updated
immunooncologysoftwaresinglecellonechanneltwochannelmultichannelspatialdataimportbioconductorimagingshinyvisualization
6.08 score 7 stars 57 scripts 368 downloadsMOGAMUN - MOGAMUN: A Multi-Objective Genetic Algorithm to Find Active Modules in Multiplex Biological Networks
MOGAMUN is a multi-objective genetic algorithm that identifies active modules in a multiplex biological network. This allows analyzing different biological networks at the same time. MOGAMUN is based on NSGA-II (Non-Dominated Sorting Genetic Algorithm, version II), which we adapted to work on networks.
Last updated
systemsbiologygraphandnetworkdifferentialexpressionbiomedicalinformaticstranscriptomicsclusteringnetwork
6.07 score 13 stars 10 scripts 298 downloadsiSEEtree - Interactive visualisation for microbiome data
iSEEtree is an extension of iSEE for the TreeSummarizedExperiment data container. It provides interactive panel designs to explore hierarchical datasets, such as the microbiome and cell lines.
Last updated
softwarevisualizationmicrobiomeguishinyappsdataimportshiny-appsvisualisation
6.03 score 3 stars 1 dependents 7 scripts 304 downloadsGeoDiff - Count model based differential expression and normalization on GeoMx RNA data
A series of statistical models using count generating distributions for background modelling, feature and sample QC, normalization and differential expression analysis on GeoMx RNA data. The application of these methods are demonstrated by example data analysis vignette.
Last updated
geneexpressiondifferentialexpressionnormalizationopenblascppopenmp
6.03 score 9 stars 24 scripts 384 downloadsBiocFHIR - Illustration of FHIR ingestion and transformation using R
FHIR R4 bundles in JSON format are derived from https://synthea.mitre.org/downloads. Transformation inspired by a kaggle notebook published by Dr Alexander Scarlat, https://www.kaggle.com/code/drscarlat/fhir-starter-parse-healthcare-bundles-into-tables. This is a very limited illustration of some basic parsing and reorganization processes. Additional tooling will be required to move beyond the Synthea data illustrations.
Last updated
infrastructuredataimportdatarepresentationfhir
6.00 score 4 stars 25 scripts 306 downloadsdemuxmix - Demultiplexing oligo-barcoded scRNA-seq data using regression mixture models
A package for demultiplexing single-cell sequencing experiments of pooled cells labeled with barcode oligonucleotides. The package implements methods to fit regression mixture models for a probabilistic classification of cells, including multiplet detection. Demultiplexing error rates can be estimated, and methods for quality control are provided.
Last updated
singlecellsequencingpreprocessingclassificationregression
6.00 score 5 stars 1 dependents 33 scripts 386 downloads
visiumStitched - Enable downstream analysis of Visium capture areas stitched together with Fiji
This package provides helper functions for working with multiple Visium capture areas that overlap each other. This package was developed along with the companion example use case data available from https://github.com/LieberInstitute/visiumStitched_brain. visiumStitched prepares SpaceRanger (10x Genomics) output files so you can stitch the images from groups of capture areas together with Fiji. Then visiumStitched builds a SpatialExperiment object with the stitched data and makes an artificial hexagonal grid enabling the seamless use of spatial clustering methods that rely on such grid to identify neighboring spots, such as PRECAST and BayesSpace. The SpatialExperiment objects created by visiumStitched are compatible with spatialLIBD, which can be used to build interactive websites for stitched SpatialExperiment objects. visiumStitched also enables casting SpatialExperiment objects as Seurat objects.
Last updated
softwarespatialtranscriptomicstranscriptiongeneexpressionvisualizationdataimport10xgenomicsbioconductorspatial-transcriptomicsspatialexperimentspatiallibdvisium
5.98 score 4 stars 7 scriptsgDRstyle - A package with style requirements for the gDR suite
Package fills a helper package role for whole gDR suite. It helps to support good development practices by keeping style requirements and style tests for other packages. It also contains build helpers to make all package requirements met.
Last updated
softwareinfrastructure
5.98 score 3 stars 2 scriptsgatom - Finding an Active Metabolic Module in Atom Transition Network
This package implements a metabolic network analysis pipeline to identify an active metabolic module based on high throughput data. The pipeline takes as input transcriptional and/or metabolic data and finds a metabolic subnetwork (module) most regulated between the two conditions of interest. The package further provides functions for module post-processing, annotation and visualization.
Last updated
geneexpressiondifferentialexpressionpathwaysnetwork
5.98 score 8 stars 16 scripts 412 downloadsbettr - A Better Way To Explore What Is Best
bettr provides a set of interactive visualization methods to explore the results of a benchmarking study, where typically more than a single performance measures are computed. The user can weight the performance measures according to their preferences. Performance measures can also be grouped and aggregated according to additional annotations.
Last updated
visualizationshinyappsgui
5.98 score 6 stars 20 scripts 241 downloadscrisprViz - Visualization Functions for CRISPR gRNAs
Provides functionalities to visualize and contextualize CRISPR guide RNAs (gRNAs) on genomic tracks across nucleases and applications. Works in conjunction with the crisprBase and crisprDesign Bioconductor packages. Plots are produced using the Gviz framework.
Last updated
crisprfunctionalgenomicsgenetargetbioconductorbioconductor-packagecrispr-analysiscrispr-designgrnagrna-sequencegrna-sequencessgrnasgrna-designvisualization
5.98 score 8 stars 2 dependents 10 scripts 408 downloads
scifer - Scifer: Single-Cell Immunoglobulin Filtering of Sanger Sequences
Have you ever index sorted cells in a 96 or 384-well plate and then sequenced using Sanger sequencing? If so, you probably had some struggles to either check the electropherogram of each cell sequenced manually, or when you tried to identify which cell was sorted where after sequencing the plate. Scifer was developed to solve this issue by performing basic quality control of Sanger sequences and merging flow cytometry data from probed single-cell sorted B cells with sequencing data. scifer can export summary tables, 'fasta' files, electropherograms for visual inspection, and generate reports.
Last updated
preprocessingqualitycontrolsangerseqsequencingsoftwareflowcytometrysinglecell
5.98 score 7 stars 30 scripts 362 downloadscrisprBowtie - Bowtie-based alignment of CRISPR gRNA spacer sequences
Provides a user-friendly interface to map on-targets and off-targets of CRISPR gRNA spacer sequences using bowtie. The alignment is fast, and can be performed using either commonly-used or custom CRISPR nucleases. The alignment can work with any reference or custom genomes. Both DNA- and RNA-targeting nucleases are supported.
Last updated
crisprfunctionalgenomicsalignmentalignerbioconductorbioconductor-packagebowtiecrispr-analysiscrispr-cas9crispr-designcrispr-targetgrnagrna-sequencegrna-sequencessgrnasgrna-design
5.97 score 3 stars 4 dependents 13 scripts 433 downloadsspeckle - Statistical methods for analysing single cell RNA-seq data
The speckle package contains functions for the analysis of single cell RNA-seq data. The speckle package currently contains functions to analyse differences in cell type proportions. There are also functions to estimate the parameters of the Beta distribution based on a given counts matrix, and a function to normalise a counts matrix to the median library size. There are plotting functions to visualise cell type proportions and the mean-variance relationship in cell type proportions and counts. As our research into specialised analyses of single cell data continues we anticipate that the package will be updated with new functions.
Last updated
singlecellrnaseqregressiongeneexpression
5.95 score 597 scripts 652 downloadsbenchdamic - Benchmark of differential abundance methods on microbiome data
Starting from a microbiome dataset (16S or WMS with absolute count values) it is possible to perform several analysis to assess the performances of many differential abundance detection methods. A basic and standardized version of the main differential abundance analysis methods is supplied but the user can also add his method to the benchmark. The analyses focus on 4 main aspects: i) the goodness of fit of each method's distributional assumptions on the observed count data, ii) the ability to control the false discovery rate, iii) the within and between method concordances, iv) the truthfulness of the findings if any apriori knowledge is given. Several graphical functions are available for result visualization.
Last updated
metagenomicsmicrobiomedifferentialexpressionmultiplecomparisonnormalizationpreprocessingsoftwarebenchmarkdifferential-abundance-methods
5.94 score 10 stars 11 scripts 298 downloadssimpleSeg - A package to perform simple cell segmentation
Image segmentation is the process of identifying the borders of individual objects (in this case cells) within an image. This allows for the features of cells such as marker expression and morphology to be extracted, stored and analysed. simpleSeg provides functionality for user friendly, watershed based segmentation on multiplexed cellular images in R based on the intensity of user specified protein marker channels. simpleSeg can also be used for the normalization of single cell data obtained from multiple images.
Last updated
classificationsurvivalsinglecellnormalizationspatialspatial-statistics
5.94 score 1 stars 2 dependents 29 scripts 502 downloadsmetabCombiner - Method for Combining LC-MS Metabolomics Feature Measurements
This package aligns LC-HRMS metabolomics datasets acquired from biologically similar specimens analyzed under similar, but not necessarily identical, conditions. Peak-picked and simply aligned metabolomics feature tables (consisting of m/z, rt, and per-sample abundance measurements, plus optional identifiers & adduct annotations) are accepted as input. The package outputs a combined table of feature pair alignments, organized into groups of similar m/z, and ranked by a similarity score. Input tables are assumed to be acquired using similar (but not necessarily identical) analytical methods.
Last updated
softwaremassspectrometrymetabolomicsmass-spectrometry
5.93 score 13 stars 22 scripts 413 downloadsUPDhmm - Detecting Uniparental Disomy through NGS trio data
Uniparental disomy (UPD) is a genetic condition where an individual inherits both copies of a chromosome or part of it from one parent, rather than one copy from each parent. This package contains a HMM for detecting UPDs through HTS (High Throughput Sequencing) data from trio assays. By analyzing the genotypes in the trio, the model infers a hidden state (normal, father isodisomy, mother isodisomy, father heterodisomy and mother heterodisomy).
Last updated
softwarehiddenmarkovmodelgenetics
5.92 score 4 stars 14 scripts 263 downloadsQTLExperiment - S4 classes for QTL summary statistics and metadata
QLTExperiment defines an S4 class for storing and manipulating summary statistics from QTL mapping experiments in one or more states. It is based on the 'SummarizedExperiment' class and contains functions for creating, merging, and subsetting objects. 'QTLExperiment' also stores experiment metadata and has checks in place to ensure that transformations apply correctly.
Last updated
functionalgenomicsdataimportdatarepresentationinfrastructuresequencingsnpsoftware
5.92 score 3 stars 1 dependents 31 scripts 270 downloadsSingleCellAlleleExperiment - S4 Class for Single Cell Data with Allele and Functional Levels for Immune Genes
Defines a S4 class that is based on SingleCellExperiment. In addition to the usual gene layer the object can also store data for immune genes such as HLAs, Igs and KIRs at allele and functional level. The package is part of a workflow named single-cell ImmunoGenomic Diversity (scIGD), that firstly incorporates allele-aware quantification data for immune genes. This new data can then be used with the here implemented data structure and functionalities for further data handling and data analysis.
Last updated
datarepresentationinfrastructuresinglecelltranscriptomicsgeneexpressiongeneticsimmunooncologydataimport
5.91 score 8 stars 17 scripts 334 downloads
CellBarcode - Cellular DNA Barcode Analysis toolkit
The package CellBarcode performs Cellular DNA Barcode analysis. It can handle all kinds of DNA barcodes, as long as the barcode is within a single sequencing read and has a pattern that can be matched by a regular expression. \code{CellBarcode} can handle barcodes with flexible lengths, with or without UMI (unique molecular identifier). This tool also can be used for pre-processing some amplicon data such as CRISPR gRNA screening, immune repertoire sequencing, and metagenome data.
Last updated
preprocessingqualitycontrolsequencingcrisprampliconamplicon-sequencingcellular-barcodecpp
5.91 score 3 stars 45 scripts 423 downloads
shiny.gosling - A Grammar-based Toolkit for Scalable and Interactive Genomics Data Visualization for R and Shiny
A Grammar-based Toolkit for Scalable and Interactive Genomics Data Visualization. http://gosling-lang.org/. This R package is based on gosling.js. It uses R functions to create gosling plots that could be embedded onto R Shiny apps.
Last updated
shinyappsgeneticsvisualization
5.88 score 1 dependents 51 scripts 226 downloadsSEraster - Rasterization Preprocessing Framework for Scalable Spatial Omics Data Analysis
SEraster is a rasterization preprocessing framework that aggregates cellular information into spatial pixels to reduce resource requirements for spatial omics data analysis. SEraster reduces the number of spatial points in spatial omics datasets for downstream analysis through a process of rasterization where single cells’ gene expression or cell-type labels are aggregated into equally sized pixels based on a user-defined resolution. SEraster is built on an R/Bioconductor S4 class called SpatialExperiment. SEraster can be incorporated with other packages to conduct downstream analyses for spatial omics datasets, such as detecting spatially variable genes.
Last updated
softwarespatialgeneexpressiontranscriptomicssinglecellpreprocessingspatial-analysisspatial-data-analysisspatial-omicsspatial-transcriptomics
5.88 score 19 stars 20 scripts 266 downloadsGeoTcgaData - Processing Various Types of Data on GEO and TCGA
Gene Expression Omnibus(GEO) and The Cancer Genome Atlas (TCGA) provide us with a wealth of data, such as RNA-seq, DNA Methylation, SNP and Copy number variation data. It's easy to download data from TCGA using the gdc tool, but processing these data into a format suitable for bioinformatics analysis requires more work. This R package was developed to handle these data.
Last updated
geneexpressiondifferentialexpressionrnaseqcopynumbervariationmicroarraysoftwarednamethylationdifferentialmethylationsnpatacseqmethylationarray
5.88 score 28 stars 27 scripts 334 downloadsphantasusLite - Loading and annotation RNA-seq counts matrices
PhantasusLite – a lightweight package with helper functions of general interest extracted from phantasus package. In parituclar it simplifies working with public RNA-seq datasets from GEO by providing access to the remote HSDS repository with the precomputed gene counts from ARCHS4 and DEE2 projects.
Last updated
geneexpressiontranscriptomicsrnaseq
5.86 score 11 stars 1 dependents 11 scripts 268 downloadsdandelionR - Single-cell Immune Repertoire Trajectory Analysis in R
dandelionR is an R package for performing single-cell immune repertoire trajectory analysis, based on the original python implementation. It provides the necessary functions to interface with scRepertoire and a custom implementation of an absorbing Markov chain for pseudotime inference, inspired by the Palantir Python package.
Last updated
softwareimmunooncologysinglecell
5.86 score 12 stars 10 scripts 259 downloadsGeDi - Defining and visualizing the distances between different genesets
The package provides different distances measurements to calculate the difference between genesets. Based on these scores the genesets are clustered and visualized as graph. This is all presented in an interactive Shiny application for easy usage.
Last updated
guigenesetenrichmentsoftwaretranscriptionrnaseqvisualizationclusteringpathwaysreportwritinggokeggreactomeshinyapps
5.85 score 2 stars 78 scripts 230 downloadsCRISPRball - Shiny Application for Interactive CRISPR Screen Visualization, Exploration, Comparison, and Filtering
A Shiny application for visualization, exploration, comparison, and filtering of CRISPR screens analyzed with MAGeCK RRA or MLE. Features include interactive plots with on-click labeling, full customization of plot aesthetics, data upload and/or download, and much more. Quickly and easily explore your CRISPR screen results and generate publication-quality figures in seconds.
Last updated
softwareshinyappscrisprqualitycontrolvisualizationguicrispr-screendata-visualizationinteractive-visualizationsmageckplotlyscreeningshiny
5.83 score 14 stars 24 scripts 266 downloadsscRNAseqApp - A single-cell RNAseq Shiny app-package
The scRNAseqApp is a Shiny app package designed for interactive visualization of single-cell data. It is an enhanced version derived from the ShinyCell, repackaged to accommodate multiple datasets. The app enables users to visualize data containing various types of information simultaneously, facilitating comprehensive analysis. Additionally, it includes a user management system to regulate database accessibility for different users.
Last updated
visualizationsinglecellrnaseqinteractive-visualizationsmultiple-usersshiny-appssingle-cell-rna-seq
5.80 score 6 stars 9 scripts 375 downloadsrhinotypeR - Rhinovirus genotyping
"rhinotypeR" is designed to automate the comparison of sequence data against prototype strains, streamlining the genotype assignment process. By implementing predefined pairwise distance thresholds, this package makes genotype assignment accessible to researchers and public health professionals. This tool enhances our epidemiological toolkit by enabling more efficient surveillance and analysis of rhinoviruses (RVs) and other viral pathogens with complex genomic landscapes. Additionally, "rhinotypeR" supports comprehensive visualization and analysis of single nucleotide polymorphisms (SNPs) and amino acid substitutions, facilitating in-depth genetic and evolutionary studies.
Last updated
sequencinggeneticsphylogeneticsvisualizationmultiplesequencealignmentmultiplecomparison
5.78 score 4 stars 4 scripts 228 downloadsTDbasedUFE - Tensor Decomposition Based Unsupervised Feature Extraction
This is a comprehensive package to perform Tensor decomposition based unsupervised feature extraction. It can perform unsupervised feature extraction. It uses tensor decomposition. It is applicable to gene expression, DNA methylation, and histone modification etc. It can perform multiomics analysis. It is also potentially applicable to single cell omics data sets.
Last updated
geneexpressionfeatureextractionmethylationarraysinglecellbioinformaticsdna-methylationgene-expression-profileshistone-modificationsmultiomicstensor-decomposition
5.78 score 5 stars 1 dependents 20 scripts 302 downloadsOmicsMLRepoR - Search harmonized metadata created under the OmicsMLRepo project
This package provides functions to browse the harmonized metadata for large omics databases. This package also supports data navigation if the metadata incorporates ontology.
Last updated
softwareinfrastructuredatarepresentationu24ca289073
5.77 score 2 stars 22 scripts 209 downloads
beer - Bayesian Enrichment Estimation in R
BEER implements a Bayesian model for analyzing phage-immunoprecipitation sequencing (PhIP-seq) data. Given a PhIPData object, BEER returns posterior probabilities of enriched antibody responses, point estimates for the relative fold-change in comparison to negative control samples, and more. Additionally, BEER provides a convenient implementation for using edgeR to identify enriched antibody responses.
Last updated
softwarestatisticalmethodbayesiansequencingcoveragejagscpp
5.77 score 11 stars 18 scripts 365 downloads
MsBackendSql - SQL-based Mass Spectrometry Data Backend
SQL-based mass spectrometry (MS) data backend supporting also storange and handling of very large data sets. Objects from this package are supposed to be used with the Spectra Bioconductor package. Through the MsBackendSql with its minimal memory footprint, this package thus provides an alternative MS data representation for very large or remote MS data sets.
Last updated
infrastructuremassspectrometrymetabolomicsdataimportproteomics
5.76 score 4 stars 36 scripts 386 downloadscfTools - Informatics Tools for Cell-Free DNA Study
The cfTools R package provides methods for cell-free DNA (cfDNA) methylation data analysis to facilitate cfDNA-based studies. Given the methylation sequencing data of a cfDNA sample, for each cancer marker or tissue marker, we deconvolve the tumor-derived or tissue-specific reads from all reads falling in the marker region. Our read-based deconvolution algorithm exploits the pervasiveness of DNA methylation for signal enhancement, therefore can sensitively identify a trace amount of tumor-specific or tissue-specific cfDNA in plasma. cfTools provides functions for (1) cancer detection: sensitively detect tumor-derived cfDNA and estimate the tumor-derived cfDNA fraction (tumor burden); (2) tissue deconvolution: infer the tissue type composition and the cfDNA fraction of multiple tissue types for a plasma cfDNA sample. These functions can serve as foundations for more advanced cfDNA-based studies, including cancer diagnosis and disease monitoring.
Last updated
softwarebiomedicalinformaticsepigeneticssequencingmethylseqdnamethylationdifferentialmethylationcpp
5.69 score 11 stars 2 scripts 252 downloadshoodscanR - Spatial cellular neighbourhood scanning in R
hoodscanR is an user-friendly R package providing functions to assist cellular neighborhood analysis of any spatial transcriptomics data with single-cell resolution. All functions in the package are built based on the SpatialExperiment object, allowing integration into various spatial transcriptomics-related packages from Bioconductor. The package can result in cell-level neighborhood annotation output, along with funtions to perform neighborhood colocalization analysis and neighborhood-based cell clustering.
Last updated
spatialtranscriptomicssinglecellclusteringcpp
5.69 score 14 stars 35 scriptsMetMashR - Metabolite Mashing with R
A package to merge, filter sort, organise and otherwise mash together metabolite annotation tables. Metabolite annotations can be imported from multiple sources (software) and combined using workflow steps based on S4 class templates derived from the `struct` package. Other modular workflow steps such as filtering, merging, splitting, normalisation and rest-api queries are included.
Last updated
workflowstepmetabolomicskegg
5.68 score 3 stars 10 scripts 298 downloads
plyinteractions - Extending tidy verbs to genomic interactions
Operate on `GInteractions` objects as tabular data using `dplyr`-like verbs. The functions and methods in `plyinteractions` provide a grammatical approach to manipulate `GInteractions`, to facilitate their integration in genomic analysis workflows.
Last updated
softwareinfrastructure
5.68 score 1 dependents 20 scripts 300 downloadsmastR - Markers Automated Screening Tool in R
mastR is an R package designed for automated screening of signatures of interest for specific research questions. The package is developed for generating refined lists of signature genes from multiple group comparisons based on the results from edgeR and limma differential expression (DE) analysis workflow. It also takes into account the background noise of tissue-specificity, which is often ignored by other marker generation tools. This package is particularly useful for the identification of group markers in various biological and medical applications, including cancer research and developmental biology.
Last updated
softwaregeneexpressiontranscriptomicsdifferentialexpressionvisualization
5.68 score 6 stars 7 scripts 379 downloadsscBubbletree - Quantitative visual exploration of scRNA-seq data
scBubbletree is a quantitative method for the visual exploration of scRNA-seq data, preserving key biological properties such as local and global cell distances and cell density distributions across samples. It effectively resolves overplotting and enables the visualization of diverse cell attributes from multiomic single-cell experiments. Additionally, scBubbletree is user-friendly and integrates seamlessly with popular scRNA-seq analysis tools, facilitating comprehensive and intuitive data interpretation.
Last updated
visualizationclusteringsinglecelltranscriptomicsrnaseqbig-databigdatascrna-seqscrna-seq-analysisvisualvisual-exploration
5.68 score 8 stars 30 scripts 378 downloadsompBAM - C++ Library for OpenMP-based multi-threaded sequential profiling of Binary Alignment Map (BAM) files
This packages provides C++ header files for developers wishing to create R packages that processes BAM files. ompBAM automates file access, memory management, and handling of multiple threads 'behind the scenes', so developers can focus on creating domain-specific functionality. The included vignette contains detailed documentation of this API, including quick-start instructions to create a new ompBAM-based package, and step-by-step explanation of the functionality behind the example packaged included within ompBAM.
Last updated
alignmentdataimportrnaseqsoftwaresequencingtranscriptomicssinglecell
5.68 score 4 stars 2 dependents 5 scripts 342 downloadsknowYourCG - Functional analysis of DNA methylome datasets
KnowYourCG (KYCG) is a supervised learning framework designed for the functional analysis of DNA methylation data. Unlike existing tools that focus on genes or genomic intervals, KnowYourCG directly targets CpG dinucleotides, featuring automated supervised screenings of diverse biological and technical influences, including sequence motifs, transcription factor binding, histone modifications, replication timing, cell-type-specific methylation, and trait-epigenome associations. KnowYourCG addresses the challenges of data sparsity in various methylation datasets, including low-pass Nanopore sequencing, single-cell DNA methylomes, 5-hydroxymethylation profiles, spatial DNA methylation maps, and array-based datasets for epigenome-wide association studies and epigenetic clocks (<doi:10.1126/sciadv.adw3027>).
Last updated
epigeneticsdnamethylationsequencingsinglecellspatialtranscriptionmethylationarrayzlib
5.67 score 7 stars 28 scripts 330 downloads
raer - RNA editing tools in R
Toolkit for identification and statistical testing of RNA editing signals from within R. Provides support for identifying sites from bulk-RNA and single cell RNA-seq datasets, and general methods for extraction of allelic read counts from alignment files. Facilitates annotation and exploratory analysis of editing signals using Bioconductor packages and resources.
Last updated
multiplecomparisonrnaseqsinglecellsequencingcoverageepitranscriptomicsfeatureextractionannotationalignmentbioconductor-packagerna-seq-analysissingle-cell-analysissingle-cell-rna-seqcurlbzip2xz-utilszlib
5.64 score 10 stars 11 scripts 368 downloadseasylift - An R package to perform genomic liftover
The easylift package provides a convenient tool for genomic liftover operations between different genome assemblies. It seamlessly works with Bioconductor's GRanges objects and chain files from the UCSC Genome Browser, allowing for straightforward handling of genomic ranges across various genome versions. One noteworthy feature of easylift is its integration with the BiocFileCache package. This integration automates the management and caching of chain files necessary for liftover operations. Users no longer need to manually specify chain file paths in their function calls, reducing the complexity of the liftover process.
Last updated
softwareworkflowstepsequencingcoveragegenomeassemblydataimport
5.64 score 8 stars 27 scripts 292 downloadsrgoslin - Lipid Shorthand Name Parsing and Normalization
The R implementation for the Grammar of Succint Lipid Nomenclature parses different short hand notation dialects for lipid names. It normalizes them to a standard name. It further provides calculated monoisotopic masses and sum formulas for each successfully parsed lipid name and supplements it with LIPID MAPS Category and Class information. Also, the structural level and further structural details about the head group, fatty acyls and functional groups are returned, where applicable.
Last updated
softwarelipidomicsmetabolomicspreprocessingnormalizationmassspectrometrycpp
5.64 score 6 stars 36 scripts 446 downloadsAnVILWorkflow - Run workflows implemented in Terra/AnVIL workspace
The AnVIL is a cloud computing resource developed in part by the National Human Genome Research Institute. The main cloud-based genomics platform deported by the AnVIL project is Terra. The AnVILWorkflow package allows remote access to Terra implemented workflows, enabling end-user to utilize Terra/ AnVIL provided resources - such as data, workflows, and flexible/scalble computing resources - through the conventional R functions.
Last updated
infrastructuresoftwareanvilgcpterrau24hg010263workflows
5.62 score 7 stars 7 scripts 223 downloadsStabMap - Stabilised mosaic single cell data integration using unshared features
StabMap performs single cell mosaic data integration by first building a mosaic data topology, and for each reference dataset, traverses the topology to project and predict data onto a common embedding. Mosaic data should be provided in a list format, with all relevant features included in the data matrices within each list object. The output of stabMap is a joint low-dimensional embedding taking into account all available relevant features. Expression imputation can also be performed using the StabMap embedding and any of the original data matrices for given reference and query cell lists.
Last updated
singlecelldimensionreductionsoftware
5.62 score 92 scripts 288 downloadsmethyLImp2 - Missing value estimation of DNA methylation data
This package allows to estimate missing values in DNA methylation data. methyLImp method is based on linear regression since methylation levels show a high degree of inter-sample correlation. Implementation is parallelised over chromosomes since probes on different chromosomes are usually independent. Mini-batch approach to reduce the runtime in case of large number of samples is available.
Last updated
dnamethylationmicroarraysoftwaremethylationarrayregressionimputationmethylationmissing-value-imputation
5.61 score 9 stars 15 scripts 265 downloadsMSPrep - Package for Summarizing, Filtering, Imputing, and Normalizing Metabolomics Data
Package performs summarization of replicates, filtering by frequency, several different options for imputing missing data, and a variety of options for transforming, batch correcting, and normalizing data.
Last updated
metabolomicsmassspectrometrypreprocessing
5.60 score 10 stars 6 scriptsAnVILGCP - The GCP R Client for the AnVIL
The package provides a set of functions to interact with the Google Cloud Platform (GCP) services on the AnVIL platform. The package is designed to use the API calls from the AnVIL package. It coordinates AnVIL workspace functionality with native GCP tools.
Last updated
softwareinfrastructurethirdpartyclientdataimportu24hg010263
5.60 score 5 dependents 38 scripts 306 downloadsGCPtools - Tools for working with gcloud and gsutil
Lower-level functionality to interface with Google Cloud Platform tools. 'gcloud' and 'gsutil' are both supported. The functionality provided centers around utilities for the AnVIL platform.
Last updated
softwareinfrastructurethirdpartyclientdataimportbioconductor-packageu24hg010263
5.60 score 15 dependents 22 scripts 428 downloads
VDJdive - Analysis Tools for 10X V(D)J Data
This package provides functions for handling and analyzing immune receptor repertoire data, such as produced by the CellRanger V(D)J pipeline. This includes reading the data into R, merging it with paired single-cell data, quantifying clonotype abundances, calculating diversity metrics, and producing common plots. It implements the E-M Algorithm for clonotype assignment, along with other methods, which makes use of ambiguous cells for improved quantification.
Last updated
softwareimmunooncologysinglecellannotationrnaseqtargetedresequencingcpp
5.59 score 13 stars 6 scripts 288 downloadsgDR - Umbrella package for R packages in the gDR suite
Package is a part of the gDR suite. It reexports functions from other packages in the gDR suite that contain critical processing functions and utilities. The vignette walks through the full processing pipeline for drug response analyses that the gDR suite offers.
Last updated
softwaredataimportshinyapps
5.58 score 2 stars 21 scripts 277 downloadsNewWave - Negative binomial model for scRNA-seq
A model designed for dimensionality reduction and batch effect removal for scRNA-seq data. It is designed to be massively parallelizable using shared objects that prevent memory duplication, and it can be used with different mini-batch approaches in order to reduce time consumption. It assumes a negative binomial distribution for the data with a dispersion parameter that can be both commonwise across gene both genewise.
Last updated
softwaregeneexpressiontranscriptomicssinglecellbatcheffectsequencingcoverageregressionbatch-effectsdimensionality-reductionnegative-binomialscrna-seq
5.57 score 5 stars 37 scripts 334 downloadspreciseTAD - preciseTAD: A machine learning framework for precise TAD boundary prediction
preciseTAD provides functions to predict the location of boundaries of topologically associated domains (TADs) and chromatin loops at base-level resolution. As an input, it takes BED-formatted genomic coordinates of domain boundaries detected from low-resolution Hi-C data, and coordinates of high-resolution genomic annotations from ENCODE or other consortia. preciseTAD employs several feature engineering strategies and resampling techniques to address class imbalance, and trains an optimized random forest model for predicting low-resolution domain boundaries. Translated on a base-level, preciseTAD predicts the probability for each base to be a boundary. Density-based clustering and scalable partitioning techniques are used to detect precise boundary regions and summit points. Compared with low-resolution boundaries, preciseTAD boundaries are highly enriched for CTCF, RAD21, SMC3, and ZNF143 signal and more conserved across cell lines. The pre-trained model can accurately predict boundaries in another cell line using CTCF, RAD21, SMC3, and ZNF143 annotation data for this cell line.
Last updated
softwarehicsequencingclusteringclassificationfunctionalgenomicsfeatureextraction
5.57 score 8 stars 23 scripts 365 downloadsCaMutQC - An R Package for Comprehensive Filtration and Selection of Cancer Somatic Mutations
CaMutQC is able to filter false positive mutations generated due to technical issues, as well as to select candidate cancer mutations through a series of well-structured functions by labeling mutations with various flags. And a detailed and vivid filter report will be offered after completing a whole filtration or selection section. Also, CaMutQC integrates serveral methods and gene panels for Tumor Mutational Burden (TMB) estimation.
Last updated
softwarequalitycontrolgenetargetcancer-genomicssomatic-mutations
5.56 score 8 stars 7 scripts 234 downloadsINTACT - Integrate TWAS and Colocalization Analysis for Gene Set Enrichment Analysis
This package integrates colocalization probabilities from colocalization analysis with transcriptome-wide association study (TWAS) scan summary statistics to implicate genes that may be biologically relevant to a complex trait. The probabilistic framework implemented in this package constrains the TWAS scan z-score-based likelihood using a gene-level colocalization probability. Given gene set annotations, this package can estimate gene set enrichment using posterior probabilities from the TWAS-colocalization integration step.
Last updated
bayesiangenesetenrichment
5.55 score 17 stars 21 scripts 272 downloadsGSgalgoR - An Evolutionary Framework for the Identification and Study of Prognostic Gene Expression Signatures in Cancer
A multi-objective optimization algorithm for disease sub-type discovery based on a non-dominated sorting genetic algorithm. The 'Galgo' framework combines the advantages of clustering algorithms for grouping heterogeneous 'omics' data and the searching properties of genetic algorithms for feature selection. The algorithm search for the optimal number of clusters determination considering the features that maximize the survival difference between sub-types while keeping cluster consistency high.
Last updated
geneexpressiontranscriptionclusteringclassificationsurvival
5.52 score 15 stars 11 scripts 342 downloadsBatChef - Single-cell RNA-seq batch effects correction methods interface
This package implements a variety of methods for batch correction in single-cell RNA sequencing (scRNA-seq) data. It incorporates quantitative metrics (e.g. Wasserstein distance, Adjusted Rand Index) to evaluate their performance. Furthermore, the package assists users in identifying and applying the optimal method for specific datasets.
Last updated
batcheffectsinglecellsequencingcpp
5.51 score 4 stars 2 scriptsgypsum - Interface to the gypsum REST API
Client for the gypsum REST API (https://gypsum.artifactdb.com), a cloud-based file store in the ArtifactDB ecosystem. This package provides functions for uploads, downloads, and various adminstrative and management tasks. Check out the documentation at https://github.com/ArtifactDB/gypsum-worker for more details.
Last updated
dataimport
5.50 score 1 stars 1 dependents 20 scripts 5.3k downloadsMotifPeeker - Benchmarking Epigenomic Profiling Methods Using Motif Enrichment
MotifPeeker is used to compare and analyse datasets from epigenomic profiling methods with motif enrichment as the key benchmark. The package outputs an HTML report consisting of three sections: (1. General Metrics) Overview of peaks-related general metrics for the datasets (FRiP scores, peak widths and motif-summit distances). (2. Known Motif Enrichment Analysis) Statistics for the frequency of user-provided motifs enriched in the datasets. (3. Motif Discovery Enrichment Analysis) Statistics for the frequency of ab-initio discovered motifs enriched in the datasets and compared with known motifs.
Last updated
epigeneticsgeneticsqualitycontrolchipseqmultiplecomparisonfunctionalgenomicsmotifdiscoverysequencematchingsoftwarealignmentbioconductorbioconductor-packagechip-seqepigenomicsinteractive-reportmotif-enrichment-analysis
5.48 score 3 stars 7 scripts 197 downloadssketchR - An R interface for python subsampling/sketching algorithms
Provides an R interface for various subsampling algorithms implemented in python packages. Currently, interfaces to the geosketch and scSampler python packages are implemented. In addition it also provides diagnostic plots to evaluate the subsampling.
Last updated
singlecell
5.48 score 3 stars 9 scripts 308 downloadsenrichViewNet - From functional enrichment results to biological networks
This package enables the visualization of functional enrichment results as network graphs. First the package enables the visualization of enrichment results, in a format corresponding to the one generated by gprofiler2, as a customizable Cytoscape network. In those networks, both gene datasets (GO terms/pathways/protein complexes) and genes associated to the datasets are represented as nodes. While the edges connect each gene to its dataset(s). The package also provides the option to create enrichment maps from functional enrichment results. Enrichment maps enable the visualization of enriched terms into a network with edges connecting overlapping genes.
Last updated
biologicalquestionsoftwarenetworknetworkenrichmentgocystocapefunctional-enrichment
5.48 score 6 stars 6 scripts
TREG - Tools for finding Total RNA Expression Genes in single nucleus RNA-seq data
RNA abundance and cell size parameters could improve RNA-seq deconvolution algorithms to more accurately estimate cell type proportions given the different cell type transcription activity levels. A Total RNA Expression Gene (TREG) can facilitate estimating total RNA content using single molecule fluorescent in situ hybridization (smFISH). We developed a data-driven approach using a measure of expression invariance to find candidate TREGs in postmortem human brain single nucleus RNA-seq. This R package implements the method for identifying candidate TREGs from snRNA-seq data.
Last updated
softwaresinglecellrnaseqgeneexpressiontranscriptomicstranscriptionsequencingbioconductordeconvolutionrnascopescrna-seqsmfishsnrna-seqtreg
5.48 score 5 stars 5 scripts 334 downloadstxcutr - Transcriptome CUTteR
Various mRNA sequencing library preparation methods generate sequencing reads specifically from the transcript ends. Analyses that focus on quantification of isoform usage from such data can be aided by using truncated versions of transcriptome annotations, both at the alignment or pseudo-alignment stage, as well as in downstream analysis. This package implements some convenience methods for readily generating such truncated annotations and their corresponding sequences.
Last updated
alignmentannotationrnaseqsequencingtranscriptomics
5.48 score 5 stars 10 scripts 380 downloadsReUseData - Reusable and reproducible Data Management
ReUseData is an _R/Bioconductor_ software tool to provide a systematic and versatile approach for standardized and reproducible data management. ReUseData facilitates transformation of shell or other ad hoc scripts for data preprocessing into workflow-based data recipes. Evaluation of data recipes generate curated data files in their generic formats (e.g., VCF, bed). Both recipes and data are cached using database infrastructure for easy data management and reuse. Prebuilt data recipes are available through ReUseData portal ("https://rcwl.org/dataRecipes/") with full annotation and user instructions. Pregenerated data are available through ReUseData cloud bucket that is directly downloadable through "getCloudData()".
Last updated
softwareinfrastructuredataimportpreprocessingimmunooncology
5.46 score 4 stars 12 scripts 255 downloadsHuBMAPR - Interface to 'HuBMAP'
'HuBMAP' provides an open, global bio-molecular atlas of the human body at the cellular level. The `datasets()`, `samples()`, `donors()`, `publications()`, and `collections()` functions retrieves the information for each of these entity types. `*_details()` are available for individual entries of each entity type. `*_derived()` are available for retrieving derived datasets or samples for individual entries of each entity type. Data files can be accessed using `bulk_data_transfer()`.
Last updated
softwaresinglecelldataimportthirdpartyclientspatialinfrastructurebioconductor-packageclienthubmaprstudio
5.43 score 3 stars 1 scripts 286 downloadslimpca - An R package for the linear modeling of high-dimensional designed data based on ASCA/APCA family of methods
This package has for objectives to provide a method to make Linear Models for high-dimensional designed data. limpca applies a GLM (General Linear Model) version of ASCA and APCA to analyse multivariate sample profiles generated by an experimental design. ASCA/APCA provide powerful visualization tools for multivariate structures in the space of each effect of the statistical model linked to the experimental design and contrarily to MANOVA, it can deal with mutlivariate datasets having more variables than observations. This method can handle unbalanced design.
Last updated
statisticalmethodprincipalcomponentregressionvisualizationexperimentaldesignmultiplecomparisongeneexpressionmetabolomics
5.43 score 2 stars 7 scripts 246 downloadsiSEEindex - iSEE extension for a landing page to a custom collection of data sets
This package provides an interface to any collection of data sets within a single iSEE web-application. The main functionality of this package is to define a custom landing page allowing app maintainers to list a custom collection of data sets that users can selected from and directly load objects into an iSEE web-application.
Last updated
softwareinfrastructurebioconductorhacktoberfest
5.43 score 2 stars 15 scripts 306 downloadsgcatest - Genotype Conditional Association TEST
GCAT is an association test for genome wide association studies that controls for population structure under a general class of trait models. This test conditions on the trait, which makes it immune to confounding by unmodeled environmental factors. Population structure is modeled via logistic factors, which are estimated using the `lfa` package.
Last updated
snpdimensionreductionprincipalcomponentgenomewideassociation
5.43 score 6 stars 7 scripts 400 downloadsTEKRABber - An R package estimates the correlations of orthologs and transposable elements between two species
TEKRABber is made to provide a user-friendly pipeline for comparing orthologs and transposable elements (TEs) between two species. It considers the orthology confidence between two species from BioMart to normalize expression counts and detect differentially expressed orthologs/TEs. Then it provides one to one correlation analysis for desired orthologs and TEs. There is also an app function to have a first insight on the result. Users can prepare orthologs/TEs RNA-seq expression data by their own preference to run TEKRABber following the data structure mentioned in the vignettes.
Last updated
differentialexpressionnormalizationtranscriptiongeneexpressionbioconductorcpp
5.43 score 3 stars 20 scripts 408 downloadsGenomicPlot - Plot profiles of next generation sequencing data in genomic features
Visualization of next generation sequencing (NGS) data is essential for interpreting high-throughput genomics experiment results. 'GenomicPlot' facilitates plotting of NGS data in various formats (bam, bed, wig and bigwig); both coverage and enrichment over input can be computed and displayed with respect to genomic features (such as UTR, CDS, enhancer), and user defined genomic loci or regions. Statistical tests on signal intensity within user defined regions of interest can be performed and represented as boxplots or bar graphs. Parallel processing is used to speed up computation on multicore platforms. In addition to genomic plots which is suitable for displaying of coverage of genomic DNA (such as ChIPseq data), metagenomic (without introns) plots can also be made for RNAseq or CLIPseq data as well.
Last updated
alternativesplicingchipseqcoveragegeneexpressionrnaseqsequencingsoftwaretranscriptionvisualizationannotation
5.42 score 6 stars 11 scripts 346 downloadscytoMEM - Marker Enrichment Modeling (MEM)
MEM, Marker Enrichment Modeling, automatically generates and displays quantitative labels for cell populations that have been identified from single-cell data. The input for MEM is a dataset that has pre-clustered or pre-gated populations with cells in rows and features in columns. Labels convey a list of measured features and the features' levels of relative enrichment on each population. MEM can be applied to a wide variety of data types and can compare between MEM labels from flow cytometry, mass cytometry, single cell RNA-seq, and spectral flow cytometry using RMSD.
Last updated
proteomicssystemsbiologyclassificationflowcytometrydatarepresentationdataimportcellbiologysinglecellclustering
5.42 score 4 stars 1 dependents 22 scripts 420 downloadsalabaster.sce - Load and Save SingleCellExperiment from File
Save SingleCellExperiment into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
5.42 score 4 dependents 11 scripts 2.0k downloadssmoppix - Analyze Single Molecule Spatial Omics Data Using the Probabilistic Index
Test for univariate and bivariate spatial patterns in spatial omics data with single-molecule resolution. The tests implemented allow for analysis of nested designs and are automatically calibrated to different biological specimens. Tests for aggregation, colocalization, gradients and vicinity to cell edge or centroid are provided.
Last updated
transcriptomicsspatialsinglecellcpp
5.41 score 1 stars 1 dependents 5 scripts 322 downloadsalabaster.sfe - Language agnostic on disk serialization of SpatialFeatureExperiment
Builds upon the existing ArtifactDB project, expending alabaster.spatial for language agnostic on disk serialization of SpatialFeatureExperiment.
Last updated
datarepresentationspatialopenjdk
5.40 score 252 scripts 186 downloadsCDI - Clustering Deviation Index (CDI)
Single-cell RNA-sequencing (scRNA-seq) is widely used to explore cellular variation. The analysis of scRNA-seq data often starts from clustering cells into subpopulations. This initial step has a high impact on downstream analyses, and hence it is important to be accurate. However, there have not been unsupervised metric designed for scRNA-seq to evaluate clustering performance. Hence, we propose clustering deviation index (CDI), an unsupervised metric based on the modeling of scRNA-seq UMI counts to evaluate clustering of cells.
Last updated
singlecellsoftwareclusteringvisualizationsequencingrnaseqcellbasedassays
5.40 score 5 stars 8 scripts 300 downloadsTMSig - Tools for Molecular Signatures
The TMSig package contains tools to prepare, analyze, and visualize named lists of sets, with an emphasis on molecular signatures (such as gene or kinase sets). It includes fast, memory efficient functions to construct sparse incidence and similarity matrices and filter, cluster, invert, and decompose sets. Additionally, bubble heatmaps can be created to visualize the results of any differential or molecular signatures analysis.
Last updated
clusteringgenesetenrichmentgraphandnetworkpathwaysvisualizationgene-setsmolecular-signatures
5.39 score 5 stars 49 scripts 264 downloadsPhosR - A set of methods and tools for comprehensive analysis of phosphoproteomics data
PhosR is a package for the comprenhensive analysis of phosphoproteomic data. There are two major components to PhosR: processing and downstream analysis. PhosR consists of various processing tools for phosphoproteomics data including filtering, imputation, normalisation, and functional analysis for inferring active kinases and signalling pathways.
Last updated
softwareresearchfieldproteomics
5.39 score 122 scripts 494 downloads
demuxSNP - scRNAseq demultiplexing using cell hashing and SNPs
This package assists in demultiplexing scRNAseq data using both cell hashing and SNPs data. The SNP profile of each group os learned using high confidence assignments from the cell hashing data. Cells which cannot be assigned with high confidence from the cell hashing data are assigned to their most similar group based on their SNPs. We also provide some helper function to optimise SNP selection, create training data and merge SNP data into the SingleCellExperiment framework.
Last updated
classificationsinglecell
5.39 score 9 stars 27 scripts 277 downloadsjazzPanda - Finding spatially relevant marker genes in image based spatial transcriptomics data
This package contains the function to find marker genes for image-based spatial transcriptomics data. There are functions to create spatial vectors from the cell and transcript coordiantes, which are passed as inputs to find marker genes. Marker genes are detected for every cluster by two approaches. The first approach is by permtuation testing, which is implmented in parallel for finding marker genes for one sample study. The other approach is to build a linear model for every gene. This approach can account for multiple samples and backgound noise.
Last updated
spatialgeneexpressiondifferentialexpressionstatisticalmethodtranscriptomicscorrelationlinear-modelsmarker-genesspatial-transcriptomics
5.38 score 4 stars 3 scripts 274 downloadsDuplexDiscovereR - Analysis of the data from RNA duplex probing experiments
DuplexDiscovereR is a package designed for analyzing data from RNA cross-linking and proximity ligation protocols such as SPLASH, PARIS, LIGR-seq, and others. DuplexDiscovereR accepts input in the form of chimerically or split-aligned reads. It includes procedures for alignment classification, filtering, and efficient clustering of individual chimeric reads into duplex groups (DGs). Once DGs are identified, the package predicts RNA duplex formation and their hybridization energies. Additional metrics, such as p-values for random ligation hypothesis or mean DG alignment scores, can be calculated to rank final set of RNA duplexes. Data from multiple experiments or replicates can be processed separately and further compared to check the reproducibility of the experimental method.
Last updated
sequencingtranscriptomicsstructuralpredictionclusteringsplicedalignment
5.38 score 3 stars 10 scripts 219 downloadsmosdef - MOSt frequently used and useful Differential Expression Functions
This package provides functionality to run a number of tasks in the differential expression analysis workflow. This encompasses the most widely used steps, from running various enrichment analysis tools with a unified interface to creating plots and beautifying table components linking to external websites and databases. This streamlines the generation of comprehensive analysis reports.
Last updated
geneexpressionsoftwaretranscriptiontranscriptomicsdifferentialexpressionvisualizationreportwritinggenesetenrichmentgo
5.38 score 4 dependents 1 scripts 518 downloadsCardinalIO - Read and write mass spectrometry imaging files
Fast and efficient reading and writing of mass spectrometry imaging data files. Supports imzML and Analyze 7.5 formats. Provides ontologies for mass spectrometry imaging.
Last updated
softwareinfrastructuredataimportmassspectrometryimagingmassspectrometrycpp
5.38 score 4 stars 2 dependents 4 scripts 470 downloadsretrofit - RETROFIT: Reference-free deconvolution of cell mixtures in spatial transcriptomics
RETROFIT is a Bayesian non-negative matrix factorization framework to decompose cell type mixtures in ST data without using external single-cell expression references. RETROFIT outperforms existing reference-based methods in estimating cell type proportions and reconstructing gene expressions in simulations with varying spot size and sample heterogeneity, irrespective of the quality or availability of the single-cell reference. RETROFIT recapitulates known cell-type localization patterns in a Slide-seq dataset of mouse cerebellum without using any single-cell data.
Last updated
transcriptomicsvisualizationrnaseqbayesianspatialsoftwaregeneexpressiondimensionreductionfeatureextractionsinglecellcpp
5.38 score 3 stars 20 scripts 294 downloadsSiPSiC - Calculate Pathway Scores for Each Cell in scRNA-Seq Data
Infer biological pathway activity of cells from single-cell RNA-sequencing data by calculating a pathway score for each cell (pathway genes are specified by the user). It is recommended to have the data in Transcripts-Per-Million (TPM) or Counts-Per-Million (CPM) units for best results. Scores may change when adding cells to or removing cells off the data. SiPSiC stands for Single Pathway analysis in Single Cells.
Last updated
softwaredifferentialexpressiongenesetenrichmentbiomedicalinformaticscellbiologytranscriptomicsrnaseqsinglecelltranscriptionsequencingimmunooncologydataimport
5.38 score 8 stars 1 dependents 9 scripts 267 downloadsRnaSeqSampleSize - RnaSeqSampleSize
RnaSeqSampleSize package provides a sample size calculation method based on negative binomial model and the exact test for assessing differential expression analysis of RNA-seq data. It controls FDR for multiple testing and utilizes the average read count and dispersion distributions from real data to estimate a more reliable sample size. It is also equipped with several unique features, including estimation for interested genes or pathway, power curve visualization, and parameter optimization.
Last updated
immunooncologyexperimentaldesignsequencingrnaseqgeneexpressiondifferentialexpressioncpp
5.38 score 24 scripts 468 downloadsrprimer - Design Degenerate Oligos from a Multiple DNA Sequence Alignment
Functions, workflow, and a Shiny application for visualizing sequence conservation and designing degenerate primers, probes, and (RT)-(q/d)PCR assays from a multiple DNA sequence alignment. The results can be presented in data frame format and visualized as dashboard-like plots. For more information, please see the package vignette.
Last updated
alignmentddpcrcoveragemultiplesequencealignmentsequencematchingqpcr
5.36 score 4 stars 19 scripts 430 downloadsmsImpute - Imputation of label-free mass spectrometry peptides
MsImpute is a package for imputation of peptide intensity in proteomics experiments. It additionally contains tools for MAR/MNAR diagnosis and assessment of distortions to the probability distribution of the data post imputation. The missing values are imputed by low-rank approximation of the underlying data matrix if they are MAR (method = "v2"), by Barycenter approach if missingness is MNAR ("v2-mnar"), or by Peptide Identity Propagation (PIP).
Last updated
massspectrometryproteomicssoftwareimputation-algorithmlabel-free-proteomicslow-rank-approximation
5.35 score 15 stars 15 scripts 414 downloadsBioCartaImage - BioCarta Pathway Images
The core functionality of the package is to provide coordinates of genes on the BioCarta pathway images and to provide methods to add self-defined graphics to the genes of interest.
Last updated
softwarepathwaysbiocartavisualization
5.33 score 11 stars 13 scripts 258 downloadsggtreeSpace - Visualizing Phylomorphospaces using 'ggtree'
This package is a comprehensive visualization tool specifically designed for exploring phylomorphospace. It not only simplifies the process of generating phylomorphospace, but also enhances it with the capability to add graphic layers to the plot with grammar of graphics to create fully annotated phylomorphospaces. It also provide some utilities to help interpret evolutionary patterns.
Last updated
annotationvisualizationphylogeneticssoftware
5.32 score 5 stars 14 scripts 232 downloadsqsvaR - Generate Quality Surrogate Variable Analysis for Degradation Correction
The qsvaR package contains functions for removing the effect of degration in rna-seq data from postmortem brain tissue. The package is equipped to help users generate principal components associated with degradation. The components can be used in differential expression analysis to remove the effects of degradation.
Last updated
softwareworkflowstepnormalizationbiologicalquestiondifferentialexpressionsequencingcoveragebioconductorbraindegradationhumanqsva
5.32 score 6 scripts 350 downloadsPedixplorer - Pedigree Functions
Routines to handle family data with a Pedigree object. The initial purpose was to create correlation structures that describe family relationships such as kinship and identity-by-descent, which can be used to model family data in mixed effects models, such as in the coxme function. Also includes a tool for Pedigree drawing which is focused on producing compact layouts without intervention. Recent additions include utilities to trim the Pedigree object with various criteria, and kinship for the X chromosome.
Last updated
softwaredatarepresentationgeneticsgraphandnetworkvisualizationkinshippedigree
5.30 score 7 stars 1 dependents 16 scripts 271 downloads
SpectraQL - MassQL support for Spectra
The Mass Spec Query Language (MassQL) is a domain-specific language enabling to express a query and retrieve mass spectrometry (MS) data in a more natural and understandable way for MS users. It is inspired by SQL and is by design programming language agnostic. The SpectraQL package adds support for the MassQL query language to R, in particular to MS data represented by Spectra objects. Users can thus apply MassQL expressions to analyze and retrieve specific data from Spectra objects.
Last updated
infrastructureproteomicsmassspectrometrymetabolomics
5.30 score 10 stars 8 scripts 240 downloadsHiCDOC - A/B compartment detection and differential analysis
HiCDOC normalizes intrachromosomal Hi-C matrices, uses unsupervised learning to predict A/B compartments from multiple replicates, and detects significant compartment changes between experiment conditions. It provides a collection of functions assembled into a pipeline to filter and normalize the data, predict the compartments and visualize the results. It accepts several type of data: tabular `.tsv` files, Cooler `.cool` or `.mcool` files, Juicer `.hic` files or HiC-Pro `.matrix` and `.bed` files.
Last updated
hicdna3dstructurenormalizationsequencingsoftwareclusteringzlibcpp
5.30 score 5 stars 6 scripts 400 downloadsSplicingFactory - Splicing Diversity Analysis for Transcriptome Data
The SplicingFactory R package uses transcript-level expression values to analyze splicing diversity based on various statistical measures, like Shannon entropy or the Gini index. These measures can quantify transcript isoform diversity within samples or between conditions. Additionally, the package analyzes the isoform diversity data, looking for significant changes between conditions.
Last updated
transcriptomicsrnaseqdifferentialsplicingalternativesplicingtranscriptomevariantgini-indexrna-seqshannon-entropysimpson-indexsplicing
5.30 score 4 stars 1 scriptsBERT - High Performance Data Integration for Large-Scale Analyses of Incomplete Omic Profiles Using Batch-Effect Reduction Trees (BERT)
Provides efficient batch-effect adjustment of data with missing values. BERT orders all batch effect correction to a tree of pairwise computations. BERT allows parallelization over sub-trees.
Last updated
batcheffectpreprocessingexperimentaldesignqualitycontrolbatch-effectbioconductor-packagebioinformaticsdata-integrationdata-sciencenature-communications
5.26 score 4 stars 23 scripts 256 downloadsHiCParser - Parser for HiC data in R
This package is a parser to import HiC data into R. It accepts several type of data: tabular files, Cooler `.cool` or `.mcool` files, Juicer `.hic` files or HiC-Pro `.matrix` and `.bed` files. The HiC data can be several files, for several replicates and conditions. The data is formated in an InteractionSet object.
Last updated
softwarehicdataimportzlibcpp
5.26 score 1 stars 2 dependents 1 scripts 249 downloadsbroadSeq - broadSeq : for streamlined exploration of RNA-seq data
This package helps user to do easily RNA-seq data analysis with multiple methods (usually which needs many different input formats). Here the user will provid the expression data as a SummarizedExperiment object and will get results from different methods. It will help user to quickly evaluate different methods.
Last updated
geneexpressiondifferentialexpressionrnaseqtranscriptomicssequencingcoveragegenesetenrichmentgo
5.26 score 9 stars 10 scripts 288 downloads
HybridExpress - Comparative analysis of RNA-seq data for hybrids and their progenitors
HybridExpress can be used to perform comparative transcriptomics analysis of hybrids (or allopolyploids) relative to their progenitor species. The package features functions to perform exploratory analyses of sample grouping, identify differentially expressed genes in hybrids relative to their progenitors, classify genes in expression categories (N = 12) and classes (N = 5), and perform functional analyses. We also provide users with graphical functions for the seamless creation of publication-ready figures that are commonly used in the literature.
Last updated
softwarefunctionalgenomicsgeneexpressiontranscriptomicsrnaseqclassificationdifferentialexpressiongene-expressionhybridpolyploidyrna-seq
5.26 score 18 stars 8 scripts 287 downloadsSurfR - Surface Protein Prediction and Identification
Identify Surface Protein coding genes from a list of candidates. Systematically download data from GEO and TCGA or use your own data. Perform DGE on bulk RNAseq data. Perform Meta-analysis. Descriptive enrichment analysis and plots.
Last updated
softwaresequencingrnaseqgeneexpressiontranscriptiondifferentialexpressionprincipalcomponentgenesetenrichmentpathwaysbatcheffectfunctionalgenomicsvisualizationdataimportfunctionalpredictiongenepredictiongodgeenrichment-analysismetaanalysisplotsproteinspublic-datasurfacesurfaceome
5.26 score 6 stars 5 scriptsiSEEhub - iSEE for the Bioconductor ExperimentHub
This package defines a custom landing page for an iSEE app interfacing with the Bioconductor ExperimentHub. The landing page allows users to browse the ExperimentHub, select a data set, download and cache it, and import it directly into a Bioconductor iSEE app.
Last updated
dataimportimmunooncology infrastructureshinyappssinglecellsoftwarebioconductorbioconductor-packagehacktoberfestisee
5.26 score 3 stars 4 scripts 334 downloadsCNVMetrics - Copy Number Variant Metrics
The CNVMetrics package calculates similarity metrics to facilitate copy number variant comparison among samples and/or methods. Similarity metrics can be employed to compare CNV profiles of genetically unrelated samples as well as those with a common genetic background. Some metrics are based on the shared amplified/deleted regions while other metrics rely on the level of amplification/deletion. The data type used as input is a plain text file containing the genomic position of the copy number variations, as well as the status and/or the log2 ratio values. Finally, a visualization tool is provided to explore resulting metrics.
Last updated
biologicalquestionsoftwarecopynumbervariationcnvcopy-number-variationmetricsr-language
5.26 score 4 stars 8 scriptsMouseFM - In-silico methods for genetic finemapping in inbred mice
This package provides methods for genetic finemapping in inbred mice by taking advantage of their very high homozygosity rate (>95%).
Last updated
geneticssnpgenetargetvariantannotationgenomicvariationmultiplecomparisonsystemsbiologymathematicalbiologypatternlogicgenepredictionbiomedicalinformaticsfunctionalgenomicsfinemapgene-candidatesinbred-miceinbred-strainsmouseqtlqtl-mapping
5.26 score 10 scripts 317 downloadsdecontX - Decontamination of single cell genomics data
This package contains implementation of DecontX (Yang et al. 2020), a decontamination algorithm for single-cell RNA-seq, and DecontPro (Yin et al. 2023), a decontamination algorithm for single cell protein expression data. DecontX is a novel Bayesian method to computationally estimate and remove RNA contamination in individual cells without empty droplet information. DecontPro is a Bayesian method that estimates the level of contamination from ambient and background sources in CITE-seq ADT dataset and decontaminate the dataset.
Last updated
singlecellbayesianonetbbcpp
5.25 score 89 scripts 858 downloadsTSAR - Thermal Shift Analysis in R
This package automates analysis workflow for Thermal Shift Analysis (TSA) data. Processing, analyzing, and visualizing data through both shiny applications and command lines. Package aims to simplify data analysis and offer front to end workflow, from raw data to multiple trial analysis.
Last updated
softwareshinyappsvisualizationqpcr
5.23 score 14 scripts 272 downloadsiSEEde - iSEE extension for panels related to differential expression analysis
This package contains diverse functionality to extend the usage of the iSEE package, including additional classes for the panels or modes facilitating the analysis of differential expression results. This package does not perform differential expression. Instead, it provides methods to embed precomputed differential expression results in a SummarizedExperiment object, in a manner that is compatible with interactive visualisation in iSEE applications.
Last updated
softwareinfrastructuredifferentialexpressionbioconductorhacktoberfestiseeu
5.23 score 1 stars 21 scripts 312 downloadsspaSim - Spatial point data simulator for tissue images
A suite of functions for simulating spatial patterns of cells in tissue images. Output images are multitype point data in SingleCellExperiment format. Each point represents a cell, with its 2D locations and cell type. Potential cell patterns include background cells, tumour/immune cell clusters, immune rings, and blood/lymphatic vessels.
Last updated
statisticalmethodspatialbiomedicalinformatics
5.21 score 2 stars 27 scripts 344 downloadsCleanUpRNAseq - Detect and Correct Genomic DNA Contamination in RNA-seq Data
RNA-seq data generated by some library preparation methods, such as rRNA-depletion-based method and the SMART-seq method, might be contaminated by genomic DNA (gDNA), if DNase I disgestion is not performed properly during RNA preparation. CleanUpRNAseq is developed to check if RNA-seq data is suffered from gDNA contamination. If so, it can perform correction for gDNA contamination and reduce false discovery rate of differentially expressed genes.
Last updated
qualitycontrolsequencinggeneexpression
5.20 score 8 stars 6 scripts 299 downloadsBioNAR - Biological Network Analysis in R
the R package BioNAR, developed to step by step analysis of PPI network. The aim is to quantify and rank each protein’s simultaneous impact into multiple complexes based on network topology and clustering. Package also enables estimating of co-occurrence of diseases across the network and specific clusters pointing towards shared/common mechanisms.
Last updated
softwaregraphandnetworknetwork
5.20 score 3 stars 35 scripts 380 downloadsaggregateBioVar - Differential Gene Expression Analysis for Multi-subject scRNA-seq
For single cell RNA-seq data collected from more than one subject (e.g. biological sample or technical replicates), this package contains tools to summarize single cell gene expression profiles at the level of subject. A SingleCellExperiment object is taken as input and converted to a list of SummarizedExperiment objects, where each list element corresponds to an assigned cell type. The SummarizedExperiment objects contain aggregate gene-by-subject count matrices and inter-subject column metadata for individual subjects that can be processed using downstream bulk RNA-seq tools.
Last updated
softwaresinglecellrnaseqtranscriptomicstranscriptiongeneexpressiondifferentialexpression
5.20 score 5 stars 21 scripts 490 downloadsrhdf5client - Access HDF5 content from HDF Scalable Data Service
This package provides functionality for reading data from HDF Scalable Data Service from within R. The HSDSArray function bridges from HSDS to the user via the DelayedArray interface. Bioconductor manages an open HSDS instance graciously provided by John Readey of the HDF Group.
Last updated
dataimportsoftwareinfrastructure
5.18 score 2 dependents 42 scripts 463 downloads
vmrseq - Probabilistic Modeling of Single-cell Methylation Heterogeneity
High-throughput single-cell measurements of DNA methylation allows studying inter-cellular epigenetic heterogeneity, but this task faces the challenges of sparsity and noise. We present vmrseq, a statistical method that overcomes these challenges and identifies variably methylated regions accurately and robustly.
Last updated
softwareimmunooncologydnamethylationepigeneticssinglecellsequencingwholegenomecomputational-biologydimensionality-reductionepigenomics-workflowhidden-markov-modelprobabilistic-models
5.18 score 10 stars 5 scriptsscQTLtools - scQTLtools: an R/Bioconductor package for comprehensive identification and visualization of single-cell eQTLs
scQTLtools is a comprehensive R/Bioconductor package that facilitates end-to-end single-cell eQTL analysis, from preprocessing to visualization
Last updated
softwaregeneexpressiongeneticvariabilitysnpdifferentialexpressiongenomicvariationvariantdetectiongeneticsfunctionalgenomicssystemsbiologyregressionsinglecellnormalizationvisualizationpreprocessingrna-seqsc-eqtl
5.18 score 6 stars 6 scripts 242 downloads
PIUMA - Phenotypes Identification Using Mapper from topological data Analysis
The PIUMA package offers a tidy pipeline of Topological Data Analysis frameworks to identify and characterize communities in high and heterogeneous dimensional data.
Last updated
clusteringgraphandnetworkdimensionreductionnetworkclassification
5.18 score 5 stars 3 scripts 272 downloadscrisprVerse - Easily install and load the crisprVerse ecosystem for CRISPR gRNA design
The crisprVerse is a modular ecosystem of R packages developed for the design and manipulation of CRISPR guide RNAs (gRNAs). All packages share a common language and design principles. This package is designed to make it easy to install and load the crisprVerse packages in a single step. To learn more about the crisprVerse, visit <https://www.github.com/crisprVerse>.
Last updated
crisprfunctionalgenomicsgenetargetcrispr-analysiscrispr-designcrispr-targetgrnagrna-sequencegrna-sequences
5.18 score 15 stars 9 scripts 344 downloadsepistasisGA - An R package to identify multi-snp effects in nuclear family studies using the GADGETS method
This package runs the GADGETS method to identify epistatic effects in nuclear family studies. It also provides functions for permutation-based inference and graphical visualization of the results.
Last updated
geneticssnpgeneticvariabilityopenblascpp
5.17 score 1 stars 11 scripts 276 downloadsCytoMDS - Low Dimensions projection of cytometry samples
This package implements a low dimensional visualization of a set of cytometry samples, in order to visually assess the 'distances' between them. This, in turn, can greatly help the user to identify quality issues like batch effects or outlier samples, and/or check the presence of potential sample clusters that might align with the exeprimental design. The CytoMDS algorithm combines, on the one hand, the concept of Earth Mover's Distance (EMD), a.k.a. Wasserstein metric and, on the other hand, the Multi Dimensional Scaling (MDS) algorithm for the low dimensional projection. Also, the package provides some diagnostic tools for both checking the quality of the MDS projection, as well as tools to help with the interpretation of the axes of the projection.
Last updated
flowcytometryqualitycontroldimensionreductionmultidimensionalscalingsoftwarevisualization
5.16 score 1 stars 1 dependents 12 scripts 319 downloadslineagespot - Detection of SARS-CoV-2 lineages in wastewater samples using next-generation sequencing
Lineagespot is a framework written in R, and aims to identify SARS-CoV-2 related mutations based on a single (or a list) of variant(s) file(s) (i.e., variant calling format). The method can facilitate the detection of SARS-CoV-2 lineages in wastewater samples using next generation sequencing, and attempts to infer the potential distribution of the SARS-CoV-2 lineages.
Last updated
variantdetectionvariantannotationsequencing
5.15 score 2 stars 6 scripts 313 downloadsmarr - Maximum rank reproducibility
marr (Maximum Rank Reproducibility) is a nonparametric approach that detects reproducible signals using a maximal rank statistic for high-dimensional biological data. In this R package, we implement functions that measures the reproducibility of features per sample pair and sample pairs per feature in high-dimensional biological replicate experiments. The user-friendly plot functions in this package also plot histograms of the reproducibility of features per sample pair and sample pairs per feature. Furthermore, our approach also allows the users to select optimal filtering threshold values for the identification of reproducible features and sample pairs based on output visualization checks (histograms). This package also provides the subset of data filtered by reproducible features and/or sample pairs.
Last updated
qualitycontrolmetabolomicsmassspectrometryrnaseqchipseqcpp
5.12 score 3 stars 22 scripts 322 downloadsomXplore - Vizualization tools for 'omics' datasets with R
This package contains a collection of functions (written as shiny modules) for the visualisation and the statistical analysis of omics data. These plots can be displayed individually or embedded in a global Shiny module. Additionaly, it is possible to integrate third party modules to the main interface of the package omXplore.
Last updated
softwareshinyappsmassspectrometrydatarepresentationguiqualitycontrolprostar2
5.10 score 25 scripts 231 downloads
mobileRNA - mobileRNA: Investigate the RNA mobilome & population-scale changes
Genomic analysis can be utilised to identify differences between RNA populations in two conditions, both in production and abundance. This includes the identification of RNAs produced by multiple genomes within a biological system. For example, RNA produced by pathogens within a host or mobile RNAs in plant graft systems. The mobileRNA package provides methods to pre-process, analyse and visualise the sRNA and mRNA populations based on the premise of mapping reads to all genotypes at the same time.
Last updated
visualizationrnaseqsequencingsmallrnagenomeassemblyclusteringexperimentaldesignqualitycontrolworkflowstepalignmentpreprocessingbioinformaticsplant-science
5.08 score 4 stars 2 scriptsDELocal - Identifies differentially expressed genes with respect to other local genes
The goal of DELocal is to identify DE genes compared to their neighboring genes from the same chromosomal location. It has been shown that genes of related functions are generally very far from each other in the chromosome. DELocal utilzes this information to identify DE genes comparing with their neighbouring genes.
Last updated
geneexpressiondifferentialexpressionrnaseqtranscriptomics
5.08 score 2 stars 1 dependents 8 scripts 296 downloadsiSEEhex - iSEE extension for summarising data points in hexagonal bins
This package provides panels summarising data points in hexagonal bins for `iSEE`. It is part of `iSEEu`, the iSEE universe of panels that extend the `iSEE` package.
Last updated
softwareinfrastructurebioconductoriseeushiny-r
5.08 score 2 dependents 7 scripts 372 downloadsPanomiR - Detection of miRNAs that regulate interacting groups of pathways
PanomiR is a package to detect miRNAs that target groups of pathways from gene expression data. This package provides functionality for generating pathway activity profiles, determining differentially activated pathways between user-specified conditions, determining clusters of pathways via the PCxN package, and generating miRNAs targeting clusters of pathways. These function can be used separately or sequentially to analyze RNA-Seq data.
Last updated
geneexpressiongenesetenrichmentgenetargetmirnapathways
5.08 score 3 stars 20 scripts 374 downloadsshinyepico - ShinyÉPICo
ShinyÉPICo is a graphical pipeline to analyze Illumina DNA methylation arrays (450k or EPIC). It allows to calculate differentially methylated positions and differentially methylated regions in a user-friendly interface. Moreover, it includes several options to export the results and obtain files to perform downstream analysis.
Last updated
differentialmethylationdnamethylationmicroarraypreprocessingqualitycontrol
5.08 score 6 stars 6 scripts 356 downloadsscatterHatch - Creates hatched patterns for scatterplots
The objective of this package is to efficiently create scatterplots where groups can be distinguished by color and texture. Visualizations in computational biology tend to have many groups making it difficult to distinguish between groups solely on color. Thus, this package is useful for increasing the accessibility of scatterplot visualizations to those with visual impairments such as color blindness.
Last updated
visualizationsinglecellcellbiologysoftwarespatial
5.05 score 7 stars 16 scripts 293 downloadsMeLSI - Metric Learning for Statistical Inference in Microbiome Analysis
MeLSI (Metric Learning for Statistical Inference) is a novel machine learning method for microbiome data analysis that learns optimal distance metrics to improve statistical power in detecting group differences. Unlike traditional distance metrics (Bray-Curtis, Euclidean, Jaccard), MeLSI adapts to the specific characteristics of your dataset to maximize separation between groups. The method uses an ensemble of weak learners to identify which microbial features drive group differences, providing both improved statistical power and biological interpretability through feature importance weights.
Last updated
softwarestatisticalmethodmicrobiomecpp
5.04 score 1 stars 17 scripts 253 downloadsDifferentialRegulation - Differentially regulated genes from scRNA-seq data
DifferentialRegulation is a method for detecting differentially regulated genes between two groups of samples (e.g., healthy vs. disease, or treated vs. untreated samples), by targeting differences in the balance of spliced and unspliced mRNA abundances, obtained from single-cell RNA-sequencing (scRNA-seq) data. From a mathematical point of view, DifferentialRegulation accounts for the sample-to-sample variability, and embeds multiple samples in a Bayesian hierarchical model. Furthermore, our method also deals with two major sources of mapping uncertainty: i) 'ambiguous' reads, compatible with both spliced and unspliced versions of a gene, and ii) reads mapping to multiple genes. In particular, ambiguous reads are treated separately from spliced and unsplced reads, while reads that are compatible with multiple genes are allocated to the gene of origin. Parameters are inferred via Markov chain Monte Carlo (MCMC) techniques (Metropolis-within-Gibbs).
Last updated
differentialsplicingbayesiangeneticsrnaseqsequencingdifferentialexpressiongeneexpressionmultiplecomparisonsoftwaretranscriptionstatisticalmethodvisualizationsinglecellgenetargetopenblascpp
5.04 score 11 stars 5 scripts 375 downloadsRega - R Interface to European Genome-Phenome Archive
The European Genome-phenome Archive (EGA) provides long-term storage and controlled sharing of personally identifiable genetic data. The Rega package offers a streamlined and extensible R interface to the EGA API, facilitating the programmatic upload of metadata. GEO-like Excel submission template is provided as a default method of organizing submission metadata.
Last updated
softwareinfrastructurethirdpartyclient
5.02 score 3 scripts 252 downloadsiSEEpathways - iSEE extension for panels related to pathway analysis
This package contains diverse functionality to extend the usage of the iSEE package, including additional classes for the panels or modes facilitating the analysis of pathway analysis results. This package does not perform pathway analysis. Instead, it provides methods to embed precomputed pathway analysis results in a SummarizedExperiment object, in a manner that is compatible with interactive visualisation in iSEE applications.
Last updated
softwareinfrastructuredifferentialexpressiongeneexpressionguivisualizationpathwaysgenesetenrichmentgoshinyappsbioconductorhacktoberfestiseeiseeu
5.01 score 1 stars 17 scripts 276 downloads
TrIdent - TrIdent - Transduction Identification
The `TrIdent` R package automates the analysis of transductomics data by detecting, classifying, and characterizing read coverage patterns associated with potential transduction events. Transductomics is a DNA sequencing-based method for the detection and characterization of transduction events in pure cultures and complex communities. Transductomics relies on mapping sequencing reads from a viral-like particle (VLP)-fraction of a sample to contigs assembled from the metagenome (whole-community) of the same sample. Reads from bacterial DNA carried by VLPs will map back to the bacterial contigs of origin creating read coverage patterns indicative of ongoing transduction.
Last updated
coveragemetagenomicspatternlogicclassificationsequencingbacteriophagehorizontal-gene-transferpattern-matchingphagesequencing-coveragetransductiontransductomicsvirus-like-particle
5.00 score 2 stars 8 scripts 202 downloadsCytoPipelineGUI - GUI's for visualization of flow cytometry data analysis pipelines
This package is the companion of the `CytoPipeline` package. It provides GUI's (shiny apps) for the visualization of flow cytometry data analysis pipelines that are run with `CytoPipeline`. Two shiny applications are provided, i.e. an interactive flow frame assessment and comparison tool and an interactive scale transformations visualization and adjustment tool.
Last updated
flowcytometrypreprocessingqualitycontrolworkflowstepimmunooncologysoftwarevisualizationguishinyapps
5.00 score 2 stars 4 scriptskatdetectr - Detection, Characterization and Visualization of Kataegis in Sequencing Data
Kataegis refers to the occurrence of regional hypermutation and is a phenomenon observed in a wide range of malignancies. Using changepoint detection katdetectr aims to identify putative kataegis foci from common data-formats housing genomic variants. Katdetectr has shown to be a robust package for the detection, characterization and visualization of kataegis.
Last updated
wholegenomesoftwaresnpsequencingclassificationvariantannotation
5.00 score 5 stars 4 scripts 362 downloadsRolDE - RolDE: Robust longitudinal Differential Expression
RolDE detects longitudinal differential expression between two conditions in noisy high-troughput data. Suitable even for data with a moderate amount of missing values.RolDE is a composite method, consisting of three independent modules with different approaches to detecting longitudinal differential expression. The combination of these diverse modules allows RolDE to robustly detect varying differences in longitudinal trends and expression levels in diverse data types and experimental settings.
Last updated
statisticalmethodsoftwaretimecourseregressionproteomicsdifferentialexpression
5.00 score 5 stars 3 scripts 356 downloadscrisprBwa - BWA-based alignment of CRISPR gRNA spacer sequences
Provides a user-friendly interface to map on-targets and off-targets of CRISPR gRNA spacer sequences using bwa. The alignment is fast, and can be performed using either commonly-used or custom CRISPR nucleases. The alignment can work with any reference or custom genomes. Currently not supported on Windows machines.
Last updated
crisprfunctionalgenomicsalignmentalignerbioconductorbioconductor-packagebwacrispr-analysiscrispr-cas9crispr-designcrispr-targetgrnagrna-sequencegrna-sequencessgrnasgrna-design
5.00 score 2 stars 11 scripts 312 downloadsSCArray - Large-scale single-cell omics data manipulation with GDS files
Provides large-scale single-cell omics data manipulation using Genomic Data Structure (GDS) files. It combines dense and sparse matrices stored in GDS files and the Bioconductor infrastructure framework (SingleCellExperiment and DelayedArray) to provide out-of-memory data storage and large-scale manipulation using the R programming language.
Last updated
infrastructuredatarepresentationdataimportsinglecellrnaseqcpp
5.00 score 1 stars 1 dependents 11 scripts 403 downloadsCPSM - CPSM: Cancer patient survival model
CPSM provides a comprehensive computational pipeline for predicting survival probability and risk groups in cancer patients. The package includes steps for data preprocessing, training/test split, and normalization. It enables feature selection using univariate survival analysis and computes a LASSO-based prognostic index (PI) score. CPSM supports the development of predictive models using various feature sets and offers a suite of visualization tools, including survival curves based on predicted probabilities, barplots for predicted mean and median survival times, KM plots overlaid with individual survival predictions, and nomograms for estimating 1-, 3-, 5-, and 10-year survival probabilities. This makes CPSM a versatile tool for survival analysis in cancer research.
Last updated
normalizationsurvivalgeneexpressionpreprocessingfeatureextractionsoftwarevisualization
4.95 score 2 stars 6 scripts 264 downloadsMSstatsBig - MSstats Preprocessing for Larger than Memory Data
MSstats package provide tools for preprocessing, summarization and differential analysis of mass spectrometry (MS) proteomics data. Recently, some MS protocols enable acquisition of data sets that result in larger than memory quantitative data. MSstats functions are not able to process such data. MSstatsBig package provides additional converter functions that enable processing larger than memory data sets.
Last updated
massspectrometryproteomicssoftware
4.95 score 1 dependents 10 scripts 270 downloadscfdnakit - Fragmen-length analysis package from high-throughput sequencing of cell-free DNA (cfDNA)
This package provides basic functions for analyzing shallow whole-genome sequencing (~0.3X or more) of cell-free DNA (cfDNA). The package basically extracts the length of cfDNA fragments and aids the vistualization of fragment-length information. The package also extract fragment-length information per non-overlapping fixed-sized bins and used it for calculating ctDNA estimation score (CES).
Last updated
copynumbervariationsequencingwholegenome
4.95 score 9 stars 10 scripts 337 downloadsCTdata - Data companion to CTexploreR
Data from publicly available databases (GTEx, CCLE, TCGA and ENCODE) that go with CTexploreR in order to re-define a comprehensive and thoroughly curated list of CT genes and their main characteristics.
Last updated
transcriptomicsepigeneticsgeneexpressiondataimportexperimenthubsoftware
4.95 score 1 stars 1 dependents 7 scripts 316 downloadsalabaster.spatial - Save and Load Spatial 'Omics Data to/from File
Save SpatialExperiment objects and their images into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
4.95 score 2 dependents 10 scripts 328 downloadsscDotPlot - Cluster a Single-cell RNA-seq Dot Plot
Dot plots of single-cell RNA-seq data allow for an examination of the relationships between cell groupings (e.g. clusters) and marker gene expression. The scDotPlot package offers a unified approach to perform a hierarchical clustering analysis and add annotations to the columns and/or rows of a scRNA-seq dot plot. It works with SingleCellExperiment and Seurat objects as well as data frames.
Last updated
softwarevisualizationdifferentialexpressiongeneexpressiontranscriptionrnaseqsinglecellsequencingclustering
4.92 score 7 stars 12 scripts 293 downloadsmotifTestR - Perform key tests for binding motifs in sequence data
Taking a set of sequence motifs as PWMs, test a set of sequences for over-representation of these motifs, as well as any positional features within the set of motifs. Enrichment analysis can be undertaken using multiple statistical approaches. The package also contains core functions to prepare data for analysis, and to visualise results.
Last updated
motifannotationchipseqchiponchipsequencematchingsoftware
4.90 score 1 stars 3 scripts 337 downloadsPICB - piRNA Cluster Builder
piRNAs (short for PIWI-interacting RNAs) and their PIWI protein partners play a key role in fertility and maintaining genome integrity by restricting mobile genetic elements (transposons) in germ cells. piRNAs originate from genomic regions known as piRNA clusters. The piRNA Cluster Builder (PICB) is a versatile toolkit designed to identify genomic regions with a high density of piRNAs. It constructs piRNA clusters through a stepwise integration of unique and multimapping piRNAs and offers wide-ranging parameter settings, supported by an optimization function that allows users to test different parameter combinations to tailor the analysis to their specific piRNA system. The output includes extensive metadata columns, enabling researchers to rank clusters and extract cluster characteristics.
Last updated
geneticsgenomeannotationsequencingfunctionalpredictioncoveragetranscriptomics
4.90 score 8 stars 5 scripts 246 downloadsAnVILAz - R / Bioconductor Support for the AnVIL Azure Platform
The AnVIL is a cloud computing resource developed in part by the National Human Genome Research Institute. The AnVILAz package supports end-users and developers using the AnVIL platform in the Azure cloud. The package provides a programmatic interface to AnVIL resources, including workspaces, notebooks, tables, and workflows. The package also provides utilities for managing resources, including copying files to and from Azure Blob Storage, and creating shared access signatures (SAS) for secure access to Azure resources.
Last updated
softwareinfrastructurethirdpartyclientu24hg010263
4.90 score 5 scripts 260 downloadsgDNAx - Diagnostics for assessing genomic DNA contamination in RNA-seq data
Provides diagnostics for assessing genomic DNA contamination in RNA-seq data, as well as plots representing these diagnostics. Moreover, the package can be used to get an insight into the strand library protocol used and, in case of strand-specific libraries, the strandedness of the data. Furthermore, it provides functionality to filter out reads of potential gDNA origin.
Last updated
transcriptiontranscriptomicsrnaseqsequencingpreprocessingsoftwaregeneexpressioncoveragedifferentialexpressionfunctionalgenomicssplicedalignmentalignment
4.90 score 2 stars 6 scripts 332 downloadsCTSV - Identification of cell-type-specific spatially variable genes accounting for excess zeros
The R package CTSV implements the CTSV approach developed by Jinge Yu and Xiangyu Luo that detects cell-type-specific spatially variable genes accounting for excess zeros. CTSV directly models sparse raw count data through a zero-inflated negative binomial regression model, incorporates cell-type proportions, and performs hypothesis testing based on R package pscl. The package outputs p-values and q-values for genes in each cell type, and CTSV is scalable to datasets with tens of thousands of genes measured on hundreds of spots. CTSV can be installed in Windows, Linux, and Mac OS.
Last updated
geneexpressionstatisticalmethodregressionspatialgenetics
4.90 score 4 stars 20 scripts 285 downloadsTOP - TOP Constructs Transferable Model Across Gene Expression Platforms
TOP constructs a transferable model across gene expression platforms for prospective experiments. Such a transferable model can be trained to make predictions on independent validation data with an accuracy that is similar to a re-substituted model. The TOP procedure also has the flexibility to be adapted to suit the most common clinical response variables, including linear response, binomial and Cox PH models.
Last updated
softwaresurvivalgeneexpression
4.90 score 79 scripts 297 downloadsHTqPCR - Automated analysis of high-throughput qPCR data
Analysis of Ct values from high throughput quantitative real-time PCR (qPCR) assays across multiple conditions or replicates. The input data can be from spatially-defined formats such ABI TaqMan Low Density Arrays or OpenArray; LightCycler from Roche Applied Science; the CFX plates from Bio-Rad Laboratories; conventional 96- or 384-well plates; or microfluidic devices such as the Dynamic Arrays from Fluidigm Corporation. HTqPCR handles data loading, quality assessment, normalization, visualization and parametric or non-parametric testing for statistical significance in Ct values between features (e.g. genes, microRNAs).
Last updated
microtitreplateassaydifferentialexpressiongeneexpressiondataimportqualitycontrolpreprocessingvisualizationmultiplecomparisonqpcr
4.89 score 1 dependents 13 scripts 562 downloadsFuseSOM - A Correlation Based Multiview Self Organizing Maps Clustering For IMC Datasets
A correlation-based multiview self-organizing map for the characterization of cell types in highly multiplexed in situ imaging cytometry assays (`FuseSOM`) is a tool for unsupervised clustering. `FuseSOM` is robust and achieves high accuracy by combining a `Self Organizing Map` architecture and a `Multiview` integration of correlation based metrics. This allows FuseSOM to cluster highly multiplexed in situ imaging cytometry assays.
Last updated
singlecellcellbasedassaysclusteringspatial
4.89 score 1 stars 26 scripts 348 downloadsscoup - Simulate Codons with Darwinian Selection Modelled as an OU Process
An elaborate molecular evolutionary framework that facilitates straightforward simulation of codon genetic sequences subjected to different degrees and/or patterns of Darwinian selection. The model is built upon the fitness landscape paradigm of Sewall Wright, as popularised by the mutation-selection model of Halpern and Bruno. This enables realistic evolutionary process of living organisms to be reproducible seamlessly. For example, an Ornstein-Uhlenbeck fitness update algorithm is incorporated herein. Consequently, otherwise complex biological processes, such as the effect of the interplay between genetic drift and fitness landscape fluctuations on the inference of diversifying selection, may now be investigated with minimal effort. Frequency-dependent and stochastic fitness landscape update techniques are available.
Last updated
alignmentclassificationcomparativegenomicsdataimportgeneticsmathematicalbiologyresearchfieldsequencingsequencematchingsoftwarestatisticalmethodworkflowstepcomputational-biologyevolutionary-biologymolecular-biologysimulation
4.89 score 11 scripts 238 downloadsAnVILPublish - Publish Packages and Other Resources to AnVIL Workspaces
Use this package to create or update AnVIL workspaces from resources such as R / Bioconductor packages. The metadata about the package (e.g., select information from the package DESCRIPTION file and from vignette YAML headings) are used to populate the 'DASHBOARD'. Vignettes are translated to python notebooks ready for evaluation in AnVIL.
Last updated
infrastructuresoftware
4.88 score 1 dependents 2 scripts 308 downloadstransmogR - Modify a set of reference sequences using a set of variants
transmogR provides the tools needed to crate a new reference genome or reference transcriptome, using a set of variants. Variants can be any combination of SNPs, Insertions and Deletions. The intended use-case is to enable creation of variant-modified reference transcriptomes for incorporation into transcriptomic pseudo-alignment workflows, such as salmon.
Last updated
alignmentgenomicvariationsequencingtranscriptomevariantvariantannotationzlib
4.85 score 5 scripts 296 downloadsGNOSIS - Genomics explorer using statistical and survival analysis in R
GNOSIS incorporates a range of R packages enabling users to efficiently explore and visualise clinical and genomic data obtained from cBioPortal. GNOSIS uses an intuitive GUI and multiple tab panels supporting a range of functionalities. These include data upload and initial exploration, data recoding and subsetting, multiple visualisations, survival analysis, statistical analysis and mutation analysis, in addition to facilitating reproducible research.
Last updated
softwareshinyappssurvivalgui
4.85 score 7 stars 7 scripts 272 downloadsBiocBook - Write, containerize, publish and version Quarto books with Bioconductor
A BiocBook can be created by authors (e.g. R developers, but also scientists, teachers, communicators, ...) who wish to 1) write (compile a body of biological and/or bioinformatics knowledge), 2) containerize (provide Docker images to reproduce the examples illustrated in the compendium), 3) publish (deploy an online book to disseminate the compendium), and 4) version (automatically generate specific online book versions and Docker images for specific Bioconductor releases).
Last updated
infrastructurereportwritingsoftwarequarto
4.85 score 7 stars 8 scripts 320 downloadseds - eds: Low-level reader for Alevin EDS format
This packages provides a single function, readEDS. This is a low-level utility for reading in Alevin EDS format into R. This function is not designed for end-users but instead the package is predominantly for simplifying package dependency graph for other Bioconductor packages.
Last updated
sequencingrnaseqgeneexpressionsinglecellcpp
4.84 score 1 dependents 23 scripts 540 downloadsmultiWGCNA - multiWGCNA
An R package for deeping mining gene co-expression networks in multi-trait expression data. Provides functions for analyzing, comparing, and visualizing WGCNA networks across conditions. multiWGCNA was designed to handle the common case where there are multiple biologically meaningful sample traits, such as disease vs wildtype across development or anatomical region.
Last updated
sequencingrnaseqgeneexpressiondifferentialexpressionregressionclustering
4.82 score 11 scripts 332 downloadsvsclust - Feature-based variance-sensitive quantitative clustering
Feature-based variance-sensitive clustering of omics data. Optimizes cluster assignment by taking into account individual feature variance. Includes several modules for statistical testing, clustering and enrichment analysis.
Last updated
clusteringannotationprincipalcomponentdifferentialexpressionvisualizationproteomicsmetabolomicscpp
4.82 score 11 scripts 327 downloadsEpiMix - EpiMix: an integrative tool for the population-level analysis of DNA methylation
EpiMix is a comprehensive tool for the integrative analysis of high-throughput DNA methylation data and gene expression data. EpiMix enables automated data downloading (from TCGA or GEO), preprocessing, methylation modeling, interactive visualization and functional annotation.To identify hypo- or hypermethylated CpG sites across physiological or pathological conditions, EpiMix uses a beta mixture modeling to identify the methylation states of each CpG probe and compares the methylation of the experimental group to the control group.The output from EpiMix is the functional DNA methylation that is predictive of gene expression. EpiMix incorporates specialized algorithms to identify functional DNA methylation at various genetic elements, including proximal cis-regulatory elements of protein-coding genes, distal enhancers, and genes encoding microRNAs and lncRNAs.
Last updated
softwareepigeneticspreprocessingdnamethylationgeneexpressiondifferentialmethylation
4.82 score 2 stars 1 dependents 11 scripts 332 downloads
signifinder - Collection and implementation of public transcriptional cancer signatures
signifinder is an R package for computing and exploring a compendium of tumor signatures. It allows to compute a variety of signatures coming from public literature, based on gene expression values, and return single-sample (-cell/-spot) scores. Currently, signifinder collects more than 70 distinct signatures, relating to multiple tumors and multiple cancer processes.
Last updated
geneexpressiongenetargetimmunooncologybiomedicalinformaticsrnaseqmicroarrayreportwritingvisualizationsinglecellspatialgenesignaling
4.81 score 9 stars 18 scripts 311 downloadssurvClust - Identification Of Clinically Relevant Genomic Subtypes Using Outcome Weighted Learning
survClust is an outcome weighted integrative clustering algorithm used to classify multi-omic samples on their available time to event information. The resulting clusters are cross-validated to avoid over overfitting and output classification of samples that are molecularly distinct and clinically meaningful. It takes in binary (mutation) as well as continuous data (other omic types).
Last updated
softwareclusteringsurvivalclassificationcpp
4.81 score 16 stars 20 scripts 242 downloadschevreulProcess - Tools for managing SingleCellExperiment objects as projects
Tools for analyzing SingleCellExperiment objects as projects. for input into the chevreulShiny app downstream. Includes functions for analysis of single cell RNA sequencing data. Supported by NIH grants R01CA137124 and R01EY026661 to David Cobrinik.
Last updated
coveragernaseqsequencingvisualizationgeneexpressiontranscriptionsinglecelltranscriptomicsnormalizationpreprocessingqualitycontroldimensionreductiondataimport
4.78 score 2 dependents 4 scripts 263 downloadschevreulPlot - Plots used in the chevreulPlot package
Tools for plotting SingleCellExperiment objects in the chevreulPlot package. Includes functions for analysis and visualization of single-cell data. Supported by NIH grants R01CA137124 and R01EY026661 to David Cobrinik.
Last updated
coveragernaseqsequencingvisualizationgeneexpressiontranscriptionsinglecelltranscriptomicsnormalizationpreprocessingqualitycontroldimensionreductiondataimport
4.78 score 1 dependents 3 scriptsmist - Differential Methylation Analysis for scDNAm Data
mist (Methylation Inference for Single-cell along Trajectory) is a hierarchical Bayesian framework for modeling DNA methylation trajectories and performing differential methylation (DM) analysis in single-cell DNA methylation (scDNAm) data. It estimates developmental-stage-specific variations, identifies genomic features with drastic changes along pseudotime, and, for two phenotypic groups, detects features with distinct temporal methylation patterns. mist uses Gibbs sampling to estimate parameters for temporal changes and stage-specific variations.
Last updated
epigeneticsdifferentialmethylationdnamethylationsinglecellsoftware
4.78 score 2 stars 12 scripts 244 downloadsMOSClip - Multi Omics Survival Clip
Topological pathway analysis tool able to integrate multi-omics data. It finds survival-associated modules or significant modules for two-class analysis. This tool have two main methods: pathway tests and module tests. The latter method allows the user to dig inside the pathways itself.
Last updated
softwarestatisticalmethodgraphandnetworksurvivalregressiondimensionreductionpathwaysreactome
4.78 score 1 stars 5 scripts 256 downloads
seahtrue - Seahtrue revives XF data for structured data analysis
Seahtrue organizes oxygen consumption and extracellular acidification analysis data from experiments performed on an XF analyzer into structured nested tibbles.This allows for detailed processing of raw data and advanced data visualization and statistics. Seahtrue introduces an open and reproducible way to analyze these XF experiments. It uses file paths to .xlsx files. These .xlsx files are supplied by the userand are generated by the user in the Wave software from Agilent from the assay result files (.asyr). The .xlsx file contains different sheets of important data for the experiment; 1. Assay Information - Details about how the experiment was set up. 2. Rate Data - Information about the OCR and ECAR rates. 3. Raw Data - The original raw data collected during the experiment. 4. Calibration Data - Data related to calibrating the instrument. Seahtrue focuses on getting the specific data needed for analysis. Once this data is extracted, it is prepared for calculations through preprocessing. To make sure everything is accurate, both the initial data and the preprocessed data go through thorough checks.
Last updated
cellbasedassaysfunctionalpredictiondatarepresentationdataimportcellbiologycheminformaticsmetabolomicsmicrotitreplateassayvisualizationqualitycontrolbatcheffectexperimentaldesignpreprocessinggo
4.78 score 2 stars 4 scripts 272 downloadsfunOmics - Aggregating Omics Data into Higher-Level Functional Representations
The 'funOmics' package ggregates or summarizes omics data into higher level functional representations such as GO terms gene sets or KEGG metabolic pathways. The aggregated data matrix represents functional activity scores that facilitate the analysis of functional molecular sets while allowing to reduce dimensionality and provide easier and faster biological interpretations. Coordinated functional activity scores can be as informative as single molecules!
Last updated
softwaretranscriptomicsmetabolomicsproteomicspathwaysgokegg
4.78 score 6 stars 3 scripts 216 downloadscrisprShiny - Exploring curated CRISPR gRNAs via Shiny
Provides means to interactively visualize guide RNAs (gRNAs) in GuideSet objects via Shiny application. This GUI can be self-contained or as a module within a larger Shiny app. The content of the app reflects the annotations present in the passed GuideSet object, and includes intuitive tools to examine, filter, and export gRNAs, thereby making gRNA design more user-friendly.
Last updated
crisprfunctionalgenomicsgenetargetguicrispr-analysiscrispr-designshiny
4.78 score 2 stars 9 scripts 274 downloads
MIRit - Integrate microRNA and gene expression to decipher pathway complexity
MIRit is an R package that provides several methods for investigating the relationships between miRNAs and genes in different biological conditions. In particular, MIRit allows to explore the functions of dysregulated miRNAs, and makes it possible to identify miRNA-gene regulatory axes that control biological pathways, thus enabling the users to unveil the complexity of miRNA biology. MIRit is an all-in-one framework that aims to help researchers in all the central aspects of an integrative miRNA-mRNA analyses, from differential expression analysis to network characterization.
Last updated
softwaregeneregulationnetworkenrichmentnetworkinferenceepigeneticsfunctionalgenomicssystemsbiologynetworkpathwaysgeneexpressiondifferentialexpressionmirnamirna-mrna-interactionmirna-seqmirnaseq-analysiscpp
4.78 score 2 stars 6 scripts 416 downloadstidyCoverage - Extract and aggregate genomic coverage over features of interest
`tidyCoverage` framework enables tidy manipulation of collections of genomic tracks and features using `tidySummarizedExperiment` methods. It facilitates the extraction, aggregation and visualization of genomic coverage over individual or thousands of genomic loci, relying on `CoverageExperiment` and `AggregatedCoverage` classes. This accelerates the integration of genomic track data in genomic analysis workflows.
Last updated
softwaresequencingcoverage
4.78 score 24 stars 10 scripts 268 downloadsbeachmat.hdf5 - beachmat bindings for HDF5-backed matrices
Extends beachmat to support initialization of tatami matrices from HDF5-backed arrays. This allows C++ code in downstream packages to directly call the HDF5 C/C++ library to access array data, without the need for block processing via DelayedArray. Some utilities are also provided for direct creation of an in-memory tatami matrix from a HDF5 file.
Last updated
datarepresentationdataimportinfrastructurecurlopensslcpp
4.78 score 9 scripts 383 downloadsscreenCounter - Counting Reads in High-Throughput Sequencing Screens
Provides functions for counting reads from high-throughput sequencing screen data (e.g., CRISPR, shRNA) to quantify barcode abundance. Currently supports single barcodes in single- or paired-end data, and combinatorial barcodes in paired-end data.
Last updated
crispralignmentfunctionalgenomicsfunctionalpredictionzlibcpp
4.78 score 4 stars 15 scripts 324 downloadsalabaster.string - Save and Load Biostrings to/from File
Save Biostrings objects to file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
4.78 score 2 dependents 5 scriptsclevRvis - Visualization Techniques for Clonal Evolution
clevRvis provides a set of visualization techniques for clonal evolution. These include shark plots, dolphin plots and plaice plots. Algorithms for time point interpolation as well as therapy effect estimation are provided. Phylogeny-aware color coding is implemented. A shiny-app for generating plots interactively is additionally provided.
Last updated
softwareshinyappsvisualization
4.78 score 6 stars 7 scripts 270 downloadsOGRE - Calculate, visualize and analyse overlap between genomic regions
OGRE calculates overlap between user defined genomic region datasets. Any regions can be supplied i.e. genes, SNPs, or reads from sequencing experiments. Key numbers help analyse the extend of overlaps which can also be visualized at a genomic level.
Last updated
softwareworkflowstepbiologicalquestionannotationmetagenomicsvisualizationsequencing
4.78 score 2 stars 4 scripts 318 downloadsDNAfusion - Identification of gene fusions using paired-end sequencing
DNAfusion can identify gene fusions such as EML4-ALK based on paired-end sequencing results. This package was developed using position deduplicated BAM files generated with the AVENIO Oncology Analysis Software. These files are made using the AVENIO ctDNA surveillance kit and Illumina Nextseq 500 sequencing. This is a targeted hybridization NGS approach and includes ALK-specific but not EML4-specific probes.
Last updated
targetedresequencinggeneticsgenefusiondetectionsequencingbioconductor-packagecirculating-tumor-dnagene-fusionliquid-biopsynext-generation-sequencingtargeted-sequencingvariant-calling
4.75 score 4 stars 14 scripts 293 downloadsGEOfastq - Downloads ENA Fastqs With GEO Accessions
GEOfastq is used to download fastq files from the European Nucleotide Archive (ENA) starting with an accession from the Gene Expression Omnibus (GEO). To do this, sample metadata is retrieved from GEO and the Sequence Read Archive (SRA). SRA run accessions are then used to construct FTP and aspera download links for fastq files generated by the ENA.
Last updated
rnaseqdataimportbioinformaticsfastqgene-expressiongeorna-seq
4.75 score 4 stars 14 scripts 319 downloadsepiregulon.extra - Companion package to epiregulon with additional plotting, differential and graph functions
Gene regulatory networks model the underlying gene regulation hierarchies that drive gene expression and observed phenotypes. Epiregulon infers TF activity in single cells by constructing a gene regulatory network (regulons). This is achieved through integration of scATAC-seq and scRNA-seq data and incorporation of public bulk TF ChIP-seq data. Links between regulatory elements and their target genes are established by computing correlations between chromatin accessibility and gene expressions.
Last updated
generegulationnetworkgeneexpressiontranscriptionchiponchipdifferentialexpressiongenetargetnormalizationgraphandnetwork
4.73 score 18 scripts 312 downloadsregioneReloaded - RegioneReloaded: Multiple Association for Genomic Region Sets
RegioneReloaded is a package that allows simultaneous analysis of associations between genomic region sets, enabling clustering of data and the creation of ready-to-publish graphs. It takes over and expands on all the features of its predecessor regioneR. It also incorporates a strategy to improve p-value calculations and normalize z-scores coming from multiple analysis to allow for their direct comparison. RegioneReloaded builds upon regioneR by adding new plotting functions for obtaining publication-ready graphs.
Last updated
geneticschipseqdnaseqmethylseqcopynumbervariationclusteringmultiplecomparison
4.73 score 5 stars 27 scripts 330 downloadsmultistateQTL - Toolkit for the analysis of multi-state QTL data
A collection of tools for doing various analyses of multi-state QTL data, with a focus on visualization and interpretation. The package 'multistateQTL' contains functions which can remove or impute missing data, identify significant associations, as well as categorise features into global, multi-state or unique. The analysis results are stored in a 'QTLExperiment' object, which is based on the 'SummarisedExperiment' framework.
Last updated
functionalgenomicsgeneexpressionsequencingvisualizationsnpsoftware
4.71 score 2 stars 17 scripts 278 downloads
GRaNIE - GRaNIE: Reconstruction cell type specific gene regulatory networks including enhancers using single-cell or bulk chromatin accessibility and RNA-seq data
Genetic variants associated with diseases often affect non-coding regions, thus likely having a regulatory role. To understand the effects of genetic variants in these regulatory regions, identifying genes that are modulated by specific regulatory elements (REs) is crucial. The effect of gene regulatory elements, such as enhancers, is often cell-type specific, likely because the combinations of transcription factors (TFs) that are regulating a given enhancer have cell-type specific activity. This TF activity can be quantified with existing tools such as diffTF and captures differences in binding of a TF in open chromatin regions. Collectively, this forms a gene regulatory network (GRN) with cell-type and data-specific TF-RE and RE-gene links. Here, we reconstruct such a GRN using single-cell or bulk RNAseq and open chromatin (e.g., using ATACseq or ChIPseq for open chromatin marks) and optionally (Capture) Hi-C data. Our network contains different types of links, connecting TFs to regulatory elements, the latter of which is connected to genes in the vicinity or within the same chromatin domain (TAD). We use a statistical framework to assign empirical FDRs and weights to all links using a permutation-based approach.
Last updated
softwaregeneexpressiongeneregulationnetworkinferencegenesetenrichmentbiomedicalinformaticsgeneticstranscriptomicsatacseqrnaseqgraphandnetworkregressiontranscriptionchipseq
4.71 score 17 scripts 400 downloadsDegCre - Probabilistic association of DEGs to CREs from differential data
DegCre generates associations between differentially expressed genes (DEGs) and cis-regulatory elements (CREs) based on non-parametric concordance between differential data. The user provides GRanges of DEG TSS and CRE regions with differential p-value and optionally log-fold changes and DegCre returns an annotated Hits object with associations and their calculated probabilities. Additionally, the package provides functionality for visualization and conversion to other formats.
Last updated
geneexpressiongeneregulationatacseqchipseqdnaseseqrnaseq
4.70 score 5 stars 5 scripts 262 downloadsconsICA - consensus Independent Component Analysis
consICA implements a data-driven deconvolution method – consensus independent component analysis (ICA) to decompose heterogeneous omics data and extract features suitable for patient diagnostics and prognostics. The method separates biologically relevant transcriptional signals from technical effects and provides information about the cellular composition and biological processes. The implementation of parallel computing in the package ensures efficient analysis of modern multicore systems.
Last updated
technologystatisticalmethodsequencingrnaseqtranscriptomicsclassificationfeatureextraction
4.70 score 3 scriptsRAREsim - Simulation of Rare Variant Genetic Data
Haplotype simulations of rare variant genetic data that emulates real data can be performed with RAREsim. RAREsim uses the expected number of variants in MAC bins - either as provided by default parameters or estimated from target data - and an abundance of rare variants as simulated HAPGEN2 to probabilistically prune variants. RAREsim produces haplotypes that emulate real sequencing data with respect to the total number of variants, allele frequency spectrum, haplotype structure, and variant annotation.
Last updated
geneticssoftwarevariantannotationsequencing
4.70 score 5 stars 10 scripts 274 downloadsADAPT - Analysis of Microbiome Differential Abundance by Pooling Tobit Models
ADAPT carries out differential abundance analysis for microbiome metagenomics data in phyloseq format. It has two innovations. One is to treat zero counts as left censored and use Tobit models for log count ratios. The other is an innovative way to find non-differentially abundant taxa as reference, then use the reference taxa to find the differentially abundant ones.
Last updated
differentialexpressionmicrobiomenormalizationsequencingmetagenomicssoftwaremultiplecomparisonopenblascpp
4.67 score 47 scripts 221 downloadsRbwa - R wrapper for BWA-backtrack and BWA-MEM aligners
Provides an R wrapper for BWA alignment algorithms. Both BWA-backtrack and BWA-MEM are available. Convenience function to build a BWA index from a reference genome is also provided. Currently not supported for Windows machines.
Last updated
sequencingalignmentbwabwa-memdna-sequencesdna-sequencing
4.67 score 1 stars 1 dependents 13 scripts 313 downloadsmiaDash - Dashboard for the interactive analysis and exploration of microbiome data
miaDash provides a Graphical User Interface for the exploration of microbiome data. This way, no knowledge of programming is required to perform analyses. Datasets can be imported, manipulated, analysed and visualised with a user-friendly interface.
Last updated
microbiomesoftwarevisualizationguishinyappsdataimportbioinformaticsdashboardiseemiashinyvisualisationwebapp
4.65 score 1 stars 9 scripts 250 downloadsCTexploreR - Explores Cancer Testis Genes
The CTexploreR package re-defines the list of Cancer Testis/Germline (CT) genes. It is based on publicly available RNAseq databases (GTEx, CCLE and TCGA) and summarises CT genes' main characteristics. Several visualisation functions allow to explore their expression in different types of tissues and cancer cells, or to inspect the methylation status of their promoters in normal tissues.
Last updated
transcriptomicsepigeneticsdifferentialexpressiongeneexpressiondnamethylationexperimenthubsoftwaredataimportbioconductor
4.65 score 10 scripts 258 downloadshicVennDiagram - Venn Diagram for genomic interaction data
A package to generate high-resolution Venn and Upset plots for genomic interaction data from HiC, ChIA-PET, HiChIP, PLAC-Seq, Hi-TrAC, HiCAR and etc. The package generates plots specifically crafted to eliminate the deceptive visual representation caused by the counts method.
Last updated
dna3dstructurehicvisualization
4.65 score 15 scripts 279 downloadsDCATS - Differential Composition Analysis Transformed by a Similarity matrix
Methods to detect the differential composition abundances between conditions in singel-cell RNA-seq experiments, with or without replicates. It aims to correct bias introduced by missclaisification and enable controlling of confounding covariates. To avoid the influence of proportion change from big cell types, DCATS can use either total cell number or specific reference group as normalization term.
Last updated
singlecellnormalization
4.65 score 45 scripts 304 downloadssSNAPPY - Single Sample directioNAl Pathway Perturbation analYsis
A single sample pathway perturbation testing method for RNA-seq data. The method propagates changes in gene expression down gene-set topologies to compute single-sample directional pathway perturbation scores that reflect potential direction of change. Perturbation scores can be used to test significance of pathway perturbation at both individual-sample and treatment levels.
Last updated
softwaregeneexpressiongenesetenrichmentgenesignaling
4.65 score 1 stars 15 scripts 378 downloadsrnaEditr - Statistical analysis of RNA editing sites and hyper-editing regions
RNAeditr analyzes site-specific RNA editing events, as well as hyper-editing regions. The editing frequencies can be tested against binary, continuous or survival outcomes. Multiple covariate variables as well as interaction effects can also be incorporated in the statistical models.
Last updated
genetargetepigeneticsdimensionreductionfeatureextractionregressionsurvivalrnaseq
4.65 score 3 stars 9 scripts 363 downloadsmagpie - MeRIP-Seq data Analysis for Genomic Power Investigation and Evaluation
This package aims to perform power analysis for the MeRIP-seq study. It calculates FDR, FDC, power, and precision under various study design parameters, including but not limited to sample size, sequencing depth, and testing method. It can also output results into .xlsx files or produce corresponding figures of choice.
Last updated
epitranscriptomicsdifferentialmethylationsequencingrnaseqsoftware
4.63 score 43 scripts 356 downloadsSpatialOmicsOverlay - Spatial Overlay for Omic Data from Nanostring GeoMx Data
Tools for NanoString Technologies GeoMx Technology. Package to easily graph on top of an OME-TIFF image. Plotting annotations can range from tissue segment to gene expression.
Last updated
geneexpressiontranscriptioncellbasedassaysdataimporttranscriptomicsproteomicsproprietaryplatformsrnaseqspatialdatarepresentationvisualizationopenjdk
4.62 score 21 scripts 329 downloadsTENET - R package for TENET (Tracing regulatory Element Networks using Epigenetic Traits) to identify key transcription factors
TENET identifies key transcription factors (TFs) and regulatory elements (REs) linked to a specific cell type by finding significantly correlated differences in gene expression and RE DNA methylation between case and control input datasets, and identifying the top genes by number of significant RE DNA methylation site links. It also includes many tools for visualization and analysis of the results, including plots displaying and comparing methylation and expression data and methylation site link counts, survival analysis, TF motif searching in the vicinity of linked RE DNA methylation sites, custom TAD and peak overlap analysis, and UCSC Genome Browser track file generation. A utility function is also provided to download methylation, expression, and patient survival data from The Cancer Genome Atlas (TCGA) for use in TENET or other analyses.
Last updated
softwarebiomedicalinformaticscellbiologygeneticsepigeneticsmultiplecomparisongeneexpressiondifferentialexpressiondnamethylationdifferentialmethylationmethylationarraysequencingmethylseqrnaseqfunctionalgenomicsgeneregulationgenetargethistonemodificationtranscriptiontranscriptomicssurvivalvisualization
4.60 score 1 stars 20 scripts 274 downloadsXeniumIO - Import and represent Xenium data from the 10X Xenium Analyzer
The package allows users to readily import spatial data obtained from the 10X Xenium Analyzer pipeline. Supported formats include 'parquet', 'h5', and 'mtx' files. The package mainly represents data as SpatialExperiment objects.
Last updated
softwareinfrastructuredataimportsinglecellspatialu24ca289073
4.60 score 3 scripts 222 downloadsELViS - An R Package for Estimating Copy Number Levels of Viral Genome Segments Using Base-Resolution Read Depth Profile
Base-resolution copy number analysis of viral genome. Utilizes base-resolution read depth data over viral genome to find copy number segments with two-dimensional segmentation approach. Provides publish-ready figures, including histograms of read depths, coverage line plots over viral genome annotated with copy number change events and viral genes, and heatmaps showing multiple types of data with integrative clustering of samples.
Last updated
copynumbervariationcoveragegenomicvariationbiomedicalinformaticssequencingnormalizationvisualizationclustering
4.60 score 8 scripts 236 downloadsSplineDV - Differential Variability (DV) analysis for single-cell RNA sequencing data. (e.g. Identify Differentially Variable Genes across two experimental conditions)
A spline based scRNA-seq method for identifying differentially variable (DV) genes across two experimental conditions. Spline-DV constructs a 3D spline from 3 key gene statistics: mean expression, coefficient of variance, and dropout rate. This is done for both conditions. The 3D spline provides the “expected” behavior of genes in each condition. The distance of the observed mean, CV and dropout rate of each gene from the expected 3D spline is used to measure variability. As the final step, the spline-DV method compares the variabilities of each condition to identify differentially variable (DV) genes.
Last updated
softwaresinglecellsequencingdifferentialexpressionrnaseqgeneexpressiontranscriptomicsfeatureextraction
4.60 score 4 stars 2 scripts 239 downloadsEpipwR - Efficient Power Analysis for EWAS with Continuous or Binary Outcomes
A quasi-simulation based approach to performing power analysis for EWAS (Epigenome-wide association studies) with continuous or binary outcomes. 'EpipwR' relies on empirical EWAS datasets to determine power at specific sample sizes while keeping computational cost low. EpipwR can be run with a variety of standard statistical tests, controlling for either a false discovery rate or a family-wise type I error rate.
Last updated
epigeneticsexperimentaldesign
4.60 score 2 stars 2 scripts 222 downloadsHoloFoodR - R interface to EBI HoloFood resource
Utility package to facilitate integration and analysis of EBI HoloFood data in R. This package streamlines access to the resource, allowing for direct loading of data into formats optimized for downstream analytics.
Last updated
softwareinfrastructuredataimportmicrobiomemicrobiomedata
4.60 score 2 stars 5 scripts 258 downloadssaseR - Scalable Aberrant Splicing and Expression Retrieval
saseR is a highly performant and fast framework for aberrant expression and splicing analyses. The main functions are: \itemize{ \item \code{\link{BamtoAspliCounts}} - Process BAM files to ASpli counts \item \code{\link{convertASpli}} - Get gene, bin or junction counts from ASpli SummarizedExperiment \item \code{\link{calculateOffsets}} - Create an offsets assays for aberrant expression or splicing analysis \item \code{\link{saseRfindEncodingDim}} - Estimate the optimal number of latent factors to include when estimating the mean expression \item \code{\link{saseRfit}} - Parameter estimation of the negative binomial distribution and compute p-values for aberrant expression and splicing } For information upon how to use these functions, check out our vignette at \url{https://github.com/statOmics/saseR/blob/main/vignettes/Vignette.Rmd} and the saseR paper: Segers, A. et al. (2023). Juggling offsets unlocks RNA-seq tools for fast scalable differential usage, aberrant splicing and expression analyses. bioRxiv. \url{https://doi.org/10.1101/2023.06.29.547014}.
Last updated
differentialexpressiondifferentialsplicingregressiongeneexpressionalternativesplicingrnaseqsequencingsoftware
4.60 score 4 stars 7 scripts 276 downloadsregionalpcs - Summarizing Regional Methylation with Regional Principal Components Analysis
Functions to summarize DNA methylation data using regional principal components. Regional principal components are computed using principal components analysis within genomic regions to summarize the variability in methylation levels across CpGs. The number of principal components is chosen using either the Marcenko-Pasteur or Gavish-Donoho method to identify relevant signal in the data.
Last updated
dnamethylationdifferentialmethylationstatisticalmethodsoftwaremethylationarray
4.60 score 4 stars 8 scripts 278 downloadspairedGSEA - Paired DGE and DGS analysis for gene set enrichment analysis
pairedGSEA makes it simple to run a paired Differential Gene Expression (DGE) and Differencital Gene Splicing (DGS) analysis. The package allows you to store intermediate results for further investiation, if desired. pairedGSEA comes with a wrapper function for running an Over-Representation Analysis (ORA) and functionalities for plotting the results.
Last updated
differentialexpressionalternativesplicingdifferentialsplicinggeneexpressionimmunooncologygenesetenrichmentpathwaysrnaseqsoftwaretranscription
4.60 score 4 stars 5 scripts 282 downloadscytofQC - Labels normalized cells for CyTOF data and assigns probabilities for each label
cytofQC is a package for initial cleaning of CyTOF data. It uses a semi-supervised approach for labeling cells with their most likely data type (bead, doublet, debris, dead) and the probability that they belong to each label type. This package does not remove data from the dataset, but provides labels and information to aid the data user in cleaning their data. Our algorithm is able to distinguish between doublets and large cells.
Last updated
softwaresinglecellannotation
4.60 score 2 stars 8 scripts 296 downloadsHiCool - HiCool
HiCool provides an R interface to process and normalize Hi-C paired-end fastq reads into .(m)cool files. .(m)cool is a compact, indexed HDF5 file format specifically tailored for efficiently storing HiC-based data. On top of processing fastq reads, HiCool provides a convenient reporting function to generate shareable reports summarizing Hi-C experiments and including quality controls.
Last updated
hicdna3dstructuredataimport
4.60 score 2 stars 8 scripts 234 downloadsterraTCGAdata - OpenAccess TCGA Data on Terra as MultiAssayExperiment
Leverage the existing open access TCGA data on Terra with well-established Bioconductor infrastructure. Make use of the Terra data model without learning its complexities. With a few functions, you can copy / download and generate a MultiAssayExperiment from the TCGA example workspaces provided by Terra.
Last updated
softwareinfrastructuredataimportbioconductor-packageu24hg010263
4.60 score 9 scripts 290 downloadsqmtools - Quantitative Metabolomics Data Processing Tools
The qmtools (quantitative metabolomics tools) package provides basic tools for processing quantitative metabolomics data with the standard SummarizedExperiment class. This includes functions for imputation, normalization, feature filtering, feature clustering, dimension-reduction, and visualization to help users prepare data for statistical analysis. This package also offers a convenient way to compute empirical Bayes statistics for which metabolic features are different between two sets of study samples. Several functions in this package could also be used in other types of omics data.
Last updated
metabolomicspreprocessingnormalizationdimensionreductionmassspectrometry
4.60 score 2 stars 8 scripts 350 downloadsMBECS - Evaluation and correction of batch effects in microbiome data-sets
The Microbiome Batch Effect Correction Suite (MBECS) provides a set of functions to evaluate and mitigate unwated noise due to processing in batches. To that end it incorporates a host of batch correcting algorithms (BECA) from various packages. In addition it offers a correction and reporting pipeline that provides a preliminary look at the characteristics of a data-set before and after correcting for batch effects.
Last updated
batcheffectmicrobiomereportwritingvisualizationnormalizationqualitycontrol
4.60 score 4 stars 8 scripts 393 downloadsCAEN - Category encoding method for selecting feature genes for the classification of single-cell RNA-seq
With the development of high-throughput techniques, more and more gene expression analysis tend to replace hybridization-based microarrays with the revolutionary technology.The novel method encodes the category again by employing the rank of samples for each gene in each class. We then consider the correlation coefficient of gene and class with rank of sample and new rank of category. The highest correlation coefficient genes are considered as the feature genes which are most effective to classify the samples.
Last updated
differentialexpressionsequencingclassificationrnaseqatacseqsinglecellgeneexpressionripseq
4.60 score 7 scripts 312 downloadsnempi - Inferring unobserved perturbations from gene expression data
Takes as input an incomplete perturbation profile and differential gene expression in log odds and infers unobserved perturbations and augments observed ones. The inference is done by iteratively inferring a network from the perturbations and inferring perturbations from the network. The network inference is done by Nested Effects Models.
Last updated
softwaregeneexpressiondifferentialexpressiondifferentialmethylationgenesignalingpathwaysnetworkclassificationneuralnetworknetworkinferenceatacseqdnaseqrnaseqpooledscreenscrisprsinglecellsystemsbiology
4.60 score 2 stars 2 scripts 321 downloadsbnem - Training of logical models from indirect measurements of perturbation experiments
bnem combines the use of indirect measurements of Nested Effects Models (package mnem) with the Boolean networks of CellNOptR. Perturbation experiments of signalling nodes in cells are analysed for their effect on the global gene expression profile. Those profiles give evidence for the Boolean regulation of down-stream nodes in the network, e.g., whether two parents activate their child independently (OR-gate) or jointly (AND-gate).
Last updated
pathwayssystemsbiologynetworkinferencenetworkgeneexpressiongeneregulationpreprocessing
4.60 score 2 stars 7 scripts 382 downloadsprofileplyr - Visualization and annotation of read signal over genomic ranges with profileplyr
Quick and straightforward visualization of read signal over genomic intervals is key for generating hypotheses from sequencing data sets (e.g. ChIP-seq, ATAC-seq, bisulfite/methyl-seq). Many tools both inside and outside of R and Bioconductor are available to explore these types of data, and they typically start with a bigWig or BAM file and end with some representation of the signal (e.g. heatmap). profileplyr leverages many Bioconductor tools to allow for both flexibility and additional functionality in workflows that end with visualization of the read signal.
Last updated
chipseqdataimportsequencingchiponchipcoverage
4.59 score 97 scripts 506 downloadsCCAFE - Case Control Allele Frequency Estimation
Functions to reconstruct case and control AFs from summary statistics. One function uses OR, NCase, NControl, and SE(log(OR)). The second function uses OR, NCase, NControl, and AF for the whole sample.
Last updated
genomewideassociationcomparativegenomicsgeneticspreprocessingsnpsoftwarewholegenome
4.56 score 1 stars 18 scripts 276 downloadsDNABarcodes - A tool for creating and analysing DNA barcodes used in Next Generation Sequencing multiplexing experiments
The package offers a function to create DNA barcode sets capable of correcting insertion, deletion, and substitution errors. Existing barcodes can be analysed regarding their minimal, maximal and average distances between barcodes. Finally, reads that start with a (possibly mutated) barcode can be demultiplexed, i.e., assigned to their original reference barcode.
Last updated
preprocessingsequencingcppopenmp
4.55 score 59 scripts 556 downloads
GraphExperiment - S4 Class for Quantitative Data and Associated Networks
GraphExperiment provides users and developers with an S4 class that extends `SingleCellExperiment` by offering infrastructure to store and retrieve networks (`igraph` objects) representing how assay features and/or observations are associated with each other. The class was designed to store networks inferred from high-dimensional quantitative data, with feature-feature networks including gene coexpression networks (GCNs), gene regulatory networks (GRNs), and co-abundance networks (from proteomics and metabolomics), and observation-observation network including cell-cell distances, species-species relationships, and sample-sample similarities.
Last updated
datarepresentationdataimportinfrastructuregeneexpressiontranscriptomicsnetworksinglecellbioconductorbioinformaticsoop
4.54 score 1 stars 2 scripts 262 downloadsClusterFoldSimilarity - Calculate similarity of clusters from different single cell samples using foldchanges
This package calculates a similarity coefficient using the fold changes of shared features (e.g. genes) among clusters of different samples/batches/datasets. The similarity coefficient is calculated using the dot-product (Hadamard product) of every pairwise combination of Fold Changes between a source cluster i of sample/dataset n and all the target clusters j in sample/dataset m
Last updated
singlecellclusteringfeatureextractiongraphandnetworkgenetargetrnaseq
4.54 score 23 scripts 275 downloadsmbQTL - mbQTL: A package for SNP-Taxa mGWAS analysis
mbQTL is a statistical R package for simultaneous 16srRNA,16srDNA (microbial) and variant, SNP, SNV (host) relationship, correlation, regression studies. We apply linear, logistic and correlation based statistics to identify the relationships of taxa, genus, species and variant, SNP, SNV in the infected host. We produce various statistical significance measures such as P values, FDR, BC and probability estimation to show significance of these relationships. Further we provide various visualization function for ease and clarification of the results of these analysis. The package is compatible with dataframe, MRexperiment and text formats.
Last updated
snpmicrobiomewholegenomemetagenomicsstatisticalmethodregression
4.53 score 1 stars 34 scripts 279 downloadsgeyser - Gene Expression displaYer of SummarizedExperiment in R
Lightweight Expression displaYer (plotter / viewer) of SummarizedExperiment object in R. This package provides a quick and easy Shiny-based GUI to empower a user to use a SummarizedExperiment object to view
Last updated
softwareshinyappsguigeneexpression
4.52 score 22 scripts 220 downloadstadar - Transcriptome Analysis of Differential Allelic Representation
This package provides functions to standardise the analysis of Differential Allelic Representation (DAR). DAR compromises the integrity of Differential Expression analysis results as it can bias expression, influencing the classification of genes (or transcripts) as being differentially expressed. DAR analysis results in an easy-to-interpret value between 0 and 1 for each genetic feature of interest, where 0 represents identical allelic representation and 1 represents complete diversity. This metric can be used to identify features prone to false-positive calls in Differential Expression analysis, and can be leveraged with statistical methods to alleviate the impact of such artefacts on RNA-seq data.
Last updated
sequencingrnaseqsnpgenomicvariationvariantannotationdifferentialexpression
4.52 score 1 stars 11 scripts 272 downloadsalabaster.vcf - Save and Load Variant Data to/from File
Save variant calling SummarizedExperiment to file and load them back as VCF objects. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
4.52 score 1 dependents 11 scripts 288 downloadsNetActivity - Compute gene set scores from a deep learning framework
#' NetActivity enables to compute gene set scores from previously trained sparsely-connected autoencoders. The package contains a function to prepare the data (`prepareSummarizedExperiment`) and a function to compute the gene set scores (`computeGeneSetScores`). The package `NetActivityData` contains different pre-trained models to be directly applied to the data. Alternatively, the users might use the package to compute gene set scores using custom models.
Last updated
rnaseqmicroarraytranscriptionfunctionalgenomicsgogeneexpressionpathwayssoftware
4.52 score 33 scripts 301 downloadstidyFlowCore - tidyFlowCore: Bringing flowCore to the tidyverse
tidyFlowCore bridges the gap between flow cytometry analysis using the flowCore Bioconductor package and the tidy data principles advocated by the tidyverse. It provides a suite of dplyr-, ggplot2-, and tidyr-like verbs specifically designed for working with flowFrame and flowSet objects as if they were tibbles; however, your data remain flowCore data structures under this layer of abstraction. tidyFlowCore enables intuitive and streamlined analysis workflows that can leverage both the Bioconductor and tidyverse ecosystems for cytometry data.
Last updated
singlecellflowcytometryinfrastructure
4.51 score 2 stars 16 scripts 225 downloadsHarmonizR - Handles missing values and makes more data available
An implementation, which takes input data and makes it available for proper batch effect removal by ComBat or Limma. The implementation appropriately handles missing values by dissecting the input matrix into smaller matrices with sufficient data to feed the ComBat or limma algorithm. The adjusted data is returned to the user as a rebuild matrix. The implementation is meant to make as much data available as possible with minimal data loss.
Last updated
batcheffect
4.51 score 32 scripts 300 downloadsADImpute - Adaptive Dropout Imputer (ADImpute)
Single-cell RNA sequencing (scRNA-seq) methods are typically unable to quantify the expression levels of all genes in a cell, creating a need for the computational prediction of missing values (‘dropout imputation’). Most existing dropout imputation methods are limited in the sense that they exclusively use the scRNA-seq dataset at hand and do not exploit external gene-gene relationship information. Here we propose two novel methods: a gene regulatory network-based approach using gene-gene relationships learnt from external data and a baseline approach corresponding to a sample-wide average. ADImpute can implement these novel methods and also combine them with existing imputation methods (currently supported: DrImpute, SAVER). ADImpute can learn the best performing method per gene and combine the results from different methods into an ensemble.
Last updated
geneexpressionnetworkpreprocessingsequencingsinglecelltranscriptomics
4.51 score 16 scripts 479 downloadsReducedExperiment - Containers and tools for dimensionally-reduced -omics representations
Provides SummarizedExperiment-like containers for storing and manipulating dimensionally-reduced assay data. The ReducedExperiment classes allow users to simultaneously manipulate their original dataset and their decomposed data, in addition to other method-specific outputs like feature loadings. Implements utilities and specialised classes for the application of stabilised independent component analysis (sICA) and weighted gene correlation network analysis (WGCNA).
Last updated
geneexpressioninfrastructuredatarepresentationsoftwaredimensionreductionnetworkbioconductor-packagebioinformaticsdimensionality-reduction
4.48 score 3 stars 8 scripts 230 downloadsbeachmat.tiledb - beachmat bindings for TileDB-backed matrices
Extends beachmat to initialize tatami matrices from TileDB-backed arrays. This allows C++ code in downstream packages to directly call the TileDB C/C++ library to access array data, without the need for block processing via DelayedArray. Developers only need to import this package to automatically extend the capabilities of beachmat::initializeCpp to TileDBArray instances.
Last updated
datarepresentationdataimportinfrastructurecpp
4.48 score 4 scriptsPolySTest - PolySTest: Detection of differentially regulated features. Combined statistical testing for data with few replicates and missing values
The complexity of high-throughput quantitative omics experiments often leads to low replicates numbers and many missing values. We implemented a new test to simultaneously consider missing values and quantitative changes, which we combined with well-performing statistical tests for high confidence detection of differentially regulated features. The package contains functions to run the test and to visualize the results.
Last updated
massspectrometryproteomicssoftwaredifferentialexpressioncsvdifferential-protein-expression-profilingdsvexpression-datagene-expression-profileheat-maphistogramp-valueproteomics-experimentq-valuestatistical-modellingstatistics-and-probabilitysvg
4.48 score 10 scripts 239 downloadssquallms - Speedy quality assurance via lasso labeling for LC-MS data
squallms is a Bioconductor R package that implements a "semi-labeled" approach to untargeted mass spectrometry data. It pulls in raw data from mass-spec files to calculate several metrics that are then used to label MS features in bulk as high or low quality. These metrics of peak quality are then passed to a simple logistic model that produces a fully-labeled dataset suitable for downstream analysis.
Last updated
massspectrometrymetabolomicsproteomicslipidomicsshinyappsclassificationclusteringfeatureextractionprincipalcomponentregressionpreprocessingqualitycontrolvisualization
4.48 score 3 stars 7 scripts 207 downloads
HicAggR - Set of 3D genomic interaction analysis tools
This package provides a set of functions useful in the analysis of 3D genomic interactions. It includes the import of standard HiC data formats into R and HiC normalisation procedures. The main objective of this package is to improve the visualization and quantification of the analysis of HiC contacts through aggregation. The package allows to import 1D genomics data, such as peaks from ATACSeq, ChIPSeq, to create potential couples between features of interest under user-defined parameters such as distance between pairs of features of interest. It allows then the extraction of contact values from the HiC data for these couples and to perform Aggregated Peak Analysis (APA) for visualization, but also to compare normalized contact values between conditions. Overall the package allows to integrate 1D genomics data with 3D genomics data, providing an easy access to HiC contact values.
Last updated
softwarehicdataimportdatarepresentationnormalizationvisualizationdna3dstructureatacseqchipseqdnaseseqrnaseq
4.48 score 1 stars 4 scripts 289 downloadsalabaster.bumpy - Save and Load BumpyMatrices to/from file
Save BumpyMatrix objects into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
4.48 score 1 dependents 8 scripts 294 downloadsalabaster.mae - Load and Save MultiAssayExperiments
Save MultiAssayExperiments into file artifacts, and load them back into memory. This is a more portable alternative to serialization of such objects into RDS files. Each artifact is associated with metadata for further interpretation; downstream applications can enrich this metadata with context-specific properties.
Last updated
dataimportdatarepresentation
4.48 score 1 dependents 10 scripts 348 downloads
magrene - Motif Analysis In Gene Regulatory Networks
magrene allows the identification and analysis of graph motifs in (duplicated) gene regulatory networks (GRNs), including lambda, V, PPI V, delta, and bifan motifs. GRNs can be tested for motif enrichment by comparing motif frequencies to a null distribution generated from degree-preserving simulated GRNs. Motif frequencies can be analyzed in the context of gene duplications to explore the impact of small-scale and whole-genome duplications on gene regulatory networks. Finally, users can calculate interaction similarity for gene pairs based on the Sorensen-Dice similarity index.
Last updated
softwaremotifdiscoverynetworkenrichmentsystemsbiologygraphandnetworkgene-regulatory-networkmotif-analysisnetwork-motifsnetwork-science
4.48 score 3 stars 7 scripts 283 downloadsoncoscanR - Secondary analyses of CNV data (HRD and more)
The software uses the copy number segments from a text file and identifies all chromosome arms that are globally altered and computes various genome-wide scores. The following HRD scores (characteristic of BRCA-mutated cancers) are included: LST, HR-LOH, nLST and gLOH. the package is tailored for the ThermoFisher Oncoscan assay analyzed with their Chromosome Alteration Suite (ChAS) but can be adapted to any input.
Last updated
copynumbervariationmicroarraysoftware
4.48 score 3 stars 9 scripts 282 downloadsGenomicInteractionNodes - A R/Bioconductor package to detect the interaction nodes from HiC/HiChIP/HiCAR data
The GenomicInteractionNodes package can import interactions from bedpe file and define the interaction nodes, the genomic interaction sites with multiple interaction loops. The interaction nodes is a binding platform regulates one or multiple genes. The detected interaction nodes will be annotated for downstream validation.
Last updated
hicsequencingsoftware
4.48 score 3 scripts 292 downloadscomapr - Crossover analysis and genetic map construction
comapr detects crossover intervals for single gametes from their haplotype states sequences and stores the crossovers in GRanges object. The genetic distances can then be calculated via the mapping functions using estimated crossover rates for maker intervals. Visualisation functions for plotting interval-based genetic map or cumulative genetic distances are implemented, which help reveal the variation of crossovers landscapes across the genome and across individuals.
Last updated
softwaresinglecellvisualizationgenetics
4.48 score 4 scripts 308 downloadsramr - Detection of Rare Aberrantly Methylated Regions in Array and NGS Data
ramr is an R package for detection of epimutations (i.e., infrequent aberrant DNA methylation events) in large data sets obtained by methylation profiling using array or high-throughput methylation sequencing. In addition, package provides functions to visualize found aberrantly methylated regions (AMRs), to generate sets of all possible regions to be used as reference sets for enrichment analysis, and to generate biologically relevant test data sets for performance evaluation of AMR/DMR search algorithms.
Last updated
dnamethylationdifferentialmethylationepigeneticsmethylationarraymethylseqaberrant-methylationbioconductordna-methylationepimutationmethylation-microarraysnext-generation-sequencingcppopenmp
4.48 score 7 scripts 374 downloadsRLassoCox - A reweighted Lasso-Cox by integrating gene interaction information
RLassoCox is a package that implements the RLasso-Cox model proposed by Wei Liu. The RLasso-Cox model integrates gene interaction information into the Lasso-Cox model for accurate survival prediction and survival biomarker discovery. It is based on the hypothesis that topologically important genes in the gene interaction network tend to have stable expression changes. The RLasso-Cox model uses random walk to evaluate the topological weight of genes, and then highlights topologically important genes to improve the generalization ability of the Lasso-Cox model. The RLasso-Cox model has the advantage of identifying small gene sets with high prognostic performance on independent datasets, which may play an important role in identifying robust survival biomarkers for various cancer types.
Last updated
survivalregressiongeneexpressiongenepredictionnetwork
4.48 score 3 stars 2 scripts 268 downloadsSCFA - SCFA: Subtyping via Consensus Factor Analysis
Subtyping via Consensus Factor Analysis (SCFA) can efficiently remove noisy signals from consistent molecular patterns in multi-omics data. SCFA first uses an autoencoder to select only important features and then repeatedly performs factor analysis to represent the data with different numbers of factors. Using these representations, it can reliably identify cancer subtypes and accurately predict risk scores of patients.
Last updated
survivalclusteringclassification
4.48 score 3 stars 7 scripts 344 downloadscustomCMPdb - Customize and Query Compound Annotation Database
This package serves as a query interface for important community collections of small molecules, while also allowing users to include custom compound collections.
Last updated
softwarecheminformaticsannotationhubsoftware
4.48 score 1 stars 5 scripts 340 downloadsInformeasure - R implementation of information measures
This package consolidates a comprehensive set of information measurements, encompassing mutual information, conditional mutual information, interaction information, partial information decomposition, and part mutual information.
Last updated
geneexpressionnetworkinferencenetworksoftware
4.48 score 3 stars 6 scripts 296 downloadseasier - Estimate Systems Immune Response from RNA-seq data
This package provides a workflow for the use of EaSIeR tool, developed to assess patients' likelihood to respond to ICB therapies providing just the patients' RNA-seq data as input. We integrate RNA-seq data with different types of prior knowledge to extract quantitative descriptors of the tumor microenvironment from several points of view, including composition of the immune repertoire, and activity of intra- and extra-cellular communications. Then, we use multi-task machine learning trained in TCGA data to identify how these descriptors can simultaneously predict several state-of-the-art hallmarks of anti-cancer immune response. In this way we derive cancer-specific models and identify cancer-specific systems biomarkers of immune response. These biomarkers have been experimentally validated in the literature and the performance of EaSIeR predictions has been validated using independent datasets form four different cancer types with patients treated with anti-PD1 or anti-PDL1 therapy.
Last updated
geneexpressionsoftwaretranscriptionsystemsbiologypathwaysgenesetenrichmentimmunooncologyepigeneticsclassificationbiomedicalinformaticsregressionexperimenthubsoftware
4.45 score 28 scripts 416 downloadsclustSIGNAL - ClustSIGNAL: a spatial clustering method
clustSIGNAL: clustering of Spatially Informed Gene expression with Neighbourhood Adapted Learning. A tool for adaptively smoothing and clustering gene expression data. clustSIGNAL uses entropy to measure heterogeneity of cell neighbourhoods and performs a weighted, adaptive smoothing, where homogeneous neighbourhoods are smoothed more and heterogeneous neighbourhoods are smoothed less. This not only overcomes data sparsity but also incorporates spatial context into the gene expression data. The resulting smoothed gene expression data is used for clustering and could be used for other downstream analyses.
Last updated
clusteringsoftwaregeneexpressionspatialtranscriptomicssinglecell
4.43 score 6 stars 3 scripts 253 downloadsscDDboost - A compositional model to assess expression changes from single-cell rna-seq data
scDDboost is an R package to analyze changes in the distribution of single-cell expression data between two experimental conditions. Compared to other methods that assess differential expression, scDDboost benefits uniquely from information conveyed by the clustering of cells into cellular subtypes. Through a novel empirical Bayesian formulation it calculates gene-specific posterior probabilities that the marginal expression distribution is the same (or different) between the two conditions. The implementation in scDDboost treats gene-level expression data within each condition as a mixture of negative binomial distributions.
Last updated
singlecellsoftwareclusteringsequencinggeneexpressiondifferentialexpressionbayesiancpp
4.43 score 27 scripts 342 downloadsMacarron - Prioritization of potentially bioactive metabolic features from epidemiological and environmental metabolomics datasets
Macarron is a workflow for the prioritization of potentially bioactive metabolites from metabolomics experiments. Prioritization integrates strengths of evidences of bioactivity such as covariation with a known metabolite, abundance relative to a known metabolite and association with an environmental or phenotypic indicator of bioactivity. Broadly, the workflow consists of stratified clustering of metabolic spectral features which co-vary in abundance in a condition, transfer of functional annotations, estimation of relative abundance and differential abundance analysis to identify associations between features and phenotype/condition.
Last updated
sequencingmetabolomicscoveragefunctionalpredictionclustering
4.41 score 17 scripts 314 downloadsCARDspa - Spatially Informed Cell Type Deconvolution for Spatial Transcriptomics
CARD is a reference-based deconvolution method that estimates cell type composition in spatial transcriptomics based on cell type specific expression information obtained from a reference scRNA-seq data. A key feature of CARD is its ability to accommodate spatial correlation in the cell type composition across tissue locations, enabling accurate and spatially informed cell type deconvolution as well as refined spatial map construction. CARD relies on an efficient optimization algorithm for constrained maximum likelihood estimation and is scalable to spatial transcriptomics with tens of thousands of spatial locations and tens of thousands of genes.
Last updated
spatialsinglecelltranscriptomicsvisualizationopenblascppopenmp
4.38 score 16 scripts 414 downloadsspoon - Address the Mean-variance Relationship in Spatial Transcriptomics Data
This package addresses the mean-variance relationship in spatially resolved transcriptomics data. Precision weights are generated for individual observations using Empirical Bayes techniques. These weights are used to rescale the data and covariates, which are then used as input in spatially variable gene detection tools.
Last updated
spatialsinglecelltranscriptomicsgeneexpressionpreprocessing
4.38 score 24 scripts 237 downloadsBiocHail - basilisk and hail
Use hail via basilisk when appropriate, or via reticulate. This package can be used in terra.bio to interact with UK Biobank resources processed by hail.is.
Last updated
infrastructurebioconductorgeneticshail
4.38 score 6 stars 20 scripts 177 downloadsmumosa - Multi-Modal Single-Cell Analysis Methods
Assorted utilities for multi-modal analyses of single-cell datasets. Includes functions to combine multiple modalities for downstream analysis, perform MNN-based batch correction across multiple modalities, and to compute correlations between assay values for different modalities.
Last updated
immunooncologysinglecellrnaseq
4.38 score 16 scripts 358 downloads
epimutacions - Robust outlier identification for DNA methylation data
The package includes some statistical outlier detection methods for epimutations detection in DNA methylation data. The methods included in the package are MANOVA, Multivariate linear models, isolation forest, robust mahalanobis distance, quantile and beta. The methods compare a case sample with a suspected disease against a reference panel (composed of healthy individuals) to identify epimutations in the given case sample. It also contains functions to annotate and visualize the identified epimutations.
Last updated
dnamethylationbiologicalquestionpreprocessingstatisticalmethodnormalizationcpp
4.37 score 39 scripts 398 downloadsDepInfeR - Inferring tumor-specific cancer dependencies through integrating ex-vivo drug response assays and drug-protein profiling
DepInfeR integrates two experimentally accessible input data matrices: the drug sensitivity profiles of cancer cell lines or primary tumors ex-vivo (X), and the drug affinities of a set of proteins (Y), to infer a matrix of molecular protein dependencies of the cancers (ß). DepInfeR deconvolutes the protein inhibition effect on the viability phenotype by using regularized multivariate linear regression. It assigns a “dependence coefficient” to each protein and each sample, and therefore could be used to gain a causal and accurate understanding of functional consequences of genomic aberrations in a heterogeneous disease, as well as to guide the choice of pharmacological intervention for a specific cancer type, sub-type, or an individual patient. For more information, please read out preprint on bioRxiv: https://doi.org/10.1101/2022.01.11.475864.
Last updated
softwareregressionpharmacogeneticspharmacogenomicsfunctionalgenomics
4.36 score 1 stars 23 scripts 304 downloadsPepSetTest - Peptide Set Test
Peptide Set Test (PepSetTest) is a peptide-centric strategy to infer differentially expressed proteins in LC-MS/MS proteomics data. This test detects coordinated changes in the expression of peptides originating from the same protein and compares these changes against the rest of the peptidome. Compared to traditional aggregation-based approaches, the peptide set test demonstrates improved statistical power, yet controlling the Type I error rate correctly in most cases. This test can be valuable for discovering novel biomarkers and prioritizing drug targets, especially when the direct application of statistical analysis to protein data fails to provide substantial insights.
Last updated
differentialexpressionregressionproteomicsmassspectrometry
4.34 score 2 stars 11 scripts 236 downloadsginmappeR - Gene Identifier Mapper
Provides functionalities to translate gene or protein identifiers between state-of-art biological databases: CARD (<https://card.mcmaster.ca/>), NCBI Protein, Nucleotide and Gene (<https://www.ncbi.nlm.nih.gov/>), UniProt (<https://www.uniprot.org/>) and KEGG (<https://www.kegg.jp>). Also offers complementary functionality like NCBI identical proteins or UniProt similar genes clusters retrieval.
Last updated
annotationkegggeneticsthirdpartyclientsoftware
4.34 score 22 scripts 244 downloadsCepo - Cepo for the identification of differentially stable genes
Defining the identity of a cell is fundamental to understand the heterogeneity of cells to various environmental signals and perturbations. We present Cepo, a new method to explore cell identities from single-cell RNA-sequencing data using differential stability as a new metric to define cell identity genes. Cepo computes cell-type specific gene statistics pertaining to differential stable gene expression.
Last updated
classificationgeneexpressionsinglecellsoftwaresequencingdifferentialexpression
4.33 score 1 dependents 48 scripts 368 downloadsEnrichDO - a Global Weighted Model for Disease Ontology Enrichment Analysis
To implement disease ontology (DO) enrichment analysis, this package is designed and presents a double weighted model based on the latest annotations of the human genome with DO terms, by integrating the DO graph topology on a global scale. This package exhibits high accuracy that it can identify more specific DO terms, which alleviates the over enriched problem. The package includes various statistical models and visualization schemes for discovering the associations between genes and diseases from biological big data.
Last updated
annotationvisualizationgenesetenrichmentsoftware
4.32 score 14 scripts 248 downloadsalabaster.files - Wrappers to Save Common File Formats
Save common bioinformatics file formats within the alabaster framework. This includes BAM, BED, VCF, bigWig, bigBed, FASTQ, FASTA and so on. We save and load additional metadata for each file, and we support linkage between each file and its corresponding index.
Last updated
datarepresentationdataimport
4.32 score 21 scripts 256 downloadsshinyDSP - A Shiny App For Visualizing Nanostring GeoMx DSP Data
This package is a Shiny app for interactively analyzing and visualizing Nanostring GeoMX Whole Transcriptome Atlas data. Users have the option of exploring a sample data to explore this app's functionality. Regions of interest (ROIs) can be filtered based on any user-provided metadata. Upon taking two or more groups of interest, all pairwise and ANOVA-like testing are automatically performed. Available ouputs include PCA, Volcano plots, tables and heatmaps. Aesthetics of each output are highly customizable.
Last updated
differentialexpressiongeneexpressionshinyappsspatialtranscriptomics
4.30 score 1 stars 5 scripts 282 downloadschevreulShiny - Tools for managing SingleCellExperiment objects as projects
Tools for managing SingleCellExperiment objects as projects. Includes functions for analysis and visualization of single-cell data. Also included is a shiny app for visualization of pre-processed scRNA data. Supported by NIH grants R01CA137124 and R01EY026661 to David Cobrinik.
Last updated
coveragernaseqsequencingvisualizationgeneexpressiontranscriptionsinglecelltranscriptomicsnormalizationpreprocessingqualitycontroldimensionreductiondataimport
4.30 score 5 scripts 284 downloadstidysbml - Extract SBML's data into dataframes
Starting from one SBML file, it extracts information from each listOfCompartments, listOfSpecies and listOfReactions element by saving them into data frames. Each table provides one row for each entity (i.e. either compartment, species, reaction or speciesReference) and one set of columns for the attributes, one column for the content of the 'notes' subelement and one set of columns for the content of the 'annotation' subelement.
Last updated
graphandnetworknetworkpathwayssoftware
4.30 score 2 stars 3 scripts 237 downloadsiSEEfier - Streamlining the creation of initial states for starting an iSEE instance
iSEEfier provides a set of functionality to quickly and intuitively create, inspect, and combine initial configuration objects. These can be conveniently passed in a straightforward manner to the function call to launch iSEE() with the specified configuration. This package currently works seamlessly with the sets of panels provided by the iSEE and iSEEu packages, but can be extended to accommodate the usage of any custom panel (e.g. from iSEEde, iSEEpathways, or any panel developed independently by the user).
Last updated
cellbasedassaysclusteringdimensionreductionfeatureextractionguigeneexpressionimmunooncologyshinyappssinglecellsoftwaretranscriptiontranscriptomicsvisualization
4.30 score 7 scripts 244 downloadssmartid - Scoring and Marker Selection Method Based on Modified TF-IDF
This package enables automated selection of group specific signature, especially for rare population. The package is developed for generating specifc lists of signature genes based on Term Frequency-Inverse Document Frequency (TF-IDF) modified methods. It can also be used as a new gene-set scoring method or data transformation method. Multiple visualization functions are implemented in this package.
Last updated
softwaregeneexpressiontranscriptomics
4.30 score 1 stars 3 scripts 259 downloads
tpSVG - Thin plate models to detect spatially variable genes
The goal of `tpSVG` is to detect and visualize spatial variation in the gene expression for spatially resolved transcriptomics data analysis. Specifically, `tpSVG` introduces a family of count-based models, with generalizable parametric assumptions such as Poisson distribution or negative binomial distribution. In addition, comparing to currently available count-based model for spatially resolved data analysis, the `tpSVG` models improves computational time, and hence greatly improves the applicability of count-based models in SRT data analysis.
Last updated
spatialtranscriptomicsgeneexpressionsoftwarestatisticalmethoddimensionreductionregressionpreprocessingspatially-resolvespatially-variable-genes
4.30 score 2 stars 4 scripts 242 downloads
biocroxytest - Handle Long Tests in Bioconductor Packages
This package provides a roclet for roxygen2 that identifies and processes code blocks in your documentation marked with `@longtests`. These blocks should contain tests that take a long time to run and thus cannot be included in the regular test suite of the package. When you run `roxygen2::roxygenise` with the `longtests_roclet`, it will extract these long tests from your documentation and save them in a separate directory. This allows you to run these long tests separately from the rest of your tests, for example, on a continuous integration server that is set up to run long tests.
Last updated
softwareinfrastructure
4.30 score 2 stars 3 scripts 240 downloadsRvisdiff - Interactive Graphs for Differential Expression
Creates a muti-graph web page which allows the interactive exploration of differential analysis tests. The graphical web interface presents results as a table which is integrated with five interactive graphs: MA-plot, volcano plot, box plot, lines plot and cluster heatmap. Graphical aspect and information represented in the graphs can be customized by means of user controls. Final graphics can be exported as PNG format.
Last updated
softwarevisualizationrnaseqdatarepresentationdifferentialexpression
4.30 score 3 scripts 312 downloadsorthos - `orthos` is an R package for variance decomposition using conditional variational auto-encoders
`orthos` decomposes RNA-seq contrasts, for example obtained from a gene knock-out or compound treatment experiment, into unspecific and experiment-specific components. Original and decomposed contrasts can be efficiently queried against a large database of contrasts (derived from ARCHS4, https://maayanlab.cloud/archs4/) to identify similar experiments. `orthos` furthermore provides plotting functions to visualize the results of such a search for similar contrasts.
Last updated
rnaseqdifferentialexpressiongeneexpression
4.30 score 7 scripts 345 downloadsSVMDO - Identification of Tumor-Discriminating mRNA Signatures via Support Vector Machines Supported by Disease Ontology
It is an easy-to-use GUI using disease information for detecting tumor/normal sample discriminating gene sets from differentially expressed genes. Our approach is based on an iterative algorithm filtering genes with disease ontology enrichment analysis and wilk and wilks lambda criterion connected to SVM classification model construction. Along with gene set extraction, SVMDO also provides individual prognostic marker detection. The algorithm is designed for FPKM and RPKM normalized RNA-Seq transcriptome datasets.
Last updated
genesetenrichmentdifferentialexpressionguiclassificationrnaseqtranscriptomicssurvivalmachine-learningrna-seqshiny
4.30 score 4 scripts 263 downloadsFeatSeekR - FeatSeekR an R package for unsupervised feature selection
FeatSeekR performs unsupervised feature selection using replicated measurements. It iteratively selects features with the highest reproducibility across replicates, after projecting out those dimensions from the data that are spanned by the previously selected features. The selected a set of features has a high replicate reproducibility and a high degree of uniqueness.
Last updated
softwarestatisticalmethodfeatureextractionmassspectrometry
4.30 score 2 stars 5 scripts 266 downloadsmetabinR - Abundance and Compositional Based Binning of Metagenomes
Provide functions for performing abundance and compositional based binning on metagenomic samples, directly from FASTA or FASTQ files. Functions are implemented in Java and called via rJava. Parallel implementation that operates directly on input FASTA/FASTQ files for fast execution. Inputs may be file paths or Biostrings/ShortRead sequence objects; results are returned as a MetabinResult S4 object wrapping cluster assignments, algorithm parameters, and input metadata.
Last updated
classificationclusteringmicrobiomesequencingsoftwareopenjdk
4.30 score 2 stars 4 scripts 324 downloads
ccImpute - ccImpute: an accurate and scalable consensus clustering based approach to impute dropout events in the single-cell RNA-seq data (https://doi.org/10.1186/s12859-022-04814-8)
Dropout events make the lowly expressed genes indistinguishable from true zero expression and different than the low expression present in cells of the same type. This issue makes any subsequent downstream analysis difficult. ccImpute is an imputation algorithm that uses cell similarity established by consensus clustering to impute the most probable dropout events in the scRNA-seq datasets. ccImpute demonstrated performance which exceeds the performance of existing imputation approaches while introducing the least amount of new noise as measured by clustering performance characteristics on datasets with known cell identities.
Last updated
singlecellsequencingprincipalcomponentdimensionreductionclusteringrnaseqtranscriptomicsopenblascppopenmp
4.30 score 2 stars 7 scripts 311 downloadsATACseqTFEA - Transcription Factor Enrichment Analysis for ATAC-seq
Assay for Transpose-Accessible Chromatin using sequencing (ATAC-seq) is a technique to assess genome-wide chromatin accessibility by probing open chromatin with hyperactive mutant Tn5 Transposase that inserts sequencing adapters into open regions of the genome. ATACseqTFEA is an improvement of the current computational method that detects differential activity of transcription factors (TFs). ATACseqTFEA not only uses the difference of open region information, but also (or emphasizes) the difference of TFs footprints (cutting sites or insertion sites). ATACseqTFEA provides an easy, rigorous way to broadly assess TF activity changes between two conditions.
Last updated
sequencingdnaseqatacseqmnaseseqgeneregulation
4.30 score 1 stars 8 scripts 378 downloadsgoSorensen - Statistical inference based on the Sorensen-Dice dissimilarity and the Gene Ontology (GO)
This package implements inferential methods to compare gene lists in terms of their biological meaning as expressed in the GO. The compared gene lists are characterized by cross-tabulation frequency tables of enriched GO items. Dissimilarity between gene lists is evaluated using the Sorensen-Dice index. The fundamental guiding principle is that two gene lists are taken as similar if they share a great proportion of common enriched GO items.
Last updated
annotationgogenesetenrichmentsoftwaremicroarraypathwaysgeneexpressionmultiplecomparisongraphandnetworkreactomeclusteringkegg
4.30 score 7 scriptsfactR - Functional Annotation of Custom Transcriptomes
factR contain tools to process and interact with custom-assembled transcriptomes (GTF). At its core, factR constructs CDS information on custom transcripts and subsequently predicts its functional output. In addition, factR has tools capable of plotting transcripts, correcting chromosome and gene information and shortlisting new transcripts.
Last updated
alternativesplicingfunctionalpredictiongenepredictioncustom-transcriptomesfunctional-annotationgtfrna-seq-analysis
4.30 score 2 stars 10 scripts 328 downloadstomoseqr - R Package for Analyzing Tomo-seq Data
`tomoseqr` is an R package for analyzing Tomo-seq data. Tomo-seq is a genome-wide RNA tomography method that combines combining high-throughput RNA sequencing with cryosectioning for spatially resolved transcriptomics. `tomoseqr` reconstructs 3D expression patterns from tomo-seq data and visualizes the reconstructed 3D expression patterns.
Last updated
geneexpressionsequencingrnaseqtranscriptomicsspatialvisualizationsoftware
4.30 score 20 scripts 338 downloadsprotGear - Protein Micro Array Data Management and Interactive Visualization
A generic three-step pre-processing package for protein microarray data. This package contains different data pre-processing procedures to allow comparison of their performance.These steps are background correction, the coefficient of variation (CV) based filtering, batch correction and normalization.
Last updated
microarrayonechannelpreprocessingbiomedicalinformaticsproteomicsbatcheffectnormalizationbayesianclusteringregressionsystemsbiologyimmunooncologybackground-correctionmicroarray-datanormalisationproteomics-datashinyshinydashboard
4.30 score 1 stars 8 scripts 406 downloadsupdateObject - Find/fix old serialized S4 instances
A set of tools built around updateObject() to work with old serialized S4 instances. The package is primarily useful to package maintainers who want to update the serialized S4 instances included in their package. This is still work-in-progress.
Last updated
infrastructuredatarepresentationbioconductor-packagecore-package
4.30 score 1 stars 6 scripts 322 downloadsMotif2Site - Detect binding sites from motifs and ChIP-seq experiments, and compare binding sites across conditions
Detect binding sites using motifs IUPAC sequence or bed coordinates and ChIP-seq experiments in bed or bam format. Combine/compare binding sites across experiments, tissues, or conditions. All normalization and differential steps are done using TMM-GLM method. Signal decomposition is done by setting motifs as the centers of the mixture of normal distribution curves.
Last updated
softwaresequencingchipseqdifferentialpeakcallingepigeneticssequencematching
4.30 score 5 scripts 345 downloadscenscyt - Differential abundance analysis with a right censored covariate in high-dimensional cytometry
Methods for differential abundance analysis in high-dimensional cytometry data when a covariate is subject to right censoring (e.g. survival time) based on multiple imputation and generalized linear mixed models.
Last updated
immunooncologyflowcytometryproteomicssinglecellcellbasedassayscellbiologyclusteringfeatureextractionsoftwaresurvival
4.30 score 2 scriptsAnVILBilling - Provide functions to retrieve and report on usage expenses in NHGRI AnVIL (anvilproject.org).
AnVILBilling helps monitor AnVIL-related costs in R, using queries to a BigQuery table to which costs are exported daily. Functions are defined to help categorize tasks and associated expenditures, and to visualize and explore expense profiles over time. This package will be expanded to help users estimate costs for specific task sets.
Last updated
infrastructuresoftware
4.30 score 5 scriptsmoanin - An R Package for Time Course RNASeq Data Analysis
Simple and efficient workflow for time-course gene expression data, built on publictly available open-source projects hosted on CRAN and bioconductor. moanin provides helper functions for all the steps required for analysing time-course data using functional data analysis: (1) functional modeling of the timecourse data; (2) differential expression analysis; (3) clustering; (4) downstream analysis.
Last updated
timecoursegeneexpressionrnaseqmicroarraydifferentialexpressionclustering
4.26 score 18 scripts 310 downloadsfamat - Functional analysis of metabolic and transcriptomic data
Famat is made to collect data about lists of genes and metabolites provided by user, and to visualize it through a Shiny app. Information collected is: - Pathways containing some of the user's genes and metabolites (obtained using a pathway enrichment analysis). - Direct interactions between user's elements inside pathways. - Information about elements (their identifiers and descriptions). - Go terms enrichment analysis performed on user's genes. The Shiny app is composed of: - information about genes, metabolites, and direct interactions between them inside pathways. - an heatmap showing which elements from the list are in pathways (pathways are structured in hierarchies). - hierarchies of enriched go terms using Molecular Function and Biological Process.
Last updated
functionalpredictiongenesetenrichmentpathwaysgoreactomekeggcompoundgene-ontologygenesshiny
4.26 score 1 stars 7 scripts 384 downloadsbetaHMM - A Hidden Markov Model Approach for Identifying Differentially Methylated Sites and Regions for Beta-Valued DNA Methylation Data
A novel approach utilizing a homogeneous hidden Markov model. And effectively model untransformed beta values. To identify DMCs while considering the spatial. Correlation of the adjacent CpG sites.
Last updated
dnamethylationdifferentialmethylationimmunooncologybiomedicalinformaticsmethylationarraysoftwaremultiplecomparisonsequencingspatialcoveragegenetargethiddenmarkovmodelmicroarray
4.22 score 11 scripts 241 downloadschihaya - Save Delayed Operations to a HDF5 File
Saves the delayed operations of a DelayedArray to a HDF5 file. This enables efficient recovery of the DelayedArray's contents in other languages and analysis frameworks.
Last updated
dataimportdatarepresentationcurlopensslcpp
4.20 score 16 scripts 416 downloadsimmunogenViewer - Visualization and evaluation of protein immunogens
Plots protein properties and visualizes position of peptide immunogens within protein sequence. Allows evaluation of immunogens based on structural and functional annotations to infer suitability for antibody-based methods aiming to detect native proteins.
Last updated
featureextractionproteomicssoftwarevisualization
4.18 score 15 scripts 198 downloadsGrafGen - Classification of Helicobacter Pylori Genomes
To classify Helicobacter pylori genomes according to genetic distance from nine reference populations. The nine reference populations are hpgpAfrica, hpgpAfrica-distant, hpgpAfroamerica, hpgpEuroamerica, hpgpMediterranea, hpgpEurope, hpgpEurasia, hpgpAsia, and hpgpAklavik86-like. The vertex populations are Africa, Europe and Asia.
Last updated
geneticssoftwaregenomeannotationclassificationcpp
4.18 score 2 scripts 258 downloadsRNAdecay - Maximum Likelihood Decay Modeling of RNA Degradation Data
RNA degradation is monitored through measurement of RNA abundance after inhibiting RNA synthesis. This package has functions and example scripts to facilitate (1) data normalization, (2) data modeling using constant decay rate or time-dependent decay rate models, (3) the evaluation of treatment or genotype effects, and (4) plotting of the data and models. Data Normalization: functions and scripts make easy the normalization to the initial (T0) RNA abundance, as well as a method to correct for artificial inflation of Reads per Million (RPM) abundance in global assessments as the total size of the RNA pool decreases. Modeling: Normalized data is then modeled using maximum likelihood to fit parameters. For making treatment or genotype comparisons (up to four), the modeling step models all possible treatment effects on each gene by repeating the modeling with constraints on the model parameters (i.e., the decay rate of treatments A and B are modeled once with them being equal and again allowing them to both vary independently). Model Selection: The AICc value is calculated for each model, and the model with the lowest AICc is chosen. Modeling results of selected models are then compiled into a single data frame. Graphical Plotting: functions are provided to easily visualize decay data model, or half-life distributions using ggplot2 package functions.
Last updated
immunooncologysoftwaregeneexpressiongeneregulationdifferentialexpressiontranscriptiontranscriptomicstimecourseregressionrnaseqnormalizationworkflowstep
4.18 score 2 scripts 353 downloadsRegionalST - Investigating regions of interest and performing regional cell type-specific analysis with spatial transcriptomics data
This package analyze spatial transcriptomics data through cross-regional cell type-specific analysis. It selects regions of interest (ROIs) and identifys cross-regional cell type-specific differential signals. The ROIs can be selected using automatic algorithm or through manual selection. It facilitates manual selection of ROIs using a shiny application.
Last updated
spatialtranscriptomicsreactomekegg
4.18 score 9 scripts 263 downloadsBiocHubsShiny - View AnnotationHub and ExperimentHub Resources Interactively
A package that allows interactive exploration of AnnotationHub and ExperimentHub resources. It uses DT / DataTable to display resources for multiple organisms. It provides template code for reproducibility and for downloading resources via the indicated Hub package.
Last updated
softwareshinyapps
4.18 score 5 scripts 332 downloadsoctad - Open Cancer TherApeutic Discovery (OCTAD)
OCTAD provides a platform for virtually screening compounds targeting precise cancer patient groups. The essential idea is to identify drugs that reverse the gene expression signature of disease by tamping down over-expressed genes and stimulating weakly expressed ones. The package offers deep-learning based reference tissue selection, disease gene expression signature creation, pathway enrichment analysis, drug reversal potency scoring, cancer cell line selection, drug enrichment analysis and in silico hit validation. It currently covers ~20,000 patient tissue samples covering 50 cancer types, and expression profiles for ~12,000 distinct compounds.
Last updated
classificationgeneexpressionpharmacogeneticspharmacogenomicssoftwaregenesetenrichment
4.18 score 356 downloadsggtreeDendro - Drawing 'dendrogram' using 'ggtree'
Offers a set of 'autoplot' methods to visualize tree-like structures (e.g., hierarchical clustering and classification/regression trees) using 'ggtree'. You can adjust graphical parameters using grammar of graphic syntax and integrate external data to the tree.
Last updated
clusteringclassificationdecisiontreephylogeneticsvisualization
4.18 score 15 scripts 268 downloadspengls - Fit Penalised Generalised Least Squares models
Combine generalised least squares methodology from the nlme package for dealing with autocorrelation with penalised least squares methods from the glmnet package to deal with high dimensionality. This pengls packages glues them together through an iterative loop. The resulting method is applicable to high dimensional datasets that exhibit autocorrelation, such as spatial or temporal data.
Last updated
transcriptomicsregressiontimecoursespatial
4.18 score 9 scripts 278 downloadsflowGraph - Identifying differential cell populations in flow cytometry data accounting for marker frequency
Identifies maximal differential cell populations in flow cytometry data taking into account dependencies between cell populations; flowGraph calculates and plots SpecEnr abundance scores given cell population cell counts.
Last updated
flowcytometrystatisticalmethodimmunooncologysoftwarecellbasedassaysvisualization
4.18 score 1 stars 15 scripts 324 downloadsMultiRNAflow - An R package for integrated analysis of temporal RNA-seq data with multiple biological conditions
Our R package MultiRNAflow provides an easy to use unified framework allowing to automatically make both unsupervised and supervised (DE) analysis for datasets with an arbitrary number of biological conditions and time points. In particular, our code makes a deep downstream analysis of DE information, e.g. identifying temporal patterns across biological conditions and DE genes which are specific to a biological condition for each time.
Last updated
sequencingrnaseqgeneexpressiontranscriptiontimecoursepreprocessingvisualizationnormalizationprincipalcomponentclusteringdifferentialexpressiongenesetenrichmentpathways
4.15 score 7 stars 6 scripts 301 downloadsISLET - Individual-Specific ceLl typE referencing Tool
ISLET is a method to conduct signal deconvolution for general -omics data. It can estimate the individual-specific and cell-type-specific reference panels, when there are multiple samples observed from each subject. It takes the input of the observed mixture data (feature by sample matrix), and the cell type mixture proportions (sample by cell type matrix), and the sample-to-subject information. It can solve for the reference panel on the individual-basis and conduct test to identify cell-type-specific differential expression (csDE) genes. It also improves estimated cell type mixture proportions by integrating personalized reference panels.
Last updated
softwarernaseqtranscriptomicstranscriptionsequencinggeneexpressiondifferentialexpressiondifferentialmethylation
4.15 score 14 scripts 342 downloadsVarCon - VarCon: an R package for retrieving neighboring nucleotides of an SNV
VarCon is an R package which converts the positional information from the annotation of an single nucleotide variation (SNV) (either referring to the coding sequence or the reference genomic sequence). It retrieves the genomic reference sequence around the position of the single nucleotide variation. To asses, whether the SNV could potentially influence binding of splicing regulatory proteins VarCon calcualtes the HEXplorer score as an estimation. Besides, VarCon additionally reports splice site strengths of splice sites within the retrieved genomic sequence and any changes due to the SNV.
Last updated
functionalgenomicsalternativesplicing
4.15 score 14 scripts 344 downloadsOutSplice - Comparison of Splicing Events between Tumor and Normal Samples
An easy to use tool that can compare splicing events in tumor and normal tissue samples using either a user generated matrix, or data from The Cancer Genome Atlas (TCGA). This package generates a matrix of splicing outliers that are significantly over or underexpressed in tumors samples compared to normal denoted by chromosome location. The package also will calculate the splicing burden in each tumor and characterize the types of splicing events that occur.
Last updated
alternativesplicingdifferentialexpressiondifferentialsplicinggeneexpressionrnaseqsoftwarevariantannotation
4.11 score 1 stars 13 scripts 333 downloadsbedbaser - A BEDbase client
A client for BEDbase. bedbaser provides access to the API at api.bedbase.org. It also includes convenience functions to import BED files into GRanges objects and BEDsets into GRangesLists.
Last updated
softwaredataimportthirdpartyclientu24ca289073
4.08 score 3 stars 6 scripts 278 downloads
gINTomics - Multi-Omics data integration
gINTomics is an R package for Multi-Omics data integration and visualization. gINTomics is designed to detect the association between the expression of a target and of its regulators, taking into account also their genomics modifications such as Copy Number Variations (CNV) and methylation. What is more, gINTomics allows integration results visualization via a Shiny-based interactive app.
Last updated
geneexpressionrnaseqmicroarrayvisualizationcopynumbervariationgenetargetquarto
4.08 score 3 stars 4 scripts 198 downloadsbandle - An R package for the Bayesian analysis of differential subcellular localisation experiments
The Bandle package enables the analysis and visualisation of differential localisation experiments using mass-spectrometry data. Experimental methods supported include dynamic LOPIT-DC, hyperLOPIT, Dynamic Organellar Maps, Dynamic PCP. It provides Bioconductor infrastructure to analyse these data.
Last updated
bayesianclassificationclusteringimmunooncologyqualitycontroldataimportproteomicsmassspectrometryopenblascppopenmp
4.08 score 4 stars 4 scripts 384 downloadsPoDCall - Positive Droplet Calling for DNA Methylation Droplet Digital PCR
Reads files exported from 'QX Manager or QuantaSoft' containing amplitude values from a run of ddPCR (96 well plate) and robustly sets thresholds to determine positive droplets for each channel of each individual well. Concentration and normalized concentration in addition to other metrics is then calculated for each well. Results are returned as a table, optionally written to file, as well as optional plots (scatterplot and histogram) for both channels per well written to file. The package includes a shiny application which provides an interactive and user-friendly interface to the full functionality of PoDCall.
Last updated
classificationepigeneticsddpcrdifferentialmethylationcpgislanddnamethylation
4.08 score 12 scripts 280 downloadsMEIGOR - MEIGOR - MEtaheuristics for bIoinformatics Global Optimization
MEIGOR provides a comprehensive environment for performing global optimization tasks in bioinformatics and systems biology. It leverages advanced metaheuristic algorithms to efficiently search the solution space and is specifically tailored to handle the complexity and high-dimensionality of biological datasets. This package supports various optimization routines and is integrated with Bioconductor's infrastructure for a seamless analysis workflow.
Last updated
systemsbiologyoptimizationsoftware
4.07 score 59 scripts 457 downloadsggmanh - Visualization Tool for GWAS Result
Manhattan plot and QQ Plot are commonly used to visualize the end result of Genome Wide Association Study. The "ggmanh" package aims to keep the generation of these plots simple while maintaining customizability. Main functions include manhattan_plot, qqunif, and thinPoints.
Last updated
visualizationgenomewideassociationgenetics
4.07 score 59 scripts 412 downloads
BioGA - Bioinformatics Genetic Algorithm (BioGA)
Genetic algorithm are a class of optimization algorithms inspired by the process of natural selection and genetics. This package allows users to analyze and optimize high throughput genomic data using genetic algorithms. The functions provided are implemented in C++ for improved speed and efficiency, with an easy-to-use interface for use within R.
Last updated
experimentaldesigntechnologydata-analysisgene-expressiongenetic-algorithmsgenomicsoptimization-algorithmscpp
4.04 score 11 scripts 228 downloadsxcore - xcore expression regulators inference
xcore is an R package for transcription factor activity modeling based on known molecular signatures and user's gene expression data. Accompanying xcoredata package provides a collection of molecular signatures, constructed from publicly available ChiP-seq experiments. xcore use ridge regression to model changes in expression as a linear combination of molecular signatures and find their unknown activities. Obtained, estimates can be further tested for significance to select molecular signatures with the highest predicted effect on the observed expression changes.
Last updated
geneexpressiongeneregulationepigeneticsregressionsequencing
4.04 score 11 scripts 374 downloadsmicrobiomeExplorer - Microbiome Exploration App
The MicrobiomeExplorer R package is designed to facilitate the analysis and visualization of marker-gene survey feature data. It allows a user to perform and visualize typical microbiome analytical workflows either through the command line or an interactive Shiny application included with the package. In addition to applying common analytical workflows the application enables automated analysis report generation.
Last updated
classificationclusteringgeneticvariabilitydifferentialexpressionmicrobiomemetagenomicsnormalizationvisualizationmultiplecomparisonsequencingsoftwareimmunooncology
4.04 score 11 scripts 475 downloadsLheuristic - Detection of scatterplots with L-shaped pattern
The Lheuristic package identifies scatterpots that follow and L-shaped, negative distribution. It can be used to identify genes regulated by methylation by integration of an expression and a methylation array. The package uses two different methods to detect expression and methyaltion L- shapped scatterplots. The parameters can be changed to detect other scatterplot patterns.
Last updated
dnamethylationstatisticalmethodmethylationarrayl-shapedmethylation-expressionscatterplot-matrix
4.00 score 9 scripts 225 downloadsSite2Target - An R package to associate peaks and target genes
Statistics implemented for both peak-wise and gene-wise associations. In peak-wise associations, the p-value of the target genes of a given set of peaks are calculated. Negative binomial or Poisson distributions can be used for modeling the unweighted peaks targets and log-nromal can be used to model the weighted peaks. In gene-wise associations a table consisting of a set of genes, mapped to specific peaks, is generated using the given rules.
Last updated
annotationchipseqsoftwareepigeneticsgeneexpressiongenetarget
4.00 score 6 scripts 231 downloadsPolytect - An R package for digital data clustering
Polytect is an advanced computational tool designed for the analysis of multi-color digital PCR data. It provides automatic clustering and labeling of partitions into distinct groups based on clusters first identified by the flowPeaks algorithm. Polytect is particularly useful for researchers in molecular biology and bioinformatics, enabling them to gain deeper insights into their experimental results through precise partition classification and data visualization.
Last updated
ddpcrclusteringmultichannelclassification
4.00 score 1 stars 4 scripts 205 downloads
mspms - Tools for the analysis of MSP-MS data
This package provides functions for the analysis of data generated by the multiplex substrate profiling by mass spectrometry for proteases (MSP-MS) method. Data exported from upstream proteomics software is accepted as input and subsequently processed for analysis. Tools for statistical analysis, visualization, and interpretation of the data are provided.
Last updated
proteomicsmassspectrometrypreprocessingproteaseproteomics-data-analysis
4.00 score 1 stars 7 scripts 288 downloadsspatialSimGP - Simulate Spatial Transcriptomics Data with the Mean-variance Relationship
This packages simulates spatial transcriptomics data with the mean- variance relationship using a Gaussian Process model per gene.
Last updated
spatialtranscriptomicsgeneexpression
4.00 score 2 scripts 240 downloadsxenLite - Simple classes and methods for managing Xenium datasets
Define a relatively light class for managing Xenium data using Bioconductor. Address use of parquet for coordinates, SpatialExperiment for assay and sample data. Address serialization and use of cloud storage.
Last updated
infrastructureu24ca289073
4.00 score 1 stars 4 scripts 242 downloadsDeepTarget - Deep characterization of cancer drugs
This package predicts a drug’s primary target(s) or secondary target(s) by integrating large-scale genetic and drug screens from the Cancer Dependency Map project run by the Broad Institute. It further investigates whether the drug specifically targets the wild-type or mutated target forms. To show how to use this package in practice, we provided sample data along with step-by-step example.
Last updated
genetargetgenepredictionpathwaysgeneexpressionrnaseqimmunooncologydifferentialexpressiongenesetenrichmentreportwritingcrispr
4.00 score 7 scripts 250 downloadszitools - Analysis of zero-inflated count data
zitools allows for zero inflated count data analysis by either using down-weighting of excess zeros or by replacing an appropriate proportion of excess zeros with NA. Through overloading frequently used statistical functions (such as mean, median, standard deviation), plotting functions (such as boxplots or heatmap) or differential abundance tests, it allows a wide range of downstream analyses for zero-inflated data in a less biased manner. This becomes applicable in the context of microbiome analyses, where the data is often overdispersed and zero-inflated, therefore making data analysis extremly challenging.
Last updated
softwarestatisticalmethodmicrobiome
4.00 score 9 scripts 254 downloadsggseqalign - Minimal Visualization of Sequence Alignments
Simple visualizations of alignments of DNA or AA sequences as well as arbitrary strings. Compatible with Biostrings and ggplot2. The plots are fully customizable using ggplot2 modifiers such as theme().
Last updated
alignmentmultiplesequencealignmentsoftwarevisualizationbioinformaticsggplot2-enhancementsminimalistic
4.00 score 10 scripts 242 downloadsfindIPs - Influential Points Detection for Feature Rankings
Feature rankings can be distorted by a single case in the context of high-dimensional data. The cases exerts abnormal influence on feature rankings are called influential points (IPs). The package aims at detecting IPs based on case deletion and quantifies their effects by measuring the rank changes (DOI:10.48550/arXiv.2303.10516). The package applies a novel rank comparing measure using the adaptive weights that stress the top-ranked important features and adjust the weights to ranking properties.
Last updated
geneexpressiondifferentialexpressionregressionsurvival
4.00 score 4 scripts 261 downloadsMAPFX - MAssively Parallel Flow cytometry Xplorer (MAPFX): A Toolbox for Analysing Data from the Massively-Parallel Cytometry Experiments
MAPFX is an end-to-end toolbox that pre-processes the raw data from MPC experiments (e.g., BioLegend's LEGENDScreen and BD Lyoplates assays), and further imputes the ‘missing’ infinity markers in the wells without those measurements. The pipeline starts by performing background correction on raw intensities to remove the noise from electronic baseline restoration and fluorescence compensation by adapting a normal-exponential convolution model. Unwanted technical variation, from sources such as well effects, is then removed using a log-normal model with plate, column, and row factors, after which infinity markers are imputed using the informative backbone markers as predictors. The completed dataset can then be used for clustering and other statistical analyses. Additionally, MAPFX can be used to normalise data from FFC assays as well.
Last updated
softwareflowcytometrycellbasedassayssinglecellproteomicsclustering
4.00 score 1 stars 276 downloadshdxmsqc - An R package for quality Control for hydrogen deuterium exchange mass spectrometry experiments
The hdxmsqc package enables us to analyse and visualise the quality of HDX-MS experiments. Either as a final quality check before downstream analysis and publication or as part of a interative procedure to determine the quality of the data. The package builds on the QFeatures and Spectra packages to integrate with other mass-spectrometry data.
Last updated
qualitycontroldataimportproteomicsmassspectrometrymetabolomics
4.00 score 1 stars 7 scripts 286 downloadsBREW3R.r - R package associated to BREW3R
This R package provide functions that are used in the BREW3R workflow. This mainly contains a function that extend a gtf as GRanges using information from another gtf (also as GRanges). The process allows to extend gene annotation without increasing the overlap between gene ids.
Last updated
genomeannotation
4.00 score 10 scripts 276 downloadscypress - Cell-Type-Specific Power Assessment
CYPRESS is a cell-type-specific power tool. This package aims to perform power analysis for the cell-type-specific data. It calculates FDR, FDC, and power, under various study design parameters, including but not limited to sample size, and effect size. It takes the input of a SummarizeExperimental(SE) object with observed mixture data (feature by sample matrix), and the cell-type mixture proportions (sample by cell-type matrix). It can solve the cell-type mixture proportions from the reference free panel from TOAST and conduct tests to identify cell-type-specific differential expression (csDE) genes.
Last updated
softwaregeneexpressiondataimportrnaseqsequencing
4.00 score 1 stars 2 scripts 286 downloadsdinoR - Differential NOMe-seq analysis
dinoR tests for significant differences in NOMe-seq footprints between two conditions, using genomic regions of interest (ROI) centered around a landmark, for example a transcription factor (TF) motif. This package takes NOMe-seq data (GCH methylation/protection) in the form of a Ranged Summarized Experiment as input. dinoR can be used to group sequencing fragments into 3 or 5 categories representing characteristic footprints (TF bound, nculeosome bound, open chromatin), plot the percentage of fragments in each category in a heatmap, or averaged across different ROI groups, for example, containing a common TF motif. It is designed to compare footprints between two sample groups, using edgeR's quasi-likelihood methods on the total fragment counts per ROI, sample, and footprint category.
Last updated
nucleosomepositioningepigeneticsmethylseqdifferentialmethylationcoveragetranscriptionsequencingsoftware
4.00 score 8 scripts 302 downloadssimPIC - Flexible simulation of paired-insertion counts for single-cell ATAC-sequencing data
simPIC is a package for simulating single-cell ATAC-seq count data. It provides a user-friendly, well documented interface for data simulation. Functions are provided for parameter estimation, realistic scATAC-seq data simulation, and comparing real and simulated datasets.
Last updated
singlecellatacseqsoftwaresequencingimmunooncologydataimportbioconductorbioinformaticsscatac-seqsimulation
4.00 score 1 stars 9 scripts 255 downloadsspillR - Spillover Compensation in Mass Cytometry Data
Channel interference in mass cytometry can cause spillover and may result in miscounting of protein markers. We develop a nonparametric finite mixture model and use the mixture components to estimate the probability of spillover. We implement our method using expectation-maximization to fit the mixture model.
Last updated
flowcytometryimmunooncologymassspectrometrypreprocessingsinglecellsoftwarestatisticalmethodvisualizationregression
4.00 score 5 scripts 245 downloadsplasmut - Stratifying mutations observed in cell-free DNA and white blood cells as germline, hematopoietic, or somatic
A Bayesian method for quantifying the liklihood that a given plasma mutation arises from clonal hematopoesis or the underlying tumor. It requires sequencing data of the mutation in plasma and white blood cells with the number of distinct and mutant reads in both tissues. We implement a Monte Carlo importance sampling method to assess the likelihood that a mutation arises from the tumor relative to non-tumor origin.
Last updated
bayesiansomaticmutationgermlinemutationsequencing
4.00 score 3 scripts 244 downloadsgg4way - 4way Plots of Differential Expression
4way plots enable a comparison of the logFC values from two contrasts of differential gene expression. The gg4way package creates 4way plots using the ggplot2 framework and supports popular Bioconductor objects. The package also provides information about the correlation between contrasts and significant genes of interest.
Last updated
softwarevisualizationdifferentialexpressiongeneexpressiontranscriptionrnaseqsinglecellsequencing
4.00 score 7 scripts 265 downloadscompSPOT - compSPOT: Tool for identifying and comparing significantly mutated genomic hotspots
Clonal cell groups share common mutations within cancer, precancer, and even clinically normal appearing tissues. The frequency and location of these mutations may predict prognosis and cancer risk. It has also been well established that certain genomic regions have increased sensitivity to acquiring mutations. Mutation-sensitive genomic regions may therefore serve as markers for predicting cancer risk. This package contains multiple functions to establish significantly mutated hotspots, compare hotspot mutation burden between samples, and perform exploratory data analysis of the correlation between hotspot mutation burden and personal risk factors for cancer, such as age, gender, and history of carcinogen exposure. This package allows users to identify robust genomic markers to help establish cancer risk.
Last updated
softwaretechnologysequencingdnaseqwholegenomeclassificationsinglecellsurvivalmultiplecomparison
4.00 score 4 scripts 253 downloadsHERON - Hierarchical Epitope pROtein biNding
HERON is a software package for analyzing peptide binding array data. In addition to identifying significant binding probes, HERON also provides functions for finding epitopes (string of consecutive peptides within a protein). HERON also calculates significance on the probe, epitope, and protein level by employing meta p-value methods. HERON is designed for obtaining calls on the sample level and calculates fractions of hits for different conditions.
Last updated
microarraysoftware
4.00 score 1 stars 6 scriptsMICSQTL - MICSQTL (Multi-omic deconvolution, Integration and Cell-type-specific Quantitative Trait Loci)
Our pipeline, MICSQTL, utilizes scRNA-seq reference and bulk transcriptomes to estimate cellular composition in the matched bulk proteomes. The expression of genes and proteins at either bulk level or cell type level can be integrated by Angle-based Joint and Individual Variation Explained (AJIVE) framework. Meanwhile, MICSQTL can perform cell-type-specic quantitative trait loci (QTL) mapping to proteins or transcripts based on the input of bulk expression data and the estimated cellular composition per molecule type, without the need for single cell sequencing. We use matched transcriptome-proteome from human brain frontal cortex tissue samples to demonstrate the input and output of our tool.
Last updated
geneexpressiongeneticsproteomicsrnaseqsequencingsinglecellsoftwarevisualizationcellbasedassayscoverage
4.00 score 8 scripts 286 downloadsMultimodalExperiment - Integrative Bulk and Single-Cell Experiment Container
MultimodalExperiment is an S4 class that integrates bulk and single-cell experiment data; it is optimally storage-efficient, and its methods are exceptionally fast. It effortlessly represents multimodal data of any nature and features normalized experiment, subject, sample, and cell annotations, which are related to underlying biological experiments through maps. Its coordination methods are opt-in and employ database-like join operations internally to deliver fast and flexible management of multimodal data.
Last updated
datarepresentationinfrastructuresinglecell
4.00 score 4 scriptsEDIRquery - Query the EDIR Database For Specific Gene
EDIRquery provides a tool to search for genes of interest within the Exome Database of Interspersed Repeats (EDIR). A gene name is a required input, and users can additionally specify repeat sequence lengths, minimum and maximum distance between sequences, and whether to allow a 1-bp mismatch. Outputs include a summary of results by repeat length, as well as a dataframe of query results. Example data provided includes a subset of the data for the gene GAA (ENSG00000171298). To query the full database requires providing a path to the downloaded database files as a parameter.
Last updated
geneticssequencematching
4.00 score 3 scripts 270 downloadsseq.hotSPOT - Targeted sequencing panel design based on mutation hotspots
seq.hotSPOT provides a resource for designing effective sequencing panels to help improve mutation capture efficacy for ultradeep sequencing projects. Using SNV datasets, this package designs custom panels for any tissue of interest and identify the genomic regions likely to contain the most mutations. Establishing efficient targeted sequencing panels can allow researchers to study mutation burden in tissues at high depth without the economic burden of whole-exome or whole-genome sequencing. This tool was developed to make high-depth sequencing panels to study low-frequency clonal mutations in clinically normal and cancerous tissues.
Last updated
softwaretechnologysequencingdnaseqwholegenome
4.00 score 8 scripts 258 downloadsalabaster - Umbrella for the Alabaster Framework
Umbrella for the alabaster suite, providing a single-line import for all alabaster.* packages. Installing this package ensures that all known alabaster.* packages are also installed, avoiding problems with missing packages when a staging method or loading function is dynamically requested. Obviously, this comes at the cost of needing to install more packages, so advanced users and application developers may prefer to install the required alabaster.* packages individually.
Last updated
datarepresentationdataimport
4.00 score 5 scripts 286 downloadsrifiComparative - 'rifiComparative' compares the output of rifi from two different conditions.
'rifiComparative' is a continuation of rifi package. It compares two conditions output of rifi using half-life and mRNA at time 0 segments. As an input for the segmentation, the difference between half-life of both condtions and log2FC of the mRNA at time 0 are used. The package provides segmentation, statistics, summary table, fragments visualization and some additional useful plots for further anaylsis.
Last updated
rnaseqdifferentialexpressiongeneregulationtranscriptomicsmicroarraysoftware
4.00 score 5 scripts 272 downloadsBG2 - Performs Bayesian GWAS analysis for non-Gaussian data using BG2
This package is built to perform GWAS analysis for non-Gaussian data using BG2. The BG2 method uses penalized quasi-likelihood along with nonlocal priors in a two step manner to identify SNPs in GWAS analysis. The research related to this package was supported in part by National Science Foundation awards DMS 1853549 and DMS 2054173.
Last updated
bayesianassaydomainsnpgenomewideassociation
4.00 score 5 scripts 252 downloads
planttfhunter - Identification and classification of plant transcription factors
planttfhunter is used to identify plant transcription factors (TFs) from protein sequence data and classify them into families and subfamilies using the classification scheme implemented in PlantTFDB. TFs are identified using pre-built hidden Markov model profiles for DNA-binding domains. Then, auxiliary and forbidden domains are used with DNA-binding domains to classify TFs into families and subfamilies (when applicable). Currently, TFs can be classified in 58 different TF families/subfamilies.
Last updated
softwaretranscriptionfunctionalpredictiongenomeannotationfunctionalgenomicshiddenmarkovmodelsequencingclassificationfunctional-genomicsgene-familieshidden-markov-modelsplant-genomicsplantsprotein-domainstranscription-factors
4.00 score 1 stars 7 scripts 268 downloadsMetaPhOR - Metabolic Pathway Analysis of RNA
MetaPhOR was developed to enable users to assess metabolic dysregulation using transcriptomic-level data (RNA-sequencing and Microarray data) and produce publication-quality figures. A list of differentially expressed genes (DEGs), which includes fold change and p value, from DESeq2 or limma, can be used as input, with sample size for MetaPhOR, and will produce a data frame of scores for each KEGG pathway. These scores represent the magnitude and direction of transcriptional change within the pathway, along with estimated p-values.MetaPhOR then uses these scores to visualize metabolic profiles within and between samples through a variety of mechanisms, including: bubble plots, heatmaps, and pathway models.
Last updated
metabolomicsrnaseqpathwaysgeneexpressiondifferentialexpressionkeggsequencingmicroarray
4.00 score 1 scripts 350 downloadsCircSeqAlignTk - End-to-End Analysis of Small RNA-Seq Data from Viroids
CircSeqAlignTk is a toolkit for the analysis of RNA-Seq data derived from circular genome sequences, with a primary focus on viroids, circular RNAs typically consisting of a few hundred nucleotides. The toolkit supports an end-to-end analysis pipeline, from alignment to visualization.
Last updated
sequencingsmallrnaalignmentsoftware
4.00 score 3 scriptsphenomis - Postprocessing and univariate analysis of omics data
The 'phenomis' package provides methods to perform post-processing (i.e. quality control and normalization) as well as univariate statistical analysis of single and multi-omics data sets. These methods include quality control metrics, signal drift and batch effect correction, intensity transformation, univariate hypothesis testing, but also clustering (as well as annotation of metabolomics data). The data are handled in the standard Bioconductor formats (i.e. SummarizedExperiment and MultiAssayExperiment for single and multi-omics datasets, respectively; the alternative ExpressionSet and MultiDataSet formats are also supported for convenience). As a result, all methods can be readily chained as workflows. The pipeline can be further enriched by multivariate analysis and feature selection, by using the 'ropls' and 'biosigner' packages, which support the same formats. Data can be conveniently imported from and exported to text files. Although the methods were initially targeted to metabolomics data, most of the methods can be applied to other types of omics data (e.g., transcriptomics, proteomics).
Last updated
batcheffectclusteringcoveragekeggmassspectrometrymetabolomicsnormalizationproteomicsqualitycontrolsequencingstatisticalmethodtranscriptomics
4.00 score 7 scripts 276 downloadsEasyCellType - Annotate cell types for scRNA-seq data
We developed EasyCellType which can automatically examine the input marker lists obtained from existing software such as Seurat over the cell markerdatabases. Two quantification approaches to annotate cell types are provided: Gene set enrichment analysis (GSEA) and a modified versio of Fisher's exact test. The function presents annotation recommendations in graphical outcomes: bar plots for each cluster showing candidate cell types, as well as a dot plot summarizing the top 5 significant annotations for each cluster.
Last updated
singlecellsoftwaregeneexpressiongenesetenrichment
4.00 score 10 scripts 301 downloadsSUITOR - Selecting the number of mutational signatures through cross-validation
An unsupervised cross-validation method to select the optimal number of mutational signatures. A data set of mutational counts is split into training and validation data.Signatures are estimated in the training data and then used to predict the mutations in the validation data.
Last updated
geneticssoftwaresomaticmutation
4.00 score 2 scripts 293 downloadsExperimentSubset - Manages subsets of data with Bioconductor Experiment objects
Experiment objects such as the SummarizedExperiment or SingleCellExperiment are data containers for one or more matrix-like assays along with the associated row and column data. Often only a subset of the original data is needed for down-stream analysis. For example, filtering out poor quality samples will require excluding some columns before analysis. The ExperimentSubset object is a container to efficiently manage different subsets of the same data without having to make separate objects for each new subset.
Last updated
infrastructuresoftwaredataimportdatarepresentation
4.00 score 10 scripts 464 downloadstomoda - Tomo-seq data analysis
This package provides many easy-to-use methods to analyze and visualize tomo-seq data. The tomo-seq technique is based on cryosectioning of tissue and performing RNA-seq on consecutive sections. (Reference: Kruse F, Junker JP, van Oudenaarden A, Bakkers J. Tomo-seq: A method to obtain genome-wide expression data with spatial resolution. Methods Cell Biol. 2016;135:299-307. doi:10.1016/bs.mcb.2016.01.006) The main purpose of the package is to find zones with similar transcriptional profiles and spatially expressed genes in a tomo-seq sample. Several visulization functions are available to create easy-to-modify plots.
Last updated
geneexpressionsequencingrnaseqtranscriptomicsspatialclusteringvisualization
4.00 score 8 scripts 326 downloadsrsemmed - An interface to the Semantic MEDLINE database
A programmatic interface to the Semantic MEDLINE database. It provides functions for searching the database for concepts and finding paths between concepts. Path searching can also be tailored to user specifications, such as placing restrictions on concept types and the type of link between concepts. It also provides functions for summarizing and visualizing those paths.
Last updated
softwareannotationpathwayssystemsbiology
4.00 score 10 scripts 343 downloadsBOBaFIT - Refitting diploid region profiles using a clustering procedure
This package provides a method to refit and correct the diploid region in copy number profiles. It uses a clustering algorithm to identify pathology-specific normal (diploid) chromosomes and then use their copy number signal to refit the whole profile. The package is composed by three functions: DRrefit (the main function), ComputeNormalChromosome and PlotCluster.
Last updated
copynumbervariationclusteringvisualizationnormalizationsoftware
3.90 score 6 scripts 334 downloadsroastgsa - Rotation based gene set analysis
This package implements a variety of functions useful for gene set analysis using rotations to approximate the null distribution. It contributes with the implementation of seven test statistic scores that can be used with different goals and interpretations. Several functions are available to complement the statistical results with graphical representations.
Last updated
microarraypreprocessingnormalizationgeneexpressionsurvivaltranscriptionsequencingtranscriptomicsbayesianclusteringregressionrnaseqmicrornaarraymrnamicroarrayfunctionalgenomicssystemsbiologyimmunooncologydifferentialexpressiongenesetenrichmentbatcheffectmultiplecomparisonqualitycontroltimecoursemetabolomicsproteomicsepigeneticscheminformaticsexonarrayonechanneltwochannelproprietaryplatformscellbiologybiomedicalinformaticsalternativesplicingdifferentialsplicingdataimportpathways
3.89 score 13 scripts 298 downloadsRcollectl - Help use collectl with R in Linux, to measure resource consumption in R processes
Provide functions to obtain instrumentation data on processes in a unix environment. Parse output of a collectl run. Vizualize aspects of system usage over time, with annotation.
Last updated
softwareinfrastructure
3.89 score 3 stars 13 scripts 183 downloadsmethodical - Discovering genomic regions where methylation is strongly associated with transcriptional activity
DNA methylation is generally considered to be associated with transcriptional silencing. However, comprehensive, genome-wide investigation of this relationship requires the evaluation of potentially millions of correlation values between the methylation of individual genomic loci and expression of associated transcripts in a relatively large numbers of samples. Methodical makes this process quick and easy while keeping a low memory footprint. It also provides a novel method for identifying regions where a number of methylation sites are consistently strongly associated with transcriptional expression. In addition, Methodical enables housing DNA methylation data from diverse sources (e.g. WGBS, RRBS and methylation arrays) with a common framework, lifting over DNA methylation data between different genome builds and creating base-resolution plots of the association between DNA methylation and transcriptional activity at transcriptional start sites.
Last updated
dnamethylationmethylationarraytranscriptiongenomewideassociationsoftware
3.83 score 40 scripts 280 downloadspfamAnalyzeR - Identification of domain isotypes in pfam data
Protein domains is one of the most import annoation of proteins we have with the Pfam database/tool being (by far) the most used tool. This R package enables the user to read the pfam prediction from both webserver and stand-alone runs into R. We have recently shown most human protein domains exist as multiple distinct variants termed domain isotypes. Different domain isotypes are used in a cell, tissue, and disease-specific manner. Accordingly, we find that domain isotypes, compared to each other, modulate, or abolish the functionality of a protein domain. This R package enables the identification and classification of such domain isotypes from Pfam data.
Last updated
alternativesplicingtranscriptomevariantbiomedicalinformaticsfunctionalgenomicssystemsbiologyannotationfunctionalpredictiongenepredictiondataimport
3.78 score 1 stars 1 dependents 2 scripts 595 downloadsCONSTANd - Data normalization by matrix raking
Normalizes a data matrix `data` by raking (using the RAS method by Bacharach, see references) the Nrows by Ncols matrix such that the row means and column means equal 1. The result is a normalized data matrix `K=RAS`, a product of row mulipliers `R` and column multipliers `S` with the original matrix `A`. Missing information needs to be presented as `NA` values and not as zero values, because CONSTANd is able to ignore missing values when calculating the mean. Using CONSTANd normalization allows for the direct comparison of values between samples within the same and even across different CONSTANd-normalized data matrices.
Last updated
massspectrometrycheminformaticsnormalizationpreprocessingdifferentialexpressiongeneticstranscriptomicsproteomics
3.78 score 7 scripts 316 downloadsMethReg - Assessing the regulatory potential of DNA methylation regions or sites on gene transcription
Epigenome-wide association studies (EWAS) detects a large number of DNA methylation differences, often hundreds of differentially methylated regions and thousands of CpGs, that are significantly associated with a disease, many are located in non-coding regions. Therefore, there is a critical need to better understand the functional impact of these CpG methylations and to further prioritize the significant changes. MethReg is an R package for integrative modeling of DNA methylation, target gene expression and transcription factor binding sites data, to systematically identify and rank functional CpG methylations. MethReg evaluates, prioritizes and annotates CpG sites with high regulatory potential using matched methylation and gene expression data, along with external TF-target interaction databases based on manually curation, ChIP-seq experiments or gene regulatory network analysis.
Last updated
methylationarrayregressiongeneexpressionepigeneticsgenetargettranscription
3.75 score 5 stars 25 scripts 358 downloadsDamsel - Damsel: an end to end analysis of DamID
Damsel provides an end to end analysis of DamID data. Damsel takes bam files from Dam-only control and fusion samples and counts the reads matching to each GATC region. edgeR is utilised to identify regions of enrichment in the fusion relative to the control. Enriched regions are combined into peaks, and are associated with nearby genes. Damsel allows for IGV style plots to be built as the results build, inspired by ggcoverage, and using the functionality and layering ability of ggplot2. Damsel also conducts gene ontology testing with bias correction through goseq, and future versions of Damsel will also incorporate motif enrichment analysis. Overall, Damsel is the first package allowing for an end to end analysis with visual capabilities. The goal of Damsel was to bring all the analysis into one place, and allow for exploratory analysis within R.
Last updated
differentialmethylationpeakdetectiongenepredictiongenesetenrichment
3.70 score 1 stars 25 scripts 268 downloadsborealis - Bisulfite-seq OutlieR mEthylation At singLe-sIte reSolution
Borealis is an R library performing outlier analysis for count-based bisulfite sequencing data. It detectes outlier methylated CpG sites from bisulfite sequencing (BS-seq). The core of Borealis is modeling Beta-Binomial distributions. This can be useful for rare disease diagnoses.
Last updated
sequencingcoveragednamethylationdifferentialmethylation
3.66 score 23 scripts 343 downloadspgxRpi - R wrapper for Progenetix
The package is an R wrapper for Progenetix REST API built upon the Beacon v2 protocol. Its purpose is to provide a seamless way for retrieving genomic data from Progenetix database—an open resource dedicated to curated oncogenomic profiles. Empowered by this package, users can effortlessly access and visualize data from Progenetix.
Last updated
copynumbervariationgenomicvariationdataimportsoftware
3.65 score 3 stars 9 scripts 288 downloadsscviR - experimental inferface from R to scvi-tools
This package defines interfaces from R to scvi-tools. A vignette works through the totalVI tutorial for analyzing CITE-seq data. Another vignette compares outputs of Chapter 12 of the OSCA book with analogous outputs based on totalVI quantifications. Future work will address other components of scvi-tools, with a focus on building understanding of probabilistic methods based on variational autoencoders.
Last updated
infrastructuresinglecelldataimportbioconductorcite-seqczieoss6-biocgpuscverse
3.62 score 7 stars 12 scripts 202 downloadsGWAS.BAYES - Bayesian analysis of Gaussian GWAS data
This package is built to perform GWAS analysis using Bayesian techniques. Currently, GWAS.BAYES has functionality for the implementation of BICOSS (Williams, J., Ferreira, M. A., and Ji, T. (2022). BICOSS: Bayesian iterative conditional stochastic search for GWAS. BMC Bioinformatics), BGWAS (Williams, J., Xu, S., Ferreira, M. A.. (2023) "BGWAS: Bayesian variable selection in linear mixed models with nonlocal priors for genome-wide association studies." BMC Bioinformatics), and GINA. All methods currently are for the analysis of Gaussian phenotypes The research related to this package was supported in part by National Science Foundation awards DMS 1853549, DMS 1853556, and DMS 2054173.
Last updated
bayesianassaydomainsnpgenomewideassociation
3.60 score 8 scripts 374 downloadsASURAT - Functional annotation-driven unsupervised clustering for single-cell data
ASURAT is a software for single-cell data analysis. Using ASURAT, one can simultaneously perform unsupervised clustering and biological interpretation in terms of cell type, disease, biological process, and signaling pathway activity. Inputting a single-cell RNA-seq data and knowledge-based databases, such as Cell Ontology, Gene Ontology, KEGG, etc., ASURAT transforms gene expression tables into original multivariate tables, termed sign-by-sample matrices (SSMs).
Last updated
geneexpressionsinglecellsequencingclusteringgenesignalingcpp
3.52 score 22 scripts 441 downloadsBASiCStan - Stan implementation of BASiCS
Provides an interface to infer the parameters of BASiCS using the variational inference (ADVI), Markov chain Monte Carlo (NUTS), and maximum a posteriori (BFGS) inference engines in the Stan programming language. BASiCS is a Bayesian hierarchical model that uses an adaptive Metropolis within Gibbs sampling scheme. Alternative inference methods provided by Stan may be preferable in some situations, for example for particularly large data or posterior distributions with difficult geometries.
Last updated
immunooncologynormalizationsequencingrnaseqsoftwaregeneexpressiontranscriptomicssinglecelldifferentialexpressionbayesiancellbiologysingle-cell-rna-seqonetbbcpp
3.48 score 2 scripts 306 downloadsBiostrings - Efficient manipulation of biological strings
Memory efficient string containers, string matching algorithms, and other utilities, for fast manipulation of large biological sequences or sets of sequences.
Last updated
sequencematchingalignmentsequencinggeneticsdataimportdatarepresentationinfrastructurebioconductor-packagecore-package
17.78 score 69 stars 1.2k dependents 13k scripts 104k downloadsGSVA - Gene Set Variation Analysis for Microarray and RNA-Seq Data
Gene Set Variation Analysis (GSVA) is a non-parametric, unsupervised method for estimating variation of gene set enrichment through the samples of a expression data set. GSVA performs a change in coordinate systems, transforming the data from a gene by sample matrix to a gene-set by sample matrix, thereby allowing the evaluation of pathway enrichment for each sample. This new matrix of GSVA enrichment scores facilitates applying standard analytical methods like functional enrichment, survival analysis, clustering, CNV-pathway analysis or cross-tissue pathway analysis, in a pathway-centric manner.
Last updated
functionalgenomicsmicroarrayrnaseqpathwaysgenesetenrichmentgene-set-enrichmentgenomicspathway-enrichment-analysis
15.68 score 247 stars 21 dependents 2.8k scripts 17k downloadsMSnbase - Base Functions and Classes for Mass Spectrometry and Proteomics
MSnbase provides infrastructure for manipulation, processing and visualisation of mass spectrometry and proteomics data, ranging from raw to quantitative and annotated data.
Last updated
immunooncologyinfrastructureproteomicsmassspectrometryqualitycontroldataimportbioconductorbioinformaticsmass-spectrometryproteomics-datavisualisationcpp
14.95 score 137 stars 37 dependents 1.2k scripts 6.8k downloadsBiocGenerics - S4 generic functions used in Bioconductor
The package defines many S4 generic functions used in Bioconductor.
Last updated
infrastructurebioconductor-packagecore-package
14.77 score 14 stars 2.5k dependents 1.3k scripts 142k downloads
xcms - LC-MS and GC-MS Data Analysis
Framework for processing and visualization of chromatographically separated and single-spectra mass spectral data. Imports from AIA/ANDI NetCDF, mzXML, mzData and mzML files. Preprocesses data for high-throughput, untargeted analyte profiling.
Last updated
immunooncologymassspectrometrymetabolomicsbioconductorfeature-detectionmass-spectrometrypeak-detectioncpp
14.66 score 231 stars 14 dependents 1.4k scripts 4.6k downloadsensembldb - Utilities to create and use Ensembl-based annotation databases
The package provides functions to create and use transcript centric annotation databases/packages. The annotation for the databases are directly fetched from Ensembl using their Perl API. The functionality and data is similar to that of the TxDb packages from the GenomicFeatures package, but, in addition to retrieve all gene/transcript models and annotations from the database, ensembldb provides a filter framework allowing to retrieve annotations for specific entries like genes encoded on a chromosome region or transcript models of lincRNA genes. EnsDb databases built with ensembldb contain also protein annotations and mappings between proteins and their encoding transcripts. Finally, ensembldb provides functions to map between genomic, transcript and protein coordinates.
Last updated
geneticsannotationdatasequencingcoverageannotationbioconductorbioconductor-packagesensembl
14.55 score 36 stars 113 dependents 1.7k scripts 19k downloadslimma - Linear Models for Microarray and Omics Data
Data analysis, linear models and differential expression for omics data.
Last updated
exonarraygeneexpressiontranscriptionalternativesplicingdifferentialexpressiondifferentialsplicinggenesetenrichmentdataimportbayesianclusteringregressiontimecoursemicroarraymicrornaarraymrnamicroarrayonechannelproprietaryplatformstwochannelsequencingrnaseqbatcheffectmultiplecomparisonnormalizationpreprocessingqualitycontrolbiomedicalinformaticscellbiologycheminformaticsepigeneticsfunctionalgenomicsgeneticsimmunooncologymetabolomicsproteomicssystemsbiologytranscriptomics
13.88 score 643 dependents 24k scripts 69k downloadsGviz - Plotting data and annotation information along genomic coordinates
Genomic data analyses requires integrated visualization of known genomic information and new experimental data. Gviz uses the biomaRt and the rtracklayer packages to perform live annotation queries to Ensembl and UCSC and translates this to e.g. gene/transcript structures in viewports of the grid graphics package. This results in genomic information plotted together with your data.
Last updated
visualizationmicroarraysequencing
13.50 score 92 stars 46 dependents 2.0k scripts 6.3k downloadsedgeR - Empirical Analysis of Digital Gene Expression Data in R
Differential expression analysis of sequence count data. Implements a range of statistical methodology based on the negative binomial distributions, including empirical Bayes estimation, exact tests, generalized linear models, quasi-likelihood, and gene set enrichment. Can perform differential analyses of any type of omics data that produces read counts, including RNA-seq, ChIP-seq, ATAC-seq, Bisulfite-seq, SAGE, CAGE, metabolomics, or proteomics spectral counts. RNA-seq analyses can be conducted at the gene or isoform level, and tests can be conducted for differential exon or transcript usage.
Last updated
alternativesplicingbatcheffectbayesianbiomedicalinformaticscellbiologychipseqclusteringcoveragedifferentialexpressiondifferentialmethylationdifferentialsplicingdnamethylationepigeneticsfunctionalgenomicsgeneexpressiongenesetenrichmentgeneticsimmunooncologymultiplecomparisonnormalizationpathwaysproteomicsqualitycontrolregressionrnaseqsagesequencingsinglecellsystemsbiologytimecoursetranscriptiontranscriptomicsopenblas
13.46 score 289 dependents 28k scripts 43k downloadsSingleR - Reference-Based Single-Cell RNA-Seq Annotation
Performs unbiased cell type recognition from single-cell RNA sequencing data, by leveraging reference transcriptomic datasets of pure cell types to infer the cell of origin of each single cell independently.
Last updated
softwaresinglecellgeneexpressiontranscriptomicsclassificationclusteringannotationbioconductorsinglercpp
13.22 score 204 stars 3 dependents 3.8k scripts 7.5k downloadsExperimentHub - Client to access ExperimentHub resources
This package provides a client for the Bioconductor ExperimentHub web resource. ExperimentHub provides a central location where curated data from experiments, publications or training courses can be accessed. Each resource has associated metadata, tags and date of modification. The client creates and manages a local cache of files retrieved enabling quick and reproducible access.
Last updated
infrastructuredataimportguithirdpartyclientcore-packageu24ca289073
12.67 score 11 stars 73 dependents 1.8k scripts 13k downloadsmsa - Multiple Sequence Alignment
The 'msa' package provides a unified R/Bioconductor interface to the multiple sequence alignment algorithms ClustalW, ClustalOmega, and Muscle. All three algorithms are integrated in the package, therefore, they do not depend on any external software tools and are available for all major platforms. The multiple sequence alignment algorithms are complemented by a function for pretty-printing multiple sequence alignments using the LaTeX package TeXshade.
Last updated
multiplesequencealignmentalignmentmultiplecomparisonsequencingcpp
12.11 score 26 stars 8 dependents 1.3k scripts 3.4k downloadsuniversalmotif - Import, Modify, and Export Motifs with R
Allows for importing most common motif types into R for use by functions provided by other Bioconductor motif-related packages. Motifs can be exported into most major motif formats from various classes as defined by other Bioconductor packages. A suite of motif and sequence manipulation and analysis functions are included, including enrichment, comparison, P-value calculation, shuffling, trimming, higher-order motifs, and others.
Last updated
motifannotationmotifdiscoverydataimportgeneregulationmotif-analysismotif-enrichment-analysissequence-logocpp
11.59 score 34 stars 14 dependents 729 scripts 2.1k downloadsanndataR - AnnData interoperability in R
Bring the power and flexibility of AnnData to the R ecosystem, allowing you to effortlessly manipulate and analyse your single-cell data. This package lets you work with backed h5ad and zarr files, directly access various slots (e.g. X, obs, var), or convert the data into SingleCellExperiment and Seurat objects.
Last updated
singlecelldataimportdatarepresentationanndatah5adinteroperability
11.55 score 192 stars 3 dependents 314 scriptsmiloR - Differential neighbourhood abundance testing on a graph
Milo performs single-cell differential abundance testing. Cell states are modelled as representative neighbourhoods on a nearest neighbour graph. Hypothesis testing is performed using either a negative bionomial generalized linear model or negative binomial generalized linear mixed model.
Last updated
singlecellmultiplecomparisonfunctionalgenomicssoftwareopenblascppopenmp
11.52 score 440 stars 2 dependents 664 scripts 1.7k downloadsDECIPHER - Tools for curating, analyzing, and manipulating biological sequences
A toolset for deciphering and managing biological sequences.
Last updated
clusteringgeneticssequencingdataimportvisualizationmicroarrayqualitycontrolqpcralignmentwholegenomemicrobiomeimmunooncologygenepredictionphylogeneticscomparativegenomicsopenmp
10.68 score 21 dependents 1.5k scripts 5.8k downloadsAnVIL - Bioconductor on the AnVIL compute environment
The AnVIL is a cloud computing resource developed in part by the National Human Genome Research Institute. The AnVIL package provides programatic access to the Dockstore, Leonardo, Rawls, TDR, and Terra RESTful programming interfaces. For platform-specific user-level functionality, see either the AnVILGCP or AnVILAz package.
Last updated
infrastructureu24hg010263
10.47 score 7 stars 12 dependents 333 scripts 1.2k downloads
rhdf5filters - HDF5 Compression Filters
Provides a collection of additional compression filters for HDF5 datasets. The package is intended to provide seamless integration with rhdf5, however the compiled filters can also be used with external applications.
Last updated
infrastructuredataimportcompressionfilter-pluginhdf5
10.35 score 5 stars 231 dependents 9 scripts 46k downloadsdiffcyt - Differential discovery in high-dimensional cytometry via high-resolution clustering
Statistical methods for differential discovery analyses in high-dimensional cytometry data (including flow cytometry, mass cytometry or CyTOF, and oligonucleotide-tagged cytometry), based on a combination of high-resolution clustering and empirical Bayes moderated tests adapted from transcriptomics.
Last updated
immunooncologyflowcytometryproteomicssinglecellcellbasedassayscellbiologyclusteringfeatureextractionsoftware
10.25 score 25 stars 5 dependents 451 scripts 930 downloadsTCGAutils - TCGA utility functions for data management
A suite of helper functions for checking and manipulating TCGA data including data obtained from the curatedTCGAData experiment package. These functions aim to simplify and make working with TCGA data more manageable. Exported functions include those that import data from flat files into Bioconductor objects, convert row annotations, and identifier translation via the GDC API.
Last updated
softwareworkflowsteppreprocessingdataimportbioconductor-packagetcgau24ca289073utilities
10.20 score 31 stars 11 dependents 259 scripts 2.0k downloadsmethylumi - Handle Illumina methylation data
This package provides classes for holding and manipulating Illumina methylation data. Based on eSet, it can contain MIAME information, sample information, feature information, and multiple matrices of data. An "intelligent" import function, methylumiR can read the Illumina text files and create a MethyLumiSet. methylumIDAT can directly read raw IDAT files from HumanMethylation27 and HumanMethylation450 microarrays. Normalization, background correction, and quality control features for GoldenGate, Infinium, and Infinium HD arrays are also included.
Last updated
dnamethylationtwochannelpreprocessingqualitycontrolcpgislandquarto
10.02 score 9 stars 17 dependents 113 scripts 2.5k downloadsMassSpecWavelet - Peak Detection for Mass Spectrometry data using wavelet-based algorithms
Peak Detection in Mass Spectrometry data is one of the important preprocessing steps. The performance of peak detection affects subsequent processes, including protein identification, profile alignment and biomarker identification. Using Continuous Wavelet Transform (CWT), this package provides a reliable algorithm for peak detection that does not require any type of smoothing or previous baseline correction method, providing more consistent results for different spectra. See <doi:10.1093/bioinformatics/btl355} for further details.
Last updated
immunooncologymassspectrometryproteomicspeakdetection
9.94 score 11 stars 21 dependents 49 scripts 4.0k downloadsh5mread - A fast HDF5 reader
The main function in the h5mread package is h5mread(), which allows reading arbitrary data from an HDF5 dataset into R, similarly to what the h5read() function from the rhdf5 package does. In the case of h5mread(), the implementation has been optimized to make it as fast and memory-efficient as possible.
Last updated
infrastructuredatarepresentationdataimportu24ca289073curlopenssl
9.50 score 3 stars 162 dependents 4 scripts 18k downloadsBanksy - Spatial transcriptomic clustering
Banksy is an R package that incorporates spatial information to cluster cells in a feature space (e.g. gene expression). To incorporate spatial information, BANKSY computes the mean neighborhood expression and azimuthal Gabor filters that capture gene expression gradients. These features are combined with the cell's own expression to embed cells in a neighbor-augmented product space which can then be clustered, allowing for accurate and spatially-aware cell typing and tissue domain segmentation.
Last updated
clusteringspatialsinglecellgeneexpressiondimensionreductionclustering-algorithmsingle-cell-omicsspatial-omics
9.42 score 157 stars 698 scripts 938 downloadsassorthead - Assorted Header-Only C++ Libraries
Vendors an assortment of useful header-only C++ libraries. Bioconductor packages can use these libraries in their own C++ code by LinkingTo this package without introducing any additional dependencies. The use of a central repository avoids duplicate vendoring of libraries across multiple R packages, and enables better coordination of version updates across cohorts of interdependent C++ libraries.
Last updated
singlecellqualitycontrolnormalizationdatarepresentationdataimportdifferentialexpressionalignment
9.27 score 1 stars 217 dependents 24k downloadsBatchQC - Batch Effects Quality Control Software
Sequencing and microarray samples often are collected or processed in multiple batches or at different times. This often produces technical biases that can lead to incorrect results in the downstream analysis. BatchQC is a software tool that streamlines batch preprocessing and evaluation by providing interactive diagnostics, visualizations, and statistical analyses to explore the extent to which batch variation impacts the data. BatchQC diagnostics help determine whether batch adjustment needs to be done, and how correction should be applied before proceeding with a downstream analysis. Moreover, BatchQC interactively applies multiple common batch effect approaches to the data and the user can quickly see the benefits of each method. BatchQC is developed as a Shiny App. The output is organized into multiple tabs and each tab features an important part of the batch effect analysis and visualization of the data. The BatchQC interface has the following analysis groups: Summary, Differential Expression, Median Correlations, Heatmaps, Circular Dendrogram, PCA Analysis, Shape, ComBat and SVA.
Last updated
batcheffectgeneexpressiongraphandnetworkmicroarraynormalizationprincipalcomponentsequencingsoftwarevisualizationqualitycontrolrnaseqpreprocessingdifferentialexpressionimmunooncology
9.27 score 11 stars 68 scripts 518 downloadsUniProt.ws - R Interface to UniProt Web Services
The Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. This package provides a collection of functions for retrieving, processing, and re-packaging UniProt web services. The package makes use of UniProt's modernized REST API and allows mapping of identifiers accross different databases.
Last updated
annotationinfrastructuregokeggbiocartabioconductor-packagecore-package
9.23 score 10 stars 6 dependents 271 scripts 1.4k downloadsscrapper - Bindings to C++ Libraries for Single-Cell Analysis
Implements R bindings to C++ code for analyzing single-cell (expression) data, mostly from various libscran libraries. Each function performs an individual step in the single-cell analysis workflow, ranging from quality control to clustering and marker detection. Additional wrappers are provided for easy construction of end-to-end workflows involving Bioconductor objects like SingleCellExperiments.
Last updated
normalizationrnaseqsoftwaregeneexpressiontranscriptomicssinglecellbatcheffectqualitycontroldifferentialexpressionfeatureextractionprincipalcomponentclusteringopenblascpp
9.11 score 8 stars 14 dependents 131 scripts 2.0k downloadscigarillo - Efficient manipulation of CIGAR strings
CIGAR stands for Concise Idiosyncratic Gapped Alignment Report. CIGAR strings are found in the BAM files produced by most aligners and in the AIRR-formatted output produced by IgBLAST. The cigarillo package provides functions to parse and inspect CIGAR strings, trim them, turn them into ranges of positions relative to the "query space" or "reference space", and project positions or sequences from one space to the other. Note that these operations are low-level operations that the user rarely needs to perform directly. More typically, they are performed behind the scene by higher-level functionality implemented in other packages like Bioconductor packages GenomicAlignments and igblastr.
Last updated
infrastructurealignmentsequencematchingsequencingbioconductor-packagecore-package
9.07 score 566 dependents 8 scripts 23k downloadsbamsignals - Extract read count signals from bam files
This package allows to efficiently obtain count vectors from indexed bam files. It counts the number of reads in given genomic ranges and it computes reads profiles and coverage profiles. It also handles paired-end data.
Last updated
dataimportsequencingcoveragealignmentcurlbzip2xz-utilszlibcpp
9.06 score 15 stars 8 dependents 58 scripts 2.2k downloadsRarr - Read Zarr Files in R
The Zarr specification defines a format for chunked, compressed, N-dimensional arrays. It's design allows efficient access to subsets of the stored array, and supports both local and cloud storage systems. Rarr aims to implement this specification in R with minimal reliance on an external tools or libraries.
Last updated
dataimportbioconductorome-ngffome-zarron-diskout-of-memoryzarrc-blosclibzstd
8.97 score 54 stars 7 dependents 89 scripts 759 downloadsSeqVarTools - Tools for variant data
An interface to the fast-access storage format for VCF data provided in SeqArray, with tools for common operations and analysis.
Last updated
snpgeneticvariabilitysequencinggenetics
8.94 score 3 stars 2 dependents 570 scripts 954 downloadsBiocPkgTools - Collection of simple tools for learning about Bioconductor Packages
Bioconductor has a rich ecosystem of metadata around packages, usage, and build status. This package is a simple collection of functions to access that metadata from R. The goal is to expose metadata for data mining and value-added functionality such as package searching, text mining, and analytics on packages.
Last updated
softwareinfrastructurebioconductormetadatau24ca289073
8.80 score 22 stars 1 dependents 143 scripts 638 downloadsimcRtools - Methods for imaging mass cytometry data analysis
This R package supports the handling and analysis of imaging mass cytometry and other highly multiplexed imaging data. The main functionality includes reading in single-cell data after image segmentation and measurement, data formatting to perform channel spillover correction and a number of spatial analysis approaches. First, cell-cell interactions are detected via spatial graph construction; these graphs can be visualized with cells representing nodes and interactions representing edges. Furthermore, per cell, its direct neighbours are summarized to allow spatial clustering. Per image/grouping level, interactions between types of cells are counted, averaged and compared against random permutations. In that way, types of cells that interact more (attraction) or less (avoidance) frequently than expected by chance are detected.
Last updated
immunooncologysinglecellspatialdataimportclusteringimcsingle-cell
8.77 score 33 stars 444 scripts 680 downloadsCOTAN - COexpression Tables ANalysis
Statistical and computational method to analyze the co-expression of gene pairs at single cell level. It provides the foundation for single-cell gene interactome analysis. The basic idea is studying the zero UMI counts' distribution instead of focusing on positive counts; this is done with a generalized contingency tables framework. COTAN can effectively assess the correlated or anti-correlated expression of gene pairs. It provides a numerical index related to the correlation and an approximate p-value for the associated independence test. COTAN can also evaluate whether single genes are differentially expressed, scoring them with a newly defined global differentiation index. Moreover, this approach provides ways to plot and cluster genes according to their co-expression pattern with other genes, effectively helping the study of gene interactions and becoming a new tool to identify cell-identity marker genes.
Last updated
systemsbiologytranscriptomicsgeneexpressionsinglecelldifferentialexpressionclusteringgpu
8.29 score 17 stars 136 scripts 436 downloadsHIBAG - HLA Genotype Imputation with Attribute Bagging
Imputes HLA classical alleles using GWAS SNP data, and it relies on a training set of HLA and SNP genotypes. HIBAG can be used by researchers with published parameter estimates instead of requiring access to large training sample datasets. It combines the concepts of attribute bagging, an ensemble classifier method, with haplotype inference for SNPs and HLA types. Attribute bagging is a technique which improves the accuracy and stability of classifier ensembles using bootstrap aggregating and random variable selection.
Last updated
geneticsstatisticalmethodbioinformaticsgpuhlaimputationmhcsnpcpp
8.27 score 31 stars 66 scripts 577 downloadsbeadarray - Quality assessment and low-level analysis for Illumina BeadArray data
The package is able to read bead-level data (raw TIFFs and text files) output by BeadScan as well as bead-summary data from BeadStudio. Methods for quality assessment and low-level analysis are provided.
Last updated
microarrayonechannelqualitycontrolpreprocessing
8.11 score 3 dependents 90 scripts 1.3k downloads
SpectriPy - Enhancing Cross-Language Mass Spectrometry Data Analysis with R and Python
The SpectriPy package allows integration of Python-based MS analysis code with the Spectra package. Spectra objects can be converted into Python MS data structures. In addition, SpectriPy integrates and wraps the similarity scoring and processing/filtering functions from the Python matchms package into R.
Last updated
infrastructuremetabolomicsmassspectrometryproteomicsmass-spectrometrypythonquarto
8.10 score 13 stars 62 scripts 174 downloadsscDiagnostics - Cell type annotation diagnostics
The scDiagnostics package provides diagnostic plots to assess the quality of cell type assignments from single cell gene expression profiles. The implemented functionality allows to assess the reliability of cell type annotations, investigate gene expression patterns, and explore relationships between different cell types in query and reference datasets allowing users to detect potential misalignments between reference and query datasets. The package also provides visualization capabilities for diagnostics purposes.
Last updated
annotationclassificationclusteringgeneexpressionrnaseqsinglecellsoftwaretranscriptomics
8.07 score 14 stars 76 scripts 255 downloadsnnSVG - Scalable identification of spatially variable genes in spatially-resolved transcriptomics data
Method for scalable identification of spatially variable genes (SVGs) in spatially-resolved transcriptomics data. The method is based on nearest-neighbor Gaussian processes and uses the BRISC algorithm for model fitting and parameter estimation. Allows identification and ranking of SVGs with flexible length scales across a tissue slide or within spatial domains defined by covariates. Scales linearly with the number of spatial locations and can be applied to datasets containing thousands or more spatial locations.
Last updated
spatialsinglecelltranscriptomicsgeneexpressionpreprocessing
7.90 score 25 stars 1 dependents 352 scripts 514 downloadslemur - Latent Embedding Multivariate Regression
Fit a latent embedding multivariate regression (LEMUR) model to multi-condition single-cell data. The model provides a parametric description of single-cell data measured with treatment vs. control or more complex experimental designs. The parametric model is used to (1) align conditions, (2) predict log fold changes between conditions for all cells, and (3) identify cell neighborhoods with consistent log fold changes. For those neighborhoods, a pseudobulked differential expression test is conducted to assess which genes are significantly changed.
Last updated
transcriptomicsdifferentialexpressionsinglecelldimensionreductionregressionquartoopenblascpp
7.88 score 102 stars 92 scripts 356 downloadschromVAR - Chromatin Variation Across Regions
Determine variation in chromatin accessibility across sets of annotations or peaks. Designed primarily for single-cell or sparse chromatin accessibility data, e.g. from scATAC-seq or sparse bulk ATAC or DNAse-seq experiments.
Last updated
singlecellsequencinggeneregulationimmunooncologycpp
7.84 score 1.6k scripts 2.9k downloads
MsBackendMgf - Mass Spectrometry Data Backend for Mascot Generic Format (mgf) Files
Mass spectrometry (MS) data backend supporting import and export of MS/MS spectra data from Mascot Generic Format (mgf) files. Objects defined in this package are supposed to be used with the Spectra Bioconductor package. This package thus adds mgf file support to the Spectra package.
Last updated
infrastructureproteomicsmassspectrometrymetabolomicsdataimport
7.79 score 5 stars 2 dependents 125 scripts 1.6k downloadsPhyloProfile - PhyloProfile
PhyloProfile is a tool for exploring complex phylogenetic profiles. Phylogenetic profiles, presence/absence patterns of genes over a set of species, are commonly used to trace the functional and evolutionary history of genes across species and time. With PhyloProfile we can enrich regular phylogenetic profiles with further data like sequence/structure similarity, to make phylogenetic profiling more meaningful. Besides the interactive visualisation powered by R-Shiny, the package offers a set of further analysis features to gain insights like the gene age estimation or core gene identification.
Last updated
softwarevisualizationdatarepresentationmultiplecomparisonfunctionalpredictiondimensionreductionbioinformaticsheatmapinteractive-visualizationsorthologsphylogenetic-profileshiny
7.77 score 38 stars 16 scripts 520 downloadsDiffBind - Differential Binding Analysis of ChIP-Seq Peak Data
Compute differentially bound sites from multiple ChIP-seq experiments using affinity (quantitative) data. Also enables occupancy (overlap) analysis and plotting functions.
Last updated
sequencingchipseqatacseqdnaseseqmethylseqripseqdifferentialpeakcallingdifferentialmethylationgeneregulationhistonemodificationpeakdetectionbiomedicalinformaticscellbiologymultiplecomparisonnormalizationreportwritingepigeneticsfunctionalgenomicscurlbzip2xz-utilszlibcpp
7.74 score 2 dependents 1.2k scripts 2.6k downloadspipeComp - pipeComp pipeline benchmarking framework
A simple framework to facilitate the comparison of pipelines involving various steps and parameters. The `pipelineDefinition` class represents pipelines as, minimally, a set of functions consecutively executed on the output of the previous one, and optionally accompanied by step-wise evaluation and aggregation functions. Given such an object, a set of alternative parameters/methods, and benchmark datasets, the `runPipeline` function then proceeds through all combinations arguments, avoiding recomputing the same step twice and compiling evaluations on the fly to avoid storing potentially large intermediate data.
Last updated
geneexpressiontranscriptomicsclusteringdatarepresentationbenchmarkbioconductorpipeline-benchmarkingpipelinessingle-cell-rna-seq
7.47 score 44 stars 45 scripts 330 downloadsggspavis - Visualization functions for spatial transcriptomics data
Visualization functions for spatial transcriptomics data. Includes functions to generate several types of plots, including spot plots, feature (molecule) plots, reduced dimension plots, spot-level quality control (QC) plots, and feature-level QC plots, for datasets from the 10x Genomics Visium and other technological platforms. Datasets are assumed to be in either SpatialExperiment or SingleCellExperiment format.
Last updated
spatialsinglecelltranscriptomicsgeneexpressionqualitycontroldimensionreduction
7.43 score 5 stars 670 scripts 688 downloadsSharedObject - Sharing R objects across multiple R processes without memory duplication
This package is developed for facilitating parallel computing in R. It is capable to create an R object in the shared memory space and share the data across multiple R processes. It avoids the overhead of memory dulplication and data transfer, which make sharing big data object across many clusters possible.
Last updated
infrastructuresharedobjectcpp
7.38 score 51 stars 1 dependents 13 scripts 362 downloadsMicrobiomeProfiler - An R/shiny package for microbiome functional enrichment analysis
This is an R/shiny package to perform functional enrichment analysis for microbiome data. This package was based on clusterProfiler. Moreover, MicrobiomeProfiler support KEGG enrichment analysis, COG enrichment analysis, Microbe-Disease association enrichment analysis, Metabo-Pathway analysis.
Last updated
microbiomesoftwarevisualizationkegg
7.37 score 42 stars 45 scripts 391 downloadsATACseqQC - ATAC-seq Quality Control
ATAC-seq, an assay for Transposase-Accessible Chromatin using sequencing, is a rapid and sensitive method for chromatin accessibility analysis. It was developed as an alternative method to MNase-seq, FAIRE-seq and DNAse-seq. Comparing to the other methods, ATAC-seq requires less amount of the biological samples and time to process. In the process of analyzing several ATAC-seq dataset produced in our labs, we learned some of the unique aspects of the quality assessment for ATAC-seq data.To help users to quickly assess whether their ATAC-seq experiment is successful, we developed ATACseqQC package partially following the guideline published in Nature Method 2013 (Greenleaf et al.), including diagnostic plot of fragment size distribution, proportion of mitochondria reads, nucleosome positioning pattern, and CTCF or other Transcript Factor footprints.
Last updated
sequencingdnaseqatacseqgeneregulationqualitycontrolcoveragenucleosomepositioningimmunooncology
7.37 score 1 dependents 388 scripts 989 downloadsggsc - Visualizing Single Cell and Spatial Transcriptomics
Useful functions to visualize single cell and spatial data. It supports visualizing 'Seurat', 'SingleCellExperiment' and 'SpatialExperiment' objects through grammar of graphics syntax implemented in 'ggplot2'.
Last updated
dimensionreductiongeneexpressionsinglecellsoftwarespatialtranscriptomicsvisualizationopenblascppopenmp
7.31 score 51 stars 38 scripts 336 downloadsSpectralTAD - SpectralTAD: Hierarchical TAD detection using spectral clustering
SpectralTAD is an R package designed to identify Topologically Associated Domains (TADs) from Hi-C contact matrices. It uses a modified version of spectral clustering that uses a sliding window to quickly detect TADs. The function works on a range of different formats of contact matrices and returns a bed file of TAD coordinates. The method does not require users to adjust any parameters to work and gives them control over the number of hierarchical levels to be returned.
Last updated
softwarehicsequencingfeatureextractionclustering
7.29 score 13 stars 25 scripts 473 downloadsplyxp - Data masks for SummarizedExperiment enabling dplyr-like manipulation
The package provides `rlang` data masks for the SummarizedExperiment class. The enables the evaluation of unquoted expression in different contexts of the SummarizedExperiment object with optional access to other contexts. The goal for `plyxp` is for evaluation to feel like a data.frame object without ever needing to unwind to a rectangular data.frame.
Last updated
annotationgenomeannotationtranscriptomics
7.09 score 9 stars 2 dependents 23 scripts 402 downloadsZarrArray - Bring Zarr datasets in R as DelayedArray objects
The ZarrArray package leverages the Rarr package to bring Zarr datasets in R as DelayedArray objects. The main class in the package is the ZarrArray class. A ZarrArray object is an array-like object that represents a Zarr dataset in R. ZarrArray objects are DelayedArray derivatives and therefore support all operations (delayed or block-processed) supported by DelayedArray objects.
Last updated
infrastructuredatarepresentationdataimportbioconductor-packagecore-packageu24ca289073
7.07 score 5 stars 5 dependents 13 scripts 628 downloads
plaid - PLAID ultrafast gene set enrichment scoring
PLAID (Pathway Level Average Intensity Detection) is an ultra-fast method to compute single-sample enrichment scores for gene expression or proteomics data. For each sample, plaid computes the gene set score as the average intensity of the genes/proteins in the gene set. The output is a gene set score matrix suitable for further analyses.
Last updated
genesetenrichmentgeneexpressionproteomicsbioinformaticsenrichment-analysisomics-datarna-seq-analysis
7.06 score 24 stars 3 scripts 157 downloads
crumblr - Count ratio uncertainty modeling base linear regression
Crumblr enables analysis of count ratio data using precision weighted linear (mixed) models. It uses an asymptotic normal approximation of the variance following the centered log ration transform (CLR) that is widely used in compositional data analysis. Crumblr provides a fast, flexible alternative to GLMs and GLMM's while retaining high power and controlling the false positive rate.
Last updated
rnaseqgeneexpressiondifferentialexpressionbatcheffectqualitycontrolsinglecellregressionepigeneticsfunctionalgenomicstranscriptomicsnormalizationclusteringdimensionreductionpreprocessingsoftware
7.04 score 7 stars 70 scripts 268 downloadsMsDataHub - Mass Spectrometry Data on ExperimentHub
The MsDataHub package uses the ExperimentHub infrastructure to distribute raw mass spectrometry data files, peptide spectrum matches or quantitative data from proteomics and metabolomics experiments.
Last updated
experimenthubsoftwaremassspectrometryproteomicsmetabolomicsbioconductordatamass-spectrometry
6.96 score 2 stars 1 dependents 108 scripts 430 downloadsRedeR - Interactive visualization and manipulation of nested networks
RedeR combines an R package with a stand-alone Java application for interactive visualization and manipulation of nested networks. Graph, node, and edge attributes can be configured using either graphical or command-line methods, following igraph syntax rules.
Last updated
guigraphandnetworknetworknetworkenrichmentnetworkinferencesoftwaresystemsbiology
6.95 score 7 dependents 106 scripts 724 downloadsrols - An R interface to the Ontology Lookup Service
The rols package is an interface to the Ontology Lookup Service (OLS) to access and query hundred of ontolgies directly from R.
Last updated
immunooncologysoftwareannotationmassspectrometrygo
6.89 score 11 stars 1 dependents 106 scriptsigblastr - User-friendly R Wrapper to IgBLAST
The igblastr package provides functions to conveniently install and use a local IgBLAST installation from within R. The package also includes a set of preinstalled IgBLAST-compatible germline databases from OGRDB, the AIRR Community’s Open Germline Receptor Database, for various organisms. It provides functions to install additional IgBLAST-compatible germline databases using reference sequences retrieved from IMGT/V-QUEST or OGRDB, or from local FASTA files supplied by the user. When possible, annotations for the V and J alleles in a new germline database are automatically generated and added to the database, so they can be used as replacements for the internal and auxiliary data provided by IgBLAST. IgBLAST is described at <https://pubmed.ncbi.nlm.nih.gov/23671333/>. IgBLAST web interface: <https://www.ncbi.nlm.nih.gov/igblast/>. OGRDB: <https://ogrdb.airr-community.org/>. IMGT/V-QUEST download site: <https://www.imgt.org/download/V-QUEST/>.
Last updated
immunologyimmunogeneticsimmunooncologycellbiologybioconductor-package
6.86 score 5 stars 24 scripts 396 downloadsMSstatsConvert - Import Data from Various Mass Spectrometry Signal Processing Tools to MSstats Format
MSstatsConvert provides tools for importing reports of Mass Spectrometry data processing tools into R format suitable for statistical analysis using the MSstats and MSstatsTMT packages.
Last updated
massspectrometryproteomicssoftwaredataimportqualitycontrolcpp
6.75 score 8 dependents 42 scripts 1.0k downloadsspatialHeatmap - spatialHeatmap: Visualizing Spatial Assays in Anatomical Images and Large-Scale Data Extensions
The spatialHeatmap package offers the primary functionality for visualizing cell-, tissue- and organ-specific assay data in spatial anatomical images. Additionally, it provides extended functionalities for large-scale data mining routines and co-visualizing bulk and single-cell data. A description of the project is available here: https://spatialheatmap.org.
Last updated
spatialvisualizationmicroarraysequencinggeneexpressiondatarepresentationnetworkclusteringgraphandnetworkcellbasedassaysatacseqdnaseqtissuemicroarraysinglecellcellbiologygenetarget
6.70 score 7 stars 20 scripts 448 downloadsnormr - Normalization and difference calling in ChIP-seq data
Robust normalization and difference calling procedures for ChIP-seq and alike data. Read counts are modeled jointly as a binomial mixture model with a user-specified number of components. A fitted background estimate accounts for the effect of enrichment in certain regions and, therefore, represents an appropriate null hypothesis. This robust background is used to identify significantly enriched or depleted regions.
Last updated
bayesiandifferentialpeakcallingclassificationdataimportchipseqripseqfunctionalgenomicsgeneticsmultiplecomparisonnormalizationpeakdetectionpreprocessingalignmentcppopenmp
6.60 score 11 stars 30 scripts 552 downloadsCellMentor - Supervised Non-negative Matrix Factorization for Dimensional Reduction in Single-Cell Analysis
Implements supervised cell type-aware non-negative matrix factorization (NMF) for dimensional reduction in single-cell RNA sequencing analysis. The package provides methods for incorporating cell type information into the dimensionality reduction process, enabling improved visualization and downstream analysis of single-cell data while preserving biological structure. CellMentor employs a unique loss function that simultaneously minimizes variation within known cell populations while maximizing distinctions between different cell types, enabling effective transfer of learned patterns from labeled reference datasets to new unlabeled data.
Last updated
softwaresinglecelltranscriptomicsdimensionreduction
6.45 score 19 stars 37 scripts 195 downloadsTADCompare - TADCompare: Identification and characterization of differential TADs
TADCompare is an R package designed to identify and characterize differential Topologically Associated Domains (TADs) between multiple Hi-C contact matrices. It contains functions for finding differential TADs between two datasets, finding differential TADs over time and identifying consensus TADs across multiple matrices. It takes all of the main types of HiC input and returns simple, comprehensive, easy to analyze results.
Last updated
softwarehicsequencingfeatureextractionclustering
6.35 score 27 stars 26 scripts 446 downloadsBgeeCall - Automatic RNA-Seq present/absent gene expression calls generation
BgeeCall allows to generate present/absent gene expression calls without using an arbitrary cutoff like TPM<1. Calls are generated based on reference intergenic sequences. These sequences are generated based on expression of all RNA-Seq libraries of each species integrated in Bgee (https://bgee.org).
Last updated
softwaregeneexpressionrnaseqbiologygene-expressiongene-levelintergenic-regionspresent-absent-callsrna-seqrna-seq-librariesscrna-seq
6.28 score 5 stars 12 scripts 448 downloadsdistinct - distinct: a method for differential analyses via hierarchical permutation tests
distinct is a statistical method to perform differential testing between two or more groups of distributions; differential testing is performed via hierarchical non-parametric permutation tests on the cumulative distribution functions (cdfs) of each sample. While most methods for differential expression target differences in the mean abundance between conditions, distinct, by comparing full cdfs, identifies, both, differential patterns involving changes in the mean, as well as more subtle variations that do not involve the mean (e.g., unimodal vs. bi-modal distributions with the same mean). distinct is a general and flexible tool: due to its fully non-parametric nature, which makes no assumptions on how the data was generated, it can be applied to a variety of datasets. It is particularly suitable to perform differential state analyses on single cell data (i.e., differential analyses within sub-populations of cells), such as single cell RNA sequencing (scRNA-seq) and high-dimensional flow or mass cytometry (HDCyto) data. To use distinct one needs data from two or more groups of samples (i.e., experimental conditions), with at least 2 samples (i.e., biological replicates) per group.
Last updated
geneticsrnaseqsequencingdifferentialexpressiongeneexpressionmultiplecomparisonsoftwaretranscriptionstatisticalmethodvisualizationsinglecellflowcytometrygenetargetopenblascpp
6.22 score 13 stars 1 dependents 43 scripts 513 downloadsPirat - Precursor or Peptide Imputation under Random Truncation
Pirat enables the imputation of missing values (either MNARs or MCARs) in bottom-up LC-MS/MS proteomics data using a penalized maximum likelihood strategy. It does not require any parameter tuning, it models the instrument censorship from the data available. It accounts for sibling peptides correlations and it can leverage complementary transcriptomics measurements.
Last updated
proteomicsmassspectrometrypreprocessingsoftwareprostar2
6.17 score 3 stars 11 scripts 302 downloadsrexposome - Exposome exploration and outcome data analysis
Package that allows to explore the exposome and to perform association analyses between exposures and health outcomes.
Last updated
softwarebiologicalquestioninfrastructuredataimportdatarepresentationbiomedicalinformaticsexperimentaldesignmultiplecomparisonclassificationclustering
6.16 score 1 dependents 32 scripts 536 downloadsCoSIA - An Investigation Across Different Species and Tissues
Cross-Species Investigation and Analysis (CoSIA) is a package that provides researchers with an alternative methodology for comparing across species and tissues using normal wild-type RNA-Seq Gene Expression data from Bgee. Using RNA-Seq Gene Expression data, CoSIA provides multiple visualization tools to explore the transcriptome diversity and variation across genes, tissues, and species. CoSIA uses the Coefficient of Variation and Shannon Entropy and Specificity to calculate transcriptome diversity and variation. CoSIA also provides additional conversion tools and utilities to provide a streamlined methodology for cross-species comparison.
Last updated
softwarebiologicalquestiongeneexpressionmultiplecomparisonthirdpartyclientdataimportgui
6.08 score 12 stars 10 scripts 320 downloadssplicelogic - splicelogic: differential transcripts to splice events
Translate differential transcript usage results into discrete splice events.
Last updated
alternativesplicingdifferentialsplicingtranscriptomicsrnaseqlongreadannotationfunctionalgenomicsdtusplicing
5.98 score 3 stars 5 scripts 193 downloadsChIPQC - Quality metrics for ChIPseq data
Quality metrics for ChIPseq data.
Last updated
sequencingchipseqqualitycontrolreportwriting
5.97 score 231 scripts 832 downloadsRFLOMICS - Interactive web application for Omics-data analysis
R-package with shiny interface, provides a framework for the analysis of transcriptomics, proteomics and/or metabolomics data. The interface offers a guided experience for the user, from the definition of the experimental design to the integration of several omics table together. A report can be generated with all settings and analysis results.
Last updated
shinyappsdifferentialexpressionmetabolomicsproteomicstranscriptomics
5.96 score 34 scripts 254 downloadsmetabom8 - A High-Performance R Package for Metabolomics Modeling and Analysis
Tools for 1D NMR metabolomics workflows, including import and preprocessing of Bruker experiments, multivariate modeling (PCA, PLS, OPLS) and model analytics and validation (y-permutations, cv-anova). Performance-critical routines are implemented in C++ and use the Armadillo and Eigen linear algebra libraries to improve runtime.
Last updated
metabolomicscheminformaticspreprocessingdataimportalignmentworkflowsteparmadilloeigenrcppcpp
5.93 score 3 stars 17 scripts 204 downloadsRhisat2 - R Wrapper for HISAT2 Aligner
An R interface to the HISAT2 spliced short-read aligner by Kim et al. (2015). The package contains wrapper functions to create a genome index and to perform the read alignment to the generated index.
Last updated
alignmentsequencingsplicedalignmentcpp
5.83 score 3 stars 1 dependents 8 scripts 560 downloads
MsBackendMetaboLights - Retrieve Mass Spectrometry Data from MetaboLights
MetaboLights is one of the main public repositories for storage of metabolomics experiments, which includes analysis results as well as raw data. The MsBackendMetaboLights package provides functionality to retrieve and represent mass spectrometry (MS) data from MetaboLights. Data files are downloaded and cached locally avoiding repetitive downloads. MS data from metabolomics experiments can thus be directly and seamlessly integrated into R-based analysis workflows with the Spectra and MsBackendMetaboLights package.
Last updated
infrastructuremassspectrometrymetabolomicsdataimportproteomicsmass-spectrometrymetabolomics-data
5.83 score 2 stars 32 scripts 338 downloadsregsplice - L1-regularization based methods for detection of differential splicing
Statistical methods for detection of differential splicing (differential exon usage) in RNA-seq and exon microarray data, using L1-regularization (lasso) to improve power.
Last updated
immunooncologyalternativesplicingdifferentialexpressiondifferentialsplicingsequencingrnaseqmicroarrayexonarrayexperimentaldesignsoftware
5.73 score 3 stars 20 scripts 400 downloadscarnation - Interactive Exploration & Management of RNA-Seq Analyses
Highly interactive & modular shiny app to explore three facets of RNA-Seq analysis: differential expression (DE), functional enrichment and pattern analysis. Several visualizations are implemented to provide a wide-ranging view of data sets. For DE analysis, we provide PCA plot, MA plot, Upset plot & heatmaps, in addition to a highly customizable gene plot. Seven different visualizations are available for functional enrichment analysis, and we also support gene pattern analysis. Genes of interest can be tracked across all modules using the gene scratchpad. In addition, carnation provides an integrated platform to manage multiple projects and user access that can be run on a central server to share with collaborators.
Last updated
guigeneexpressionsoftwareshinyappsgotranscriptiontranscriptomicsvisualizationdifferentialexpressionpathwaysgenesetenrichment
5.71 score 2 stars 17 scripts 217 downloads
blase - Bulk Linking Analysis for Single-cell Experiments
BLASE is a method for finding where bulk RNA-seq data lies on a single-cell pseudotime trajectory. It uses a fast and understandable approach based on Spearman correlation, with bootstrapping to provide confidence. BLASE can be used to "date" bulk RNA-seq data, annotate cell types in scRNA-seq, and help correct for developmental phenotype differences in bulk RNA-seq experiments.
Last updated
transcriptomicssinglecellsequencinggeneexpressiontranscriptionrnaseqtimecoursecellbiologysoftwarecellbasedassays
5.61 score 1 stars 7 scripts 251 downloadsbiotmle - Targeted Learning with Moderated Statistics for Biomarker Discovery
Tools for differential expression biomarker discovery based on microarray and next-generation sequencing data that leverage efficient semiparametric estimators of the average treatment effect for variable importance analysis. Estimation and inference of the (marginal) average treatment effects of potential biomarkers are computed by targeted minimum loss-based estimation, with joint, stable inference constructed across all biomarkers using a generalization of moderated statistics for use with the estimated efficient influence function. The procedure accommodates the use of ensemble machine learning for the estimation of nuisance functions.
Last updated
regressiongeneexpressiondifferentialexpressionsequencingmicroarrayrnaseqimmunooncologybioconductorbioconductor-packagebioconductor-packagesbioinformaticsbiomarker-discoverybiostatisticscausal-inferencecomputational-biologymachine-learningstatisticstargeted-learning
5.57 score 5 stars 5 scriptsSVP - Predicting cell states and their variability in single-cell or spatial omics data
SVP uses the distance between cells and cells, features and features, cells and features in the space of MCA to build nearest neighbor graph, then uses random walk with restart algorithm to calculate the activity score of gene sets (such as cell marker genes, kegg pathway, go ontology, gene modules, transcription factor or miRNA target sets, reactome pathway, ...), which is then further weighted using the hypergeometric test results from the original expression matrix. To detect the spatially or single cell variable gene sets or (other features) and the spatial colocalization between the features accurately, SVP provides some global and local spatial autocorrelation method to identify the spatial variable features. SVP is developed based on SingleCellExperiment class, which can be interoperable with the existing computing ecosystem.
Last updated
singlecellsoftwarespatialtranscriptomicsgenetargetgeneexpressiongenesetenrichmenttranscriptiongokeggopenblascppopenmp
5.56 score 12 stars 8 scripts 331 downloadsGBScleanR - Error correction tool for noisy genotyping by sequencing (GBS) data
GBScleanR is a package for quality check, filtering, and error correction of genotype data derived from next generation sequcener (NGS) based genotyping platforms. GBScleanR takes Variant Call Format (VCF) file as input. The main function of this package is `estGeno()` which estimates the true genotypes of samples from given read counts for genotype markers using a hidden Markov model with incorporating uneven observation ratio of allelic reads. This implementation gives robust genotype estimation even in noisy genotype data usually observed in Genotyping-By-Sequnencing (GBS) and similar methods, e.g. RADseq. The current implementation accepts genotype data of a diploid population at any generation of multi-parental cross, e.g. biparental F2 from inbred parents, biparental F2 from outbred parents, and 8-way recombinant inbred lines (8-way RILs) which can be refered to as MAGIC population.
Last updated
geneticvariabilitysnpgeneticshiddenmarkovmodelsequencingqualitycontrolcpp
5.56 score 4 stars 10 scripts 380 downloadsepigraHMM - Epigenomic R-based analysis with hidden Markov models
epigraHMM provides a set of tools for the analysis of epigenomic data based on hidden Markov Models. It contains two separate peak callers, one for consensus peaks from biological or technical replicates, and one for differential peaks from multi-replicate multi-condition experiments. In differential peak calling, epigraHMM provides window-specific posterior probabilities associated with every possible combinatorial pattern of read enrichment across conditions.
Last updated
chipseqatacseqdnaseseqhiddenmarkovmodelepigeneticscurlopensslopenblascppopenmp
5.44 score 91 scripts 360 downloadsMetaboSignal - MetaboSignal: a network-based approach to overlay and explore metabolic and signaling KEGG pathways
MetaboSignal is an R package that allows merging, analyzing and customizing metabolic and signaling KEGG pathways. It is a network-based approach designed to explore the topological relationship between genes (signaling- or enzymatic-genes) and metabolites, representing a powerful tool to investigate the genetic landscape and regulatory networks of metabolic phenotypes.
Last updated
graphandnetworkgenesignalinggenetargetnetworkpathwayskeggreactomesoftware
5.28 score 16 scripts 576 downloadsrigvf - R interface to the IGVF Catalog
The IGVF Catalog provides data on the impact of genomic variants on function. The `rigvf` package provides an interface to the IGVF Catalog, allowing easy integration with Bioconductor resources.
Last updated
thirdpartyclientannotationvariantannotationfunctionalgenomicsgeneregulationgenomicvariationgenetarget
5.13 score 2 scripts 261 downloadsfenr - Fast functional enrichment for interactive applications
Perform fast functional enrichment on feature lists (like genes or proteins) using the hypergeometric distribution. Tailored for speed, this package is ideal for interactive platforms such as Shiny. It supports the retrieval of functional data from sources like GO, KEGG, Reactome, Bioplanet and WikiPathways. By downloading and preparing data first, it allows for rapid successive tests on various feature selections without the need for repetitive, time-consuming preparatory steps typical of other packages.
Last updated
functionalpredictiondifferentialexpressiongenesetenrichmentgokeggreactomeproteomics
5.10 score 1 dependents 12 scripts 360 downloadsIsoBayes - IsoBayes: Single Isoform protein inference Method via Bayesian Analyses
IsoBayes is a Bayesian method to perform inference on single protein isoforms. Our approach infers the presence/absence of protein isoforms, and also estimates their abundance; additionally, it provides a measure of the uncertainty of these estimates, via: i) the posterior probability that a protein isoform is present in the sample; ii) a posterior credible interval of its abundance. IsoBayes inputs liquid cromatography mass spectrometry (MS) data, and can work with both PSM counts, and intensities. When available, trascript isoform abundances (i.e., TPMs) are also incorporated: TPMs are used to formulate an informative prior for the respective protein isoform relative abundance. We further identify isoforms where the relative abundance of proteins and transcripts significantly differ. We use a two-layer latent variable approach to model two sources of uncertainty typical of MS data: i) peptides may be erroneously detected (even when absent); ii) many peptides are compatible with multiple protein isoforms. In the first layer, we sample the presence/absence of each peptide based on its estimated probability of being mistakenly detected, also known as PEP (i.e., posterior error probability). In the second layer, for peptides that were estimated as being present, we allocate their abundance across the protein isoforms they map to. These two steps allow us to recover the presence and abundance of each protein isoform.
Last updated
statisticalmethodbayesianproteomicsmassspectrometryalternativesplicingsequencingrnaseqgeneexpressiongeneticsvisualizationsoftwarecpp
5.08 score 8 stars 15 scripts 292 downloadsdecompTumor2Sig - Decomposition of individual tumors into mutational signatures by signature refitting
Uses quadratic programming for signature refitting, i.e., to decompose the mutation catalog from an individual tumor sample into a set of given mutational signatures (either Alexandrov-model signatures or Shiraishi-model signatures), computing weights that reflect the contributions of the signatures to the mutation load of the tumor.
Last updated
softwaresnpsequencingdnaseqgenomicvariationsomaticmutationbiomedicalinformaticsgeneticsbiologicalquestionstatisticalmethod
5.03 score 1 stars 1 dependents 12 scripts 466 downloadsHubPub - Utilities to create and use Bioconductor Hubs
HubPub provides users with functionality to help with the Bioconductor Hub structures. The package provides the ability to create a skeleton of a Hub style package that the user can then populate with the necessary information. There are also functions to help add resources to the Hub package metadata files as well as publish data to the Bioconductor S3 bucket.
Last updated
dataimportinfrastructuresoftwarethirdpartyclientbioconductor-package
5.03 score 3 stars 10 scripts 944 downloadsbetterChromVAR - Improved ChromVAR (Chromatin Variation Across Regions)
A much faster analytical implementation of chromVAR, with additional features, used to infer TF activity from (bulk or single-cell) ATAC-seq data and motif annotations (or binding probabilities). The package also includes the CVnorm normalization method based on the chromVAR logic.
Last updated
softwareatacseqnormalizationepigeneticssequencing
4.98 score 2 stars 12 scripts 238 downloadsGOaGO - Gene Ontology enrichment analysis of gene pairs
GO-a-GO annotates Gene Ontology terms that are enriched in a given set of gene pairs. The enrichment is calculated from a permutation test for overrepresentation of gene pairs that are associated with a shared term. Such gene pairs are counted for the original set of gene pairs and compared against randomized sets in which the structure of the pairs is preserved, but the gene identities (including the associated terms) are permuted.
Last updated
gogenesetenrichment
4.95 score 2 stars 5 scripts 240 downloadsBiocBuildReporter - Functions to process a bioconductor build report database
This package reads remote parquet files that have processed Bioconductor build report logs. Users may query the tables directly for specific information or use pre-defined helper functions for common queries. The logs processed are from https://bioconductor.org/checkResults/. In the future we will extend this package out to include processing of r-universe logs.
Last updated
softwareinfrastructure
4.85 score 2 scripts 215 downloadsSEMPLR - SNP Effect Matrix Pipeline in R
SEMPLR computes transcription factor binding affinity scores for genomic positions and genetic variants. Scores are computed from SNP Effect Matrices (SEMs) produced by SEMpl. 223 pre-computed SEMs are included with the package or custom sets can be provided. Enrichment can be tested among sets of genomic positions to determine if transcription factor binding events occur more often than expected. Comparing binding affinity scores between alleles can reveal differences in transcription factor binding with genetic variation. This package also includes several visualization functions to view scores both on the motif and variant/position level.
Last updated
motifannotationtranscriptionsnpgenomicvariationcpp
4.85 score 1 stars 2 scripts 248 downloadsscMitoMut - Single-cell Mitochondrial Mutation Analysis Tool
This package is designed for calling lineage-informative mitochondrial mutations using single-cell sequencing data, such as scRNASeq and scATACSeq (preferably the latter due to RNA editing issues). It includes functions for mutation calling and visualization. Mutation calling is done using beta-binomial distribution.
Last updated
preprocessingsequencingsinglecellopenblascpp
4.73 score 3 stars 12 scripts 270 downloadsBiocAzul - Programmatic Access to the Azul API
Represents the OpenAPI v2 Azul API as an R object for performing requests. The infrastructure uses the AnVIL and rapiclient packages. Users can connect to either the AnVIL or Human Cell Atlas Data Explorers.
Last updated
softwareinfrastructuredataimportthirdpartyclientu24hg010263
4.70 score 2 scripts 182 downloads
GenomAutomorphism - Compute the automorphisms between DNA's Abelian group representations
This is a R package to compute the automorphisms between pairwise aligned DNA sequences represented as elements from a Genomic Abelian group. In a general scenario, from genomic regions till the whole genomes from a given population (from any species or close related species) can be algebraically represented as a direct sum of cyclic groups or more specifically Abelian p-groups. Basically, we propose the representation of multiple sequence alignments of length N bp as element of a finite Abelian group created by the direct sum of homocyclic Abelian group of prime-power order.
Last updated
mathematicalbiologycomparativegenomicsfunctionalgenomicsmultiplesequencealignmentwholegenomegenetic-codegenetic-code-algebragenomegenome-algebra
4.68 score 12 scripts 390 downloadsTDbasedUFEadv - Advanced package of tensor decomposition based unsupervised feature extraction
This is an advanced version of TDbasedUFE, which is a comprehensive package to perform Tensor decomposition based unsupervised feature extraction. In contrast to TDbasedUFE which can perform simple the feature selection and the multiomics analyses, this package can perform more complicated and advanced features, but they are not so popularly required. Only users who require more specific features can make use of its functionality.
Last updated
geneexpressionfeatureextractionmethylationarraysinglecellsoftwarebioconductor-packagebioinformaticstensor-decomposition
4.65 score 10 scripts 326 downloadsSeqtometry - Signature scoring for single cell analysis
This package provides functions used in Seqtometry (Kousnetsov et al. 2024), a method for analyzing single cell (scRNA-seq or scATAC-seq) data via signature (gene set) enrichment scores. The Seqtometry scores may be useful for annotating or characterizing cells, either in a flow cytometry like workflow (where scores are standalone features used for progressive partitoning as described in the Seqtometry publication) or in a cluster-based workflow (as features of clusters). The exported impute function (a port of Python's MAGIC-impute, van Dijk et al. 2018), may also be useful for single cell analysis on its own.
Last updated
singlecellgenesetenrichmentgeneexpressioncpp
4.60 score 1 stars 1 scriptsflowGate - Interactive Cytometry Gating in R
flowGate adds an interactive Shiny app to allow manual GUI-based gating of flow cytometry data in R. Using flowGate, you can draw 1D and 2D span/rectangle gates, quadrant gates, and polygon gates on flow cytometry data by interactively drawing the gates on a plot of your data, rather than by specifying gate coordinates. This package is especially geared toward wet-lab cytometerists looking to take advantage of R for cytometry analysis, without necessarily having a lot of R experience.
Last updated
softwareworkflowstepflowcytometrypreprocessingimmunooncologydataimport
4.59 score 26 scripts 378 downloadsfraq - A High-Throughput and Extensible Toolkit for Processing FASTQ Data
High-throughput extensible toolkit for processing FASTQ data. The goal of this package is to empower users to quickly build out small programmatic 'kernels' to define any FASTQ processing task they may need. Builds on Intel TBB’s flow graph to orchestrate concurrent I/O and data processing; throughput can be as fast as compression and disk speed allows. The package also ships with a suite of predefined kernels for common FASTQ tasks.
Last updated
softwareinfrastructuresequencingdnaseqqualitycontrolalignmentlibzstdcpp
4.56 score 1 stars 12 scripts 186 downloadstoppgene - Gene List Enrichment Analysis using the ToppGene Suite
The ToppGene Suite is a one-stop portal for gene list enrichment analysis and candidate gene prioritization based on functional annotations and protein interactions network. Although the ToppCluster web application provides convenient graphical access to the ToppGene Suite, the OpenAPI 3.0 compliant interface of ToppGene is better suited for automation and reproducibility. This package includes Bioconductor class interfaces and biological examples.
Last updated
clusteringgeneexpressiongenesetenrichmentgeneticsmotifdiscoverynetworknetworkenrichmentpathwayspharmacogeneticsproteomicssoftwarethirdpartyclient
4.54 score 1 stars 2 scripts 62 downloadssmoothclust - smoothclust
Method for identification of spatial domains and spatially-aware clustering in spatial transcriptomics data. The method generates spatial domains with smooth boundaries by smoothing gene expression profiles across neighboring spatial locations, followed by unsupervised clustering. Spatial domains consisting of consistent mixtures of cell types may then be further investigated by applying cell type compositional analyses or differential analyses.
Last updated
spatialsinglecelltranscriptomicsgeneexpressionclustering
4.48 score 1 stars 12 scripts 292 downloadsDNAcycP2 - DNA Cyclizability Prediction
This package performs prediction of intrinsic cyclizability of of every 50-bp subsequence in a DNA sequence. The input could be a file either in FASTA or text format. The output will be the C-score, the estimated intrinsic cyclizability score for each 50 bp sequences in each entry of the sequence set.
Last updated
neuralnetworkstructuralprediction
4.48 score 3 scripts 241 downloadssingIST - comparative single-cell transcriptomics between disease models and a human condition
Provides with toolkits to implement a full singIST analysis with pseudobulked Seurat objects of disease models and human data.
Last updated
singlecellclassificationtranscriptomics
4.40 score 9 scripts 128 downloadsdmGsea - Efficient Gene Set Enrichment Analysis for DNA Methylation Data
The R package dmGsea provides efficient gene set enrichment analysis specifically for DNA methylation data. It addresses key biases, including probe dependency and varying probe numbers per gene. The package supports Illumina 450K, EPIC, and mouse methylation arrays. Users can also apply it to other omics data by supplying custom probe-to-gene mapping annotations. dmGsea is flexible, fast, and well-suited for large-scale epigenomic studies.
Last updated
genesetenrichmentpathwaysdnamethylationproteomicssequencingcopynumbervariationgeneexpressiongenomicvariationcoverage
4.40 score 3 scripts 256 downloadsHiCBricks - Framework for Storing and Accessing Hi-C Data Through HDF Files
HiCBricks is a library designed for handling large high-resolution Hi-C datasets. Over the years, the Hi-C field has experienced a rapid increase in the size and complexity of datasets. HiCBricks is meant to overcome the challenges related to the analysis of such large datasets within the R environment. HiCBricks offers user-friendly and efficient solutions for handling large high-resolution Hi-C datasets. The package provides an R/Bioconductor framework with the bricks to build more complex data analysis pipelines and algorithms. HiCBricks already incorporates example algorithms for calling domain boundaries and functions for high quality data visualization.
Last updated
dataimportinfrastructuresoftwaretechnologysequencinghic
4.35 score 15 scripts 518 downloadsCyTOFpower - Power analysis for CyTOF experiments
This package is a tool to predict the power of CyTOF experiments in the context of differential state analyses. The package provides a shiny app with two options to predict the power of an experiment: i. generation of in-sicilico CyTOF data, using users input ii. browsing in a grid of parameters for which the power was already precomputed.
Last updated
flowcytometrysinglecellcellbiologystatisticalmethodsoftware
4.30 score 4 scripts 261 downloadsRegEnrich - Gene regulator enrichment analysis
This package is a pipeline to identify the key gene regulators in a biological process, for example in cell differentiation and in cell development after stimulation. There are four major steps in this pipeline: (1) differential expression analysis; (2) regulator-target network inference; (3) enrichment analysis; and (4) regulators scoring and ranking.
Last updated
geneexpressiontranscriptomicsrnaseqtwochanneltranscriptiongenetargetnetworkenrichmentdifferentialexpressionnetworknetworkinferencegenesetenrichmentfunctionalprediction
4.08 score 30 scripts 387 downloadsncRNAtools - An R toolkit for non-coding RNA
ncRNAtools provides a set of basic tools for handling and analyzing non-coding RNAs. These include tools to access the RNAcentral database and to predict and visualize the secondary structure of non-coding RNAs. The package also provides tools to read, write and interconvert the file formats most commonly used for representing such secondary structures.
Last updated
functionalgenomicsdataimportthirdpartyclientvisualizationstructuralprediction
4.07 score 1 stars 39 scripts 350 downloadsStatescopeR - StatescopeR framework for discovery of cell states from cell type-specific gene expression profiles inferred from bulk mRNA profiles
StatescopeR is an R wrapper around Statescope, a computational framework designed to discover cell states from cell type-specific gene expression profiles inferred from bulk RNA profiles.
Last updated
geneexpressionrnaseqsinglecellbayesiantranscriptomicssoftware
3.74 score 4 scripts 71 downloadsrfaRm - An R interface to the Rfam database
rfaRm provides a client interface to the Rfam database of RNA families. Data that can be retrieved include RNA families, secondary structure images, covariance models, sequences within each family, alignments leading to the identification of a family and secondary structures in the dot-bracket format.
Last updated
functionalgenomicsdataimportthirdpartyclientvisualizationmultiplesequencealignment
2.48 score 7 scripts 340 downloads
treeio - Base Classes and Functions for Phylogenetic Tree Input and Output
'treeio' is an R package to make it easier to import and store phylogenetic tree with associated data; and to link external data from different sources to phylogeny. It also supports exporting phylogenetic tree with heterogeneous associated data to a single tree file and can be served as a platform for merging tree with associated data and converting file formats.
Last updated
softwareannotationclusteringdataimportdatarepresentationalignmentmultiplesequencealignmentphylogeneticsexporterparserphylogenetic-trees
14.86 score 104 stars 138 dependents 2.7k scripts 47k downloadsSparseArray - High-performance sparse data representation and manipulation in R
The SparseArray package provides array-like containers for efficient in-memory representation of multidimensional sparse data in R (arrays and matrices). The package defines the SparseArray virtual class and two concrete subclasses: COO_SparseArray and SVT_SparseArray. Each subclass uses its own internal representation of the nonzero multidimensional data: the "COO layout" and the "SVT layout", respectively. SVT_SparseArray objects mimic as much as possible the behavior of ordinary matrix and array objects in base R. In particular, they suppport most of the "standard matrix and array API" defined in base R and in the matrixStats package from CRAN.
Last updated
infrastructuredatarepresentationbioconductor-packagecore-packageu24ca289073openmp
12.81 score 11 stars 1.4k dependents 107 scripts 97k downloadsOmnipathR - OmniPath web service client and more
A client for the OmniPath web service (https://www.omnipathdb.org) and many other resources. It also includes functions to transform and pretty print some of the downloaded data, functions to access a number of other resources such as BioPlex, ConsensusPathDB, EVEX, Gene Ontology, Guide to Pharmacology (IUPHAR/BPS), Harmonizome, HTRIdb, Human Phenotype Ontology, InWeb InBioMap, KEGG Pathway, Pathway Commons, Ramilowski et al. 2015, RegNetwork, ReMap, TF census, TRRUST and Vinayagam et al. 2011. Furthermore, OmnipathR features a close integration with the NicheNet method for ligand activity prediction from transcriptomics data, and its R implementation `nichenetr` (available only on github).
Last updated
graphandnetworknetworkpathwayssoftwarethirdpartyclientdataimportdatarepresentationgenesignalinggeneregulationsystemsbiologytranscriptomicssinglecellannotationkeggcomplexesenzyme-ptmnetworksnetworks-biologyomnipathproteinsquarto
10.42 score 171 stars 4 dependents 390 scripts
MetaboCoreUtils - Core Utils for Metabolomics Data
MetaboCoreUtils defines metabolomics-related core functionality provided as low-level functions to allow a data structure-independent usage across various R packages. This includes functions to calculate between ion (adduct) and compound mass-to-charge ratios and masses or functions to work with chemical formulas. The package provides also a set of adduct definitions and information on some commercially available internal standard mixes commonly used in MS experiments.
Last updated
infrastructuremetabolomicsmassspectrometrymass-spectrometry
10.29 score 9 stars 71 dependents 102 scripts 3.7k downloadsSeqinfo - A simple S4 class for storing basic information about a collection of genomic sequences
The Seqinfo class stores the names, lengths, circularity flags, and genomes for a particular collection of sequences. These sequences are typically the chromosomes and/or scaffolds of a specific genome assembly of a given organism. Seqinfo objects are rarely used as standalone objects. Instead, they are used as part of higher-level objects to represent their seqinfo() component. Examples of such higher-level objects are GRanges, RangedSummarizedExperiment, VCF, GAlignments, etc... defined in other Bioconductor infrastructure packages.
Last updated
infrastructuredatarepresentationgenomeassemblyannotationgenomeannotationbioconductor-packagecore-package
10.05 score 1 stars 1.9k dependents 22 scripts 61k downloadsSpiecEasi - Sparse Inverse Covariance for Ecological Statistical Inference
Estimate networks from the precision matrix of compositional microbial abundance data.
Last updated
softwaremicrobiomemetagenomicsgraphandnetworknetworkinferenceopenblascpp
9.59 score 233 stars 788 scripts 575 downloadsorthogene - Gene mapping made easy
`orthogene` is an R package for easy mapping of orthologous genes across hundreds of species. It pulls up-to-date gene ortholog mappings across **700+ organisms**. It also provides various utility functions to aggregate/expand common objects (e.g. data.frames, gene expression matrices, lists) using **1:1**, **many:1**, **1:many** or **many:many** gene mappings, both within- and between-species.
Last updated
geneticscomparativegenomicspreprocessingphylogeneticstranscriptomicsgeneexpressionanimal-modelsbioconductorbioconductor-packagebioinformaticsbiomedicinecomparative-genomicsevolutionary-biologygenesgenomicsontologiestranslational-research
9.23 score 58 stars 3 dependents 89 scripts 1.2k downloadsnetZooR - A menagerie of methods for the inference and analysis of gene regulatory networks
netZooR unifies the implementations of several Network Zoo methods (netzoo, netzoo.github.io) into a single package by creating interfaces between network inference and network analysis methods. Currently, the package has 3 methods for network inference including PANDA and its optimized implementation OTTER (network reconstruction using mutliple lines of biological evidence), LIONESS (single-sample network inference), and EGRET (genotype-specific networks). Network analysis methods include CONDOR (community detection), ALPACA (differential community detection), CRANE (significance estimation of differential modules), MONSTER (estimation of network transition states). In addition, YARN allows to process gene expresssion data for tissue-specific analyses and SAMBAR infers missing mutation data based on pathway information.
Last updated
geneexpressiongeneregulationgraphandnetworkmicroarraynetworknetworkinferencetranscriptiongene-regulatory-networktranscription-factors
8.93 score 120 stars 143 scripts 180 downloadshypeR - An R Package For Geneset Enrichment Workflows
An R Package for Geneset Enrichment Workflows.
Last updated
genesetenrichmentannotationpathwaysbioinformaticscomputational-biologygeneset-enrichment-analysis
8.92 score 79 stars 348 scriptsClusterGVis - One-Step to Cluster and Visualize Gene Expression Data
Provides a streamlined workflow for clustering and visualizing gene expression patterns, particularly from time-series RNA-Seq and single-cell experiments. The package is designed to integrate seamlessly within the Bioconductor ecosystem by operating directly on standard data classes such as `SummarizedExperiment` and `SingleCellExperiment`. It implements common clustering algorithms (e.g., k-means, fuzzy c-means) and generates a suite of publication-ready visualizations to explore co-expressed gene modules. Functions are also included to facilitate the visualization of clustering results derived from other popular tools.
Last updated
rnaseqtranscriptomicsvisualizationsinglecellgeneexpressionclusteringcomplexheatmapgene-clusteringgene-expressionmfuzz
8.62 score 389 stars 80 scripts 175 downloadsPTMods - Managing Post-Translational Modifications in R
An interface to the community supported database for amino acid/protein modifications using mass spectrometry.
Last updated
proteomicsmassspectrometryamino-acid-modificationsmass-spectrometryprotein
8.51 score 11 stars 41 dependents 5 scripts 1.6k downloadsnotame - Workflow for non-targeted LC-MS metabolic profiling
Provides functionality for untargeted LC-MS metabolomics research as specified in the associated protocol article in the 'Metabolomics Data Processing and Data Analysis—Current Best Practices' special issue of the Metabolites journal (2020). This includes tabular data preprocessing and quality control, uni- and multivariate analysis as well as quality control visualizations, feature-wise visualizations and results visualizations. Raw data preprocessing and functionality related to biological context, such as pathway analysis, is not included.
Last updated
biomedicalinformaticsmetabolomicsdataimportmassspectrometrybatcheffectmultiplecomparisonnormalizationqualitycontrolvisualizationpreprocessing
8.45 score 6 stars 2 dependents 113 scripts 280 downloadsigvR - igvR: integrative genomics viewer
Access to igv.js, the Integrative Genomics Viewer running in a web browser.
Last updated
visualizationthirdpartyclientgenomebrowsers
8.43 score 46 stars 145 scriptsSpaceTrooper - SpaceTrooper performs Quality Control analysis of Image-Based spatial
SpaceTrooper performs Quality Control analysis using data driven GLM models of Image-Based spatial data, providing exploration plots, QC metrics computation, outlier detection. It implements a GLM strategy for the detection of low quality cells in imaging-based spatial data (Transcriptomics and Proteomics). It additionally implements several plots for the visualization of imaging based polygons through the ggplot2 package.
Last updated
softwaretranscriptomicsgeneexpressionqualitycontrolspatialsinglecelldataimportimmunooncology
7.80 score 11 stars 38 scripts 313 downloadsontoProc - processing of ontologies of anatomy, cell lines, and so on
Support harvesting of diverse bioinformatic ontologies, making particular use of the ontologyIndex package on CRAN. We provide snapshots of key ontologies for terms about cells, cell lines, chemical compounds, and anatomy, to help analyze genome-scale experiments, particularly cell x compound screens. Another purpose is to strengthen development of compelling use cases for richer interfaces to emerging ontologies.
Last updated
infrastructuregobioinformaticsgenomicsontology
7.74 score 5 stars 2 dependents 102 scripts 612 downloadsBgeeDB - Annotation and gene expression data retrieval from Bgee database. TopAnat, an anatomical entities Enrichment Analysis tool for UBERON ontology
A package for the annotation and gene expression data download from Bgee database, and TopAnat analysis: GO-like enrichment of anatomical terms, mapped to genes by expression patterns.
Last updated
softwaredataimportsequencinggeneexpressionmicroarraygogenesetenrichmentbioinformaticsenrichment-analysisrna-seqscrna-seqsingle-cell
7.63 score 16 stars 1 dependents 34 scripts 533 downloadsigvShiny - igvShiny: a wrapper of Integrative Genomics Viewer (IGV - an interactive tool for visualization and exploration integrated genomic data)
This package is a wrapper of Integrative Genomics Viewer (IGV). It comprises an htmlwidget version of IGV. It can be used as a module in Shiny apps.
Last updated
softwareshinyappssequencingcoverage
7.43 score 38 stars 1 dependents 148 scripts 370 downloadsdrugfindR - Investigate iLINCS for candidate repurposable drugs
This package provides a convenient way to access the LINCS Signatures available in the iLINCS database. These signatures include Consensus Gene Knockdown Signatures, Gene Overexpression signatures and Chemical Perturbagen Signatures. It also provides a way to enter your own transcriptomic signatures and identify concordant and discordant signatures in the LINCS database.
Last updated
lincsilincsdrug repurposingdrug discoverytranscriptomicsgene expressiongene knockdowngene overexpressionchemical perturbagendrugfindrbioinformaticsbioinformatics-pipeline
7.26 score 12 stars 159 scriptsMetaboAnnotatoR - Automated Annotation of All-Ion Fragmentation LC-MS Metabolomic Features
Performs feature annotations on LC-MS All-ion fragmentation datasets using fragment ion libraries.
Last updated
massspectrometrymetabolomics
7.14 score 14 stars 39 scripts 136 downloadsimmReferent - An Interface for Immune Receptor and HLA Gene Reference Data
Provides a consistent interface for downloading, storing, and accessing immune receptor (TCR/BCR) and HLA sequences from IMGT, IPD-IMGT/HLA, and OGRDB (AIRR-C). Supports export to popular analysis tools including MiXCR, TRUST4, Cell Ranger, and IgBLAST. This package serves as a core dependency for immunogenomics packages, ensuring reliable and high-quality sequence access with local caching for reproducibility.
Last updated
softwareannotationsequencing
7.08 score 9 stars 5 dependents 6 scripts 168 downloadsAerith - visualization and annotation of isotopic enrichment patterns of peptides and metabolites with stable isotope labeling from proteomics and metabolomics
Visualisation of peptide isotopic peaks and SIP peptide spectra match (PSM). Filtration of high quality PSM. Accurate isotopic abundance calculation of peptide and metabolites. Visualisation of SIP proteomics results.
Last updated
proteomicsmetabolomicsmassspectrometrysoftwarevisualizationqualitycontrolannotationlc-msmass-spectrometrystable-isotope-mass-spectrometrystable-isotope-tracingstable-isotopescppopenmp
7.02 score 5 stars 54 scripts 176 downloadsImageArray - A framework for on-disk and in-memory image arrays
ImageArray provides a framework for on-disk and in-memory image arrays, specifically for pyramidal images stored in HDF5, Zarr and life sciences image file formats (OME Bio-Formats).
Last updated
softwarevisualization
7.02 score 6 stars 2 dependents 18 scripts 243 downloadsscTypeEval - Evaluation of cell type classifications in single-cell transcriptomics
scTypeEval provides tools to evaluate and validate cell type classifications in single-cell transcriptomics when ground truth labels are limited or unavailable. Results are organized in an S4 object that integrates processed data, dimensional reductions, dissimilarity assays, and consistency metrics computed across samples. The workflow includes preprocessing and feature selection, principal component analysis, computation of dissimilarity matrices, internal validation metrics (for example, silhouette-based summaries), and visualization utilities to inspect heatmaps and PCA plots. Functions support common single-cell containers and enable comparison of clustering and labeling strategies across datasets.
Last updated
singlecelltranscriptomicsgeneexpressioncellbasedassaysdimensionreductionpreprocessingprincipalcomponent
6.95 score 8 stars 23 scripts 199 downloadsdominoSignal - Cell Communication Analysis for Single Cell RNA Sequencing
dominoSignal is a package developed to analyze cell signaling through ligand - receptor - transcription factor networks in scRNAseq data. It takes as input information transcriptomic data, requiring counts, z-scored counts, and cluster labels, as well as information on transcription factor activation (such as from SCENIC) and a database of ligand and receptor pairings (such as from CellPhoneDB). This package creates an object storing ligand - receptor - transcription factor linkages by cluster and provides several methods for exploring, summarizing, and visualizing the analysis.
Last updated
systemsbiologysinglecelltranscriptomicsnetwork
6.92 score 7 stars 28 scripts 264 downloadsChromatograms - Infrastructure for Chromatographic Mass Spectrometry Data
The Chromatograms packages defines an efficient infrastructure for storing and handling of chromatographic mass spectrometry data. It provides different implementations of *backends* to store and represent the data. Such backends can be optimized for small memory footprint or fast data access/processing. A lazy evaluation queue and chunk-wise processing capabilities ensure efficient analysis of also very large data sets.
Last updated
infrastructuremetabolomicsmassspectrometryproteomics
6.91 score 3 stars 1 dependents 25 scripts 270 downloadsIbex - Methods for BCR single-cell embedding
Implementation of the Ibex algorithm for single-cell embedding based on BCR sequences. The package includes a standalone function to encode BCR sequence information by amino acid properties or sequence order using tensorflow-based autoencoder. In addition, the package interacts with SingleCellExperiment or Seurat data objects.
Last updated
softwareimmunooncologysinglecellclassificationannotationsequencing
6.82 score 27 stars 18 scripts 191 downloadsmiaTime - Microbiome Time Series Analysis
The `miaTime` package provides tools for microbiome time series analysis based on (Tree)SummarizedExperiment infrastructure.
Last updated
microbiomesoftwaresequencing
6.79 score 8 stars 37 scripts 232 downloadsscLANE - Model Gene Expression Dynamics with Spline-Based NB GLMs, GEEs, & GLMMs
Our scLANE model uses truncated power basis spline models to build flexible, interpretable models of single cell gene expression over pseudotime or latent time. The modeling architectures currently supported are Negative-binomial GLMs, GEEs, & GLMMs. Downstream analysis functionalities include model comparison, dynamic gene clustering, smoothed counts generation, gene set enrichment testing, & visualization.
Last updated
rnaseqsoftwareclusteringtimecoursesequencingregressionsinglecellvisualizationgeneexpressiontranscriptomicsgenesetenrichmentdifferentialexpressiondifferential-expressionestimating-equationsgenomicsmixed-modelspseudotimerna-velocityscrna-seqsingle-celltrajectory-inferencecpp
6.79 score 17 stars 36 scripts 222 downloadsMSstatsBioNet - Network Analysis for MS-based Proteomics Experiments
A set of tools for network analysis using mass spectrometry-based proteomics data and network databases. The package takes as input the output of MSstats differential abundance analysis and provides functions to perform enrichment analysis and visualization in the context of prior knowledge from past literature. Notably, this package integrates with INDRA, which is a database of biological networks extracted from the literature using text mining techniques.
Last updated
immunooncologymassspectrometryproteomicssoftwarequalitycontrolnetworkenrichmentnetwork
6.75 score 2 stars 1 dependents 12 scripts 279 downloads
MutSeqR - Analysis of Error-Corrected Sequencing Data for Mutation Detection
Standard methods for analysis of mutation data following error- corrected sequencing (ECS) for the purpose of mutagencity assessment. Functions include importing the mutation lists provided by a variant caller, and a set of analytical tools for statistical testing and visualization of mutation data; comparison to COSMIC and/or germline signatures; etc.
Last updated
sequencingsomaticmutationvisualizationgenomicvariationdrivermutationstatisticalmethodgenetarget
6.72 score 11 stars 8 scripts 202 downloadsimageFeatureTCGA - Import features from hovernet, provgigapath into a MultiAssayExperiment
The package imports data from HoverNet, and ProvGigaPath pipelines. Pipeline output data are hosted in a self-owned online repository. Package functionality conveniently incorporates pipeline data into existing MultiAssayExperiment instances from curatedTCGAData.
Last updated
softwareinfrastructuredataimportdatarepresentation
6.58 score 2 stars 1 dependents 19 scripts 193 downloadsSuperCellCyto - SuperCell For Cytometry Data
SuperCellCyto provides the ability to summarise cytometry data into supercells by merging together cells that are similar in their marker expressions using the SuperCell package.
Last updated
cellbiologyflowcytometrysoftwaresinglecellbioinformaticscomputational-biologycytometry
6.55 score 13 stars 22 scripts 234 downloadsMetaProViz - METabolomics pre-PRocessing, functiOnal analysis and VIZualisation
MetaProViz can analyse standard metabolomics and exometabolomics data (CoRe). It performs pre-processing including feature filtering, missing value imputation, normalisation and outlier detection. It performs functional analysis including differential metabolite analysis (DMA), clustering based on regulatory rules (MCA) and contains different visualisation methods to extract biological interpretable graphs and saves them in a publication ready format.
Last updated
clusteringmetabolomicspathwaysqualitycontrolsoftwaresystemsbiologyvisualizationquarto
6.46 score 22 stars 31 scripts 144 downloadsBrowserViz - BrowserViz: interactive R/browser graphics using websockets and JSON
Interactvive graphics in a web browser from R, using websockets and JSON.
Last updated
visualizationthirdpartyclient
6.38 score 2 stars 2 dependents 25 scripts 488 downloads