Architecture

Overview of the codebase design, key components, and the patterns used throughout GeneCircuitry.

Table of contents

  1. TOC

Package layout

genecircuitry/
├── __init__.py                  # Package exports; optional deps wrapped in try/except
├── config.py                    # Central config — single source of truth for all parameters
├── preprocessing.py             # QC, normalization, clustering (Scanpy wrappers)
├── celloracle_processing.py     # GRN inference (optional dep: celloracle)
├── hotspot_processing.py        # Gene module detection (optional dep: hotspotsc)
├── grn_deep_analysis.py         # NetworkX GRN visualisation
├── atac_peaks_processing.py     # ATAC peak → TF motif matrix (optional dep: genomepy, gimmemotifs)
├── enrichment_analysis.py       # Local ORA enrichment engine (bundled gene sets, disk caching)
├── logging_utils.py             # Internal logging helpers
├── pipeline/
│   ├── __init__.py              # Exports PipelineController, log_step, log_error
│   └── controller.py           # PipelineController class — central orchestrator
├── plotting/
│   ├── __init__.py
│   ├── qc_plots.py             # Canonical QC violin/scatter plots
│   ├── grn_plots.py            # GRN network graphs, centrality heatmaps
│   ├── hotspot_plots.py        # Hotspot module heatmaps, score violins
│   └── utils.py                # save_plot(), plot_exists()
└── reporting/
    ├── __init__.py              # generate_report() entry point
    ├── generator.py            # ReportGenerator class (HTML + PDF)
    └── sections.py             # Section builders (data summary, QC, GRN, Hotspot…)

Key components

File / Directory Purpose
genecircuitry/config.py Single source of truth — all parameters live here
genecircuitry/pipeline/controller.py PipelineController class — central orchestrator (~1900 lines)
run_complete_analysis.py CLI entry point — thin wrapper around genecircuitry.pipeline.main
genecircuitry/preprocessing.py Scanpy wrappers: QC, normalization, dimensionality reduction, clustering
genecircuitry/celloracle_processing.py CellOracle GRN inference (optional dep)
genecircuitry/hotspot_processing.py Hotspot gene module analysis (optional dep)
genecircuitry/grn_deep_analysis.py NetworkX network visualization
genecircuitry/plotting/ Canonical plot generation (qc_plots.py, grn_plots.py, hotspot_plots.py)
genecircuitry/reporting/ HTML / PDF report generation
genecircuitry/atac_peaks_processing.py ATAC-seq peak motif scanning (optional)
genecircuitry/enrichment_analysis.py Fast local gene set over-representation analysis (ORA) with bundled gene sets

Data flow

AnnData (.h5ad)
    │
    ▼
[load_data()]
    │
    ▼
[preprocessing_pipeline()]         genecircuitry/preprocessing.py
  ├─ perform_qc()
  └─ perform_normalization()
    │
    ▼ (checkpoint: preprocessed_adata.h5ad)
    │
    ├─► [stratification_pipeline()]  — optional, splits by cell type
    │
    ▼
[dimensionality_reduction_clustering()]
  └─ perform_dimensionality_reduction_clustering()
    │
    ▼ (checkpoint: clustered_adata.h5ad)
    │
    ├──────────────────────────┐
    ▼                          ▼
[celloracle_pipeline()]    [hotspot_pipeline()]
  ├─ perform_grn_pre_processing   create_hotspot_object()
  ├─ create_oracle_object()       run_hotspot_analysis()
  ├─ run_PCA()
  ├─ run_KNN()
  └─ run_links()
    │                          │
    ▼ (checkpoint: oracle + links)
    │
    ▼
[grn_deep_analysis_pipeline()]   genecircuitry/grn_deep_analysis.py
  └─ process, plot, compare networks
    │
    ▼
[generate_report()]              genecircuitry/reporting/
  └─ HTML + PDF with embedded figures

PipelineController

PipelineController (in genecircuitry/pipeline/controller.py) is the central class. It:

  • Holds pipeline state (self.adata, self.adata_preprocessed, self.stratification_results)
  • Exposes one run_step_*() method per pipeline step
  • Handles checkpointing, logging, and error recovery per step
  • Supports stratified (per-cluster) runs via run_stratified_pipeline_sequential()

Adding a new step:

# In genecircuitry/pipeline/controller.py — add to PipelineController class
def run_step_my_analysis(self, adata, log_dir=None):
    """Execute my custom analysis step."""
    log_step("Controller.MyAnalysis", "STARTED")
    try:
        result = my_analysis_function(adata)
        log_step("Controller.MyAnalysis", "COMPLETED")
        return result
    except Exception as e:
        log_error("Controller.MyAnalysis", e)
        raise

Then register the step name in run_complete_pipeline().


Config as single source of truth

All numeric thresholds, figure sizes, and directory paths are constants in genecircuitry/config.py.
Functions use them as defaults, allowing per-call overrides without mutating global state:

def perform_qc(adata, min_genes=None, ...):
    if min_genes is None:
        min_genes = config.QC_MIN_GENES   # falls back to config

To change a constant globally at runtime: config.update_config(QC_MIN_GENES=500).


Optional dependencies

CellOracle and Hotspot are wrapped in try/except ImportError in genecircuitry/__init__.py:

try:
    from . import celloracle_processing
except ImportError:
    celloracle_processing = None

Downstream code checks if genecircuitry.celloracle_processing is not None before use. New optional-dep modules must follow the same pattern.


Plotting subpackage

The canonical location for all plots is genecircuitry/plotting/:

Module Contents
qc_plots.py plot_qc_violin_pre_filter(), plot_qc_scatter_pre_filter(), post-filter variants
grn_plots.py generate_all_grn_plots(), plot_enriched_tf_network(), plot_tf_shared_target_network()
hotspot_plots.py Hotspot module heatmaps and per-module gene plots
utils.py plot_exists(), save_plot() — shared helpers

Legacy inline plotting in grn_deep_analysis.py and hotspot_processing.py is kept for backward compatibility but delegates to the plotting/ subpackage.


Logging

Two log files are created in <output>/logs/:

File Contents
pipeline.log Every step with start/complete timestamp and key metrics
error.log Errors with full Python tracebacks
from genecircuitry.pipeline import log_step, log_error

log_step("MyStep", "STARTED", {"n_cells": adata.n_obs})
log_step("MyStep", "COMPLETED", {"result_count": len(results)})
log_error("MyStep", exception)