Calibrate SURF predictions with SCAN.

SCAN fits a Negative Binomial spline model between a SURF prediction and held-out coverage. Train on one genomic region, evaluate on another, then carry both latent-mean and posterior-predictive intervals into interpretation.

Calibration workflow

Use the SCAN Agent Skill as the guide.

The Skill should walk through input checks, the R fit, convergence diagnostics, held-out coverage, and interval export. The steps below mirror the repository’s CLI and notebooks so every result remains reproducible without a matched external truth set.

1. Produce held-out SURF tracks

In jzhoulab/SURF, install the locked environment and run prediction with --half. The modality is read from the conf.txt saved beside the checkpoint.

git clone https://github.com/jzhoulab/SURF.git
cd SURF
uv sync
uv run surf predict -i path/to/training_output_dir -m best.model.checkpoint -c chrom_lengths.tsv -o predictions.bw --half

The diagnostic run supplies the three tracks SCAN needs:

2. Install SCAN

SCAN requires R 4.3 or newer. Its installer adds Stan, spline, plotting, and BigWig dependencies.

git clone https://github.com/jzhoulab/SCAN.git
cd SCAN
Rscript install_deps.R

3. Fit on one held-out chromosome

Fit the spline calibration on chr8. The default model uses 3,000 log-spaced bins, a monotone I-spline for the mean, a local B-spline for dispersion, four chains, and 1,000 iterations.

Rscript scan_calibration.R --pred_bw /path/to/half.pred.bw --tar_bw /path/to/half.tar.bw --chrom chr8 --out_dir out/sub01 --prefix sub01

A complete fit writes:

4. Check the fit before trusting intervals

  1. Review Rhat and effective sample size for failed or weakly mixed chains.
  2. Check the predicted-versus-observed diagnostic across the supported signal range.
  3. Inspect sparsely represented high-signal bins rather than relying on an overall average.
  4. Keep evaluation below the calibration data’s upper prediction limit.

5. Evaluate on a different chromosome

Point the pretrained analysis notebook at the data and fitted model, then evaluate on chr10. The full-analysis notebook can instead reproduce fitting and evaluation end to end.

export SCAN_DATA_ROOT=/path/to/coverage_models
export SCAN_GT_BW=/path/to/full_gt.bw
export SCAN_MODEL_DIR="$PWD/out/sub01"

Do not score against a full-coverage track as if it were independent: it contains the reads supplied to SURF. The repository notebooks evaluate against max(full coverage − input, 0) and scale the fitted mean to the held-out depth before checking empirical coverage.

6. Read the two uncertainty bands

The second interval is usually wider because it includes read-sampling variation. Choose the interval from the question you are asking, not from which band looks more decisive.

7. Save the calibration evidence

Keep the fitted model, knot definitions, binned summary, sampler diagnostics, empirical-coverage plots, example region, SURF checkpoint, and the exact input and target tracks together. This is the evidence needed to reproduce the intervals or explain why a region was flagged.