Train the toy model, then SURF your data.

Start with the DNase-seq reference configuration in jzhoulab/SURF. Once you can explain its inputs, split, checkpoint, and prediction, use the repository’s development and prediction Skills to adapt the same safeguards to a new base-pair assay.

Two repository Skills

Let the repo supply the current workflow.

skills/develop-surf-model/ covers training and new modalities. skills/predict-surf-model/covers inference and BigWig output. Read the development Skill first so an agent works with SURF’s actual strategy interface and configuration schema.

git clone https://github.com/jzhoulab/SURF.git
cd SURF
cat skills/develop-surf-model/SKILL.md
Prompt: Read skills/develop-surf-model/SKILL.md and follow the DNase reference workflow end to end. Set up the documented environment, train the toy model without changing its configuration, and summarize the data split, outputs, and validation checks. Stop and explain any mismatch with the documented result before adapting it.

1. Check the toy inputs

The reference run is defined by confs/dnase.txt. Before training, confirm that its DNase BigWig and hg38 FASTA paths exist. The current configuration expects files under test_data/and a CUDA device; the repository’s data-download instructions must be completed when those assets are not already present.

2. Install and train the reference model

SURF uses the committed uv.lock. Install the exact environment, then launch the generic trainer with the DNase configuration.

uv sync
uv run surf train -c confs/dnase.txt

A completed run writes best.model.checkpoint, bestvalloss.model.checkpoint,last.model.checkpoint, a saved conf.txt, the log, and training-curve PNGs to the configured output directory.

Pause and verify

3. Understand the correctness gate

Direct Predictive Inference learns to predict one partition of an experiment from DNA sequence and the other partition. The in-loader 50/50 split is valid only when the track contains raw integer point counts and one unit represents one read at one position—such as a DNase cut site, CAGE 5′ end, or per-C methylation call.

Do not apply the in-loader split to normalized, smoothed, binned, or extended-fragment coverage. The same read can then leak into both halves. Reprocess to point counts, or split reads or UMIs before building separate input and target BigWigs.

4. Generate an inspectable prediction

Prediction reads the modality and data settings from the saved conf.txt. Supply a tab-separated chromosome-length file and keep the prediction beside the trained model.

uv run surf predict -i dnase.test.com -m best.model.checkpoint -c chrom_lengths.tsv -o predictions.bw

Add --half for a diagnostic run. SURF then predicts from a half-coverage input and also writes the complementary input and target tracks. View all three together before moving to your own assay.

5. Decide whether the new assay needs code

Config only

Reuse the DNase strategy when the assay is a single-channel, non-negative count track with the same input layout and exponential output transform. Point a new config at the data and choose a suitable count loss.

New strategy

Add a modality strategy when output dimensions, activation, masking, preprocessing, or metrics differ. The development Skill starts from src/surf/strategies/dnase.py, registers the new strategy, and keeps the generic trainer unchanged.

The manuscript gives three useful precedents:

6. Develop the smallest adaptation with the Skill

Adaptation prompt
Prompt: Read skills/develop-surf-model/SKILL.md. I want to train SURF on [assay and data format] at [target resolution]. Before editing code, map my data to the model inputs and training target, recommend the observation model, loss, window size, and chromosome split, and identify the smallest changes from the toy workflow. Then implement a small smoke test and show me how to validate it.

Begin with one small, inspectable region and a short smoke test. Prefer a new config over new code, and a new strategy over edits to the generic trainer. Change one modeling assumption at a time.

7. Validate before scaling up

  1. Check shapes, strand layout, coordinates, count totals, and reference-genome alignment.
  2. Confirm that the read partitions and held-out chromosomes are disjoint.
  3. Compare SURF with the observed input against the same held-out target.
  4. Run sequence-only and data-only ablations to show where the gain comes from.
  5. Inspect a signal-rich locus and a quiet locus at broad and fine scales.
  6. Test the coverage levels, cell types, or resolutions required by the intended claim.
Preflight review
Prompt: Audit this SURF run against skills/develop-surf-model/SKILL.md. Check for read or chromosome leakage, inconsistent preprocessing, an inappropriate loss, failed controls, and missing provenance. Compare the prediction with the observed input on held-out data and tell me what must be fixed before a full training run.

8. Train, predict, and preserve provenance

After the controls pass, freeze the split and config, train the full model with uv run surf train -c path/to/conf.json, and apply the selected checkpoint with the generic prediction command. If the new modality changes its output transform or layout, invoke skills/predict-surf-model/ before inference.

Keep the data version, preprocessing, chromosome split, config, seed, checkpoint, training curves, held-out metrics, and prediction track together. That bundle is the smallest reproducible model result.

9. Calibrate the result with SCAN

Use a half-coverage SURF run to produce the prediction and held-out target tracks required by SCAN. Install the SCAN R dependencies, fit its Negative Binomial spline model on chr8, and evaluate empirical coverage on a different chromosome—chr10 in the repository notebooks.

Rscript scan_calibration.R --pred_bw /path/to/half.pred.bw --tar_bw /path/to/half.tar.bw --chrom chr8 --out_dir out/toy --prefix toy

Preserve both the latent-mean and posterior-predictive intervals. The SCAN Agent Skill should guide this same fit, diagnostic, held-out evaluation, and export sequence when it is present in the repository.

Before calling the model ready