Train the toy model, then SURF your data.
Start with the DNase-seq reference configuration in jzhoulab/SURF. Once you can explain its inputs, split, checkpoint, and prediction, use the repository’s development and prediction Skills to adapt the same safeguards to a new base-pair assay.
Let the repo supply the current workflow.
skills/develop-surf-model/ covers training and new modalities. skills/predict-surf-model/covers inference and BigWig output. Read the development Skill first so an agent works with SURF’s actual strategy interface and configuration schema.
git clone https://github.com/jzhoulab/SURF.git
cd SURF
cat skills/develop-surf-model/SKILL.md1. Check the toy inputs
The reference run is defined by confs/dnase.txt. Before training, confirm that its DNase BigWig and hg38 FASTA paths exist. The current configuration expects files under test_data/and a CUDA device; the repository’s data-download instructions must be completed when those assets are not already present.
- Model: baseline DNase strategy with a 10 kb sliding window.
- Objective: pseudo-Poisson KL on a 50/50 binomial read split.
- Toy duration: 1,000 iterations with a batch size of 16.
- Genome split: chr8 validation; chr9 and chr10 test.
- Output directory:
dnase.test.com.
2. Install and train the reference model
SURF uses the committed uv.lock. Install the exact environment, then launch the generic trainer with the DNase configuration.
uv sync
uv run surf train -c confs/dnase.txtA completed run writes best.model.checkpoint, bestvalloss.model.checkpoint,last.model.checkpoint, a saved conf.txt, the log, and training-curve PNGs to the configured output directory.
Pause and verify
- The training and validation metrics remain finite and follow the pattern documented by the repo.
- The saved
conf.txtmatches the run you intended to launch. - You can identify which chromosomes and read partition were withheld from optimization.
- The selected checkpoint and every downstream prediction can be traced to this output directory.
3. Understand the correctness gate
Direct Predictive Inference learns to predict one partition of an experiment from DNA sequence and the other partition. The in-loader 50/50 split is valid only when the track contains raw integer point counts and one unit represents one read at one position—such as a DNase cut site, CAGE 5′ end, or per-C methylation call.
Do not apply the in-loader split to normalized, smoothed, binned, or extended-fragment coverage. The same read can then leak into both halves. Reprocess to point counts, or split reads or UMIs before building separate input and target BigWigs.
4. Generate an inspectable prediction
Prediction reads the modality and data settings from the saved conf.txt. Supply a tab-separated chromosome-length file and keep the prediction beside the trained model.
uv run surf predict -i dnase.test.com -m best.model.checkpoint -c chrom_lengths.tsv -o predictions.bwAdd --half for a diagnostic run. SURF then predicts from a half-coverage input and also writes the complementary input and target tracks. View all three together before moving to your own assay.
5. Decide whether the new assay needs code
Config only
Reuse the DNase strategy when the assay is a single-channel, non-negative count track with the same input layout and exponential output transform. Point a new config at the data and choose a suitable count loss.
New strategy
Add a modality strategy when output dimensions, activation, masking, preprocessing, or metrics differ. The development Skill starts from src/surf/strategies/dnase.py, registers the new strategy, and keeps the generic trainer unchanged.
The manuscript gives three useful precedents:
- DNase-seq: point counts, exponential output transform, and count likelihood.
- CAGE: 5′ ends with a longer sequence context and a Puffin-D-like architecture.
- Methylation: methylated and unmethylated channels, sigmoid output, and masked Bernoulli loss.
6. Develop the smallest adaptation with the Skill
Begin with one small, inspectable region and a short smoke test. Prefer a new config over new code, and a new strategy over edits to the generic trainer. Change one modeling assumption at a time.
7. Validate before scaling up
- Check shapes, strand layout, coordinates, count totals, and reference-genome alignment.
- Confirm that the read partitions and held-out chromosomes are disjoint.
- Compare SURF with the observed input against the same held-out target.
- Run sequence-only and data-only ablations to show where the gain comes from.
- Inspect a signal-rich locus and a quiet locus at broad and fine scales.
- Test the coverage levels, cell types, or resolutions required by the intended claim.
8. Train, predict, and preserve provenance
After the controls pass, freeze the split and config, train the full model with uv run surf train -c path/to/conf.json, and apply the selected checkpoint with the generic prediction command. If the new modality changes its output transform or layout, invoke skills/predict-surf-model/ before inference.
Keep the data version, preprocessing, chromosome split, config, seed, checkpoint, training curves, held-out metrics, and prediction track together. That bundle is the smallest reproducible model result.
9. Calibrate the result with SCAN
Use a half-coverage SURF run to produce the prediction and held-out target tracks required by SCAN. Install the SCAN R dependencies, fit its Negative Binomial spline model on chr8, and evaluate empirical coverage on a different chromosome—chr10 in the repository notebooks.
Rscript scan_calibration.R --pred_bw /path/to/half.pred.bw --tar_bw /path/to/half.tar.bw --chrom chr8 --out_dir out/toy --prefix toyPreserve both the latent-mean and posterior-predictive intervals. The SCAN Agent Skill should guide this same fit, diagnostic, held-out evaluation, and export sequence when it is present in the repository.
Before calling the model ready
- The DNase reference run is reproducible in the locked environment.
- The new data are split before any operation that could leak one observation into both halves.
- The loss, resolution, context, activation, and metrics match the assay.
- Held-out comparisons show where SURF improves over the observed input—and where it does not.
- SCAN is fit and evaluated on different held-out genomic regions.
- The config, checkpoint, prediction, uncertainty output, and evaluation all trace to one run.
