Cromwell workflow report

rMAP-GWAS

Reproducible microbial GWAS from paired-end reads for one explicitly coded binary phenotype per run, including gene presence/absence GWAS, optional SNP marker GWAS, population-structure visualization, post-GWAS reference annotation, and ranked association reporting.

CromwellDockerizedDistance-corrected pyseerMulti-pathogen readyExplicit phenotype codingPhenotype interpretation guardrailsGenBank annotation rescueOptional SNP GWAS
Automated interpretation disclaimer: The brief summaries in this report are rule-based supportive guidance for navigation and structured review only. Final biological, clinical, or public-health interpretation should be validated by a qualified expert.
GWAS phenotype being tested: This run tests one explicitly configured binary phenotype only: MSSA_vs_MRSA. The model contrast is case (MRSA) coded as 1 versus control (MSSA) coded as 0. This report does not assume the contrast is antimicrobial resistance unless the configured phenotype or group labels indicate resistance/susceptibility. Interpret candidate markers only as associations with this configured phenotype.
Association scope: General association guardrail: this is a marker-phenotype GWAS for the configured binary contrast, not automatically an AMR analysis. Gene presence/absence is the primary marker set; optional SNP GWAS adds point-mutation markers. Review sequencing depth, assembly quality, population structure, sampling design, and phenotype definition before making biological or public-health claims.
Sample-size interpretation note: This run has 4 samples (2 cases (MRSA) and 2 controls (MSSA)). Interpret association strength in relation to cohort size, phenotype balance, population structure, and independent validation before making biological, clinical, or public-health claims.
cases (MRSA)
2
controls (MSSA)
2
Total samples
4
Top gene hit
No significant gene hit
No significant gene hit at selected threshold

01 Phenotype legend and coding

Brief interpretation: This section confirms the exact binary contrast used by pyseer: case (MRSA) is coded as 1 and control (MSSA) is coded as 0. Confirm that these labels match the study metadata before interpreting associations.
Phenotype tested: Binary phenotype MSSA_vs_MRSA. Configured contrast: case (MRSA) versus control (MSSA). The report displays the metadata-derived biological contrast throughout the HTML: case (MRSA) = 1 and control (MSSA) = 0.
ClassPyseer codingMetadata columnMetadata value used for displaySamplesReport label
case1MSSA_vs_MRSAMRSA2case (MRSA)
control0MSSA_vs_MRSAMSSA2control (MSSA)

01b Phenotype association scope

Brief interpretation: This section explains how to interpret marker associations for the configured phenotype contrast: case (MRSA) versus control (MSSA).
Configured phenotype: This run tests one explicitly configured binary phenotype only: MSSA_vs_MRSA. The model contrast is case (MRSA) coded as 1 versus control (MSSA) coded as 0. This report does not assume the contrast is antimicrobial resistance unless the configured phenotype or group labels indicate resistance/susceptibility. Interpret candidate markers only as associations with this configured phenotype.
Interpretation issueCurrent workflow/report behaviorRecommended interpretation
Phenotype clarityOne binary phenotype is tested per run using the phenotype TSV supplied to pyseer.Use phenotype_name and phenotype_display_values to show the real biological contrast, for example specimen_source, HIV_status, MRSA_vs_MSSA, or rpoB_marker_status.
Gene presence/absenceThe primary GWAS branch tests pangenome gene-cluster presence/absence with population-structure correction.Interpret hits as gene-cluster associations with the configured phenotype, not as resistance genes unless the phenotype is explicitly an AMR contrast.
Point mutationsPoint mutations are assessed only when the optional SNP GWAS branch is enabled.Use do_snp_gwas=true for marker-defined or mutation-driven contrasts, then interpret SNP hits against the configured case/control labels.
Sample metadata and sampling designThe report displays the metadata-derived case/control labels, but metadata columns are not automatically modeled as covariates.For phenotypes such as specimen source, HIV status, geography, host category, or marker carriage, check whether lineage, site, or sampling design explains the association.
Sequencing depth and assembly qualityThe workflow generates FASTQ/assembly QC outputs, but read depth and assembly quality are not automatically modeled as covariates.Filter low-depth or poor-quality samples and confirm that apparent gene absence is not caused by technical dropout.
Population structure and linkageMash distances/MDS and population-structure plots are generated to help assess lineage confounding.If the configured phenotype clusters by lineage or outbreak, treat hits as candidate associations requiring validation rather than causal determinants.

02 Workflow architecture

Brief interpretation: Reads are quality-trimmed, assembled, annotated with Prokka, summarized as a Panaroo pangenome, tested with pyseer, and then annotated against the selected reference package.
01
Input
Sample-set validation

Checks sample names, paired FASTQs, group labels, and case (MRSA)/control (MSSA) balance.

02
QC
fastp trimming

Generates cleaned reads plus QC summaries.

03
Assembly
Shovill

Builds de novo genome assemblies using safe Cromwell memory handling.

04
Annotation
Prokka + Panaroo

Creates GFF annotations and pangenome gene matrices. Prokka remains the default; Bakta can be enabled with use_bakta=true.

05
GWAS
Mash + pyseer

Runs population-structure-aware gene and optional SNP association testing.

06
Rescue
GenBank annotation

Maps prioritized pangenome/SNP markers to reference GenBank features where possible.

03 Run configuration

Brief interpretation: The report records the reference package, species label, annotation engine, GWAS mode, SNP branch, Gubbins status, and container backend so that the run can be audited and repeated.
Reference name
SAUREUS_2026_06
Species
Staphylococcus aureus
Reference Docker
gmboowa/rmap-gwas-saureus-refs:2026.06
Generated UTC
2026-07-12T01:09:09Z
Phenotype coding: case (MRSA) = 1; control (MSSA) = 0. GWAS mode: gene_presence_absence; SNP branch: run; Gubbins recombination module: not run; annotation engine: Prokka; container backend recorded as docker.

04 Top-hit GenBank annotation rescue

Brief interpretation: The top gene hit is a candidate association for the case (MRSA) versus control (MSSA) contrast. Treat it as a hypothesis until supported by sample size, population-structure checks, annotation confidence, and independent validation.
Display name
No significant gene hit
Pangenome/SNP marker
NA
Reference locus / gene
not matched / not assigned
Confidence
none
Top-hit interpretation: Displayed from available pangenome/reference annotation; inspect manually. Reference identity: NA%; reference coverage: NA%. Product: not assigned.

05 Population structure

Brief interpretation: The PCoA plot shows moderate separation between the two phenotype groups (PCoA separation score 0.90; between/within Mash distance ratio 0.86). Some GWAS signals may still be influenced by lineage, so prioritize hits that remain biologically plausible and are not explained only by clustering.
Interpretation check: Review case (MRSA)/control (MSSA) clustering before interpreting top hits. Strong phenotype-lineage clustering can indicate lineage-associated markers rather than causal phenotype-associated variation.
Lay interpretation — PCoA: The PCoA plot shows moderate separation between the two phenotype groups (PCoA separation score 0.90; between/within Mash distance ratio 0.86). Some GWAS signals may still be influenced by lineage, so prioritize hits that remain biologically plausible and are not explained only by clustering.
Population structure: Mash PCoA ERR11728856 (control (MSSA)) ERR11728866 (control (MSSA)) ERR11728974 (case (MRSA)) ERR11728986 (case (MRSA)) case (MRSA) control (MSSA) PCoA1 (67.5% variance) PCoA2 (25.8% variance)
Mash distance / kinship matrix ERR11728856 vs ERR11728856 ERR11728856 vs ERR11728866 ERR11728856 vs ERR11728974 ERR11728856 vs ERR11728986 ERR11728866 vs ERR11728856 ERR11728866 vs ERR11728866 ERR11728866 vs ERR11728974 ERR11728866 vs ERR11728986 ERR11728974 vs ERR11728856 ERR11728974 vs ERR11728866 ERR11728974 vs ERR11728974 ERR11728974 vs ERR11728986 ERR11728986 vs ERR11728856 ERR11728986 vs ERR11728866 ERR11728986 vs ERR11728974 ERR11728986 vs ERR11728986 case (MRSA) samples control (MSSA) samples
metric	value
phenotype	MSSA_vs_MRSA
case_label	case (MRSA) (MRSA)
control_label	control (MSSA) (MSSA)
samples	4
pcoa1_variance_percent	67.5023
pcoa2_variance_percent	25.7896
pcoa1_plus_pcoa2_variance_percent	93.2919
case_within_mean_mash_distance	0.126974
control_within_mean_mash_distance	0.186761
between_group_mean_mash_distance	0.135431125
between_within_mash_distance_ratio	0.8633
pcoa_centroid_separation_score	0.8972
method	PCoA from square Mash distance matrix plus distance heatmap
layman_pcoa_interpretation	The PCoA plot shows moderate separation between the two phenotype groups (PCoA separation score 0.90; between/within Mash distance ratio 0.86). Some GWAS signals may still be influenced by lineage, so prioritize hits that remain biologically plausible and are not explained only by clustering.

06 Gene presence/absence GWAS plots

Brief interpretation: Gene Manhattan: Each dot is a gene or gene cluster. Taller dots mean stronger evidence of difference between the two groups. In this run, none of the 9 tested gene features crossed the alpha=0.05 line, so there is no strong gene presence/absence signal at this threshold. This can happen with small cohorts, weak phenotype separation, or true absence of strong gene-level associations. Gene QQ: The QQ plot compares observed p-values with what would be expected if there were no real association signal. The points do not show a strong upward tail, which supports the interpretation that this run has no clear gene-level association at the chosen threshold.
Plot note: The gene association plot is a feature-index GWAS plot, not a full reference-coordinate Manhattan plot. Significant points are interpreted against the case (MRSA) versus control (MSSA) phenotype.
Lay interpretation — gene Manhattan: Each dot is a gene or gene cluster. Taller dots mean stronger evidence of difference between the two groups. In this run, none of the 9 tested gene features crossed the alpha=0.05 line, so there is no strong gene presence/absence signal at this threshold. This can happen with small cohorts, weak phenotype separation, or true absence of strong gene-level associations.
Lay interpretation — gene QQ: The QQ plot compares observed p-values with what would be expected if there were no real association signal. The points do not show a strong upward tail, which supports the interpretation that this run has no clear gene-level association at the chosen threshold.
Gene presence/absence GWAS Manhattan-style association plot 1.0 0.8 0.6 0.4 0.2 0.0 Feature index -log10(p-value)
Gene presence/absence GWAS QQ plot 1.0 0.8 0.6 0.4 0.2 0.0 Expected -log10(p-value) Observed -log10(p-value)
metric	value
pvalues_detected	9
points_drawn	9
plot_label	Gene presence/absence GWAS
significant_points_at_alpha	0
top_feature	perR
top_pvalue	0.217
top_minus_log10_pvalue	0.6635
qq_median_delta_observed_minus_expected	-0.3010
qq_tail_delta_observed_minus_expected	-0.5917
manhattan_type	feature_index_not_genomic_coordinate
qq_plot	generated
layman_manhattan_interpretation	Each dot is a gene or gene cluster. Taller dots mean stronger evidence of difference between the two groups. In this run, none of the 9 tested gene features crossed the alpha=0.05 line, so there is no strong gene presence/absence signal at this threshold. This can happen with small cohorts, weak phenotype separation, or true absence of strong gene-level associations.
layman_qq_interpretation	The QQ plot compares observed p-values with what would be expected if there were no real association signal. The points do not show a strong upward tail, which supports the interpretation that this run has no clear gene-level association at the chosen threshold.

06 SNP marker GWAS

Brief interpretation: SNP Manhattan: Each dot is a SNP marker. Taller dots mean stronger evidence that the marker differs between the two phenotype groups. In this run, none of the 8960 tested SNP markers crossed the alpha=0.05 line, so there is no strong SNP signal at this threshold. SNP QQ: The SNP QQ plot does not show a strong upward tail, supporting the interpretation that no clear SNP-level signal was detected at the chosen threshold.
SNP branch: run. SNP marker GWAS is intended for mutation-mediated phenotypes, including MTBC drug-resistance traits. Review case (MRSA)/control (MSSA) population structure before interpreting SNP hits. Optional Gubbins status: not run.
Lay interpretation — SNP Manhattan: Each dot is a SNP marker. Taller dots mean stronger evidence that the marker differs between the two phenotype groups. In this run, none of the 8960 tested SNP markers crossed the alpha=0.05 line, so there is no strong SNP signal at this threshold.
Lay interpretation — SNP QQ: The SNP QQ plot does not show a strong upward tail, supporting the interpretation that no clear SNP-level signal was detected at the chosen threshold.
SNP marker GWAS Manhattan plot 1.0 0.8 0.6 0.4 0.2 0.0 Reference genomic coordinate / marker order -log10(p-value)
SNP marker GWAS QQ plot 1.0 0.8 0.6 0.4 0.2 0.0 Expected -log10(p-value) Observed -log10(p-value)

Top priority SNP marker hits

metric	value
plot_label	SNP marker GWAS
pvalues_detected	8960
significant_points_at_alpha	0
top_marker	NC_007795.1_14196_A_G
top_position	14196
top_pvalue	0.217
top_minus_log10_pvalue	0.6635
qq_median_delta_observed_minus_expected	-0.3011
qq_tail_delta_observed_minus_expected	-3.5898
points_drawn	5000
manhattan_type	reference_coordinate_when_available
qq_plot	generated
layman_manhattan_interpretation	Each dot is a SNP marker. Taller dots mean stronger evidence that the marker differs between the two phenotype groups. In this run, none of the 8960 tested SNP markers crossed the alpha=0.05 line, so there is no strong SNP signal at this threshold.
layman_qq_interpretation	The SNP QQ plot does not show a strong upward tail, supporting the interpretation that no clear SNP-level signal was detected at the chosen threshold.

SNP hit table

No prioritized hits available.

rank feature id variant id feature type contig position ref alt gene name product reference locus tag reference gene reference product case (MRSA) alt case (MRSA) total case (MRSA) frequency control (MSSA) alt control (MSSA) total control (MSSA) frequency enriched in beta odds ratio odds ratio ci95 odds ratio ci95 lower odds ratio ci95 upper pyseer pvalue q value priority score annotation source reference location qual notes

Showing 0 of 0 rows.

metric	value
snp_gwas_status	prioritized
pyseer_snp_rows	8960
ranked_snp_features	8960
significant_snp_features	0
alpha	0.05

06b Optional Gubbins recombination assessment

Brief interpretation: The recombination panel documents whether Gubbins was requested, whether filtering succeeded, and whether the SNP GWAS used the filtered or original VCF.
Recommendation: Gubbins is optional and is most useful for recombining bacterial species or datasets where recombinant blocks may inflate SNP associations. It is especially relevant for organisms such as Klebsiella pneumoniae, Escherichia coli, Salmonella, Streptococcus pneumoniae, Neisseria spp., and selected Acinetobacter datasets. For MTBC and other highly clonal organisms, keep it off by default unless there is a specific recombination or lineage-confounding question.
Requested
not requested
Run status
not run
Input SNP records
0
Filtered SNP records
0
Filtering note: Gubbins not run. When Gubbins is enabled and succeeds, SNP GWAS uses the Gubbins-filtered VCF. When it is skipped, unavailable, or fails softly, the original Snippy SNP VCF is used and the reason is recorded below.

Recombination intervals: 0. Output files: rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_SNP_gubbins_summary.tsv, rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_SNP_gubbins.filtered_polymorphic_sites.fasta, rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_SNP_gubbins.recombination_predictions.gff, and rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_SNP_gubbins.log.

metric	value
gubbins_status	not_run
recommendation	Optional. Enable for recombining bacteria or strong lineage/recombination concerns; usually not required as a default for MTBC/clonal analyses.
usage_note	Gubbins not requested. Primary SNP GWAS used the unfiltered Snippy/core VCF when SNP GWAS was enabled.

07 Top priority gene presence/absence GWAS hits

Brief interpretation: The prioritized table ranks candidate gene presence/absence associations by statistical evidence, enrichment direction, effect size, and annotation support.

No prioritized hits available.

Prioritized hit table

rank feature id feature type display name display label gene name product reference locus tag reference gene reference product annotation confidence reference identity reference coverage interpretation note annotation evidence case (MRSA) present case (MRSA) total case (MRSA) frequency control (MSSA) present control (MSSA) total control (MSSA) frequency enriched in beta odds ratio odds ratio ci95 odds ratio ci95 lower odds ratio ci95 upper pyseer pvalue q value priority score annotation source reference match type reference location annotation note cluster member ids notes display product

Showing 0 of 0 rows.

08 All significant gene presence/absence hits

Brief interpretation: This section lists all significant gene presence/absence hits at the selected threshold; use it for broader review beyond the top candidates.
rank feature id feature type display name display label gene name product reference locus tag reference gene reference product annotation confidence reference identity reference coverage interpretation note annotation evidence case (MRSA) present case (MRSA) total case (MRSA) frequency control (MSSA) present control (MSSA) total control (MSSA) frequency enriched in beta odds ratio odds ratio ci95 odds ratio ci95 lower odds ratio ci95 upper pyseer pvalue q value priority score annotation source reference match type reference location annotation note cluster member ids notes display product

Showing 0 of 0 rows.

09 Reference annotation confidence guide

Brief interpretation: The confidence guide explains how reference annotation rescue should be interpreted. Low-confidence or unmatched markers should retain their stable Panaroo/SNP identifiers.
Annotation caveat: Reference annotation rescue is intended to improve interpretability of pangenome/SNP markers. Low-confidence matches should not be treated as definitive gene calls. Candidate hits should be validated in independent datasets and interpreted in the context of case (MRSA) versus control (MSSA).
ConfidenceRuleInterpretation
high>=95% identity and >=90% coverage, or exact qualifier-level supportStrong reference-supported annotation
medium>=85% identity and >=70% coveragePlausible annotation; inspect manually
low>=60% identity and >=50% coverage, or weak/partial supportTentative annotation only
noneNo usable GenBank matchKeep the pangenome/SNP marker identifier

Reference annotation summary

metric	value
reference_cds_parsed	2767
reference_parse_warning	
panaroo_clusters_parsed	9
annotation_engine	Panaroo/Prokka
fasta_records_parsed	300
top_priority_input	rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_top_priority_hits.tsv
all_significant_input	rMAP_GWAS_MRSA_MSSA_4_tiny_lowram_local_all_significant_hits.tsv
method	GenBank qualifier matching plus pure-Python nucleotide similarity rescue when Panaroo representative sequences are available

10 Input validation

Brief interpretation: The validation report checks input consistency, sample counts, paired FASTQ arrays, and phenotype coding before downstream interpretation.
rMAP-GWAS sample-set input validation report
================================================
Cases (MRSA): 2
Controls (MSSA): 2
Total samples: 4
Case (MRSA) label: case (MRSA)
Control (MSSA) label: control (MSSA)
Phenotype display metadata values provided: yes
Unique phenotype display values: MRSA, MSSA

Warnings:
- Low case (MRSA) count: 2. Microbial GWAS may be underpowered.
- Low control (MSSA) count: 2. Microbial GWAS may be underpowered.
- Recommended minimum for stable microbial GWAS is often >=50 cases (MRSA) and >=50 controls (MSSA).

Status: PASS

11 Panaroo summary

Brief interpretation: The Panaroo summary describes the pangenome graph and gene matrix generated from Prokka-derived GFF files.
Panaroo output files:
panaroo_out/combined_DNA_CDS.fasta
panaroo_out/combined_protein_CDS.fasta
panaroo_out/combined_protein_cdhit_out.txt
panaroo_out/combined_protein_cdhit_out.txt.clstr
panaroo_out/final_graph.gml
panaroo_out/gene_data.csv
panaroo_out/gene_presence_absence.Rtab
panaroo_out/gene_presence_absence.csv
panaroo_out/gene_presence_absence_roary.csv
panaroo_out/pan_genome_reference.fa
panaroo_out/pre_filt_graph.gml
panaroo_out/struct_presence_absence.Rtab
panaroo_out/summary_statistics.txt

12 Interpretation guidance

Brief interpretation: These notes summarize major microbial GWAS caveats, including population structure, recombination, small sample size, and phenotype misclassification.
Microbial GWAS can be confounded by lineage structure, outbreak clustering, recombination, phenotype misclassification, small sample size, species background, variable sequencing depth/assembly quality, point mutations outside gene presence/absence, and co-carriage of multiple resistance genes in the same antibiotic class. Use the population-structure plots, QC outputs, optional SNP GWAS, and, for recombining species, consider enabling Gubbins to reduce recombinant-block-driven SNP associations. Interpret candidate markers as associations with case (MRSA) versus control (MSSA), not as proven causal determinants or standalone clinical resistance calls.

13 Key output files

Brief interpretation: This section records the key files needed for review, reruns, sharing, and downstream manuscript or supplementary reporting.